← Back to Main Site
YennJ12 Engineering Blog

Engineering insights, architecture deep dives, and technical solutions

Home Engineering Architecture Data All Posts FDE Series FDE Core Topics About Search
↑↓ navigate · ↵ open · esc close View all results →

#VLLM

4 posts tagged "vllm"

2

Part 2 — Hugging Face 實戰(二):用模型、跑 App、推送自己的模型

如何從兩百萬個 repo 中挑對模型、用四個抽象層載入它、用 Gradio 與 FastAPI + vLLM 把它變成服務,並把自己的模型完整推上 Hub。含量化、Spaces 部署與 Model Card 撰寫範例。

Aug 14, 2026 ·26 min
5

Part 5 — Hugging Face 實戰(五):端到端實戰 — 打造一個完整的 RAG 客服系統

把前四篇串起來:用 bge-m3 + FAISS + reranker + 微調模型,從文件切分、混合檢索、引用生成、FastAPI 服務化到 Gradio 部署與線上評估,一套可直接執行的完整專案程式碼。

Aug 17, 2026 ·30 min
22

Part 22 — AI 工程從零開始|Phase 11 Part 1:LLM 推論工程 — 從實驗到每秒千次請求

深入解析 LLM 生產推論:vLLM PagedAttention、連續批次、投機解碼、量化(GPTQ/AWQ/INT4)、推論成本優化與 SLA 設計

Jun 21, 2026 ·23 min
47

Part 47 — FDE 面試準備指南(四十七):RKK 實戰——大模型與地端微型模型的智慧混合路由與冷啟動優化

深度拆解 Edge/On-Premise 小模型與雲端大模型的雙軌路由架構:基於 Token 概率熵值的早停路由(Early-Exit Confidence Routing)、vLLM logprobs API 整合、PII 強制本地路由、冷啟動優化策略,以及三個演進階段的完整系統設計

Jun 8, 2026 ·26 min

About

  • About me
  • Blog home
  • All posts
  • Authors
  • GitHub
  • Contact

Categories

  • Engineering
  • AI & ML
  • DevOps
  • Cloud & AWS
  • Data Engineering
  • Tools & Productivity

Resources

  • All tags
  • Archives
  • Series
  • Documentation

Community

  • GitHub profile
  • Blog repository
  • Report an issue
  • RSS feed
English
Taipei

© 2026 YennJ12 Engineering Team. All rights reserved.

Built with Hugo