Back to Main Site
YennJ12 Engineering Blog

Engineering insights, architecture deep dives, and technical solutions

Home
Topics
AI & LLM 219 Building with language models — from prompt to production. Engineering 262 The craft: writing, shipping and fixing real systems. Architecture 29 Design decisions and the tradeoffs behind them. Cloud & Infrastructure 26 Where the code actually runs. Finance & Investing 41 Reading filings and pricing businesses. Business & Growth 28 Turning engineering into a business. Developer Tools 19 Sharpening the tools you use every day. Creative & Media 9 Shipping creative products, not just code.
All categories Full archive (349)
Series
FDE Interview Guide Forward-deployed engineering, end to end FDE Core Topics The recurring themes, deep-dived AI Engineering From zero to production AI systems 10-K Deep Dives Institutional-grade filing digests Browse all tags →
Archive
About
↑↓ navigate · ↵ open · esc close View all results →

#Serving

2 posts tagged "serving"

22

Part 22 — AI 工程從零開始|Phase 11 Part 1:LLM 推論工程 — 從實驗到每秒千次請求

深入解析 LLM 生產推論:vLLM PagedAttention、連續批次、投機解碼、量化(GPTQ/AWQ/INT4)、推論成本優化與 SLA 設計

Jun 21, 2026 ·23 min aiengineering
36

Part 36 — AI 工程從零開始|Phase 17 Part 1:AI 推論服務架構 — 從單機到全球部署

深入解析 AI 推論服務工程:模型服務器選型(Triton/TorchServe/vLLM)、負載均衡、自動擴縮容、GPU 共享與多租戶隔離架構

Jun 22, 2026 ·23 min aiengineering

About

  • About me
  • Blog home
  • All posts
  • Authors
  • GitHub
  • Contact

Topics

  • AI & LLM
  • Engineering
  • Architecture
  • Cloud & Infrastructure
  • Finance & Investing
  • Business & Growth
  • Developer Tools
  • Creative & Media
  • All topics

Resources

  • All tags
  • Archives
  • Calendar
  • FDE series
  • 10-K deep dives
  • Search

Community

  • GitHub profile
  • Blog repository
  • Report an issue
  • RSS feed
English
Taipei

© 2026 YennJ12 Engineering Team. All rights reserved.

Built with Hugo