← Back to Main Site
YennJ12 Engineering Blog

Engineering insights, architecture deep dives, and technical solutions

Home Engineering Architecture Data All Posts FDE Series FDE Core Topics About Search
↑↓ navigate · ↵ open · esc close View all results →

#Inference

2 posts tagged "inference"

16

Part 16 — FDE core topic - TTFT & Throughput Optimization:首字延遲與推理吞吐量的硬體級優化

深入解析 LLM 推理服務的兩大核心指標——首字時間(TTFT)與每秒 Token 吞吐量——以及 Quantization、Continuous Batching、PagedAttention、Speculative Decoding、Flash Attention 五大硬體級優化技術的原理與取捨。

Jun 8, 2026 ·18 min
22

Part 22 — AI 工程從零開始|Phase 11 Part 1:LLM 推論工程 — 從實驗到每秒千次請求

深入解析 LLM 生產推論:vLLM PagedAttention、連續批次、投機解碼、量化(GPTQ/AWQ/INT4)、推論成本優化與 SLA 設計

Jun 21, 2026 ·23 min

About

  • About me
  • Blog home
  • All posts
  • Authors
  • GitHub
  • Contact

Categories

  • Engineering
  • AI & ML
  • DevOps
  • Cloud & AWS
  • Data Engineering
  • Tools & Productivity

Resources

  • All tags
  • Archives
  • Series
  • Documentation

Community

  • GitHub profile
  • Blog repository
  • Report an issue
  • RSS feed
English
Taipei

© 2026 YennJ12 Engineering Team. All rights reserved.

Built with Hugo