Back to Main Site
YennJ12 Engineering Blog

Engineering insights, architecture deep dives, and technical solutions

Home
Topics
AI & LLM 234 Building with language models — from prompt to production. Engineering 277 The craft: writing, shipping and fixing real systems. Architecture 34 Design decisions and the tradeoffs behind them. Cloud & Infrastructure 31 Where the code actually runs. Finance & Investing 41 Reading filings and pricing businesses. Business & Growth 28 Turning engineering into a business. Developer Tools 19 Sharpening the tools you use every day. Creative & Media 9 Shipping creative products, not just code.
All categories Full archive (364)
Series
FDE Interview Guide Forward-deployed engineering, end to end FDE Core Topics The recurring themes, deep-dived AI Engineering From zero to production AI systems 10-K Deep Dives Institutional-grade filing digests Browse all tags →
Archive
About
↑↓ navigate · ↵ open · esc close View all results →

#Speculative Decoding

1 posts tagged "speculative decoding"

3

Part 3 — vLLM Intro Part 3 — 連續批次與排程器 — 決定誰在這一輪前進一格

vLLM 原始碼導讀系列第三篇:拆解連續批次的實作、V1 統一 token 預算排程器如何取代 V0 的雙軌制、Chunked Prefill 在 TTFT 與 ITL 之間的取捨、搶佔與重算機制、投機解碼的接受率數學,以及一份可直接套用的調參手冊。

Sep 11, 2026 ·27 min aiengineeringarchitecture

About

  • About me
  • Blog home
  • All posts
  • Authors
  • GitHub
  • Contact

Topics

  • AI & LLM
  • Engineering
  • Architecture
  • Cloud & Infrastructure
  • Finance & Investing
  • Business & Growth
  • Developer Tools
  • Creative & Media
  • All topics

Resources

  • All tags
  • Archives
  • Calendar
  • FDE series
  • 10-K deep dives
  • Search

Community

  • GitHub profile
  • Blog repository
  • Report an issue
  • RSS feed
English
Taipei

© 2026 YennJ12 Engineering Team. All rights reserved.

Built with Hugo