Back to Main Site
YennJ12 Engineering Blog

Engineering insights, architecture deep dives, and technical solutions

Home
Topics
AI & LLM 219 Building with language models — from prompt to production. Engineering 262 The craft: writing, shipping and fixing real systems. Architecture 29 Design decisions and the tradeoffs behind them. Cloud & Infrastructure 26 Where the code actually runs. Finance & Investing 41 Reading filings and pricing businesses. Business & Growth 28 Turning engineering into a business. Developer Tools 19 Sharpening the tools you use every day. Creative & Media 9 Shipping creative products, not just code.
All categories Full archive (349)
Series
FDE Interview Guide Forward-deployed engineering, end to end FDE Core Topics The recurring themes, deep-dived AI Engineering From zero to production AI systems 10-K Deep Dives Institutional-grade filing digests Browse all tags →
Archive
About
↑↓ navigate · ↵ open · esc close View all results →

#CostOptimization

2 posts tagged "costoptimization"

17

Part 17 — FDE core topic - Context Cache Eviction:硬體級上下文快取驅逐策略與計費陷阱

深入解析 Vertex AI Context Caching 的 KV 快取原理、三層驅逐架構設計,以及如何避免每小時 $4.50 的隱性計費陷阱。

Jun 8, 2026 ·18 min engineering
18

Part 18 — FDE core topic - Semantic Model Routing:置信度熵值驅動的智能模型分流

深入解析如何以 Shannon 熵值即時偵測模型不確定性,動態將查詢路由至最便宜的可行模型,實現隱私保護與 74% 成本節省的生產架構。

Jun 8, 2026 ·18 min engineering

About

  • About me
  • Blog home
  • All posts
  • Authors
  • GitHub
  • Contact

Topics

  • AI & LLM
  • Engineering
  • Architecture
  • Cloud & Infrastructure
  • Finance & Investing
  • Business & Growth
  • Developer Tools
  • Creative & Media
  • All topics

Resources

  • All tags
  • Archives
  • Calendar
  • FDE series
  • 10-K deep dives
  • Search

Community

  • GitHub profile
  • Blog repository
  • Report an issue
  • RSS feed
English
Taipei

© 2026 YennJ12 Engineering Team. All rights reserved.

Built with Hugo