Back to Main Site
YennJ12 Engineering Blog

Engineering insights, architecture deep dives, and technical solutions

Home
Topics
AI & LLM 219 Building with language models — from prompt to production. Engineering 262 The craft: writing, shipping and fixing real systems. Architecture 29 Design decisions and the tradeoffs behind them. Cloud & Infrastructure 26 Where the code actually runs. Finance & Investing 41 Reading filings and pricing businesses. Business & Growth 28 Turning engineering into a business. Developer Tools 19 Sharpening the tools you use every day. Creative & Media 9 Shipping creative products, not just code.
All categories Full archive (349)
Series
FDE Interview Guide Forward-deployed engineering, end to end FDE Core Topics The recurring themes, deep-dived AI Engineering From zero to production AI systems 10-K Deep Dives Institutional-grade filing digests Browse all tags →
Archive
About
↑↓ navigate · ↵ open · esc close View all results →

#Prompt-Caching

2 posts tagged "prompt-caching"

4

Part 4 — OpenWorker 深度解析(四):LLM 層 — Provider 抽象、能力降級與 Context 自動壓縮

拆解 OpenWorker 如何同時支援 OpenAI、Anthropic、Gemini、Bedrock、Vertex、Ollama 與多家轉售商:ProviderClient 契約為何刻意同步且無迴圈、能力矩陣如何驅動 vision/PDF 降級、TokenUsage 的快取拆分,以及 561 行 compaction.py 的完整壓縮演算法。

Aug 7, 2026 ·27 min aiengineering

多 Agent Token 優化系列 pt.2:Prompt Caching 實戰 — 從記憶體快取到 RAG 系統

多 Agent Token 優化系列 pt.2:深入探索 Prompt Caching 的實際應用,從 Claude API 原生快取、應用層記憶體快取、到 RAG 系統整合,提供完整程式碼範例,幫助你打造高效低成本的 AI 應用。

Mar 12, 2026 ·30 min aitools

About

  • About me
  • Blog home
  • All posts
  • Authors
  • GitHub
  • Contact

Topics

  • AI & LLM
  • Engineering
  • Architecture
  • Cloud & Infrastructure
  • Finance & Investing
  • Business & Growth
  • Developer Tools
  • Creative & Media
  • All topics

Resources

  • All tags
  • Archives
  • Calendar
  • FDE series
  • 10-K deep dives
  • Search

Community

  • GitHub profile
  • Blog repository
  • Report an issue
  • RSS feed
English
Taipei

© 2026 YennJ12 Engineering Team. All rights reserved.

Built with Hugo