← Back to Main Site
YennJ12 Engineering Blog

Engineering insights, architecture deep dives, and technical solutions

Home Engineering Architecture Data All Posts FDE Series FDE Core Topics About Search
↑↓ navigate · ↵ open · esc close View all results →

#LLM-as-Judge

2 posts tagged "llm-as-judge"

Auto Agent System - Part 2 - Harness 引擎:多模型容錯、自我修正與 LLM 評審

深入 agent_auto_system 的心臟——Harness 引擎。從第一個 PR「解析被 markdown 包住的 JSON」開始,一路講到跨模型 fallback 重試(PR #3)、驗證失敗後的自我修正、獨立 LLM 評審打分,以及每次執行的 token/成本追蹤。這是把 LLM 這匹野馬套上挽具的完整工程。

Jul 4, 2026 ·22 min

ChatPDF RAG 優化(三):可觀測性與評估 —— Langfuse 追蹤、評估歷史、即時評分

沒有量測就沒有優化。本篇拆解 chatPDF 如何補上 RAG 的可觀測性最後一塊:opt-in 零開銷的 Langfuse 追蹤、執行緒安全的 singleton、評估歷史持久化、即時答案評分(faithfulness/relevance)、relevance gate,以及無外部依賴的 SVG 趨勢圖表。

Jun 30, 2026 ·16 min

About

  • About me
  • Blog home
  • All posts
  • Authors
  • GitHub
  • Contact

Categories

  • Engineering
  • AI & ML
  • DevOps
  • Cloud & AWS
  • Data Engineering
  • Tools & Productivity

Resources

  • All tags
  • Archives
  • Series
  • Documentation

Community

  • GitHub profile
  • Blog repository
  • Report an issue
  • RSS feed
English
Taipei

© 2026 YennJ12 Engineering Team. All rights reserved.

Built with Hugo