← Back to Main Site
YennJ12 Engineering Blog

Engineering insights, architecture deep dives, and technical solutions

Home Engineering Architecture Data All Posts FDE Series FDE Core Topics About Search
↑↓ navigate · ↵ open · esc close View all results →

#LLM-as-a-Judge

2 posts tagged "llm-as-a-judge"

2

Part 2 — 如何衡量 AI 的準確度(二):大型語言模型(LLM)的評估方法

LLM 的輸出沒有唯一標準答案,該怎麼客觀評估?本文介紹 BLEU、ROUGE、Perplexity、BERTScore 及 LLM-as-a-Judge 等方法,幫助你從多個維度評估語言模型的真實能力。

May 18, 2026 ·15 min

Langfuse 入門 Part 3 — LLM 評估:Score、LLM-as-a-Judge、Dataset 與 Experiment

LLM 應用最難的問題:你怎麼知道它『答得好不好』?本篇拆解 Langfuse 的評估體系——用 Score 量化品質、用 LLM-as-a-Judge 自動評分、用人工標註校準、再用 Dataset + Experiment 在上線前做回歸測試,把『我覺得改好了』變成『數據證明改好了』。

Jun 30, 2026 ·16 min

About

  • About me
  • Blog home
  • All posts
  • Authors
  • GitHub
  • Contact

Categories

  • Engineering
  • AI & ML
  • DevOps
  • Cloud & AWS
  • Data Engineering
  • Tools & Productivity

Resources

  • All tags
  • Archives
  • Series
  • Documentation

Community

  • GitHub profile
  • Blog repository
  • Report an issue
  • RSS feed
English
Taipei

© 2026 YennJ12 Engineering Team. All rights reserved.

Built with Hugo