Back to Main Site
YennJ12 Engineering Blog

Engineering insights, architecture deep dives, and technical solutions

Home
Topics
AI & LLM 219 Building with language models — from prompt to production. Engineering 262 The craft: writing, shipping and fixing real systems. Architecture 29 Design decisions and the tradeoffs behind them. Cloud & Infrastructure 26 Where the code actually runs. Finance & Investing 41 Reading filings and pricing businesses. Business & Growth 28 Turning engineering into a business. Developer Tools 19 Sharpening the tools you use every day. Creative & Media 9 Shipping creative products, not just code.
All categories Full archive (349)
Series
FDE Interview Guide Forward-deployed engineering, end to end FDE Core Topics The recurring themes, deep-dived AI Engineering From zero to production AI systems 10-K Deep Dives Institutional-grade filing digests Browse all tags →
Archive
About
↑↓ navigate · ↵ open · esc close View all results →

#Policy Gradient

1 posts tagged "policy gradient"

18

Part 18 — AI 工程從零開始|Phase 9:強化學習基礎 — RLHF 與遊戲 AI 的根基

深入解析強化學習工程原理:MDP/Q-Learning/Policy Gradient/PPO/RLHF,理解 ChatGPT 背後的對齊訓練機制

Jun 21, 2026 ·23 min aiengineering

About

  • About me
  • Blog home
  • All posts
  • Authors
  • GitHub
  • Contact

Topics

  • AI & LLM
  • Engineering
  • Architecture
  • Cloud & Infrastructure
  • Finance & Investing
  • Business & Growth
  • Developer Tools
  • Creative & Media
  • All topics

Resources

  • All tags
  • Archives
  • Calendar
  • FDE series
  • 10-K deep dives
  • Search

Community

  • GitHub profile
  • Blog repository
  • Report an issue
  • RSS feed
English
Taipei

© 2026 YennJ12 Engineering Team. All rights reserved.

Built with Hugo