← Back to Main Site
YennJ12 Engineering Blog

Engineering insights, architecture deep dives, and technical solutions

Home Engineering Architecture Data All Posts FDE Series FDE Core Topics About Search
↑↓ navigate · ↵ open · esc close View all results →

#Quantization

2 posts tagged "quantization"

22

Part 22 — AI 工程從零開始|Phase 11 Part 1:LLM 推論工程 — 從實驗到每秒千次請求

深入解析 LLM 生產推論:vLLM PagedAttention、連續批次、投機解碼、量化(GPTQ/AWQ/INT4)、推論成本優化與 SLA 設計

Jun 21, 2026 ·23 min

ollama on mac - part 2 - 公開模型全覽與選型指南

從硬體、任務到量化,一套可落地的 Ollama 開源模型選型心智模型:看懂 Llama / Qwen / Gemma / Mistral / Phi / DeepSeek 家族,選對尺寸與量化,讓你的 Mac 跑出最佳性價比。

Jul 16, 2026 ·20 min

About

  • About me
  • Blog home
  • All posts
  • Authors
  • GitHub
  • Contact

Categories

  • Engineering
  • AI & ML
  • DevOps
  • Cloud & AWS
  • Data Engineering
  • Tools & Productivity

Resources

  • All tags
  • Archives
  • Series
  • Documentation

Community

  • GitHub profile
  • Blog repository
  • Report an issue
  • RSS feed
English
Taipei

© 2026 YennJ12 Engineering Team. All rights reserved.

Built with Hugo