AI Evals in Practice: Testing, reliability, and quality for LLM systems in production
暫譯: AI Evals 實務:生產環境中 LLM 系統的測試、可靠性與品質

Incau, Caio

相關主題

商品描述

Build evaluation systems that reveal whether your LLM applications are improving or regressing, using calibrated judges, RAG and agent metrics, CI regression tests, online evals, and cost-aware pipelines

Key Features:

- Build a complete Python eval harness from datasets and scorers to CI and dashboards

- Evaluate RAG, agents, and prompts with calibrated metrics and human feedback

- Move from offline testing to production evals, guardrails, and reliability workflows

Book Description:

LLM applications can look healthy in dashboards while their answers quietly become less accurate, less useful, or less reliable. AI Evals in Practice gives developers and AI engineers a systematic way to measure quality, catch regressions, and make evidence-based improvements before and after deployment.

You will build evalkit, a complete Python evaluation harness, while learning the core components of an eval: datasets, scorers, runners, and golden test sets. You will create deterministic scorers and LLM-as-judge evaluations, then calibrate judges and mitigate common biases. The book applies these foundations to prompt regression testing and CI, RAG retrieval and generation metrics, agent trajectories and tool calls, and human annotation workflows. You will then extend evaluation into production with online sampling, guardrails, and cost- and latency-aware pipelines, while comparing tools such as DeepEval, promptfoo, Langfuse, and Braintrust. A case study and final project bring the pieces together into an eval-driven development workflow with CI and a dashboard.

By the end, you will be able to design evaluation pipelines that help you ship LLM systems with measurable, repeatable quality.

What You Will Learn:

- Design eval datasets, scorers, runners, and golden test sets

- Build deterministic metrics for repeatable quality checks

- Calibrate LLM judges and reduce common evaluation bias

- Add prompt regression tests to continuous integration

- Measure retrieval and generation quality in RAG systems

- Evaluate agent trajectories, tool calls, and multi-turn behavior

- Run human annotation workflows and production online evals

- Control evaluation cost and latency without losing signal

Who this book is for:

This book is for AI engineers, LLM application developers, machine learning engineers, platform engineers, and technical leads who build or operate production systems using large language models. It is especially useful for teams working with prompts, RAG pipelines, agents, or AI features that need measurable quality and regression protection. Readers should be comfortable with Python and familiar with building or integrating LLM applications.

Table of Contents

- Why Evals Are the Missing Discipline

- Anatomy of an Eval: Datasets, Scorers, and Runners

- Building Golden Datasets

- Deterministic Scorers

- LLM-as-Judge: Design, Calibration, and Bias

- Prompt Regression Testing and CI Integration

- Evaluating RAG: Retrieval and Generation Metrics

- Evaluating Agents: Trajectories, Tool Calls, and Multi-Turn

- Human-in-the-Loop Annotation and Labeling Ops

- Production: Online Evals, Sampling, and Guardrails

- Cost and Latency of Eval Pipelines

- Tooling Landscape: DeepEval, promptfoo, Langfuse, Braintrust

- Eval-Driven Development Workflow

- Case Study: Taking a Flaky Agent to 99% Reliability

- Final Project: Complete Evalkit with CI and Dashboard

商品描述(中文翻譯)

建立評估系統,透過經過校準的評審、RAG 與 agent 指標、CI 回歸測試、線上評估,以及具成本意識的管線,判斷您的 LLM 應用程式是在進步還是退步

主要特色:

- 建立完整的 Python 評估工具組,涵蓋從資料集與評分器到 CI 和儀表板的各個環節
- 使用經過校準的指標與人工回饋,評估 RAG、agents 和 prompts
- 從離線測試邁向正式環境評估、護欄(guardrails)與可靠性工作流程

書籍簡介:

LLM 應用程式在儀表板上看似運作正常,但其回答的準確性、實用性或可靠性可能已悄悄降低。《AI Evals in Practice》提供開發人員與 AI 工程師一套系統化的方法,用來衡量品質、捕捉回歸問題,並在部署前後依據證據進行改善。

您將建置 evalkit,一套完整的 Python 評估工具組,同時學習評估流程的核心元件:資料集、評分器、執行器與黃金測試集(golden test sets)。您將建立確定性評分器與 LLM-as-judge 評估,接著校準評審並減輕常見偏誤。本書會將這些基礎應用於 prompt 回歸測試與 CI、RAG 的檢索與生成指標、agent 軌跡與工具呼叫,以及人工標註工作流程。接著,您將透過線上取樣、護欄(guardrails),以及考量成本與延遲的管線,將評估延伸至正式環境;同時比較 DeepEval、promptfoo、Langfuse 和 Braintrust 等工具。書中的案例研究與期末專案將各項內容整合成一套由評估驅動的開發工作流程,並包含 CI 與儀表板。

學完本書後,您將能夠設計評估管線,協助您以可衡量且可重複的品質交付 LLM 系統。

您將學會:

- 設計評估資料集、評分器、執行器與黃金測試集
- 建立確定性指標,以進行可重複的品質檢查
- 校準 LLM 評審並降低常見的評估偏誤
- 將 prompt 回歸測試加入持續整合(CI)
- 衡量 RAG 系統中的檢索與生成品質
- 評估 agent 軌跡、工具呼叫與多輪互動行為
- 執行人工標註工作流程與正式環境線上評估
- 在不失去評估訊號的情況下,控制評估成本與延遲

適合讀者:

本書適合建置或維運使用大型語言模型之正式環境系統的 AI 工程師、LLM 應用程式開發人員、機器學習工程師、平台工程師與技術主管。對於使用 prompts、RAG 管線、agents,或需要可衡量品質與回歸防護的 AI 功能團隊而言,本書尤其實用。讀者應熟悉 Python,並具備建置或整合 LLM 應用程式的經驗。

目錄:

- 為什麼評估是缺少的一環
- 評估的結構:資料集、評分器與執行器
- 建立黃金資料集
- 確定性評分器
- LLM-as-Judge:設計、校準與偏誤
- Prompt 回歸測試與 CI 整合
- 評估 RAG:檢索與生成指標
- 評估 agents:軌跡、工具呼叫與多輪互動
- 人機協作標註與標記作業
- 正式環境:線上評估、取樣與護欄
- 評估管線的成本與延遲
- 工具生態:DeepEval、promptfoo、Langfuse、Braintrust
- 評估驅動的開發工作流程
- 案例研究:將不穩定的 agent 提升至 99% 可靠性
- 期末專案:具備 CI 與儀表板的完整 Evalkit