Reliable Evaluations for Llms and AI Agents: End-To-End Evaluation Frameworks for Llms and Autonomous AI Agents
暫譯: LLM 與 AI Agent 的可靠評估:LLM 與自主式 AI Agent 的端到端評估框架

Robsky, Alexei, Lavitas, Liliya, Wang, Yueqing

商品描述

This book gives practitioners a concrete, systematic framework for designing evals that make AI systems safe, robust, and customer-ready before they reach production. Drawing on real-world failures, from chatbots that went off the rails to shopping assistants that hallucinated product information, it shows how seemingly small evaluation gaps can cascade into legal, financial, and reputational crisis, and how to close those gaps with disciplined, systematic testing.

Moving from foundational concepts to advanced practice, Reliable Evals for LLMs and AI Agents introduces the four core levers of effective evals: sets, templates, metrics, and evaluators. It then extends these to the unique challenges of autonomous AI agents, where systems perceive, reason, act, and adapt in iterative loops that demand fundamentally different eval approaches. Along the way, it guides readers through benchmark selection, custom eval set design, statistical rigor in metrics, human and LLM-as-a-judge rating strategies, and the infrastructure needed to automate evals at scale.

For engineering leaders, applied researchers, data scientists, and product teams shipping LLM- and agent-powered experiences, this volume offers a blueprint for building eval flywheels that continuously improve AI quality. It shows how to progress from ad-hoc checks to production-grade eval systems, align model metrics with real user satisfaction, integrate offline evals with online A/B testing, and design accessible interfaces that democratize rigorous testing across an organization.

商品描述(中文翻譯)

本書為實務工作者提供一套具體且系統化的框架,用於設計評估(evals),讓 AI 系統在正式上線前具備安全性、穩健性,並符合客戶使用需求。本書取材自真實世界的失敗案例,從失控的聊天機器人,到捏造產品資訊的購物助理,說明看似微小的評估缺口如何層層擴大,最終演變成法律、財務與聲譽危機;同時也介紹如何透過嚴謹且系統化的測試,補上這些缺口。

《Reliable Evals for LLMs and AI Agents》從基礎概念一路延伸至進階實務,介紹有效評估的四個核心槓桿:評估集(sets)、範本(templates)、指標(metrics)與評估器(evaluators)。接著,本書進一步探討自主式 AI 代理程式所面臨的獨特挑戰。這類系統會以反覆迴圈的方式感知、推理、行動與適應,因此需要從根本上不同的評估方法。全書也將引導讀者了解基準測試的選擇、自訂評估集的設計、指標的統計嚴謹性、人工評分與 LLM-as-a-judge 評分策略,以及大規模自動化評估所需的基礎架構。

對於負責推出由 LLM 與 AI 代理程式驅動之體驗的工程主管、應用研究人員、資料科學家與產品團隊而言,本書提供了一套建立評估飛輪(eval flywheel)的藍圖,協助持續改善 AI 品質。本書說明如何從臨時性的檢查,逐步發展為適用於正式環境的評估系統;如何讓模型指標與真實使用者滿意度保持一致;如何將離線評估與線上 A/B 測試整合;以及如何設計易於使用的介面,讓組織中的各個團隊都能採用嚴謹的測試方法。

作者簡介

Alexei Robsky is a technology leader with over 15 years of experience building production AI and Machine-Learning systems. As AI Leader at Microsoft, he heads the Fabric Real-Time Intelligence AI team, delivering cloud-scale AI agents and solutions that act on live data streams.
Previously, Alexei led Google's Gemini Core Modeling and Evals Data Science Research organization, driving evaluation research and training-data quality initiatives that shaped the performance, reliability, and safety of Gemini models.
Earlier in his career, he was a Data Science Manager at Twitter, overseeing evaluation of Home Timeline ranking and personalization, and a Principal Data Science Manager at Microsoft, where he guided cross-functional teams that shipped production-level Azure customer-experience solutions.
Alexei co-authored Machine Learning Governance for Managers (Springer, 2024), holds an MBA from Duke University's Fuqua School of Business, and earned a B.Sc. in Electrical Engineering and Computer Science from Tel Aviv University.

Liliya Lavitas is an accomplished Data Science leader with a deep background in statistical modeling and machine learning. Liliiya currently leads a Data Science team in Google Search incorporating AI features in Google Search experience. Previously, under leadership of Alexei, she has been leading a Data Science team at Google DeepMind, responsible for the rigorous evaluation and training of core Gemini capabilities, ensuring their performance and reliability. Prior to her role at Google, Liliya managed Data Science teams at Netflix and Twitter. Liliya earned her Ph.D. in Statistics from Boston University in 2017, with a thesis specializing in Time Series analysis.

Yueqing Wang is a statistician specializing in novel methodology for system evaluation. In recent years, she has developed new techniques and frameworks for Gen AI evaluation for Google Gemini and at Microsoft AI. Earlier, she worked on paid product ecosystem at YouTube, statistical evaluations of startups at Google Ventures, and Media Mix Modeling at Google Ads, among other things. She earned her PhD in Statistics from the University of California, Berkeley in 2012 with Professor Bin Yu.

作者簡介(中文翻譯)

Alexei Robsky 是一位科技領導者,擁有超過 15 年建置正式環境 AI 與 Machine Learning 系統的經驗。Alexei 目前擔任 Microsoft 的 AI Leader,領導 Fabric Real-Time Intelligence AI 團隊,負責提供雲端規模的 AI agents 與解決方案,讓系統能夠根據即時資料串流採取行動。

此前,Alexei 領導 Google 的 Gemini Core Modeling and Evals Data Science Research 組織,推動評估研究與訓練資料品質計畫,這些工作形塑了 Gemini 模型的效能、可靠性與安全性。

在職涯早期,Alexei 曾任 Twitter 的 Data Science Manager,負責 Home Timeline 排名與個人化功能的評估;也曾任 Microsoft 的 Principal Data Science Manager,帶領跨職能團隊推出正式環境等級的 Azure 客戶體驗解決方案。

Alexei 與他人共同撰寫了《Machine Learning Governance for Managers》(Springer,2024),擁有 Duke University Fuqua School of Business 的 MBA 學位,以及 Tel Aviv University 電機工程與 Computer Science 學士學位。

Liliya Lavitas 是一位成就卓著的 Data Science 領導者,在統計建模與 Machine Learning 方面擁有深厚背景。Liliya 目前在 Google Search 領導 Data Science 團隊,負責將 AI 功能整合至 Google Search 體驗中。此前,在 Alexei 的領導下,她曾於 Google DeepMind 領導 Data Science 團隊,負責嚴謹評估與訓練 Gemini 的核心能力,確保其效能與可靠性。在加入 Google 之前,Liliya 曾管理 Netflix 與 Twitter 的 Data Science 團隊。Liliya 於 2017 年取得 Boston University 的統計學博士學位,博士論文專攻 Time Series 分析。

Yueqing Wang 是一位統計學家,專精於系統評估的新方法論。近年來,她為 Google Gemini 與 Microsoft AI 開發生成式 AI(Gen AI)評估的新技術與框架。更早之前,她曾參與 YouTube 的付費產品生態系、Google Ventures 的新創公司統計評估,以及 Google Ads 的 Media Mix Modeling 等工作。她於 2012 年在 Bin Yu 教授的指導下,取得 University of California, Berkeley 的統計學博士學位。

最後瀏覽商品 (1)