Reinforcement Learning in Action: From Foundations to Frontier AI
暫譯: 強化學習實戰:從基礎到前沿 AI

Kamath, Uday, Vajre, Vedant

  • 出版商: CRC
  • 出版日期: 2026-10-14
  • 售價: $5,550
  • 貴賓價: 9.5 折 $5,272
  • 語言: 英文
  • 頁數: 292
  • 裝訂: Hardcover - also called cloth, retail trade, or trade
  • ISBN: 1041131410
  • ISBN-13: 9781041131410
  • 相關分類: Reinforcement、Machine Learning、Large language model
  • 尚未上市,無法訂購

相關主題

商品描述

Reinforcement learning (RL) has become the engine behind some of the most significant advances in modern artificial intelligence, from defeating world champions in Go to aligning large language models with human preferences. Yet despite its central role, RL remains poorly understood by many practitioners who work with these systems daily. Reinforcement Learning in Action: From Foundations to Frontiers bridges the gap between classical RL theory and the cutting-edge techniques driving today's AI breakthroughs. The book traces a complete path from Markov Decision Processes and Bellman equations through deep RL methods (DQN, REINFORCE, Actor-Critic, PPO) to the modern landscape of LLM alignment (RLHF, DPO, SimPO, KTO), reasoning optimization (GRPO, VinePPO, MCTS), and agentic systems with tool use, memory, and multi-turn planning. A distinguishing feature is the book's consistent five-layer pedagogical structure: each algorithm is presented with its key characteristics, a full mathematical derivation, an honest assessment of its advantages and limitations, a complete from-scratch Python/PyTorch implementation in which variable names match the equations, and a hands-on case study with reproducible experiments. Case studies progress from Grid World navigation and CartPole control to fine-tuning language models with DPO on the HuggingFace ecosystem, training reasoning models with GRPO on mathematical benchmarks, and building a full agentic customer support system. Written for ML engineers, researchers, and advanced students, this book provides both the conceptual depth and implementation fluency needed to understand, build, and extend the RL systems shaping the future of AI.

商品描述(中文翻譯)

強化學習(Reinforcement Learning,RL)已成為現代人工智慧最重要進展背後的核心引擎之一,從擊敗圍棋世界冠軍,到讓大型語言模型符合人類偏好,皆可見其身影。然而,儘管 RL 扮演著核心角色,許多每天使用這些系統的實務工作者,對 RL 的理解仍相當有限。

《Reinforcement Learning in Action: From Foundations to Frontiers》旨在彌合理論上的經典 RL 與推動當今 AI 突破性進展的前沿技術之間的落差。本書完整介紹從馬可夫決策過程(Markov Decision Processes)與 Bellman 方程式,到深度 RL 方法(DQN、REINFORCE、Actor-Critic、PPO),再延伸至現代大型語言模型對齊(LLM alignment)的技術版圖(RLHF、DPO、SimPO、KTO)、推理最佳化(GRPO、VinePPO、MCTS),以及具備工具使用、記憶與多回合規劃能力的代理系統(agentic systems)。

本書的一大特色,是始終採用一致的五層次教學架構:首先說明每個演算法的關鍵特性;接著提供完整的數學推導;誠實評估其優勢與限制;以從零開始的 Python/PyTorch 實作完整呈現演算法,其中變數名稱與數學方程式保持一致;最後則透過可重現實驗的實作案例,協助讀者深入理解。

案例研究的內容,從 Grid World 導航與 CartPole 控制開始,逐步進展至在 HuggingFace 生態系統中使用 DPO 微調語言模型、在數學基準測試上使用 GRPO 訓練推理模型,以及建構完整的代理式客戶支援系統。

本書專為 ML 工程師、研究人員與進階學生撰寫,兼具足夠的概念深度與實作熟練度,協助讀者理解、建構並延伸正在塑造 AI 未來的 RL 系統。

作者簡介

Uday Kamath has over 25 years of experience in AI product development with a Ph.D. in scalable machine learning. His significant contributions span numerous journals, conferences, books, and patents. Notable books include Large Language Models: A Deep Dive, Applied Causal Inference, Explainable Artificial Intelligence, Transformers for Machine Learning, Deep Learning for NLP and Speech Recognition, Mastering Java Machine Learning, and Machine Learning: End-to-End Guide for Java Developers. Currently serving as the Chief Analytics Officer at Smarsh, he spearheads data science and research in communications AI for regulated industries. He is also an active member of the Board of Advisors for entities, including commercial companies and academic institutions.

Vedant Vajre is an aspiring AI researcher with a strong interest in reinforcement learning and intelligent decision-making systems. He is graduating from Penn State University and intends to pursue doctoral research in artificial intelligence. He has authored two peer-reviewed research publications with IEEE and continues to pursue active research in machine learning. Having worked at organizations including NASA, IBM, and early-stage startups, he has gained experience applying machine learning and AI in both research and production settings. Outside of research, he loves spending time with his two Shih Tzus (Pinot and Buzz), playing tennis, and solving Sudoku puzzles.

作者簡介(中文翻譯)

Uday Kamath 擁有超過 25 年的 AI 產品開發經驗,並取得可擴展機器學習(scalable machine learning)博士學位。他的重大貢獻遍及眾多期刊、會議、書籍與專利。著作包括《大型語言模型:深入探討》(Large Language Models: A Deep Dive)、《應用因果推論》(Applied Causal Inference)、《可解釋人工智慧》(Explainable Artificial Intelligence)、《用於機器學習的 Transformers》(Transformers for Machine Learning)、《NLP 與語音辨識的深度學習》(Deep Learning for NLP and Speech Recognition)、《精通 Java 機器學習》(Mastering Java Machine Learning),以及《機器學習:Java 開發人員的端到端指南》(Machine Learning: End-to-End Guide for Java Developers)。目前他擔任 Smarsh 的 Chief Analytics Officer,負責引領受監管產業之通訊 AI 領域的資料科學與研究工作。他也是多個組織之諮詢委員會(Board of Advisors)的活躍成員,這些組織包括商業公司與學術機構。

Vedant Vajre 是一名立志成為 AI 研究人員的研究者,對強化學習(reinforcement learning)與智慧決策系統(intelligent decision-making systems)抱有濃厚興趣。他即將畢業於 Penn State University,並計畫進一步攻讀人工智慧博士,從事相關研究。他曾與 IEEE 合作發表兩篇經同儕審查的研究論文,目前持續積極投入機器學習研究。他曾任職於 NASA、IBM 及早期階段的新創公司,因而累積了在研究與正式上線環境中應用機器學習與 AI 的經驗。除了研究之外,他也喜歡與兩隻西施犬(Pinot 和 Buzz)共度時光、打網球,以及解 Sudoku。