Reinforcement Learning in Action: From Foundations to Frontier AI
暫譯: 實戰強化學習:從基礎到前沿 AI

Kamath, Uday, Vajre, Vedant

  • 出版商: CRC
  • 出版日期: 2026-10-14
  • 售價: $2,570
  • 貴賓價: 9.5 折 $2,441
  • 語言: 英文
  • 頁數: 292
  • 裝訂: Quality Paper - also called trade paper
  • ISBN: 1041130813
  • ISBN-13: 9781041130819
  • 相關分類: Reinforcement、Machine Learning、Python
  • 尚未上市,無法訂購

相關主題

商品描述

Reinforcement learning (RL) has become the engine behind some of the most significant advances in modern artificial intelligence, from defeating world champions in Go to aligning large language models with human preferences. Yet despite its central role, RL remains poorly understood by many practitioners who work with these systems daily. Reinforcement Learning in Action: From Foundations to Frontiers bridges the gap between classical RL theory and the cutting-edge techniques driving today's AI breakthroughs. The book traces a complete path from Markov Decision Processes and Bellman equations through deep RL methods (DQN, REINFORCE, Actor-Critic, PPO) to the modern landscape of LLM alignment (RLHF, DPO, SimPO, KTO), reasoning optimization (GRPO, VinePPO, MCTS), and agentic systems with tool use, memory, and multi-turn planning. A distinguishing feature is the book's consistent five-layer pedagogical structure: each algorithm is presented with its key characteristics, a full mathematical derivation, an honest assessment of its advantages and limitations, a complete from-scratch Python/PyTorch implementation in which variable names match the equations, and a hands-on case study with reproducible experiments. Case studies progress from Grid World navigation and CartPole control to fine-tuning language models with DPO on the HuggingFace ecosystem, training reasoning models with GRPO on mathematical benchmarks, and building a full agentic customer support system. Written for ML engineers, researchers, and advanced students, this book provides both the conceptual depth and implementation fluency needed to understand, build, and extend the RL systems shaping the future of AI.

商品描述(中文翻譯)

強化學習(Reinforcement Learning,RL)已成為現代人工智慧最重要進展背後的核心引擎之一,從擊敗圍棋世界冠軍,到讓大型語言模型符合人類偏好,皆可見其身影。然而,儘管 RL 扮演著核心角色,許多日常使用這些系統的實務工作者,對其仍缺乏充分理解。

《Reinforcement Learning in Action: From Foundations to Frontiers》旨在彌合理論上的經典 RL 與推動當今 AI 突破性進展的前沿技術之間的落差。本書完整介紹從馬可夫決策過程(Markov Decision Processes)與 Bellman 方程式,到深度 RL 方法(DQN、REINFORCE、Actor-Critic、PPO),再延伸至現代大型語言模型對齊(LLM alignment)的技術版圖(RLHF、DPO、SimPO、KTO)、推理最佳化(GRPO、VinePPO、MCTS),以及具備工具使用、記憶與多回合規劃能力的代理系統(agentic systems)。

本書的一大特色,是始終貫徹五層次的教學架構:每個演算法都會說明其關鍵特性、提供完整的數學推導、誠實評估其優勢與限制、以從零開始的 Python/PyTorch 實作完整呈現,並確保變數名稱與方程式一致,最後搭配可實際操作且能重現實驗結果的案例研究。

案例研究的內容循序漸進,從 Grid World 導航與 CartPole 控制,到運用 HuggingFace 生態系統以 DPO 微調語言模型、在數學基準測試上使用 GRPO 訓練推理模型,以及建構完整的代理式客戶支援系統。

本書專為 ML 工程師、研究人員與具備進階程度的學生撰寫,兼具理解、建構與延伸 RL 系統所需的概念深度與實作熟練度,協助讀者掌握正在形塑 AI 未來的強化學習技術。

作者簡介

Uday Kamath has over 25 years of experience in AI product development with a Ph.D. in scalable machine learning. His significant contributions span numerous journals, conferences, books, and patents. Notable books include Large Language Models: A Deep Dive, Applied Causal Inference, Explainable Artificial Intelligence, Transformers for Machine Learning, Deep Learning for NLP and Speech Recognition, Mastering Java Machine Learning, and Machine Learning: End-to-End Guide for Java Developers. Currently serving as the Chief Analytics Officer at Smarsh, he spearheads data science and research in communications AI for regulated industries. He is also an active member of the Board of Advisors for entities, including commercial companies and academic institutions.

Vedant Vajre is an aspiring AI researcher with a strong interest in reinforcement learning and intelligent decision-making systems. He is graduating from Penn State University and intends to pursue doctoral research in artificial intelligence. He has authored two peer-reviewed research publications with IEEE and continues to pursue active research in machine learning. Having worked at organizations including NASA, IBM, and early-stage startups, he has gained experience applying machine learning and AI in both research and production settings. Outside of research, he loves spending time with his two Shih Tzus (Pinot and Buzz), playing tennis, and solving Sudoku puzzles.

作者簡介(中文翻譯)

Uday Kamath 擁有超過 25 年的 AI 產品開發經驗,並取得可擴展機器學習(scalable machine learning)博士學位。他的重要貢獻涵蓋眾多期刊、研討會、書籍與專利。其著作包括《Large Language Models: A Deep Dive》、《Applied Causal Inference》、《Explainable Artificial Intelligence》、《Transformers for Machine Learning》、《Deep Learning for NLP and Speech Recognition》、《Mastering Java Machine Learning》,以及《Machine Learning: End-to-End Guide for Java Developers》。目前,他擔任 Smarsh 的 Chief Analytics Officer,負責引領受監管產業通訊 AI 領域的資料科學與研究工作。他也是多個組織之顧問委員會(Board of Advisors)的活躍成員,這些組織包括商業公司與學術機構。

Vedant Vajre 是一位志向遠大的 AI 研究人員,對強化學習(reinforcement learning)與智慧決策系統抱有濃厚興趣。他即將畢業於 Penn State University,並計畫進一步從事人工智慧博士研究。他曾與 IEEE 合作發表兩篇經同儕審查的研究論文,目前持續投入機器學習研究。他曾任職於 NASA、IBM 及處於早期階段的新創公司等組織,累積了在研究與正式生產環境中應用機器學習和 AI 的經驗。除了研究之外,他喜歡與自己的兩隻 Shih Tzu(Pinot 和 Buzz)共度時光、打網球,以及解 Sudoku。