The Orange Book of Machine Learning Green Edition: The essentials of making predictions using supervised regression and classification for tabular dat
暫譯: 機器學習橙皮書綠色版:使用監督式回歸和分類對表格數據進行預測的基本要素

Ellis, Carl McBride

  • 出版商: Packt Publishing
  • 出版日期: 2026-07-24
  • 售價: $1,580
  • 貴賓價: 9.5$1,501
  • 語言: 英文
  • 頁數: 238
  • 裝訂: Quality Paper - also called trade paper
  • ISBN: 1808081315
  • ISBN-13: 9781808081316
  • 相關分類: Machine Learning
  • 海外代購書籍(需單獨結帳)

相關主題

商品描述

Learn supervised machine learning for tabular data using Python, pandas, scikit-learn, CatBoost, LightGBM, XGBoost, TabPFN, and TabICL for regression, classification, and predictive modeling

Key Features:

- Explore, clean, and prepare tabular datasets for machine learning workflows

- Build regression and classification models using modern machine learning tools

- Improve predictions with calibration, conformal intervals, and optimization techniques

Book Description:

Master the essential tools and techniques for supervised machine learning on tabular data with this practical guide to regression and classification. Through clear explanations, code snippets, and hands-on notebooks, you'll learn how to use Python and leading machine learning libraries, including pandas, scikit-learn, CatBoost, LightGBM, XGBoost, TabPFN, and TabICL, to build predictive models for real-world datasets.

The book covers the complete workflow, from data exploration and cleaning to model development, evaluation, and optimization. You'll learn how to perform regression analysis for accurate point predictions and estimate uncertainty using conformal prediction intervals. For classification tasks, you'll explore probabilistic predictions and calibration techniques to improve model reliability. You'll also discover practical approaches to feature engineering, feature selection, and hyperparameter optimization to enhance model performance. In addition, the book introduces tabular foundation models and in-context learning techniques, providing insight into the latest advances in machine learning for structured data.

By the end of the book, you'll have the skills and confidence to develop, evaluate, and deploy supervised machine learning models for a wide range of tabular data applications.

What You Will Learn:

- Perform exploratory data analysis and data cleaning

- Apply cross-validation for reliable model evaluation

- Build regression models and prediction intervals

- Develop calibrated probabilistic classification models

- Optimize models through hyperparameter tuning

- Engineer and select features for improved performance

- Use ensemble learning methods effectively

- Explore tabular foundation models and in-context learning

Who this book is for:

This book is designed for motivated self-learners, university students studying applied machine learning, junior data scientists, and academic researchers looking to incorporate machine learning into their analytical workflows. Readers should have a basic familiarity with Python and data analysis concepts. Whether you're developing predictive models for business, research, or educational purposes, this book provides the practical guidance needed to apply modern machine learning techniques to structured and tabular datasets.

Table of Contents

- Introduction

- Statistics

- Exploratory data analysis (EDA)

- Data cleaning

- Cross-validation

- Interpolation and smoothing

- Regression

- Classification

- GLM and GAM

- Ensemble estimators

- Hyperparameter optimization (HPO)

- Feature engineering and selection

- Tabular foundation models (TFM)

商品描述(中文翻譯)

使用 Python、pandas、scikit-learn、CatBoost、LightGBM、XGBoost、TabPFN 和 TabICL 學習針對表格數據的監督式機器學習,應用於回歸、分類和預測建模

主要特點:
- 探索、清理並準備表格數據集以進行機器學習工作流程
- 使用現代機器學習工具構建回歸和分類模型
- 通過校準、符合區間和優化技術改善預測

書籍描述:
掌握針對表格數據的監督式機器學習的基本工具和技術,這本實用指南涵蓋回歸和分類。通過清晰的解釋、代碼片段和實作筆記本,您將學會如何使用 Python 和主要的機器學習庫,包括 pandas、scikit-learn、CatBoost、LightGBM、XGBoost、TabPFN 和 TabICL,為現實世界數據集構建預測模型。

本書涵蓋完整的工作流程,從數據探索和清理到模型開發、評估和優化。您將學會如何進行回歸分析以獲得準確的點預測,並使用符合預測區間來估計不確定性。對於分類任務,您將探索概率預測和校準技術以提高模型的可靠性。您還將發現實用的方法來進行特徵工程、特徵選擇和超參數優化,以提升模型性能。此外,本書介紹了表格基礎模型和上下文學習技術,提供對結構化數據機器學習最新進展的見解。

在書籍結束時,您將具備開發、評估和部署監督式機器學習模型的技能和信心,適用於各種表格數據應用。

您將學到的內容:
- 執行探索性數據分析和數據清理
- 應用交叉驗證以進行可靠的模型評估
- 構建回歸模型和預測區間
- 開發經過校準的概率分類模型
- 通過超參數調整優化模型
- 進行特徵工程和選擇以改善性能
- 有效使用集成學習方法
- 探索表格基礎模型和上下文學習

本書適合誰:
本書專為有動力的自學者、大學學習應用機器學習的學生、初級數據科學家和希望將機器學習納入其分析工作流程的學術研究者而設計。讀者應對 Python 和數據分析概念有基本的熟悉度。無論您是為商業、研究還是教育目的開發預測模型,本書提供了應用現代機器學習技術於結構化和表格數據集所需的實用指導。

目錄
- 介紹
- 統計
- 探索性數據分析 (EDA)
- 數據清理
- 交叉驗證
- 插值和平滑
- 回歸
- 分類
- GLM 和 GAM
- 集成估計器
- 超參數優化 (HPO)
- 特徵工程和選擇
- 表格基礎模型 (TFM)

最後瀏覽商品 (2)