Python Data Analysis - Fourth Edition: Master Python Analytics with Machine Learning, Deep Learning, GenAI, LLMs, and Data Engineering
暫譯: Python 數據分析(第四版):掌握 Python 分析技術,涵蓋機器學習、深度學習、生成式 AI、大型語言模型及數據工程
Navlani, Avinash, Wijaya, Cornellius Yudha
- 出版商: Packt Publishing
- 出版日期: 2026-06-26
- 售價: $1,780
- 貴賓價: 9.5 折 $1,691
- 語言: 英文
- 頁數: 766
- 裝訂: Quality Paper - also called trade paper
- ISBN: 1806022877
- ISBN-13: 9781806022878
-
相關分類:
Python、Machine Learning
海外代購書籍(需單獨結帳)
相關主題
商品描述
Understand data analysis pipelines using Python Data Analysis, machine learning, pandas, scikit-learn, and data visualization techniques. Build scalable workflows for time series, NLP, image analytics, and big data processing.
Key Features:
- Prepare, clean, and transform data with Python, pandas, and exploratory data analysis techniques
- Apply machine learning with Python using regression, classification, clustering, PCA, and Bayesian methods
- Scale analytics workflows using Dask, Ray, Modin, and PySpark
Book Description:
Modern data analysis goes beyond cleaning and visualizing data. Today's practitioners need to build scalable data pipelines, apply machine learning, work with text and image data, and understand emerging AI techniques such as Generative AI and Large Language Models (LLMs). This guide shows you how to tackle these challenges using Python's modern data ecosystem.
Unlike books focused on a single library or technique, this book provides an end-to-end approach to Python data analysis. You'll learn how to move from data preparation and exploratory analysis to machine learning, NLP, image analytics, scalable processing, and AI-powered workflows.
Starting with statistical foundations, you'll learn how to clean, transform, wrangle, and visualize data. You'll then explore time series analysis, signal processing, forecasting, and predictive analytics before applying machine learning techniques such as regression, classification, clustering, PCA, probabilistic methods, and Bayesian approaches.
The book also covers graph analytics, sentiment analysis, NLP, image analytics, Generative AI, and LLMs. Finally, you'll learn to scale analytics workflows using Dask, Modin, Ray, and PySpark.
By the end of the book, you'll be able to build end-to-end data analysis pipelines and apply modern data science and AI techniques to solve real-world challenges.
What You Will Learn:
- Prepare, clean, and transform data for exploratory data analysis and data wrangling
- Analyze and visualize data using Python and pandas
- Perform time series analysis, forecasting, and signal processing
- Apply machine learning with Python using scikit-learn techniques
- Use regression, classification, clustering, PCA, and Bayesian methods
- Perform sentiment analysis, NLP, graph analytics, and image analytics
- Accelerate workflows using Dask, Modin, and Ray
- Build scalable big data analytics pipelines with PySpark
Who this book is for:
This book is for data analysts, data scientists, business analysts, statisticians, students, and academic professionals who want to strengthen their Python Data Analysis skills. It is ideal for readers looking to apply data science with Python to real-world problems involving data preparation, visualization, machine learning, NLP, image analytics, and big data processing. A basic understanding of mathematics and working knowledge of Python will help you get the most from this book.
Table of Contents
- Getting Started with Python Libraries
- NumPy and Pandas
- Statistics for Data Insights
- Linear Algebra
- Data Visualization
- Retrieving, Processing, and Storing Data
- Cleaning Messy Data
- Time Series Analysis
- Supervised Learning: Regression and Classification
- Unsupervised Learning: Dimensionality Reduction, Clustering, Anomaly Detection
- Ensemble Methods: Bagging and Boosting Methods
- Artificial Neural Networks and Deep Learning
- Analyzing Text Data
- Analyzing Image Data
- LLMs and Gen AI
- Parallel Computing Using Dask, Modin, and Ray
- Big Data Analytics using PySpark
商品描述(中文翻譯)
了解使用 Python 數據分析、機器學習、pandas、scikit-learn 和數據可視化技術的數據分析管道。構建可擴展的工作流程以處理時間序列、自然語言處理 (NLP)、圖像分析和大數據處理。
主要特點:
- 使用 Python、pandas 和探索性數據分析技術準備、清理和轉換數據
- 使用 Python 應用機器學習,包括回歸、分類、聚類、主成分分析 (PCA) 和貝葉斯方法
- 使用 Dask、Ray、Modin 和 PySpark 擴展分析工作流程
書籍描述:
現代數據分析不僅僅是清理和可視化數據。當今的從業者需要構建可擴展的數據管道,應用機器學習,處理文本和圖像數據,並理解新興的人工智慧技術,如生成式人工智慧和大型語言模型 (LLMs)。本指南將向您展示如何使用 Python 的現代數據生態系統來應對這些挑戰。
與專注於單一庫或技術的書籍不同,本書提供了 Python 數據分析的端到端方法。您將學習如何從數據準備和探索性分析轉向機器學習、NLP、圖像分析、可擴展處理和人工智慧驅動的工作流程。
從統計基礎開始,您將學習如何清理、轉換、處理和可視化數據。然後,您將探索時間序列分析、信號處理、預測和預測分析,然後應用機器學習技術,如回歸、分類、聚類、PCA、概率方法和貝葉斯方法。
本書還涵蓋圖形分析、情感分析、NLP、圖像分析、生成式人工智慧和大型語言模型。最後,您將學習如何使用 Dask、Modin、Ray 和 PySpark 擴展分析工作流程。
到本書結束時,您將能夠構建端到端的數據分析管道,並應用現代數據科學和人工智慧技術來解決現實世界的挑戰。
您將學到的內容:
- 準備、清理和轉換數據以進行探索性數據分析和數據處理
- 使用 Python 和 pandas 分析和可視化數據
- 執行時間序列分析、預測和信號處理
- 使用 scikit-learn 技術應用 Python 機器學習
- 使用回歸、分類、聚類、PCA 和貝葉斯方法
- 執行情感分析、NLP、圖形分析和圖像分析
- 使用 Dask、Modin 和 Ray 加速工作流程
- 使用 PySpark 構建可擴展的大數據分析管道
本書適合誰:
本書適合數據分析師、數據科學家、商業分析師、統計學家、學生和希望加強其 Python 數據分析技能的學術專業人士。它非常適合希望將數據科學與 Python 應用於涉及數據準備、可視化、機器學習、NLP、圖像分析和大數據處理的現實問題的讀者。對數學的基本理解和 Python 的工作知識將幫助您充分利用本書。
目錄
- 開始使用 Python 庫
- NumPy 和 Pandas
- 數據洞察的統計學
- 線性代數
- 數據可視化
- 檢索、處理和存儲數據
- 清理雜亂數據
- 時間序列分析
- 監督學習:回歸和分類
- 非監督學習:降維、聚類、異常檢測
- 集成方法:袋裝和提升方法
- 人工神經網絡和深度學習
- 分析文本數據
- 分析圖像數據
- 大型語言模型和生成式人工智慧
- 使用 Dask、Modin 和 Ray 的並行計算
- 使用 PySpark 的大數據分析