An Introduction to Large Language Models
暫譯: 大型語言模型導論

Dhanith P. R., Joe, S, Geetha, Abdullah a., Sheik

相關主題

商品描述

The book offers an introduction to large language models (LLMs) that bridge foundational natural language processing (NLP) concepts with the advanced techniques underlying large language models. It offers a structured exploration of NLP evolution, from rule- based approaches to transformer architectures. Covering key principles such as tokenization, attention mechanisms, and model architectures (BERT, GPT, T5), the book explains pretraining objectives like masked and causal language modeling. It also addresses optimization techniques such as LoRA, pruning, and quantization for efficient LLM deployment. Multimodal models, including GPT- 4 and PaLM- E, are explored alongside retrieval- augmented generation and AI- powered agents.

Features:

- Discusses foundational NLP concepts, theoretical depth, advanced techniques, and real- world applications.

- Covers perplexity, BLEU, ROUGE, and datasets like SuperGLUE and SQuAD for assessing LLM performance, discusses LoRA, pruning, and quantization to optimize LLM deployment in resource- constrained settings.

- Explores GPT- 4, PaLM- E, and retrieval- augmented generation, expanding beyond traditional NLP models.

- Provides Python implementations for fine- tuning, classification, summarization, and conversational AI tasks.

- Highlights use cases in text generation, code generation, sentiment analysis, and multimodal AI.

This book is an invaluable textbook for students, researchers, and industry professionals seeking a deep technical understanding of LLMs and their applications.

商品描述(中文翻譯)

本書介紹大型語言模型(large language models,LLMs),將自然語言處理(natural language processing,NLP)的基礎概念,與大型語言模型背後的進階技術串聯起來。全書有系統地探討 NLP 的演進,從以規則為基礎的方法一路發展到 Transformer 架構。書中涵蓋 tokenization、attention mechanisms,以及 BERT、GPT、T5 等模型架構等重要原理,並說明 masked language modeling 與 causal language modeling 等預訓練目標。此外,也介紹 LoRA、pruning 與 quantization 等最佳化技術,以實現高效率的 LLM 部署。本書同時探討 GPT-4 與 PaLM-E 等多模態模型,以及 retrieval-augmented generation 和 AI-powered agents。

特色:

- 討論 NLP 的基礎概念、理論深度、進階技術與實際應用。
- 涵蓋 perplexity、BLEU、ROUGE,以及 SuperGLUE 和 SQuAD 等用於評估 LLM 效能的資料集;同時介紹 LoRA、pruning 與 quantization,協助在資源受限的環境中最佳化 LLM 部署。
- 探索 GPT-4、PaLM-E 與 retrieval-augmented generation,拓展傳統 NLP 模型以外的應用範疇。
- 提供 Python 實作,涵蓋 fine-tuning、分類、摘要與 conversational AI 等任務。
- 強調文字生成、程式碼生成、情感分析與多模態 AI 等使用案例。

對於希望深入理解 LLM 及其應用技術的學生、研究人員與業界專業人士而言,本書是不可或缺的教材。

作者簡介

Joe Dhanith P. R. is an Assistant Professor at the Vellore Institute of Technology (VIT), Chennai, with a strong academic and research background in the field of computer science. He has made notable contributions to high- impact journals, particularly in advancing the domains of natural language processing (NLP), web mining, and information retrieval. His research focuses on developing intelligent systems that enhance text understanding, knowledge extraction, and data- driven decision- making. Dr. Joe has been actively involved in mentoring students, guiding research projects, and delivering courses related to programming, machine learning, and deep learning. His work reflects a commitment to bridging theoretical foundations with practical applications, contributing to both academic research and real- world problem solving.

Geetha S. is a Professor and Associate Dean (Research) in School of Computer Science and Engineering, VIT University, Chennai Campus, India. She has received the B.E. and M.E. degrees in Computer Science and Engineering from Madurai Kamaraj University, India and Anna University of Chennai, India, and Ph.D. degree from Anna University, respectively. She has 23 years of rich teaching and research experience. She has published more than 100 papers in reputed International conferences and refereed journals. Her h- Index is 21 and i- 10 index is 46. Her research interests include image processing, deep learning, feature selection, steganography, steganalysis, multimedia security, intrusion detection systems, malware detection, machine learning paradigms, and information forensics. She has given many expert lectures and keynote addresses in international and national conferences. She has organized many workshops, conferences, and FDPs. She is a recipient of University Rank and Academic Topper Award in B.E. and M.E, respectively. She is also the proud recipient of ASDF Best Academic Researcher Award 2013, ASDF Best Professor Award 2014, Research Award 2016- 2020, High Performer Award 2016, from VIT University and DST - ISCA Best Poster Award 2018, Innovation Award 2022 - Cybersecurity Roadshow conducted in IISc.

Sheik Abdullah A. is working as an Associate Professor in the School of Computer Science Engineering, Vellore Institute of Technology, Chennai. He completed his Ph.D. from Anna University. He is a visiting faculty at The Institute of Mathematical Sciences IMSC Chennai and contributed and worked in computational biology, mathematical decision support models in clinical informatics, data analytics, and statistics. He is also a visiting researcher at Chennai Mathematical Institute (CMI) and is engaged in works corresponding to Timed automata and its applications. More recently, he has contributed his novelty in assessing risk factors that contribute to type II diabetes with swarm intelligence and machine learning approaches. His works correspond to real- time analysis of medical data and medical experts with the development of clinical decision support models for hospitals in rural areas. He also contributed his research intelligence in NLP, big data, knowledge- based systems, E- governance, learning analytics, and probabilistic planning algorithms. He has published over 45+ archival research papers to his credit, 15 book chapters, and a book. Being a Gold medalist (PG), he has been awarded the honorable chief minister award for the best project in E- governance.

作者簡介(中文翻譯)

Joe Dhanith P. R. 是 Vellore Institute of Technology (VIT) Chennai 分校的助理教授,在電腦科學領域具備深厚的學術與研究背景。他對高影響力期刊做出了重要貢獻,尤其致力於推動自然語言處理(Natural Language Processing, NLP)、網頁探勘(web mining)與資訊檢索(information retrieval)等領域的發展。他的研究重點是開發智慧型系統,以提升文字理解、知識擷取及資料驅動的決策能力。Joe 博士積極參與學生指導、研究計畫輔導,以及程式設計、機器學習與深度學習相關課程的教學。他的研究工作展現了將理論基礎與實務應用相結合的承諾,同時為學術研究及現實世界的問題解決貢獻心力。

Geetha S. 是印度 VIT University Chennai 校區電腦科學與工程學院的教授兼研究副院長(Associate Dean, Research)。她分別於印度 Madurai Kamaraj University 及 Chennai Anna University 取得電腦科學與工程學士(B.E.)及碩士(M.E.)學位,並取得 Anna University 的博士學位。她擁有 23 年豐富的教學與研究經驗,已在知名國際會議及同儕評審期刊發表超過 100 篇論文。她的 h-index 為 21,i10-index 為 46。

她的研究興趣包括影像處理、深度學習、特徵選擇、資訊隱藏(steganography)、隱寫分析(steganalysis)、多媒體安全、入侵偵測系統、惡意軟體偵測、機器學習典範,以及資訊鑑識。她曾在許多國際及國內會議發表專家演講與主題演講,也曾主辦多場工作坊、會議及 FDP(Faculty Development Programs,教師發展計畫)。她分別於 B.E. 及 M.E. 學業期間獲得大學排名及學業優異獎。她也是多項獎項的榮獲者,包括 VIT University 頒發的 ASDF Best Academic Researcher Award 2013、ASDF Best Professor Award 2014、Research Award 2016–2020 及 High Performer Award 2016,以及 DST-ISCA Best Poster Award 2018;此外,她也獲得在 IISc 舉辦的 Cybersecurity Roadshow 所頒發的 Innovation Award 2022。

Sheik Abdullah A. 目前任職於 Vellore Institute of Technology Chennai 分校電腦科學工程學院,擔任副教授。他於 Anna University 取得博士學位,並擔任 The Institute of Mathematical Sciences(IMSc)Chennai 分校的客座教師。他曾參與計算生物學、臨床資訊學中的數學決策支援模型、資料分析與統計學等領域的研究與工作。他也是 Chennai Mathematical Institute(CMI)的客座研究員,目前從事時間自動機(Timed automata)及其應用相關研究。

近年來,他運用群體智慧(swarm intelligence)與機器學習方法,創新地評估導致第二型糖尿病的風險因素。他的研究涵蓋醫療資料的即時分析,以及為偏鄉醫院開發臨床決策支援模型,以協助醫療專業人員進行決策。他也在自然語言處理(NLP)、大數據、知識型系統、電子治理(E-governance)、學習分析及機率規劃演算法等領域做出研究貢獻。他已發表超過 45 篇典藏研究論文、15 篇書籍章節,並出版 1 本書。作為研究所金牌得主,他曾因在電子治理領域的最佳專案而獲頒榮譽首席部長獎。