An edition of AI Engineering (2024)

AI Engineering

Building Applications with Foundation Models

  • 4.3 (3 ratings)
  • 88 Want to read
  • 2 Currently reading
  • 3 Have read
Locate

My Reading Lists:

Create a new list

  • 88 Want to read
  • 2 Currently reading
  • 3 Have read

Buy this book

Last edited by sentientai
November 14, 2025 | History
An edition of AI Engineering (2024)

AI Engineering

Building Applications with Foundation Models

  • 4.3 (3 ratings)
  • 88 Want to read
  • 2 Currently reading
  • 3 Have read

Recent breakthroughs in AI have not only increased demand for AI products, they've also lowered the barriers to entry for those who want to build AI products. The model-as-a-service approach has transformed AI from an esoteric discipline into a powerful development tool that anyone can use. Everyone, including those with minimal or no prior AI experience, can now leverage AI models to build applications. In this book, author Chip Huyen discusses AI engineering: the process of building applications with readily available foundation models.

The book starts with an overview of AI engineering, explaining how it differs from traditional ML engineering and discussing the new AI stack. The more AI is used, the more opportunities there are for catastrophic failures, and therefore, the more important evaluation becomes. This book discusses different approaches to evaluating open-ended models, including the rapidly growing AI-as-a-judge approach.

AI application developers will discover how to navigate the AI landscape, including models, datasets, evaluation benchmarks, and the seemingly infinite number of use cases and application patterns. You'll learn a framework for developing an AI application, starting with simple techniques and progressing toward more sophisticated methods, and discover how to efficiently deploy these applications.

  • Understand what AI engineering is and how it differs from traditional machine learning engineering
  • Learn the process for developing an AI application, the challenges at each step, and approaches to address them
  • Explore various model adaptation techniques, including prompt engineering, RAG, fine-tuning, agents, and dataset engineering, and understand how and why they work
  • Examine the bottlenecks for latency and cost when serving foundation models and learn how to overcome them
  • Choose the right model, dataset, evaluation benchmarks, and metrics for your needs

Chip Huyen works to accelerate data analytics on GPUs at Voltron Data. Previously, she was with Snorkel AI and NVIDIA, founded an AI infrastructure startup, and taught Machine Learning Systems Design at Stanford. She's the author of the book Designing Machine Learning Systems, an Amazon bestseller in AI.

AI Engineering builds upon and is complementary to Designing Machine Learning Systems (O'Reilly).

Publish Date
Language
English
Pages
350

Buy this book

Edition Availability
Cover of: Ingegneria Dell'IA
Ingegneria Dell'IA: Costruire Applicazioni con I Modelli Di Base
2026, O'Reilly Media, Incorporated
in Italian
Cover of: AI Engineering
AI Engineering: Building Applications with Foundation Models
2024, O'Reilly Media, Incorporated
in English

Add another edition?

Book Details


First Sentence

"If I could use only one word to describe AI post-2020, it’d be scale."

Table of Contents

Preface. 10
What This Book Is About. 12
What This Book Is Not. 16
Who This Book Is For. 18
Navigating This Book. 20
Conventions Used in This Book. 24
Using Code Examples. 26
O’Reilly Online Learning. 27
How to Contact Us. 28
Acknowledgments. 29
1. Introduction to Building AI Applications with Foundation Models. 34
The Rise of AI Engineering. 36
From Language Models to Large Language Models. 36
From Large Language Models to Foundation Models. 48
From Foundation Models to AI Engineering. 55
Foundation Model Use Cases. 63
Coding. 71
Image and Video Production. 74
Writing. 75
Education. 78
Conversational Bots. 80
Information Aggregation. 82
Data Organization. 83
Workflow Automation. 85
Planning AI Applications. 86
Use Case Evaluation. 86
Setting Expectations. 94
Milestone Planning. 95
Maintenance. 97
The AI Engineering Stack. 100
Three Layers of the AI Stack. 102
AI Engineering Versus ML Engineering. 106
AI Engineering Versus Full-Stack Engineering. 123
Summary. 124
2. Understanding Foundation Models. 130
Training Data. 132
Multilingual Models. 135
Domain-Specific Models. 144
Modeling. 148
Model Architecture. 148
Model Size. 165
Post-Training. 184
Supervised Finetuning. 189
Preference Finetuning. 195
Sampling. 205
Sampling Fundamentals. 205
Sampling Strategies. 209
Test Time Compute. 219
Structured Outputs. 226
The Probabilistic Nature of AI. 235
Summary. 248
3. Evaluation Methodology. 256
Challenges of Evaluating Foundation Models. 258
Understanding Language Modeling Metrics. 264
Entropy. 266
Cross Entropy. 268
Bits-per-Character and Bits-per-Byte. 269
Perplexity. 270
Perplexity Interpretation and Use Cases. 272
Exact Evaluation. 277
Functional Correctness. 278
Similarity Measurements Against Reference Data. 281
Introduction to Embedding. 292
AI as a Judge. 297
Why AI as a Judge?. 297
How to Use AI as a Judge. 300
Limitations of AI as a Judge. 306
What Models Can Act as Judges?. 315
Ranking Models with Comparative Evaluation. 320
Challenges of Comparative Evaluation. 327
The Future of Comparative Evaluation. 333
Summary. 335
4. Evaluate AI Systems. 340
Evaluation Criteria. 341
Domain-Specific Capability. 345
Generation Capability. 350
Instruction-Following Capability. 368
Cost and Latency. 381
Model Selection. 385
Model Selection Workflow. 386
Model Build Versus Buy. 389
Navigate Public Benchmarks. 411
Design Your Evaluation Pipeline. 428
Step 1. Evaluate All Components in a System. 428
Step 2. Create an Evaluation Guideline. 431
Step 3. Define Evaluation Methods and Data. 436
Summary. 446
5. Prompt Engineering. 453
Introduction to Prompting. 454
In-Context Learning: Zero-Shot and Few-Shot. 457
System Prompt and User Prompt. 461
Context Length and Context Efficiency. 467
Prompt Engineering Best Practices. 471
Write Clear and Explicit Instructions. 472
Provide Sufficient Context. 480
Break Complex Tasks into Simpler Subtasks. 482
Give the Model Time to Think. 488
Iterate on Your Prompts. 494
Evaluate Prompt Engineering Tools. 495
Organize and Version Prompts. 501
Defensive Prompt Engineering. 505
Proprietary Prompts and Reverse Prompt Engineering. 507
Jailbreaking and Prompt Injection. 511
Information Extraction. 521
Defenses Against Prompt Attacks. 528
Summary. 536
6. RAG and Agents. 541
RAG. 542
RAG Architecture. 546
Retrieval Algorithms. 548
Retrieval Optimization. 572
RAG Beyond Texts. 583
Agents. 588
Agent Overview. 590
Tools. 594
Planning. 602
Agent Failure Modes and Evaluation. 635
Memory. 641
Summary. 650
7. Finetuning. 655
Finetuning Overview. 657
When to Finetune. 663
Reasons to Finetune. 664
Reasons Not to Finetune. 666
Finetuning and RAG. 674
Memory Bottlenecks. 681
Backpropagation and Trainable Parameters. 684
Memory Math. 687
Numerical Representations. 692
Quantization. 697
Finetuning Techniques. 705
Parameter-Efficient Finetuning. 706
Model Merging and Multi-Task Finetuning. 734
Finetuning Tactics. 751
Summary. 760
8. Dataset Engineering. 768
Data Curation. 771
Data Quality. 778
Data Coverage. 782
Data Quantity. 787
Data Acquisition and Annotation. 796
Data Augmentation and Synthesis. 801
Why Data Synthesis. 803
Traditional Data Synthesis Techniques. 807
AI-Powered Data Synthesis. 814
Model Distillation. 832
Data Processing. 834
Inspect Data. 835
Deduplicate Data. 838
Clean and Filter Data. 842
Format Data. 844
Summary. 847
9. Inference Optimization. 852
Understanding Inference Optimization. 854
Inference Overview. 854
Inference Performance Metrics. 864
AI Accelerators. 878
Inference Optimization. 891
Model Optimization. 892
Inference Service Optimization. 919
Summary. 933
10. AI Engineering Architecture and User Feedback. 941
AI Engineering Architecture. 942
Step 1. Enhance Context. 944
Step 2. Put in Guardrails. 946
Step 3. Add Model Router and Gateway. 954
Step 4. Reduce Latency with Caches. 962
Step 5. Add Agent Patterns. 967
Monitoring and Observability. 970
AI Pipeline Orchestration. 984
User Feedback. 989
Extracting Conversational Feedback. 990
Feedback Design. 1003
Feedback Limitations. 1017
Summary. 1023
Epilogue. 1027
Index. 1029
About the Author. 1094

The Physical Object

Number of pages
350
Weight
0.666

Edition Identifiers

Open Library
OL54058212M
ISBN 13
9781098166304

Work Identifiers

Work ID
OL39671094W

Source records

Links outside Open Library

Community Reviews (0)

No community reviews have been submitted for this work.

Lists

Download catalog record: RDF / JSON / OPDS | Wikipedia citation