Fundamentals of Deep Learning: Designing Next-Generation Machine Intelligence Algorithms

  • 11 Want to read
  • 1 Currently reading
  • 1 Have read
Preview

My Reading Lists:

Create a new list

  • 11 Want to read
  • 1 Currently reading
  • 1 Have read

Buy this book

Last edited by Drini
September 22, 2025 | History

Fundamentals of Deep Learning: Designing Next-Generation Machine Intelligence Algorithms

  • 11 Want to read
  • 1 Currently reading
  • 1 Have read

With the reinvigoration of neural networks in the 2000s, deep learning has become an extremely active area of research, one that’s paving the way for modern machine learning. In this practical book, author Nikhil Buduma provides examples and clear explanations to guide you through major concepts of this complicated field.

Companies such as Google, Microsoft, and Facebook are actively growing in-house deep-learning teams. For the rest of us, however, deep learning is still a pretty complex and difficult subject to grasp. If you’re familiar with Python, and have a background in calculus, along with a basic understanding of machine learning, this book will get you started.

Examine the foundations of machine learning and neural networks. Learn how to train feed-forward neural networks. Use TensorFlow to implement your first neural network. Manage problems that arise as you begin to make networks deeper. Build neural networks that analyze complex images. Perform effective dimensionality reduction using autoencoders. Dive deep into sequence analysis to examine language. Learn the fundamentals of reinforcement learning.

Publish Date
Publisher
O'Reilly Media
Pages
298

Buy this book

Book Details


Table of Contents

Preface
Page ix
1. The Neural Network
Page 1
Building Intelligent Machines
Page 1
The Limits of Traditional Computer Programs
Page 2
The Mechanics of Machine Learning
Page 3
The Neuron
Page 7
Expressing Linear Perceptrons as Neurons
Page 8
Feed-Forward Neural Networks
Page 9
Linear Neurons and Their Limitations
Page 12
Sigmoid, Tanh, and ReLU Neurons
Page 13
Softmax Output Layers
Page 15
Looking Forward
Page 15
2. Training Feed-Forward Neural Networks
Page 17
The Fast-Food Problem
Page 17
Gradient Descent
Page 19
The Delta Rule and Learning Rates
Page 21
Gradient Descent with Sigmoidal Neurons
Page 22
The Backpropagation Algorithm
Page 23
Stochastic and Minibatch Gradient Descent
Page 25
Test Sets, Validation Sets, and Overfitting
Page 27
Preventing Overfitting in Deep Neural Networks
Page 34
Summary
Page 37
3. Implementing Neural Networks in TensorFlow
Page 39
What Is TensorFlow?
Page 39
How Does TensorFlow Compare to Alternatives?
Page 40
Installing TensorFlow
Page 41
Creating and Manipulating TensorFlow Variables
Page 43
TensorFlow Operations
Page 45
Placeholder Tensors
Page 45
Sessions in TensorFlow
Page 46
Navigating Variable Scopes and Sharing Variables
Page 48
Managing Models over the CPU and GPU
Page 51
Specifying the Logistic Regression Model in TensorFlow
Page 52
Logging and Training the Logistic Regression Model
Page 55
Leveraging TensorBoard to Visualize Computation Graphs and Learning
Page 58
Building a Multilayer Model for MNIST in TensorFlow
Page 59
Summary
Page 62
4. Beyond Gradient Descent
Page 63
The Challenges with Gradient Descent
Page 63
Local Minima in the Error Surfaces of Deep Networks
Page 64
Model Identifiability
Page 65
How Pesky Are Spurious Local Minima in Deep Networks?
Page 66
Flat Regions in the Error Surface
Page 69
When the Gradient Points in the Wrong Direction
Page 71
Momentum-Based Optimization
Page 74
A Brief View of Second-Order Methods
Page 77
Learning Rate Adaptation
Page 78
AdaGrad — Accumulating Historical Gradients
Page 79
RMSProp — Exponentially Weighted Moving Average of Gradients
Page 80
Adam — Combining Momentum and RMSProp
Page 81
The Philosophy Behind Optimizer Selection
Page 83
Summary
Page 83
5. Convolutional Neural Networks
Page 85
Neurons in Human Vision
Page 85
The Shortcomings of Feature Selection
Page 86
Vanilla Deep Neural Networks Don't Scale
Page 89
Filters and Feature Maps
Page 90
Full Description of the Convolutional Layer
Page 95
Max Pooling
Page 98
Full Architectural Description of Convolution Networks
Page 99
Closing the Loop on MNIST with Convolutional Networks
Page 101
Image Preprocessing Pipelines Enable More Robust Models
Page 103
Accelerating Training with Batch Normalization
Page 104
Building a Convolutional Network for CIFAR-10
Page 107
Visualizing Learning in Convolutional Networks
Page 109
Leveraging Convolutional Filters to Replicate Artistic Styles
Page 113
Learning Convolutional Filters for Other Problem Domains
Page 114
Summary
Page 115
6. Embedding and Representation Learning
Page 117
Learning Lower-Dimensional Representations
Page 117
Principal Component Analysis
Page 118
Motivating the Autoencoder Architecture
Page 120
Implementing an Autoencoder in TensorFlow
Page 121
Denoising to Force Robust Representations
Page 134
Sparsity in Autoencoders
Page 137
When Context Is More Informative than the Input Vector
Page 140
The Word2Vec Framework
Page 143
Implementing the Skip-Gram Architecture
Page 146
Summary
Page 152
7. Models for Sequence Analysis
Page 153
Analyzing Variable-Length Inputs
Page 153
Tackling seq2seq with Neural N-Grams
Page 155
Implementing a Part-of-Speech Tagger
Page 156
Dependency Parsing and SyntaxNet
Page 164
Beam Search and Global Normalization
Page 168
A Case for Stateful Deep Learning Models
Page 172
Recurrent Neural Networks
Page 173
The Challenges with Vanishing Gradients
Page 176
Long Short-Term Memory (LSTM) Units
Page 178
TensorFlow Primitives for RNN Models
Page 183
Implementing a Sentiment Analysis Model
Page 185
Solving seq2seq Tasks with Recurrent Neural Networks
Page 189
Augmenting Recurrent Networks with Attention
Page 191
Dissecting a Neural Translation Network
Page 194
Summary
Page 217
8. Memory Augmented Neural Networks
Page 219
Neural Turing Machines
Page 219
Attention-Based Memory Access
Page 221
NTM Memory Addressing Mechanisms
Page 223
Differentiable Neural Computers
Page 226
Interference-Free Writing in DNCs
Page 229
DNC Memory Reuse
Page 230
Temporal Linking of DNC Writes
Page 231
Understanding the DNC Read Head
Page 232
The DNC Controller Network
Page 232
Visualizing the DNC in Action
Page 234
Implementing the DNC in TensorFlow
Page 237
Teaching a DNC to Read and Comprehend
Page 242
Summary
Page 244
9. Deep Reinforcement Learning
Page 245
Deep Reinforcement Learning Masters Atari Games
Page 245
What Is Reinforcement Learning?
Page 247
Markov Decision Processes (MDP)
Page 248
Policy
Page 249
Future Return
Page 250
Discounted Future Return
Page 251
Explore Versus Exploit
Page 251
Policy Versus Value Learning
Page 253
Policy Learning via Policy Gradients
Page 254
Pole-Cart with Policy Gradients
Page 254
OpenAI Gym
Page 254
Creating an Agent
Page 255
Building the Model and Optimizer
Page 257
Sampling Actions
Page 257
Keeping Track of History
Page 257
Policy Gradient Main Function
Page 258
PGAgent Performance on Pole-Cart
Page 260
Q-Learning and Deep Q-Networks
Page 261
The Bellman Equation
Page 261
Issues with Value Iteration
Page 262
Approximating the Q-Function
Page 262
Deep Q-Network (DQN)
Page 263
Training DQN
Page 263
Learning Stability
Page 263
Target Q-Network
Page 264
Experience Replay
Page 264
From Q-Function to Policy
Page 264
DQN and the Markov Assumption
Page 265
DQN's Solution to the Markov Assumption
Page 265
Playing Breakout with DQN
Page 265
Building Our Architecture
Page 268
Stacking Frames
Page 268
Setting Up Training Operations
Page 268
Updating Our Target Q-Network
Page 269
Implementing Experience Replay
Page 269
DQN Main Loop
Page 270
DQNAgent Results on Breakout
Page 272
Improving and Moving Beyond DQN
Page 273
Deep Recurrent Q-Networks (DRQN)
Page 273
Asynchronous Advantage Actor-Critic Agent (A3C)
Page 274
Unsupervised Reinforcement and Auxiliary Learning (UNREAL)
Page 275
Summary
Page 276
Index
Page 277

Classifications

Library of Congress
Q325.5, TA347.A78, TA347.A78 B83 2017

Edition Identifiers

Open Library
OL26836319M
Internet Archive
fundamentalsofde0000budu
ISBN 10
1491925612
ISBN 13
9781491925614
OCLC/WorldCat
989166788, 992798385

Work Identifiers

Work ID
OL19545719W

Community Reviews (0)

No community reviews have been submitted for this work.

Lists

Download catalog record: RDF / JSON / OPDS | Wikipedia citation