1D convolution for neural networks, part 5: Backpropagation

Brandon Rohrer · Intermediate ·📐 ML Fundamentals ·6y ago

Key Takeaways

The video discusses backpropagation for 1D convolutional neural networks, covering the mathematical equations and differentiation required for training, with a focus on calculating the partial derivative of the loss with respect to the inputs.

Full Transcript

having equations for convolution is a great start they are specific enough that we can turn around and implement them in code the other thing that we have to be able to do in a neural network that uses back propagation to learn is we have to be able to differentiate these equations specifically with respect to the error gradient the lost gradient that's going to be passed back down through the layers if these concepts are unfamiliar there are some links in the text up above that you can use to get familiar with them for now I'm going to assume that they're not totally new and kind of jump in back propagation is how we take the sensitivity of the loss function to changes in each of our layers output values so this is the partial derivative of the loss with respect to the outputs Y and we want to be able to propagate that back and calculate the partial derivative of the loss with respect to the inputs X this is one link in our chain and we back propagate this sensitivity of the loss to each set of outputs and inputs all the way back through the network by the chain rule of calculus the way that we get the partial derivative of loss with respect to X is we take the partial derivative of loss with respect to Y and multiply it by the partial derivative of Y with respect to X now in our case x and y are arrays they're not single values they're a whole long line of values the list of values so really what we want to calculate is the partial derivative of the loss with respect to each one of the inputs individually and the way we calculate that is by the chain rule the partial derivative of the loss with respect to Y times the partial derivative of Y with respect to each input element and each of our outputs are so individual values not one element and so to get to really break this down to individual elements we have the partial of loss with respect to each input element is the sum of the partial of loss to each output element times the partial of that output element to the input element added up over all of the output elements so because it's so verbose we'll use as a shorthand for a partial derivative of loss with respect to the output we'll just call that the output gradient and the partial derivative of the loss with respect to the input will call the input gradient so to go from the output gradient to the input gradient we have to know this quantity here the partial derivative of each output element with respect to each input element that's what we mean when we say that the layer is differentiable we have to be able to calculate this bit to be able to do back propagation so in order to do that we can go back to our definition of convolution we can explode it back out and now say each y sub J is equal to X sub J minus P times W sub P etc etc we expand that summation out and we expand out for all of our potential elements J so this is just a longhand way to represent that summation sign and to represent the various elements J then we can take in each of these cases we can take the derivative of that Y sub J with respect to each of the X's that occur in it so in this case the very first element if we take the derivative of Y sub J with respect to X sub J minus P the derivative if J is W sub P the very next term the derivative of Y sub J with respect to X sub J minus P plus one is w sub P plus one and so what we get is a bunch of small partial derivatives of individual output values with respect to individual input values so this is starting to look really good this is the type of thing that we're looking for

Original Description

Part of an 9-part series on 1D convolution for neural networks. Catch the rest at https://e2eml.school/321
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Brandon Rohrer · Brandon Rohrer · 51 of 60

1 Robot Learning with a Biologically-Inspired Brain (BECCA)
Robot Learning with a Biologically-Inspired Brain (BECCA)
Brandon Rohrer
2 BECCA talk at AGI 2011
BECCA talk at AGI 2011
Brandon Rohrer
3 Robot Learning with a Biologically-Inspired Brain (BECCA), The Sequel
Robot Learning with a Biologically-Inspired Brain (BECCA), The Sequel
Brandon Rohrer
4 BECCA listens to The Hobbit
BECCA listens to The Hobbit
Brandon Rohrer
5 Learning the building blocks of speech: BECCA extracts a hierarchy of audio features
Learning the building blocks of speech: BECCA extracts a hierarchy of audio features
Brandon Rohrer
6 BECCA listens for sound effects in The Hobbit
BECCA listens for sound effects in The Hobbit
Brandon Rohrer
7 BECCA finds movie trailers while watching the Big Bang Theory
BECCA finds movie trailers while watching the Big Bang Theory
Brandon Rohrer
8 Listening for unexpected sounds: BECCA detects anomalies in audio data
Listening for unexpected sounds: BECCA detects anomalies in audio data
Brandon Rohrer
9 Learning the building blocks of vision: BECCA extracts a spatio-temporal hierarchy of features
Learning the building blocks of vision: BECCA extracts a spatio-temporal hierarchy of features
Brandon Rohrer
10 Watching for the unexpected: BECCA detects anomalies in video data
Watching for the unexpected: BECCA detects anomalies in video data
Brandon Rohrer
11 BECCA finds a stationary target
BECCA finds a stationary target
Brandon Rohrer
12 BECCA finds a stationary target at 3X speed
BECCA finds a stationary target at 3X speed
Brandon Rohrer
13 BECCA watches the X-men and Bruce Lee
BECCA watches the X-men and Bruce Lee
Brandon Rohrer
14 BECCA plays Quidditch
BECCA plays Quidditch
Brandon Rohrer
15 BECCA chases a ball
BECCA chases a ball
Brandon Rohrer
16 BECCA chases a ball, part 2
BECCA chases a ball, part 2
Brandon Rohrer
17 Becca chases a ball, part 3
Becca chases a ball, part 3
Brandon Rohrer
18 BECCA creates features from MNIST
BECCA creates features from MNIST
Brandon Rohrer
19 How reinforcement learning works in Becca 7
How reinforcement learning works in Becca 7
Brandon Rohrer
20 Deep Learning Demystified
Deep Learning Demystified
Brandon Rohrer
21 How Data Science Works
How Data Science Works
Brandon Rohrer
22 How Convolutional Neural Networks work
How Convolutional Neural Networks work
Brandon Rohrer
23 How Bayes Theorem works
How Bayes Theorem works
Brandon Rohrer
24 How Deep Neural Networks Work
How Deep Neural Networks Work
Brandon Rohrer
25 Recurrent Neural Networks (RNN) and Long Short-Term Memory (LSTM)
Recurrent Neural Networks (RNN) and Long Short-Term Memory (LSTM)
Brandon Rohrer
26 How Support Vector Machines work / How to open a black box
How Support Vector Machines work / How to open a black box
Brandon Rohrer
27 How autocorrelation works
How autocorrelation works
Brandon Rohrer
28 Getting closer to human intelligence through robotics
Getting closer to human intelligence through robotics
Brandon Rohrer
29 A minimalist's guide to slicing and indexing pandas DataFrames
A minimalist's guide to slicing and indexing pandas DataFrames
Brandon Rohrer
30 How decision trees work
How decision trees work
Brandon Rohrer
31 Data scientist archetypes
Data scientist archetypes
Brandon Rohrer
32 How to use python's datetime package
How to use python's datetime package
Brandon Rohrer
33 How optimization for machine learning works, part 1
How optimization for machine learning works, part 1
Brandon Rohrer
34 How optimization for machine learning works, part 2
How optimization for machine learning works, part 2
Brandon Rohrer
35 How optimization for machine learning works, part 3
How optimization for machine learning works, part 3
Brandon Rohrer
36 How optimization for machine learning works, part 4
How optimization for machine learning works, part 4
Brandon Rohrer
37 How convolutional neural networks work, in depth
How convolutional neural networks work, in depth
Brandon Rohrer
38 How to pick a machine learning model 4: Splitting the data
How to pick a machine learning model 4: Splitting the data
Brandon Rohrer
39 How to pick a machine learning model 3: Choosing a loss function
How to pick a machine learning model 3: Choosing a loss function
Brandon Rohrer
40 How to pick a machine learning model 2: Separating signal from noise
How to pick a machine learning model 2: Separating signal from noise
Brandon Rohrer
41 How to pick a machine learning model 1: Choosing between models
How to pick a machine learning model 1: Choosing between models
Brandon Rohrer
42 How to pick a machine learning model 5: Navigating assumptions
How to pick a machine learning model 5: Navigating assumptions
Brandon Rohrer
43 What do neural networks learn?
What do neural networks learn?
Brandon Rohrer
44 Interview with iRobot's Director of Data Science Angela Bassa
Interview with iRobot's Director of Data Science Angela Bassa
Brandon Rohrer
45 How Backpropagation Works
How Backpropagation Works
Brandon Rohrer
46 Evolutionary Powell's method: A discrete optimizer for hyperparameter optimization
Evolutionary Powell's method: A discrete optimizer for hyperparameter optimization
Brandon Rohrer
47 1D convolution for neural networks, part 1: Sliding dot product
1D convolution for neural networks, part 1: Sliding dot product
Brandon Rohrer
48 1D convolution for neural networks, part 2: Convolution copies the kernel
1D convolution for neural networks, part 2: Convolution copies the kernel
Brandon Rohrer
49 1D convolution for neural networks, part 3: Sliding dot product equations longhand
1D convolution for neural networks, part 3: Sliding dot product equations longhand
Brandon Rohrer
50 1D convolution for neural networks, part 4: Convolution equation
1D convolution for neural networks, part 4: Convolution equation
Brandon Rohrer
1D convolution for neural networks, part 5: Backpropagation
1D convolution for neural networks, part 5: Backpropagation
Brandon Rohrer
52 1D convolution for neural networks, part 6: Input gradient
1D convolution for neural networks, part 6: Input gradient
Brandon Rohrer
53 1D convolution for neural networks, part 7: Weight gradient
1D convolution for neural networks, part 7: Weight gradient
Brandon Rohrer
54 1D convolution for neural networks, part 8: Padding
1D convolution for neural networks, part 8: Padding
Brandon Rohrer
55 1D convolution for neural networks, part 9: Stride
1D convolution for neural networks, part 9: Stride
Brandon Rohrer
56 The Four Grand Challenges of Robots in the Home
The Four Grand Challenges of Robots in the Home
Brandon Rohrer
57 How Convolution Works
How Convolution Works
Brandon Rohrer
58 The Softmax neural network layer
The Softmax neural network layer
Brandon Rohrer
59 Batch normalization
Batch normalization
Brandon Rohrer
60 Getting ready to learn Python, Mac edition #1: Files and directories
Getting ready to learn Python, Mac edition #1: Files and directories
Brandon Rohrer

This video teaches how to apply backpropagation to 1D convolutional neural networks, covering the mathematical concepts and differentiation required for training. By understanding how to calculate the partial derivative of the loss with respect to the inputs, viewers can learn how to train their own neural networks.

Key Takeaways
  1. Define the convolutional neural network architecture
  2. Calculate the output of the convolutional layer
  3. Apply the chain rule to calculate the partial derivative of the loss with respect to the outputs
  4. Calculate the partial derivative of the loss with respect to the inputs
  5. Use backpropagation to update the model weights
💡 The key to backpropagation is calculating the partial derivative of the loss with respect to the inputs, which requires applying the chain rule and understanding the convolutional neural network architecture.

Related Reads

📰
One H100, Many Models: What I Learned Sharing GPUs on Kubernetes
Learn how to share GPUs on Kubernetes using time-slicing, MPS, and MIG, and discover the pros and cons of each approach
Medium · Machine Learning
📰
AFM-Jev: A Local Decision Engine for Resource-Constrained Banking Workflows
Learn about AFM-Jev, a local decision engine for resource-constrained banking workflows, designed for Apple silicon with data control and regulated deployment
Medium · Machine Learning
📰
Normal Distribution Explained with a Simple Example
Learn the basics of normal distribution with a simple example and understand its importance in statistics
Medium · AI
📰
MLOps Best Practices 2026
Learn MLOps best practices for 2026 to improve your machine learning workflow efficiency and reliability
Dev.to AI
Up next
How Neural Networks Actually Work: The Perceptron Explained
Insightforge | AI & Data Science
Watch →