What Is the Vanishing Gradient Problem?

Overview

The vanishing gradient problem is a difficulty that arises when training artificial neural networks with gradient-based methods such as backpropagation.

It makes the parameters of the earlier layers particularly difficult to train and tune. Backpropagation proceeds from the end of the network toward the beginning, so if gradients vanish along the way, the earlier layers cannot be updated effectively. The problem becomes worse as the network gains more layers.

This is not an inherent flaw in neural networks themselves. It arises when gradient-based training is combined with certain activation functions.

Let us build an intuitive understanding of the problem and examine what it causes.

 

Problem

Gradient-based methods learn parameter values by measuring how a small change in a parameter affects the network's output.

If a change in a parameter produces only a tiny change in the output, the network cannot learn that parameter effectively. That is the problem.

A gradient is ultimately a derivative—a measure of change. If that change is extremely small, the network cannot train effectively and may converge before its error has been sufficiently reduced.

This is what happens in the vanishing gradient problem: the gradient of the network's output with respect to each parameter in an early layer becomes extremely small. Put another way, even a large change to a parameter in an early layer barely affects the output.

So when does this happen, and why?

 

Cause

The vanishing gradient problem depends on the choice of activation function. Common activation functions such as sigmoid and tanh squash their inputs into a very small output range in a highly nonlinear way.

For example, sigmoid maps real numbers into the interval [0, 1]. As a result, a very large region of input space is mapped into an extremely small range.

Within that region, even a large change in the input produces only a small change in the output, because the gradient is small.

This effect becomes worse when we stack multiple layers of these nonlinear functions on top of one another.

For instance, the first layer maps a broad input region into a small output region. The second and third layers can compress that region even further.

Consequently, even a very large change to the first layer's input may barely change the final output.

To address this, we can use activation functions that do not have the same squashing behavior.

ReLU—the rectified linear unit, max(0, x)—is a common choice.

Reference

 
 

Read next