Batch normalization
Normalization technique used to make training faster and more stable by adjusting the inputs to each layer, recentering them around zero and rescaling them to a standard size
In artificial neural networks, batch normalization (also known as batch norm) is a normalization technique used to make training faster and more stable by adjusting the inputs to each layer—re-centering them around zero and re-scaling them to a standard size. It was introduced by Sergey Ioffe and Christian Szegedy in 2015.
Nº Q55080248 ★
Common · Knowledge
Batch normalization
Normalization technique used to make training faster and more stable by adjusting the inputs to each layer, recentering them around zero and rescaling them to a standard size
In artificial neural networks, batch normalization (also known as batch norm) is a normalization technique used to make training faster and more stable by adjusting the inputs to each layer—re-centering them around zero and re-scaling them to a standard size. It was introduced by Sergey Ioffe and Christian Szegedy in 2015.
From Wikipedia
In artificial neural networks, batch normalization (also known as batch norm) is a normalization technique used to make training faster and more stable by adjusting the inputs to each layer—re-centering them around zero and re-scaling them to a standard size. It was introduced by Sergey Ioffe and Christian Szegedy in 2015. Experts still debate why batch normalization works so well. It was initially thought to tackle internal covariate shift, a problem where parameter initialization and changes in the distribution of the inputs of each layer affect the learning rate of the network. However, newer research suggests it does not fix this shift but instead smooths the objective function—a mathematical guide the network follows to improve—enhancing performance. In very deep networks, batch normalization can initially cause a severe gradient explosion—where updates to the network grow uncontrollably large—but this is managed with shortcuts called skip connections in residual networks. Another theory is that batch normalization adjusts data by handling its size and path separately, speeding up training.
Text: Wikipédia, CC BY-SA 4.0. ·
Related cards
-
L
Linearization
Finding linear approximation of function at given point
Nº Q1520713 ★
Not listed
-
P
Platt scaling
Machine learning calibration technique
Nº Q17146653 ★
Not listed
-
Gram–Schmidt process
Method for orthonormalising a set of vectors
Nº Q475239 ★★★
Not listed
-
T
Third normal form
Normalizing a database design to reduce the duplication of data and ensure referential integrity
Nº Q311585 ★★
Not listed
-
Neural network (machine learning)
Computational model used in machine learning, based on connected, hierarchical functions
Nº Q192776 ★★★★★
Not listed
-
Renormalization
Process of assuring meaningful mathematical results in quantum field theory and related disciplines
Nº Q1047702 ★★
Not listed
-
Total variation denoising
Noise removal process during image processing
Nº Q7828156 ★
Not listed
-
Artificial neuron
Mathematical function conceived as a crude model
Nº Q177058 ★★
Not listed
-
R
Randomized algorithm
Algorithm designed to use randomness from auxiliary inputs as part of its logic
Nº Q583461 ★
Not listed
-
T
Types of artificial neural networks
Overview about the types of artificial neural networks
Nº Q7860946 ★
Not listed
-
Supervised learning
Machine learning task of learning a function that maps an input to an output based on example input-output pairs
Nº Q334384 ★★
Not listed
-
H
History of artificial neural networks
Aspect of history
Nº Q85766763 ★
Not listed
-
R
Renormalization group
Method for using scale changes to understand physical theories such as quantum field theories
Nº Q1203669 ★★
Not listed
-
K
Kabsch algorithm
Type of algorithm
Nº Q6344361 ★
Not listed
-
Convolutional neural network
Regularized type of feed-forward neural network that learns features by itself via filter (or kernel) optimization
Nº Q17084460 ★★★
Not listed
-
R
Radial basis function kernel
Machine learning kernel function
Nº Q7280263 ★
Not listed
-
L
Learning rate
Tuning parameter (hyperparameter) in optimization
Nº Q65121812 ★
Not listed
-
R
Remez algorithm
Algorithm to approximate functions
Nº Q2835816 ★
Not listed