Común · Saberes
Batch normalization
Normalization technique used to make training faster and more stable by adjusting the inputs to each layer, recentering them around zero and rescaling them to a standard size
In artificial neural networks, batch normalization (also known as batch norm) is a normalization technique used to make training faster and more stable by adjusting the inputs to each layer—re-centering them around zero and re-scaling them to a standard size. It was introduced by Sergey Ioffe and Christian Szegedy in 2015.
En Wikipedia
Texto en inglés Aún no hay artículo en tu idioma: extracto en inglés.
In artificial neural networks, batch normalization (also known as batch norm) is a normalization technique used to make training faster and more stable by adjusting the inputs to each layer—re-centering them around zero and re-scaling them to a standard size. It was introduced by Sergey Ioffe and Christian Szegedy in 2015. Experts still debate why batch normalization works so well. It was initially thought to tackle internal covariate shift, a problem where parameter initialization and changes in the distribution of the inputs of each layer affect the learning rate of the network. However, newer research suggests it does not fix this shift but instead smooths the objective function—a mathematical guide the network follows to improve—enhancing performance. In very deep networks, batch normalization can initially cause a severe gradient explosion—where updates to the network grow uncontrollably large—but this is managed with shortcuts called skip connections in residual networks. Another theory is that batch normalization adjusts data by handling its size and path separately, speeding up training.
Texto: Wikipedia en inglés, CC BY-SA 4.0. ·
Cartas cercanas
-
L★
LINPACK benchmarks
Software
-
C★
Consistent Overhead Byte Stuffing
Algorithm for encoding data bytes
-
T★
Teoría algorítmica de la información
-
B★
Brent's method
Root-finding algorithm
-
T★
Transformación Box-cox
Family of functions to transform data
-
★★
Optimización por enjambre de partículas
-
★
Biological neuron model
Mathematical description of the properties of certain cells in the nervous system that generate sharp electrical potentials across their cell membrane, roughly one millisecond in duration
-
R★
Reduced chi-squared statistic
Type of statistic
-
★★★
Forma canónica de Jordan
Definición particular de una matriz con su diagonal formada por bloques de Jordan.
-
★★★
Análisis de la regresión
Conjunto de procesos estadísticos para estimar las relaciones entre variables
-
B★
Bochner integral
Generalization of the Lebesgue integral to Banach-space valued functions
-
★
Binary GCD algorithm
Algorithm that computes the greatest common divisor of two integers using only arithmetic shifts, comparisons, and subtraction
-
C★★
CMA-ES
Algorithm
-
S★
Slab
-
D★★
Dynamic frequency scaling
Technique in computer architecture whereby the frequency of a microprocessor can be automatically adjusted "on the fly", either to conserve power or to reduce the amount of generated heat
-
★★
Programación en enteros
-
★★
MACD
-
★★
DBZ (meteorology)
Unit of measure used in weather radar