Commune · Savoirs
Batch normalization
Normalization technique used to make training faster and more stable by adjusting the inputs to each layer, recentering them around zero and rescaling them to a standard size
In artificial neural networks, batch normalization (also known as batch norm) is a normalization technique used to make training faster and more stable by adjusting the inputs to each layer—re-centering them around zero and re-scaling them to a standard size. It was introduced by Sergey Ioffe and Christian Szegedy in 2015.
Sur Wikipédia
Texte en anglais Pas encore d'article dans ta langue : extrait en anglais.
In artificial neural networks, batch normalization (also known as batch norm) is a normalization technique used to make training faster and more stable by adjusting the inputs to each layer—re-centering them around zero and re-scaling them to a standard size. It was introduced by Sergey Ioffe and Christian Szegedy in 2015. Experts still debate why batch normalization works so well. It was initially thought to tackle internal covariate shift, a problem where parameter initialization and changes in the distribution of the inputs of each layer affect the learning rate of the network. However, newer research suggests it does not fix this shift but instead smooths the objective function—a mathematical guide the network follows to improve—enhancing performance. In very deep networks, batch normalization can initially cause a severe gradient explosion—where updates to the network grow uncontrollably large—but this is managed with shortcuts called skip connections in residual networks. Another theory is that batch normalization adjusts data by handling its size and path separately, speeding up training.
Texte : Wikipédia en anglais, CC BY-SA 4.0. ·
Cartes voisines
-
V★
Validated numerics
Numerics including mathematically strict error evaluation
-
R★
Réseau neuronal siamois
-
S★
Solomonoff's theory of inductive inference
Mathematical formalization of Occam's razor that, assuming the world is generated by a computer program, the most likely one is the shortest, using Bayesian inference
-
U★★
U-Net
-
★
Jackknife
-
A★
AdaBoost
Algorithme de boosting
-
★
Bootstrap aggregating
-
★
Seau percé
Algorithme pour réseau informatique
-
A★★
Algorithme de Markov
-
A★
Apprentissage avec erreurs
Problème algorithmique reposant sur les réseaux euclidiens dont la difficulté supposée résisterait à un ordinateur quantique.
-
E★★★
Espace de Banach
Espace vectoriel normé sur un corps de nombres et complet pour la norme
-
★★★
Méthode de Newton
Algorithme de calcul d'un zéro d'une fonction réelle d'une variable réelle
-
★★★★
Transformeur
Architecture d'apprentissage automatique
-
P★★
Predictive coding
Psychological term
-
S★★
Standard part function
In non-standard analysis, the standard part function is a function from the limited (finite) hyperreal numbers to the real numbers.
-
N★★
Nombre normal
Nombre réel tel que, quelle que soit la base de numération choisie pour l'écrire, en recherchant une séquence finie de chiffres dans son développement, on a autant de chance de la trouver que n'importe laquelle des séquences de même longueur
-
C★
Changement de variable (simplification algébrique)
-
B★
Basic Linear Algebra Subprograms
Nsemble de fonctions standardisées (interface de programmation) réalisant des opérations de base de l'algèbre linéaire