Commune · Savoirs
Batch normalization
Normalization technique used to make training faster and more stable by adjusting the inputs to each layer, recentering them around zero and rescaling them to a standard size
In artificial neural networks, batch normalization (also known as batch norm) is a normalization technique used to make training faster and more stable by adjusting the inputs to each layer—re-centering them around zero and re-scaling them to a standard size. It was introduced by Sergey Ioffe and Christian Szegedy in 2015.
Sur Wikipédia
Texte en anglais Pas encore d'article dans ta langue : extrait en anglais.
In artificial neural networks, batch normalization (also known as batch norm) is a normalization technique used to make training faster and more stable by adjusting the inputs to each layer—re-centering them around zero and re-scaling them to a standard size. It was introduced by Sergey Ioffe and Christian Szegedy in 2015. Experts still debate why batch normalization works so well. It was initially thought to tackle internal covariate shift, a problem where parameter initialization and changes in the distribution of the inputs of each layer affect the learning rate of the network. However, newer research suggests it does not fix this shift but instead smooths the objective function—a mathematical guide the network follows to improve—enhancing performance. In very deep networks, batch normalization can initially cause a severe gradient explosion—where updates to the network grow uncontrollably large—but this is managed with shortcuts called skip connections in residual networks. Another theory is that batch normalization adjusts data by handling its size and path separately, speeding up training.
Texte : Wikipédia en anglais, CC BY-SA 4.0. ·
Cartes voisines
-
A★
Algorithme de Thompson
-
T★
Top-p sampling
Language model technique
-
B★★★
Blahut–Arimoto algorithm
Class of algorithms in information theory
-
★
Weisfeiler Leman graph isomorphism test
Heuristic algorithm for testing whether two graphs are isomorphic
-
★★
Théorème de Norton
-
★★
Matrice échelonnée
Matrice dont le nombre de 0 est croissant par les lignes ou par les colonnes
-
L★★
Lottery ticket hypothesis
Machine learning hypothesis
-
★★
Algorithme espérance-maximisation
Algorithme d'optimisation
-
★
Auto-GPT
Logiciel d'automatisation
-
é★
équations d’estimation généralisées
Méthode d'estimation statistique
-
N★
Newey–West estimator
Statistical tool
-
★★
Algorithme de Gauss-Newton
-
L★★
Limited-memory BFGS
Optimization algorithm
-
p★
programmation différentiable
Programming paradigm in which a numeric computer program can be differentiated throughout via automatic differentiation, allowing for machine learning based on gradient descent etc.
-
★★
Recuit simulé
Méthode d'optimisation
-
★
Tokenisation (sécurité informatique)
Transformation en unités numériques, appelées jetons, qui peuvent être utilisées ou échangées de manière numérique
-
H★
Hiérarchie de Borel
-
★
apprentissage sans échantillon
L'apprentissage sans échantillon est une méthode d'apprentissage automatique qui permet à un modèle de reconnaître des catégories non vues lors de son entraînement en utilisant des connaissances sémantiques ou des attributs associés.