Commune · Savoirs
Batch normalization
Normalization technique used to make training faster and more stable by adjusting the inputs to each layer, recentering them around zero and rescaling them to a standard size
In artificial neural networks, batch normalization (also known as batch norm) is a normalization technique used to make training faster and more stable by adjusting the inputs to each layer—re-centering them around zero and re-scaling them to a standard size. It was introduced by Sergey Ioffe and Christian Szegedy in 2015.
Sur Wikipédia
Texte en anglais Pas encore d'article dans ta langue : extrait en anglais.
In artificial neural networks, batch normalization (also known as batch norm) is a normalization technique used to make training faster and more stable by adjusting the inputs to each layer—re-centering them around zero and re-scaling them to a standard size. It was introduced by Sergey Ioffe and Christian Szegedy in 2015. Experts still debate why batch normalization works so well. It was initially thought to tackle internal covariate shift, a problem where parameter initialization and changes in the distribution of the inputs of each layer affect the learning rate of the network. However, newer research suggests it does not fix this shift but instead smooths the objective function—a mathematical guide the network follows to improve—enhancing performance. In very deep networks, batch normalization can initially cause a severe gradient explosion—where updates to the network grow uncontrollably large—but this is managed with shortcuts called skip connections in residual networks. Another theory is that batch normalization adjusts data by handling its size and path separately, speeding up training.
Texte : Wikipédia en anglais, CC BY-SA 4.0. ·
Cartes voisines
-
M★
Méthode de Welch
-
★★★
Algorithme LLL
Algorithme de réduction de réseau qui s'exécute en temps polynomial
-
N★★
Norm-referenced test
Yields an estimate of the testee's position in population
-
★★
A-weighting
Curves used to weigh sound pressure level
-
★
Algorithme de Lloyd-Max
-
★★
NURBS
Modèle mathématique
-
★
Droite de Henry
Droite d'ajustement d'un nuage de point à une loi normale
-
★★
Algorithme de colonies de fourmis
Algorithmes inspirés du comportement des fourmis et qui constituent une famille de métaheuristiques d’optimisation
-
K★
Knowledge cutoff
Temporal limit of a model's training data
-
★★★
Rétropropagation du gradient
Méthode pour calculer le gradient de l'erreur pour chaque neurone d'un réseau de neurones
-
G★★
Gradient boosting
Une technique d’apprentissage qui consiste à combiner des modèles faibles, le plus souvent des arbres de décision, pour former un modèle prédictif performant en corrigeant les erreurs résiduelles de manière itérative
-
★
Méthode des surfaces de réponses
-
★★
Validation croisée
-
C★
Constant folding
-
★
Arithmétique d'intervalles
-
★
Transformation de Fisher
Statistical transformation
-
A★★
Analyse discriminante linéaire
Outil statistique
-
★
Codage par intervalle