Common · Knowledge
Decision tree pruning
Algorithm improvement technique where unnecessary nodes are removed
Pruning is a data compression technique in machine learning and search algorithms that reduces the size of decision trees by removing sections of the tree that are non-critical and redundant to classify instances. Pruning reduces the complexity of the final classifier, and hence improves predictive accuracy by the reduction of overfitting.
From Wikipedia
Pruning is a data compression technique in machine learning and search algorithms that reduces the size of decision trees by removing sections of the tree that are non-critical and redundant to classify instances. Pruning reduces the complexity of the final classifier, and hence improves predictive accuracy by the reduction of overfitting. One of the questions that arises in a decision tree algorithm is the optimal size of the final tree. A tree that is too large risks overfitting the training data and poorly generalizing to new samples. A small tree might not capture important structural information about the sample space. However, it is hard to tell when a tree algorithm should stop because it is impossible to tell if the addition of a single extra node will dramatically decrease error. This problem is known as the horizon effect. A common strategy is to grow the tree until each node contains a small number of instances then use pruning to remove nodes that do not provide additional information. Pruning should reduce the size of a learning tree without reducing predictive accuracy as measured by a cross-validation set. There are many techniques for tree pruning that differ in the measurement that is used to optimize performance.
Text: Wikipédia, CC BY-SA 4.0. · Image: CrumpledBenito (CC BY-SA 4.0) ·
Related cards
-
★★
Clock gating
Technique used in synchronous circuits for reducing dynamic power dissipation, by adding more logic to a circuit to prune the clock tree (disabling portions of the circuitry so that the flip-flops in them do not have to switch states)
-
★★
Binary search tree
Data structure in tree form with 0, 1, or 2 children per node, sorted for fast lookup
-
★★★
Red–black tree
Self-balancing binary search tree data structure
-
T★
Tree topping
Practice of removing entire tops of trees
-
★★
Decision tree learning
Algorithm that recursively splits data on attributes to build a tree for classification or regression
-
A★
Algorithmic curation
The process of selecting, organizing, and presenting digital content to users on online platforms (social media, search engines) mediated by recommendation algorithms, aiming for engagement and personalization.