Burstsort
Cache-efficient algorithm for sorting strings
Burstsort and its variants are cache-efficient algorithms for sorting strings. They are variants of the traditional radix sort but faster for large data sets of common strings, first published in 2003, with some optimizing versions published in later years.
Nº Q5000665 ★
Common · Knowledge
Burstsort
Cache-efficient algorithm for sorting strings
Burstsort and its variants are cache-efficient algorithms for sorting strings. They are variants of the traditional radix sort but faster for large data sets of common strings, first published in 2003, with some optimizing versions published in later years.
From Wikipedia
Burstsort and its variants are cache-efficient algorithms for sorting strings. They are variants of the traditional radix sort but faster for large data sets of common strings, first published in 2003, with some optimizing versions published in later years. Burstsort algorithms use a trie to store prefixes of strings, with growable arrays of pointers as end nodes containing sorted, unique, suffixes (referred to as buckets). Some variants copy the string tails into the buckets. As the buckets grow beyond a predetermined threshold, the buckets are "burst" into tries, giving the sort its name. A more recent variant uses a bucket index with smaller sub-buckets to reduce memory usage. Most implementations delegate to multikey quicksort, an extension of three-way radix quicksort, to sort the contents of the buckets. By dividing the input into buckets with common prefixes, the sorting can be done in a cache-efficient manner. Burstsort was introduced as a sort that is similar to MSD radix sort, but is faster due to being aware of caching and related radixes being stored closer to each other due to specifics of trie structure. It exploits specifics of strings that are usually encountered in real world. And although asymptotically it is the same as radix sort, with time complexity of O(wn) (w – word length and n – number of strings to be sorted), but due to better memory distribution it tends to be twice as fast on big data sets of strings. It has been billed as the "fastest known algorithm to sort large sets of strings".
Text: Wikipédia, CC BY-SA 4.0. ·
Related cards
-
Radix sort
Non-comparative sorting algorithm
Nº Q830223 ★★
Not listed
-
Binary search tree
Data structure in tree form with 0, 1, or 2 children per node, sorted for fast lookup
Nº Q623818 ★★
Not listed
-
Substring
Subsequence of the symbols in a string, where the order of the elements is preserved
Nº Q2626534 ★
Not listed
-
S
String interning
Data structure for reusing strings
Nº Q4383552 ★
Not listed
-
P
Powersort
Sorting algorithm
Nº Q136399159 ★
Not listed
-
Suffix automaton
Minimal DFA accepting set of all suffixes of particular string
Nº Q19599738 ★
Not listed