U

Unicode equivalence

Aspect of the Unicode Standard

Unicode equivalence is the specification by the Unicode character encoding standard that some sequences of code points represent essentially the same character. The feature was introduced in the standard to allow compatibility with pre-existing standard character sets, which often included similar or identical characters.

Nº Q250798 ★

Comum · Saberes

Unicode equivalence

Aspect of the Unicode Standard

Texto em inglês

Unicode equivalence is the specification by the Unicode character encoding standard that some sequences of code points represent essentially the same character. The feature was introduced in the standard to allow compatibility with pre-existing standard character sets, which often included similar or identical characters.

Na Wikipédia

Texto em inglês Ainda não há artigo no seu idioma: trecho em inglês.

Unicode equivalence is the specification by the Unicode character encoding standard that some sequences of code points represent essentially the same character. The feature was introduced in the standard to allow compatibility with pre-existing standard character sets, which often included similar or identical characters. Unicode provides two such notions, canonical equivalence and compatibility. Code point sequences that are defined as canonically equivalent are assumed to have the same appearance and meaning when printed or displayed. For example, the code point U+006E n LATIN SMALL LETTER N followed by U+0303 ◌̃ COMBINING TILDE is defined by Unicode to be canonically equivalent to the single code point U+00F1 ñ LATIN SMALL LETTER N WITH TILDE. Therefore, those sequences should be displayed in the same manner, should be treated in the same way by applications such as alphabetizing names or searching, and may be substituted for each other. Similarly, each Hangul syllable block that is encoded as a single character may be equivalently encoded as a combination of a leading conjoining jamo; a vowel conjoining jamo; and, if appropriate, a trailing conjoining jamo. Sequences that are defined as compatible are assumed to have possibly distinct appearances but the same meaning in some contexts. Thus, for example, U+FB00 ff LATIN SMALL LIGATURE FF, a typographic ligature, is defined to be compatible with, but not canonically equivalent to, the sequence U+0066 U+0066 (two Latin "f" letters). Compatible sequences may be treated the same way in some applications (such as sorting and indexing) but not in others, and they may be substituted for each other in some situations, but not in others. Sequences that are canonically equivalent are also compatible, but the opposite is not necessarily true. The standard also defines a text normalization procedure, called Unicode normalization, which replaces equivalent sequences of characters so that...

Texto: Wikipédia em inglês, CC BY-SA 4.0. ·

Cartas próximas

Abrir

…

Confirmação