UTF-32/UCS-4
Sign in to saveAlso known as 32-bit Unicode Transformation Format, Unicode Transformation Format – 32-bit, Unicode Transformation Format - 32-bit, UTF32, UTF_32, UCS-4, UCS4
UTF-32 (32-bit Unicode Transformation Format), sometimes called UCS-4, is a fixed-length encoding used to encode Unicode code points that uses exactly 32 bits (four bytes) per code point (but a number of leading bits must be zero as there are far fewer than 232 Unicode code points, needing actually only 21 bits). In contrast, all other Unicode transformation formats are variable-length encodings. Each 32-bit value in UTF-32 represents one Unicode code point and is exactly equal to that code point's numerical value.
In the Vinony graph
Within Vinony's link graph, UTF-32/UCS-4 is referenced by 417 other articles, and connects out to Unicode, Universal Character Set and ISO/IEC 646.
Vinony files it under Character encoding and Unicode Transformation Formats.
Its subject is documented across 17 Wikipedia language editions.
Described at

Note: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support. Note: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support. Note: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support. There has been an error sending your feedback to the team. Your comment was saved locally, if not in an incognito browser, and will be available when attempting to submit feedback again. The IBM® i operating system does not support UTF-32 encoding with a CCSID value. Unicode was originally designed as a pure 16-bit encoding, aimed at representing all modern scripts. Over time, and especially after the addition of over 14 500 composite characters for compatibility with established sets, it became clear that 16 bits were not sufficient for many users. Out of this arose UTF-32. Arrow rightDifferent encodings of Unicode is the algorithmic mapping from every Unicode value to a unique byte sequence.") While IBM values the use of inclusive language, terms that are outside of IBM's direct influence, for the sake of maintaining user understanding, are sometimes required. As other industry leaders join IBM in embracing the use of inclusive language, IBM will continue to update the documentation to reflect those changes.
Excerpt from a page describing this subject · 12,470 chars · not written by Vinony
Wikidata facts
Show 4 more facts
- Stack Exchange tag
- stackoverflow.com/tags/utf-32
- different from
- Unicode
- described at URL
- www.ibm.com/docs/en/i/7.1?topic=unicode-utf-32
- data size
- 32
Sources (1)
via Wikidata · CC0
Article · Polski
UTF-32 (ang. 32-bit unicode transformation format) – jeden ze sposobów kodowania znaków standardu Unicode. Sposób ten wymaga użycia 32-bitowych słów. Zestaw znaków jest też zdefiniowany w standardzie ISO 10646 jako UCS-4. Kody obejmują zakres od 0 do 0x7FFFFFFF. Kod znaku zawsze ma długość 4 bajtów i w zapisie big endian przedstawia po prostu numer znaku w tabeli Unikodu. Możliwa jest również odwrotna kolejność – w zapisie little endian, co nakłada obowiązek używania znacznika kierunku BOM. Stała długość kodu każdego znaku (w przeciwieństwie do m.in. UTF-8) jest dużą zaletą tego kodowania. Kodowanie to jest jednak bardzo nieefektywne - zakodowane ciągi znaków są dwa do czterech razy dłuższe niż ciągi tych samych znaków zapisanych w innych kodowaniach. Kodowanie to z tego powodu jest zwykle stosowane tylko w pamięci operacyjnej w celu ułatwienia obsługi i przetwarzania (np. obliczenie długości czy wycinanie ciągu znaków jest bardzo proste), na innych nośnikach (takich jak połączenia sieciowe czy dysk twardy) stosuje się zwykle bardziej efektywne UTF-8 lub UTF-16. W systemach uniksowych kodowanie to jest najczęściej używane do wewnętrznego przechowywania napisów Unicode.
Abstract from DBpedia / Wikipedia · CC BY-SA