Skip to content
EntityQ736068· pop 17· linked from 417 articles

Also known as 32-bit Unicode Transformation Format, Unicode Transformation Format – 32-bit, Unicode Transformation Format - 32-bit, UTF32, UTF_32, UCS-4, UCS4

UTF-32 (32-bit Unicode Transformation Format), sometimes called UCS-4, is a fixed-length encoding used to encode Unicode code points that uses exactly 32 bits (four bytes) per code point (but a number of leading bits must be zero as there are far fewer than 232 Unicode code points, needing actually only 21 bits). In contrast, all other Unicode transformation formats are variable-length encodings. Each 32-bit value in UTF-32 represents one Unicode code point and is exactly equal to that code point's numerical value.

In the Vinony graph

Vinony's link graph records 417 inbound references to UTF-32, and connects out to Unicode, Universal Character Set and ISO/IEC 646.

Vinony files it under Character encoding and Unicode Transformation Formats.

Vinony links it to 17 Wikipedia language editions.

Described at

UTF-32 - IBM Documentation

IBM Documentation.

ibm.com →

Note: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support. Note: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support. Note: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support. There has been an error sending your feedback to the team. Your comment was saved locally, if not in an incognito browser, and will be available when attempting to submit feedback again. The IBM® i operating system does not support UTF-32 encoding with a CCSID value. Unicode was originally designed as a pure 16-bit encoding, aimed at representing all modern scripts. Over time, and especially after the addition of over 14 500 composite characters for compatibility with established sets, it became clear that 16 bits were not sufficient for many users. Out of this arose UTF-32. Arrow rightDifferent encodings of Unicode is the algorithmic mapping from every Unicode value to a unique byte sequence.") While IBM values the use of inclusive language, terms that are outside of IBM's direct influence, for the sake of maintaining user understanding, are sometimes required. As other industry leaders join IBM in embracing the use of inclusive language, IBM will continue to update the documentation to reflect those changes.

Excerpt from a page describing this subject · 12,470 chars · not written by Vinony

Wikidata facts

Show 4 more facts
different from
Unicode
data size
32
Sources (1)

via Wikidata · CC0

Article · Русский

UTF-32 (англ. Unicode Transformation Format) или UCS-4 (универсальный набор символов, англ. Universal Character Set) в информатике — один из способов кодирования символов Юникода, использующий для кодирования любого символа ровно 32 бита. Остальные кодировки, UTF-8 и UTF-16, используют для представления символов переменное число байтов. Символ UTF-32 является прямым представлением его кодовой позиции (Code point). Главное преимущество UTF-32 перед кодировками переменной длины заключается в том, что символы Юникод непосредственно индексируемы. Получение n-ой кодовой позиции является операцией, занимающей одинаковое время. Напротив, коды с переменной длиной требует последовательного доступа к n-ой кодовой позиции. Это делает замену символов в строках UTF-32 простой, для этого используется целое число в качестве индекса, как обычно делается для строк ASCII. Главный недостаток UTF-32 — это неэффективное использование пространства, так как для хранения символа используется четыре байта. Символы, лежащие за пределами нулевой (базовой) плоскости кодового пространства, редко используются в большинстве текстов. Поэтому удвоение, в сравнении с UTF-16, занимаемого строками в UTF-32 пространства, не оправдано. Хотя использование неменяющегося числа байтов на символ удобно, но не настолько, как кажется. Операция усечения строк реализуется легче в сравнении с UTF-8 и UTF-16. Но это не делает более быстрым нахождение конкретного смещения в строке, так как смещение может вычисляться и для кодировок фиксированного размера. Это не облегчает вычисление отображаемой ширины строки, за исключением ограниченного числа случаев, так как даже символ «фиксированной ширины» может быть получен комбинированием обычного символа с модифицирующим, который не имеет ширины. Например, буква «й» может быть получена из буквы «и» и диакритического знака «крючок над буквой». Сочетание таких знаков означает, что текстовые редакторы не могут рассматривать 32-битный код как единицу редактирования. Редакторы, которые ограничиваются работой с языками с письмом слева направо и составными символами (англ. Precomposed character), могут использовать символы фиксированного размера. Но такие редакторы вряд ли поддержат символы, лежащие за пределами нулевой (базовой) плоскости кодового пространства и вряд ли смогут работать одинаково хорошо с символами UTF-16.

Abstract from DBpedia / Wikipedia · CC BY-SA

Connections

Categories