File:UTF-8_Encoding_Scheme.png · Wikimedia Commons · See Wikimedia Commons
UTF-8
Sign in to saveAlso known as Unicode Transformation Format – 8-bit, UCS Transformation Format – 8-bit, UTF-2, FSS-UTF, filesystem safe UTF, UTF 8
UTF-8 is a character encoding standard used for electronic communication. Defined by the Unicode Standard, the name is derived from Unicode Transformation Format 8-bit. As of 2026, almost every webpage (99%) is transmitted as UTF-8.
In the Vinony graph
Within Vinony's link graph, UTF-8 is referenced by 1,483 other articles, and connects out to Unicode, Plan 9 and byte order mark.
It sits within the topics Character encoding, Computer-related introductions in 1993 and Encodings.
Its subject is documented across 47 Wikipedia language editions.
Key facts
- Character encoding.name
- UTF-8
- Character encoding.standard
- Unicode Standard
- Character encoding.classification
- Unicode Transformation Format, extended ASCII, variable-length encoding
- Character encoding.encodes
- ISO/IEC 10646 (Unicode)
- Character encoding.extends
- ASCII
- Character encoding.prev
- UTF-1
via Wikipedia infobox
Described at
There are several possible representations of Unicode data, including UTF-8 , UTF-16 and UTF-32 . They are all able to represent all of Unicode, but they differ for example in the number of bits for their constituent code units . A Unicode transformation format (UTF ) is an algorithmic mapping from every Unicode code point (except surrogate code points ) to a unique byte sequence. The ISO/IEC 10646 standard uses the term “UCS transformation format ” for UTF; the two terms are merely synonyms for the same concept. Each UTF is reversible, thus every UTF supports lossless round tripping : mapping from any Unicode coded character sequence S to a sequence of bytes and back will produce S again. To ensure round tripping, a UTF mapping must have a mapping for all code points (except surrogate code points). This includes reserved or unassigned code points and the 66 noncharacters (including U+FFFE and U+FFFF). In addition to being lossless, UTFs are unique: any given coded character sequence will always result in the same sequence of bytes for a given UTF. The SCSU compression method, even though it is reversible, is not a UTF because the same string can map to very many different byte sequences, depending on the particular SCSU compressor. [[AF]]( For the formal definition of UTFs see Section 3.9, Unicode Encoding Forms in The Unicode Standard . For more information on encoding forms see UTR 17: Unicode Character Encoding Model . [[AF]]( The freely available open source project International Components for Unicode (ICU) has UTF conversion built into it. The latest version may be downloaded from the ICU Project GitHub site. Many other libraries may have built-in converters, so you may not have to write your own. [[AF]]( Q: Are there any byte sequences that are not generated by a UTF? How should I interpret them? None of the UTFs can generate every arbitrary byte sequence. For example, in UTF-8 every byte of the form 110xxxxx 2 must be followed with a byte of the form 10xxxxxx 2 . A sequence such as is illegal, and must never be generated. When faced with this illegal byte sequence while transforming or interpreting, a UTF-8 conformant process must treat the first byte 110xxxxx 2 as an illegal termination error: for example, either signaling an error, filtering the byte out, or representing the byte with a marker such as U+FFFD REPLACEMENT CHARACTER. In the latter two cases, it will continue processing at the second byte 0xxxxxxx 2 . A conformant process must not interpret illegal or ill-formed byte sequences as characters, however, it may take error recovery actions. No conformant process may use irregular byte sequences to encode out-of-band information. UTF-8 is most common on the web. UTF-16 is used by Java and Windows (.Net). UTF-8 and UTF-32 are used by Linux and various Unix systems. The conversions between all of them are algorithmically based, fast and lossless. This makes it easy to support data input or output in multiple formats, while using a particular UTF for internal storage or processing. [[AF]]( UTF-16 and UTF-32 use code units that are two and four bytes long respectively. For these UTFs , there are three sub-flavors: BE, LE and unmarked. The BE form uses big-endian byte serialization (most significant byte first), the LE form uses little-endian byte serialization (least significant byte first) and the unmarked form uses big-endian byte serialization by default, but may include a byte order mark at the beginning to indicate the actual byte serialization used. [[AF]]( Use UTF-8 . This preserves ASCII , but not Latin-1, because the characters 127 are different from Latin-1. UTF-8 uses the bytes in the ASCII only for ASCII characters. Therefore, it works well in any environment where ASCII characters have a significance as syntax characters, e.g. file name syntaxes, markup languages, etc., but where the all other characters may use arbitrary bytes. Use Java or C style escapes, of the form uXXXX or xXXXX. Thi
Excerpt from a page describing this subject · 40,000 chars · not written by Vinony
Wikidata facts
- Creator
- Ken Thompson
Show 9 more facts
- described at URL
- www.ibm.com/docs/en/i/7.1?topic=unicode-utf-8
- data size
- 8
- last update
- 2003-11-00
- Commons category
- UTF-8
- Stack Exchange tag
- stackoverflow.com/tags/utf-8
- language of work or name
- multiple languages
- derivative work
- CESU-8
- different from
- Unicode
- time of discovery or invention
- 1992-09-02
Sources (3)
via Wikidata · CC0
Article · Português
UTF-8 (8-bit Unicode Transformation Format) é um tipo de codificação binária (Unicode) de comprimento variável criado por Ken Thompson e Rob Pike. Pode representar qualquer caractere universal padrão do Unicode, sendo também compatível com o ASCII. Por esta razão, está lentamente a ser adaptado como tipo de codificação padrão para e-mail, páginas web, e outros locais onde os caracteres são armazenados. UTF-8 usa de um a quatro bytes (estritamente, octetos) por caractere, dependendo do símbolo Unicode que representa. É necessário apenas um byte para codificar os 128 caracteres ASCII (Unicode U+0000 a U+007F). São necessários dois bytes para caracteres Latinos com diacríticos. São também usados dois bytes para representar caracteres dos alfabetos Grego, Cirílico, Armênio, Hebraico, Sírio e Thaana (Unicode U+0080 a U+07FF). São necessários três bytes para o resto do (que contém praticamente todos os caracteres comuns utilizados). Existem ainda outros caracteres que necessitam de quatro bytes. Quatro bytes pode parecer muito para um caractere ("code point"), mas muito raramente são utilizados. Além disso, UTF-16 (a principal alternativa ao UTF-8) necessita também de quatro bytes para estes "code points". A definição de qual dos dois é mais eficiente (UTF-8 ou UTF-16) depende da variedade de "code points" usados. Contudo, as diferenças entre os vários tipos de codificação tornam-se irrelevantes com o uso de sistemas de compressão como o DEFLATE. Para textos curtos nos quais os tradicionais algoritmos não funcionam bem e se faz necessário ter o tamanho em consideração, é geralmente usado o Esquema Padrão de Compressão para Unicode (Standard Compression Scheme for Unicode). O "Internet Engineering Task Force" (IETF) requer que todos os protocolos utilizados na Internet suportem, pelo menos, o UTF-8. O "Internet Mail Consortium" (IMC) [1] recomenda que todos os clientes de e-mail consigam ler e criar mails usando o UTF-8.
Abstract from DBpedia / Wikipedia · CC BY-SA