File:UTF-8_Encoding_Scheme.png · Wikimedia Commons · See Wikimedia Commons
UTF-8
Sign in to saveAlso known as Unicode Transformation Format – 8-bit, UCS Transformation Format – 8-bit, UTF-2, FSS-UTF, filesystem safe UTF, UTF 8
符号単位を既定したUnicode文字符号化形式及び文字符号化スキーム
In the Vinony graph
Within Vinony's link graph, UTF-8 is referenced by 1,483 other articles, and connects out to Unicode, Plan 9 and byte order mark.
It sits within the topics Character encoding, Computer-related introductions in 1993 and Encodings.
Its subject is documented across 47 Wikipedia language editions.
Key facts
- Character encoding.name
- UTF-8
- Character encoding.standard
- Unicode Standard
- Character encoding.classification
- Unicode Transformation Format, extended ASCII, variable-length encoding
- Character encoding.encodes
- ISO/IEC 10646 (Unicode)
- Character encoding.extends
- ASCII
- Character encoding.prev
- UTF-1
via Wikipedia infobox
Described at
There are several possible representations of Unicode data, including UTF-8 , UTF-16 and UTF-32 . They are all able to represent all of Unicode, but they differ for example in the number of bits for their constituent code units . A Unicode transformation format (UTF ) is an algorithmic mapping from every Unicode code point (except surrogate code points ) to a unique byte sequence. The ISO/IEC 10646 standard uses the term “UCS transformation format ” for UTF; the two terms are merely synonyms for the same concept. Each UTF is reversible, thus every UTF supports lossless round tripping : mapping from any Unicode coded character sequence S to a sequence of bytes and back will produce S again. To ensure round tripping, a UTF mapping must have a mapping for all code points (except surrogate code points). This includes reserved or unassigned code points and the 66 noncharacters (including U+FFFE and U+FFFF). In addition to being lossless, UTFs are unique: any given coded character sequence will always result in the same sequence of bytes for a given UTF. The SCSU compression method, even though it is reversible, is not a UTF because the same string can map to very many different byte sequences, depending on the particular SCSU compressor. [[AF]]( For the formal definition of UTFs see Section 3.9, Unicode Encoding Forms in The Unicode Standard . For more information on encoding forms see UTR 17: Unicode Character Encoding Model . [[AF]]( The freely available open source project International Components for Unicode (ICU) has UTF conversion built into it. The latest version may be downloaded from the ICU Project GitHub site. Many other libraries may have built-in converters, so you may not have to write your own. [[AF]]( Q: Are there any byte sequences that are not generated by a UTF? How should I interpret them? None of the UTFs can generate every arbitrary byte sequence. For example, in UTF-8 every byte of the form 110xxxxx 2 must be followed with a byte of the form 10xxxxxx 2 . A sequence such as is illegal, and must never be generated. When faced with this illegal byte sequence while transforming or interpreting, a UTF-8 conformant process must treat the first byte 110xxxxx 2 as an illegal termination error: for example, either signaling an error, filtering the byte out, or representing the byte with a marker such as U+FFFD REPLACEMENT CHARACTER. In the latter two cases, it will continue processing at the second byte 0xxxxxxx 2 . A conformant process must not interpret illegal or ill-formed byte sequences as characters, however, it may take error recovery actions. No conformant process may use irregular byte sequences to encode out-of-band information. UTF-8 is most common on the web. UTF-16 is used by Java and Windows (.Net). UTF-8 and UTF-32 are used by Linux and various Unix systems. The conversions between all of them are algorithmically based, fast and lossless. This makes it easy to support data input or output in multiple formats, while using a particular UTF for internal storage or processing. [[AF]]( UTF-16 and UTF-32 use code units that are two and four bytes long respectively. For these UTFs , there are three sub-flavors: BE, LE and unmarked. The BE form uses big-endian byte serialization (most significant byte first), the LE form uses little-endian byte serialization (least significant byte first) and the unmarked form uses big-endian byte serialization by default, but may include a byte order mark at the beginning to indicate the actual byte serialization used. [[AF]]( Use UTF-8 . This preserves ASCII , but not Latin-1, because the characters 127 are different from Latin-1. UTF-8 uses the bytes in the ASCII only for ASCII characters. Therefore, it works well in any environment where ASCII characters have a significance as syntax characters, e.g. file name syntaxes, markup languages, etc., but where the all other characters may use arbitrary bytes. Use Java or C style escapes, of the form uXXXX or xXXXX. Thi
Excerpt from a page describing this subject · 40,000 chars · not written by Vinony
Wikidata facts
- Creator
- Ken Thompson
Show 9 more facts
- described at URL
- www.ibm.com/docs/en/i/7.1?topic=unicode-utf-8
- data size
- 8
- last update
- 2003-11-00
- Commons category
- UTF-8
- Stack Exchange tag
- stackoverflow.com/tags/utf-8
- language of work or name
- multiple languages
- derivative work
- CESU-8
- different from
- Unicode
- time of discovery or invention
- 1992-09-02
Sources (3)
via Wikidata · CC0
Article · 日本語
UTF-8(ユーティーエフはち、ユーティーエフエイト)はISO/IEC 10646 (UCS) とUnicodeで使える8ビット符号単位(1〜4バイトの可変長)の文字符号化形式および文字符号化スキーム。 正式名称は、ISO/IEC 10646では “UCS Transformation Format 8”、Unicodeでは “Unicode Transformation Format-8” という。両者はISO/IEC 10646とUnicodeのコード重複範囲で互換性がある。RFCにも仕様がある。 2バイト目以降に「/」などのASCII文字が現れないように工夫されていることから、UTF-FSS (File System Safe) ともいわれる。旧名称はUTF-2。 UTF-8は、データ交換方式・ファイル形式として一般的に使われる傾向にある。 当初は、ベル研究所においてPlan 9で用いるエンコードとして、ロブ・パイクによる設計指針のもと、ケン・トンプソンによって考案された。
Abstract from DBpedia / Wikipedia · CC BY-SA