File:New_Unicode_logo.svg · Wikimedia Commons · See Wikimedia Commons
Unicode
Sign in to saveAlso known as The Unicode Standard, Uni-code, Unicode Standard
Unicode (also known as The Unicode Standard and TUS) is a character encoding standard maintained by the Unicode Consortium designed to support the use of text in all of the world's writing systems that can be digitized. Version 17.0 defines 159,801 characters and 172 scripts used in various ordinary, literary, academic and technical contexts.
Unicode is a standardized system that assigns unique numerical codes to characters used in all the world's writing systems, allowing computers to store and display text from any language or script. As of its latest version, it encompasses over 159,000 characters representing 172 different scripts, making it possible for digital devices to handle text reliably across languages and cultures.
AI-generated from the Wikipedia summary — may contain errors.
Key facts
- Character encoding.name
- Unicode
- Character encoding.image
- New Unicode logo.svg
- Character encoding.caption
- Logo of the Unicode Consortium
- Character encoding.standard
- Unicode Standard
- Character encoding.lang
- 168 scripts (list)
- Character encoding.encodings
- (uncommon) (obsolete)
- Character encoding.prev
- ISO/IEC 8859, among others
via Wikipedia infobox
Wikidata facts
- Official website
- unicode.org
Show 10 more facts
- Commons category
- Unicode
- Stack Exchange tag
- askubuntu.com/tags/unicode
- Commons gallery
- Unicode
- Unicode range
- U+0000-10FFFF
- ACM Classification Code (2012)
- 10011594
- software version identifier
- 17.0.0
- issue tracker URL
- unicode-org.atlassian.net
- publication date
- 1996-07-00
- official blog URL
- blog.unicode.org
- P13411
- Sagas of Icelanders
via Wikidata · CC0
~62 min read
Article
38 sectionsContents
- Origin and development
- {{anchor|Unicode 88}}History
- Unicode Consortium
- Scripts covered
- Proposals for adding scripts
- Versions
- Architecture and terminology<span class="anchor" id="Upluslink"></span><!-- Template:U+ links to this paragraph -->
- Codespace and code points
- Code planes and blocks
- General Category property
- Abstract characters <span class="anchor" id="Alias"></span>
- Ready-made versus composite characters
- Ligatures
- Standardized subsets
- {{anchor|UTF|UCS}}Mapping and encodings
- Adoption
- Operating systems
- Input methods
- Web
- Fonts
- Newlines
- Issues
- Character unification
- Han unification
- Italic or cursive characters in Cyrillic
- Localised case pairs
- Diacritics on lowercase {{serif|I}}
- Security<span class="anchor" id="Security issues"></span>
- Mapping to legacy character sets
- Indic scripts
- Combining characters
- Anomalies
- See also
- Notes
- References
- Further reading
- External links
Unicode (also known as The Unicode Standard and TUS) is a character encoding standard maintained by the Unicode Consortium designed to support the use of text in all of the world's writing systems that can be digitized. Version 17.0 defines 159,801 characters and 172 scripts used in various ordinary, literary, academic and technical contexts.
Unicode has largely supplanted the previous environment of myriad incompatible character sets used within different locales and on different computer architectures. The entire repertoire of these sets, plus many additional characters, were merged into the single Unicode set. Unicode is used to encode the vast majority of text on the Internet, including most web pages, and relevant Unicode support has become a common consideration in contemporary software development. Unicode is ultimately capable of encoding more than 1.1 million characters.
Gallery (14)
Available in 122 languages
- Español
- Français
- Deutsch
- 中文
- 日本語
- Русский
- Português
- Italiano
- العربية
- हिन्दी
- Afrikaans
- Albanian
- Albanian
- Amharic
- Armenian
- Assamese
- Asturian
- Azerbaijani
Show 103 more
- Bahasa Indonesia
- Bangla
- Basque
- Bavarian
- be_x_old
- Belarusian
- Bhojpuri
- blk
- Bosnian
- Breton
- Bulgarian
- Burmese
- Catalan
- cdo
- Central Kurdish
- Cherokee
- Chuvash
- Croatian
- Czech
- Danish
- Esperanto
- Estonian
- Filipino
- Finnish
- Galician
- Georgian
- Greek
- Gujarati
- Hakka Chinese
- Hebrew
- Hungarian
- Icelandic
- Iloko
- Interlingua
- Irish
- Javanese
- Kannada
- Kashmiri
- Kazakh
- Kurdish
- Kyrgyz
- Latin
- Latvian
- Lingua Franca Nova
- Lithuanian
- Low German
- Macedonian
- Maithili
- Malay
- Malayalam
- Marathi
- Mari
- Mingrelian
- mnw
- Mongolian
- Nederlands
- Nepali
- Newari
- Norwegian
- Norwegian Nynorsk
- Occitan
- Papiamento
- Pashto
- Polski
- Punjabi
- Romanian
- Sanskrit
- Santali
- Scots
- Serbian
- Serbian (Latin)
- simple
- Sindhi
- Sinhala
- Slovak
- Slovenian
- Sundanese
- Svenska
- Swahili
- Tajik
- Tamil
- Telugu
- Tiếng Việt
- Toki Pona
- Türkçe
- Tuvinian
- Ukrainian
- Urdu
- Uyghur
- Uzbek
- Walloon
- Welsh
- Western Panjabi
- Wu Chinese
- Yakut
- Yiddish
- Yoruba
- zh_classical
- zh_min_nan
- zh_yue
- فارسی
- ไทย
- 한국어
via Wikidata sitelinks · CC0