Skip to content
Unicode

File:New_Unicode_logo.svg · Wikimedia Commons · See Wikimedia Commons

EntityQ8819· pop 130· linked from 5,522 articles

Also known as The Unicode Standard, Uni-code, Unicode Standard

Unicode (also known as The Unicode Standard and TUS) is a character encoding standard maintained by the Unicode Consortium designed to support the use of text in all of the world's writing systems that can be digitized. Version 17.0 defines 159,801 characters and 172 scripts used in various ordinary, literary, academic and technical contexts.

AI overview

Unicode is a standardized system that assigns unique numerical codes to characters used in all the world's writing systems, allowing computers to store and display text from any language or script. As of its latest version, it encompasses over 159,000 characters representing 172 different scripts, making it possible for digital devices to handle text reliably across languages and cultures.

AI-generated from the Wikipedia summary — may contain errors.

Key facts

Character encoding.name
Unicode
Character encoding.image
New Unicode logo.svg
Character encoding.caption
Logo of the Unicode Consortium
Character encoding.standard
Unicode Standard
Character encoding.lang
168 scripts (list)
Character encoding.encodings
(uncommon) (obsolete)
Character encoding.prev
ISO/IEC 8859, among others

via Wikipedia infobox

Wikidata facts

Official website
unicode.org
Show 10 more facts
Commons category
Unicode
Stack Exchange tag
askubuntu.com/tags/unicode
Commons gallery
Unicode
Unicode range
U+0000-10FFFF
ACM Classification Code (2012)
10011594
software version identifier
17.0.0
issue tracker URL
unicode-org.atlassian.net
publication date
1996-07-00
official blog URL
blog.unicode.org
Sources (6)

via Wikidata · CC0

~62 min read

Article

38 sections
Contents
  • Origin and development
  • {{anchor|Unicode 88}}History
  • Unicode Consortium
  • Scripts covered
  • Proposals for adding scripts
  • Versions
  • Architecture and terminology<span class="anchor" id="Upluslink"></span><!-- Template:U+ links to this paragraph -->
  • Codespace and code points
  • Code planes and blocks
  • General Category property
  • Abstract characters <span class="anchor" id="Alias"></span>
  • Ready-made versus composite characters
  • Ligatures
  • Standardized subsets
  • {{anchor|UTF|UCS}}Mapping and encodings
  • Adoption
  • Operating systems
  • Input methods
  • Email
  • Web
  • Fonts
  • Newlines
  • Issues
  • Character unification
  • Han unification
  • Italic or cursive characters in Cyrillic
  • Localised case pairs
  • Diacritics on lowercase {{serif|I}}
  • Security<span class="anchor" id="Security issues"></span>
  • Mapping to legacy character sets
  • Indic scripts
  • Combining characters
  • Anomalies
  • See also
  • Notes
  • References
  • Further reading
  • External links

Unicode (also known as The Unicode Standard and TUS) is a character encoding standard maintained by the Unicode Consortium designed to support the use of text in all of the world's writing systems that can be digitized. Version 17.0 defines 159,801 characters and 172 scripts used in various ordinary, literary, academic and technical contexts.

Unicode has largely supplanted the previous environment of myriad incompatible character sets used within different locales and on different computer architectures. The entire repertoire of these sets, plus many additional characters, were merged into the single Unicode set. Unicode is used to encode the vast majority of text on the Internet, including most web pages, and relevant Unicode support has become a common consideration in contemporary software development. Unicode is ultimately capable of encoding more than 1.1 million characters.

Gallery (14)

Connections

Categories