Topic 6.8
Characters, Unicode and Encodings
In one line
Unicode gives every character in every script a number called a code point; a Java char is a 16-bit UTF-16 code unit, so characters beyond U+FFFF (like most emoji) take two chars. When text becomes bytes (files, network), an encoding such as UTF-8 decides the bytes, and using the wrong one on either side garbles the text.
Think of it like this
A giant numbered sticker book. Every sticker in the world (every letter, digit, symbol and emoji, in every language) has its own number in one big catalogue. To send a sticker to a friend by post, you don't send the sticker; you write its number on paper. Different countries use different ways of writing numbers on the paper, and if your friend reads your paper using the wrong system, they look up the wrong sticker. The catalogue is Unicode, the number is a code point, and the way of writing numbers on paper is an encoding like UTF-8.
Words you'll meet
New words in this topic, in plain English. Come back here whenever one feels fuzzy.
- Unicode
- The worldwide standard that gives every character in every writing system its own number.
- Code point
- A character's Unicode number, written like U+0041 for
A. - Code unit
- The basic chunk an encoding uses. In UTF-16 (Java Strings) it's 16 bits, one
char. - Surrogate pair
- Two
chars that together stand for one character above U+FFFF, such as an emoji. - Encoding / charset
- The rules for turning characters into bytes and back, like UTF-8 or ISO-8859-1.
- UTF-8
- An encoding that uses 1 to 4 bytes per character and is identical to ASCII for English letters. The standard on the web.
- Mojibake
- Garbled text caused by reading bytes with the wrong encoding, like
caféinstead ofcafé. - Normalization
- Converting text so that characters that look the same are stored the same way, such as NFC form.
- Grapheme cluster
- What a reader sees as one character, which may be made of several code points (a letter plus accents, or a family emoji).
Step by step
01From characters to numbers
Computers store numbers, not letters. Early systems used ASCII, 128 codes for English letters, digits and punctuation. Every country then invented its own 256-code extension, and they clashed: the same byte meant é in one and a Greek letter in another.
Unicode ended that by giving every character one universal code point. Java was designed in the 1990s when Unicode fitted in 16 bits, which is why char is 16 bits. Unicode later grew past U+FFFF, and Java 5 (JSR 204) added surrogate-pair support and the code point APIs.
02char, int and code points
A char is an unsigned 16-bit number, so char arithmetic works: (char) ('a' + 1) is 'b', and '7' - '0' is the int 7 (Topic 1.5). Many methods take an int code point instead of a char so they can handle every character.
Casting a code point above U+FFFF to char silently keeps only the low 16 bits: (char) 0x1F600 is U+F600, an unrelated private-use character. Use Character.toChars(cp) or Character.toString(cp) to get the right one or two chars.
int cp = 0x1F600;
char[] units = Character.toChars(cp); // {'\uD83D', '\uDE00'}
String face = Character.toString(cp); // Java 11: same two chars as a String
System.out.println(face.length()); // 2
System.out.println(face.codePointCount(0, face.length())); // 103Surrogate pairs and why length() lies
"hi \uD83D\uDE00" holds h, i, space and one emoji, but length() is 5 because the emoji takes two chars. charAt(3) returns only the high surrogate, which isn't a character on its own.
Cutting a String between the two halves (with substring, or by truncating to a maximum length) leaves a lone surrogate. When encoded to UTF-8 it becomes ?, and in a database it may be rejected or stored as garbage. Truncate on code-point boundaries: s.substring(0, s.offsetByCodePoints(0, n)).
Remember compact strings (Topic 6.1): a String with only Latin-1 characters uses 1 byte per char internally, but its API still behaves as UTF-16. Add one emoji and the whole String switches to 2 bytes per char.
04Encoding: from String to bytes and back
Whenever text leaves the JVM (a file, a socket, an HTTP body, a database column) it becomes bytes, and something must choose the charset. getBytes(StandardCharsets.UTF_8) and new String(bytes, StandardCharsets.UTF_8) make the choice explicit. Readers and writers take one too: Files.readString(path, UTF_8), new InputStreamReader(in, UTF_8) (Phase 12).
é is 1 byte in Latin-1 (E9) but 2 bytes in UTF-8 (C3 A9). Decode those two UTF-8 bytes as Latin-1 and you get two characters, Ã and ©: that's mojibake. Decoding invalid bytes as UTF-8 doesn't throw; each bad sequence becomes U+FFFD, the replacement character �. Use a CharsetDecoder with CodingErrorAction.REPORT when you need to reject bad input.
05Java 18: UTF-8 by default
Before Java 18, Charset.defaultCharset() came from the OS locale. A program that wrote a file with new FileWriter(f) on Linux (UTF-8) and read it on Windows (windows-1252) garbled every non-ASCII character. JEP 400 made UTF-8 the default for those APIs on every OS from Java 18.
Two things remain platform-dependent: System.out's encoding (the stdout.encoding property, matching the console) and file names. And code that must run on Java 17 still needs explicit charsets. The rule never changes: always pass a charset.
06Same look, different code points: normalization
"caf\u00e9" and "cafe\u0301" both display as café, but they are different code point sequences, so equals is false and they hash to different HashMap buckets. Text typed on macOS, copied from PDFs or entered on phones often arrives decomposed.
Normalize at the boundary, when text enters your system: Normalizer.normalize(s, Normalizer.Form.NFC) for storage and comparison. NFKC additionally folds look-alike "compatibility" characters (like the ligature fi to fi), which is useful for search and usernames.
07Unicode escapes are processed before everything else
\uXXXX in Java source is translated by the compiler before it even splits the code into tokens, everywhere: in strings, in identifiers and in comments. That's why // notes are in C:\users\asha fails to compile: \users starts a Unicode escape with invalid hex digits.
It's also why \u000a (a newline) inside a comment can end the comment early and turn the rest of the line into code, a known trick for hiding code in reviews. Write \n in strings, and avoid \u escapes for control characters.
Try it yourself
- 1
Count what a person sees
In the first example, add the String
"\uD83D\uDC4D\uD83C\uDFFD"(thumbs up with a skin-tone modifier). Predict itslength()and code point count, then run. A person sees one symbol: why do both numbers disagree with that? - 2
Make your own mojibake
In the encodings example, encode
"\u0915"(Hindi ka) as UTF-8 and decode it as ISO-8859-1. Predict how many characters you get back before running, using the byte count from the output. - 3
Truncate safely
Write a method that cuts a String to at most
ncode points usingoffsetByCodePoints. Test it on"hi \uD83D\uDE00!"withn = 4, and compare withsubstring(0, 4). Print both with theescapehelper.
Code & diagrams
The program prints only ASCII so the output looks the same on every console. Character.getName needs Java 7.
Expected output
plain: length 4, code points 4
accented: length 4, code points 4
emoji: length 5, code points 4
codePointAt(3): U+1F600, chars needed: 2
charAt(3) is a high surrogate: true
charAt(4) is a low surrogate: true
code points: U+0068 U+0069 U+0020 U+1F600
name: GRINNING FACE
(char) cast gives U+F600
first 4 code points: 5 charsUTF-16 output starts with FE FF, a byte-order mark. Unencodable characters become ? (byte 3F) when encoding, and malformed bytes become U+FFFD when decoding.
Expected output
UTF-8 (5 bytes): 63 61 66 C3 A9
Latin-1 (4 bytes): 63 61 66 E9
UTF-16 (10 bytes): FE FF 00 63 00 61 00 66 00 E9
emoji in UTF-8: F0 9F 98 80
Hindi ka in UTF-8: E0 A4 95
decoded as UTF-8: caf\u00E9, equals original: true
decoded as Latin-1: caf\u00C3\u00A9, length 5
cut-off UTF-8 decodes to: c\uFFFD
emoji to Latin-1: 3FThe naive reverse puts the low surrogate before the high one, which is no longer a valid character. (?U) is the inline UNICODE_CHARACTER_CLASS flag.
Expected output
lengths: 4 vs 5
equals: false
after NFC: equals true, length 4
NFD of composed: cafe\u0301
naive reverse: b\uDE00\uD83Da
builder reverse: b\uD83D\uDE00a
isLetter(e-acute): true
isDigit(Arabic-Indic 3): true
parseInt of Arabic-Indic 34, plus 1: 35
regex \d on it: false, with (?U): true
Character.toString(0x1F600).length(): 2
next letter after a: b, digit value of '7': 7import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
Path path = Path.of("menu.txt");
Files.writeString(path, text, StandardCharsets.UTF_8); // Java 11
String back = Files.readString(path, StandardCharsets.UTF_8);
// Risky before Java 18: uses the OS default charset
// new FileReader("menu.txt") -> new FileReader("menu.txt", StandardCharsets.UTF_8) (Java 11)
// See what the JVM is using
System.out.println(java.nio.charset.Charset.defaultCharset()); // UTF-8 on Java 18+Break it on purpose
Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.
Break #1
A Windows path in a comment
Add the comment // notes are in C:\users\asha to main.
Break #2
Decode UTF-8 bytes with the wrong charset
Write a file with Files.writeString(p, "café", UTF_8) and read it with new String(Files.readAllBytes(p), StandardCharsets.ISO_8859_1).
Break #3
Cut an emoji in half
Truncate "hi \uD83D\uDE00" with substring(0, 4) and write it as UTF-8.
Myth vs fact
Myth
A char is a character.
Fact
A char is a UTF-16 code unit. Characters above U+FFFF (most emoji, many rare CJK characters, historic scripts) need two chars, and a visible character can be several code points.
Myth
UTF-8 means every character is one byte.
Fact
Only ASCII is one byte. UTF-8 uses 2 bytes for most European accented letters, 3 for Indian and most East Asian scripts, and 4 for emoji.
Myth
Since Java 18 encodings are no longer my problem.
Fact
Java 18 made UTF-8 the default for file APIs, but other systems (databases, old services, Windows consoles) may still use something else, and Java 17 code doesn't get the new default. Always pass a charset explicitly.
Myth
\d matches any digit in any script.
Fact
By default \d is [0-9] only. Character.isDigit and Integer.parseInt accept Unicode digits (like Arabic-Indic), and the regex needs UNICODE_CHARACTER_CLASS ((?U)) to do the same.
Pro corner
Extra depth for experienced readers. New to this? Skip it for now and come back later.
- ▸
Java's UTF-16 choice dates from Java 1.0 (1996), when Unicode was a 16-bit code. Since Java 9, compact strings store Latin-1 text in 1 byte per char internally, but every public API still speaks UTF-16 indexes, so
charAtandlengthare O(1) and code point operations likecodePointCountare O(n). - ▸
The JVM's class files store string constants in modified UTF-8: the null character is two bytes (
C0 80) and supplementary characters are stored as two 3-byte surrogates (CESU-8 style). That's whyDataOutputStream.writeUTFoutput isn't standard UTF-8 and other languages can misread it. - ▸
Integer.parseIntaccepting non-ASCII digits (viaCharacter.digit) has caused validation bypasses: a field checked with\\d+rejects\u0663, while another path parses it happily. Validate and parse with the same rules, and normalize (NFKC) identifiers before comparing them, to block look-alike (homoglyph) tricks. - ▸
Case conversion isn't always one-to-one:
"ß".toUpperCase()is"SS", andCharacter.toUpperCase(int)can't express that, so use the String methods for text.equalsIgnoreCasecompares char by char and won't treatßandSSas equal; full case folding needs ICU4J or aCollator.
Remember this
- 1
Unicode assigns each character a code point, written
U+plus hex:Ais U+0041,éis U+00E9, the Hindi letterकis U+0915, the grinning-face emoji is U+1F600. Code points run from U+0000 to U+10FFFF (about 1.1 million). The first 65,536 (U+0000 to U+FFFF) are the Basic Multilingual Plane (BMP); everything above is supplementary. - 2
A Java **
charis 16 bits, so it can hold only U+0000 to U+FFFF. Java strings are sequences of UTF-16 code units**: a BMP character is onechar, and a supplementary character is a surrogate pair of twochars (a high surrogate in U+D800 to U+DBFF followed by a low surrogate in U+DC00 to U+DFFF). That's why"hi \uD83D\uDE00".length()is 5, not 4:length()countschars, not characters. - 3
For full-Unicode work use the code point APIs (Java 5):
codePointAt(i),codePointCount(begin, end),codePoints()(anIntStream, Java 8),offsetByCodePoints,Character.charCount(cp),Character.toChars(cp),Character.toString(int)(Java 11) and theintoverloads likeCharacter.isLetter(int).StringBuilder.reverse()keeps surrogate pairs together; reversing achar[]by hand breaks them. - 4
An encoding (Java calls it a charset) turns code points into bytes and back. UTF-8 uses 1 byte for ASCII, 2 for most European letters, 3 for most Asian scripts and 4 for emoji; it's the web's standard. UTF-16 uses 2 or 4 bytes. ISO-8859-1 (Latin-1) uses exactly 1 byte but only covers U+0000 to U+00FF. Convert with
s.getBytes(StandardCharsets.UTF_8)andnew String(bytes, StandardCharsets.UTF_8);StandardCharsetsarrived in Java 7. - 5
Mojibake (garbled text like
café) happens when bytes are decoded with a different charset than they were encoded with. Before Java 18 the default charset came from the operating system (often windows-1252 on Windows), so code that calledgetBytes()ornew FileReader(file)without a charset behaved differently per machine. JEP 400 made UTF-8 the default in Java 18, but always passing a charset explicitly is still the safe habit. - 6
One visible character can be several code points.
écan be the single code point U+00E9 (precomposed) ore+ U+0301 (a combining accent). They look identical but aren'tequals.java.text.Normalizer.normalize(s, Normalizer.Form.NFC)converts both to the same form. A grapheme cluster (what a person calls "one character", including flag emoji and family emoji built from several code points) is found withBreakIterator.getCharacterInstance()or the regex\X(Java 9).
Explain it without notes
What is the difference between a code point, a char and a byte?
Why does "\uD83D\uDE00".length() return 2, and how do you count real characters?
What is mojibake, how does it happen, and how do you prevent it?
What did JEP 400 change in Java 18, and why should you still pass a charset?
Two Strings look identical on screen but equals returns false. What could cause this and how do you fix it?
Practice
Write static String codePoints(String s) that returns the code points of s as U+XXXX values separated by spaces. Print it for "Hi" and for "\u0915\u093E" (the Hindi syllable kaa).
Write static int utf8Length(String s) that returns how many bytes s takes in UTF-8, and print it for "chai", "caf\u00e9" and "\uD83D\uDE00".
Write static boolean sameText(String a, String b) that compares two Strings after NFC normalization. Test it with "caf\u00e9" and "cafe\u0301".
Trade-offs
- ↔
UTF-16 Strings give O(1)
charAtandlength, but those count code units, not characters. Code-point-correct processing is O(n) and more verbose; use it wherever text can contain emoji or rare scripts (which, for user input, is everywhere). - ↔
UTF-8 is compact for ASCII-heavy text and universal on the web; UTF-16 can be smaller for mostly Asian-script text. In practice the interoperability of UTF-8 wins almost every time.
- ↔
Normalizing input makes comparison and search reliable, but changes the exact bytes users sent. Normalize for keys and comparisons; keep the original when you must reproduce it exactly (signatures, legal text).
Done when you can
Done when you can explain code point,
char, code unit and byte, and convert between them.Done when you count and iterate characters with the code point APIs.
Done when you always pass a charset when turning text into bytes or back.
Done when you can diagnose mojibake from its symptoms and fix the decoding side.
Done when you normalize text before comparing it and truncate on code-point boundaries.
Done when you know what changed in Java 18 (JEP 400) and what didn't.