Command Palette

Search for a command to run...

PHASE 6Intermediate ~33 min· topic 8 of 8

Topic 6.8

Characters, Unicode and Encodings

In one line

Unicode gives every character in every script a number called a code point; a Java char is a 16-bit UTF-16 code unit, so characters beyond U+FFFF (like most emoji) take two chars. When text becomes bytes (files, network), an encoding such as UTF-8 decides the bytes, and using the wrong one on either side garbles the text.

Think of it like this

A giant numbered sticker book. Every sticker in the world (every letter, digit, symbol and emoji, in every language) has its own number in one big catalogue. To send a sticker to a friend by post, you don't send the sticker; you write its number on paper. Different countries use different ways of writing numbers on the paper, and if your friend reads your paper using the wrong system, they look up the wrong sticker. The catalogue is Unicode, the number is a code point, and the way of writing numbers on paper is an encoding like UTF-8.

Words you'll meet

New words in this topic, in plain English. Come back here whenever one feels fuzzy.

Unicode
The worldwide standard that gives every character in every writing system its own number.
Code point
A character's Unicode number, written like U+0041 for A.
Code unit
The basic chunk an encoding uses. In UTF-16 (Java Strings) it's 16 bits, one char.
Surrogate pair
Two chars that together stand for one character above U+FFFF, such as an emoji.
Encoding / charset
The rules for turning characters into bytes and back, like UTF-8 or ISO-8859-1.
UTF-8
An encoding that uses 1 to 4 bytes per character and is identical to ASCII for English letters. The standard on the web.
Mojibake
Garbled text caused by reading bytes with the wrong encoding, like café instead of café.
Normalization
Converting text so that characters that look the same are stored the same way, such as NFC form.
Grapheme cluster
What a reader sees as one character, which may be made of several code points (a letter plus accents, or a family emoji).

Step by step

01From characters to numbers

Computers store numbers, not letters. Early systems used ASCII, 128 codes for English letters, digits and punctuation. Every country then invented its own 256-code extension, and they clashed: the same byte meant é in one and a Greek letter in another.

Unicode ended that by giving every character one universal code point. Java was designed in the 1990s when Unicode fitted in 16 bits, which is why char is 16 bits. Unicode later grew past U+FFFF, and Java 5 (JSR 204) added surrogate-pair support and the code point APIs.

From characters to numbersdiagram
Rendering diagram…

02char, int and code points

A char is an unsigned 16-bit number, so char arithmetic works: (char) ('a' + 1) is 'b', and '7' - '0' is the int 7 (Topic 1.5). Many methods take an int code point instead of a char so they can handle every character.

Casting a code point above U+FFFF to char silently keeps only the low 16 bits: (char) 0x1F600 is U+F600, an unrelated private-use character. Use Character.toChars(cp) or Character.toString(cp) to get the right one or two chars.

Main.javawhole filejava
int cp = 0x1F600;
char[] units = Character.toChars(cp);        // {'\uD83D', '\uDE00'}
String face = Character.toString(cp);        // Java 11: same two chars as a String
System.out.println(face.length());           // 2
System.out.println(face.codePointCount(0, face.length()));   // 1

03Surrogate pairs and why length() lies

"hi \uD83D\uDE00" holds h, i, space and one emoji, but length() is 5 because the emoji takes two chars. charAt(3) returns only the high surrogate, which isn't a character on its own.

Cutting a String between the two halves (with substring, or by truncating to a maximum length) leaves a lone surrogate. When encoded to UTF-8 it becomes ?, and in a database it may be rejected or stored as garbage. Truncate on code-point boundaries: s.substring(0, s.offsetByCodePoints(0, n)).

Remember compact strings (Topic 6.1): a String with only Latin-1 characters uses 1 byte per char internally, but its API still behaves as UTF-16. Add one emoji and the whole String switches to 2 bytes per char.

04Encoding: from String to bytes and back

Whenever text leaves the JVM (a file, a socket, an HTTP body, a database column) it becomes bytes, and something must choose the charset. getBytes(StandardCharsets.UTF_8) and new String(bytes, StandardCharsets.UTF_8) make the choice explicit. Readers and writers take one too: Files.readString(path, UTF_8), new InputStreamReader(in, UTF_8) (Phase 12).

é is 1 byte in Latin-1 (E9) but 2 bytes in UTF-8 (C3 A9). Decode those two UTF-8 bytes as Latin-1 and you get two characters, Ã and ©: that's mojibake. Decoding invalid bytes as UTF-8 doesn't throw; each bad sequence becomes U+FFFD, the replacement character �. Use a CharsetDecoder with CodingErrorAction.REPORT when you need to reject bad input.

Encoding: from String to bytes and backdiagram
Rendering diagram…

05Java 18: UTF-8 by default

Before Java 18, Charset.defaultCharset() came from the OS locale. A program that wrote a file with new FileWriter(f) on Linux (UTF-8) and read it on Windows (windows-1252) garbled every non-ASCII character. JEP 400 made UTF-8 the default for those APIs on every OS from Java 18.

Two things remain platform-dependent: System.out's encoding (the stdout.encoding property, matching the console) and file names. And code that must run on Java 17 still needs explicit charsets. The rule never changes: always pass a charset.

terminal
$ java -XshowSettings:properties -version 2>&1 | grep -E "encoding"
── expected output ──
file.encoding = UTF-8
native.encoding = Cp1252
stderr.encoding = Cp1252
stdout.encoding = Cp1252
sun.io.unicode.encoding = UnicodeLittle
sun.jnu.encoding = Cp1252

06Same look, different code points: normalization

"caf\u00e9" and "cafe\u0301" both display as café, but they are different code point sequences, so equals is false and they hash to different HashMap buckets. Text typed on macOS, copied from PDFs or entered on phones often arrives decomposed.

Normalize at the boundary, when text enters your system: Normalizer.normalize(s, Normalizer.Form.NFC) for storage and comparison. NFKC additionally folds look-alike "compatibility" characters (like the ligature fi to fi), which is useful for search and usernames.

07Unicode escapes are processed before everything else

\uXXXX in Java source is translated by the compiler before it even splits the code into tokens, everywhere: in strings, in identifiers and in comments. That's why // notes are in C:\users\asha fails to compile: \users starts a Unicode escape with invalid hex digits.

It's also why \u000a (a newline) inside a comment can end the comment early and turn the rest of the line into code, a known trick for hiding code in reviews. Write \n in strings, and avoid \u escapes for control characters.

terminal
$ javac Main.java
── expected output ──
Main.java:3: error: illegal unicode escape
// notes are in C:\users\asha
^
1 error

Try it yourself

  1. 1

    Count what a person sees

    In the first example, add the String "\uD83D\uDC4D\uD83C\uDFFD" (thumbs up with a skin-tone modifier). Predict its length() and code point count, then run. A person sees one symbol: why do both numbers disagree with that?

  2. 2

    Make your own mojibake

    In the encodings example, encode "\u0915" (Hindi ka) as UTF-8 and decode it as ISO-8859-1. Predict how many characters you get back before running, using the byte count from the output.

  3. 3

    Truncate safely

    Write a method that cuts a String to at most n code points using offsetByCodePoints. Test it on "hi \uD83D\uDE00!" with n = 4, and compare with substring(0, 4). Print both with the escape helper.

Code & diagrams

length() vs code points, and surrogate pairs Java 7+ New tab

The program prints only ASCII so the output looks the same on every console. Character.getName needs Java 7.

Sign in to run this example in your browser.

Expected output

plain: length 4, code points 4
accented: length 4, code points 4
emoji: length 5, code points 4
codePointAt(3): U+1F600, chars needed: 2
charAt(3) is a high surrogate: true
charAt(4) is a low surrogate: true
code points: U+0068 U+0069 U+0020 U+1F600
name: GRINNING FACE
(char) cast gives U+F600
first 4 code points: 5 chars
Encodings: bytes, mojibake and the replacement character Java 7+ New tab

UTF-16 output starts with FE FF, a byte-order mark. Unencodable characters become ? (byte 3F) when encoding, and malformed bytes become U+FFFD when decoding.

Sign in to run this example in your browser.

Expected output

UTF-8   (5 bytes): 63 61 66 C3 A9
Latin-1 (4 bytes): 63 61 66 E9
UTF-16  (10 bytes): FE FF 00 63 00 61 00 66 00 E9
emoji in UTF-8: F0 9F 98 80
Hindi ka in UTF-8: E0 A4 95
decoded as UTF-8:   caf\u00E9, equals original: true
decoded as Latin-1: caf\u00C3\u00A9, length 5
cut-off UTF-8 decodes to: c\uFFFD
emoji to Latin-1: 3F
Normalization, reversing safely and Unicode digits Java 11+ New tab

The naive reverse puts the low surrogate before the high one, which is no longer a valid character. (?U) is the inline UNICODE_CHARACTER_CLASS flag.

Sign in to run this example in your browser.

Expected output

lengths: 4 vs 5
equals: false
after NFC: equals true, length 4
NFD of composed: cafe\u0301
naive reverse:   b\uDE00\uD83Da
builder reverse: b\uD83D\uDE00a
isLetter(e-acute): true
isDigit(Arabic-Indic 3): true
parseInt of Arabic-Indic 34, plus 1: 35
regex \d on it: false, with (?U): true
Character.toString(0x1F600).length(): 2
next letter after a: b, digit value of '7': 7
Always name the charset for I/O (fragment)java
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;

Path path = Path.of("menu.txt");
Files.writeString(path, text, StandardCharsets.UTF_8);       // Java 11
String back = Files.readString(path, StandardCharsets.UTF_8);

// Risky before Java 18: uses the OS default charset
// new FileReader("menu.txt")   ->   new FileReader("menu.txt", StandardCharsets.UTF_8)  (Java 11)

// See what the JVM is using
System.out.println(java.nio.charset.Charset.defaultCharset());   // UTF-8 on Java 18+

Break it on purpose

Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.

Break #1

A Windows path in a comment

Add the comment // notes are in C:\users\asha to main.

terminal
$ javac Main.java
── what you'll see ──
Main.java:3: error: illegal unicode escape
// notes are in C:\users\asha
^
1 error

Break #2

Decode UTF-8 bytes with the wrong charset

Write a file with Files.writeString(p, "café", UTF_8) and read it with new String(Files.readAllBytes(p), StandardCharsets.ISO_8859_1).

terminal
$ java Main.java
── what you'll see ──
café

Break #3

Cut an emoji in half

Truncate "hi \uD83D\uDE00" with substring(0, 4) and write it as UTF-8.

terminal
$ java Main.java
── what you'll see ──
bytes: 68 69 20 3F

Myth vs fact

Myth

A char is a character.

Fact

A char is a UTF-16 code unit. Characters above U+FFFF (most emoji, many rare CJK characters, historic scripts) need two chars, and a visible character can be several code points.

Myth

UTF-8 means every character is one byte.

Fact

Only ASCII is one byte. UTF-8 uses 2 bytes for most European accented letters, 3 for Indian and most East Asian scripts, and 4 for emoji.

Myth

Since Java 18 encodings are no longer my problem.

Fact

Java 18 made UTF-8 the default for file APIs, but other systems (databases, old services, Windows consoles) may still use something else, and Java 17 code doesn't get the new default. Always pass a charset explicitly.

Myth

\d matches any digit in any script.

Fact

By default \d is [0-9] only. Character.isDigit and Integer.parseInt accept Unicode digits (like Arabic-Indic), and the regex needs UNICODE_CHARACTER_CLASS ((?U)) to do the same.

Pro corner

Extra depth for experienced readers. New to this? Skip it for now and come back later.

  • ▸

    Java's UTF-16 choice dates from Java 1.0 (1996), when Unicode was a 16-bit code. Since Java 9, compact strings store Latin-1 text in 1 byte per char internally, but every public API still speaks UTF-16 indexes, so charAt and length are O(1) and code point operations like codePointCount are O(n).

  • ▸

    The JVM's class files store string constants in modified UTF-8: the null character is two bytes (C0 80) and supplementary characters are stored as two 3-byte surrogates (CESU-8 style). That's why DataOutputStream.writeUTF output isn't standard UTF-8 and other languages can misread it.

  • ▸

    Integer.parseInt accepting non-ASCII digits (via Character.digit) has caused validation bypasses: a field checked with \\d+ rejects \u0663, while another path parses it happily. Validate and parse with the same rules, and normalize (NFKC) identifiers before comparing them, to block look-alike (homoglyph) tricks.

  • ▸

    Case conversion isn't always one-to-one: "ß".toUpperCase() is "SS", and Character.toUpperCase(int) can't express that, so use the String methods for text. equalsIgnoreCase compares char by char and won't treat ß and SS as equal; full case folding needs ICU4J or a Collator.

Remember this

  1. 1

    Unicode assigns each character a code point, written U+ plus hex: A is U+0041, é is U+00E9, the Hindi letter क is U+0915, the grinning-face emoji is U+1F600. Code points run from U+0000 to U+10FFFF (about 1.1 million). The first 65,536 (U+0000 to U+FFFF) are the Basic Multilingual Plane (BMP); everything above is supplementary.

  2. 2

    A Java **char is 16 bits, so it can hold only U+0000 to U+FFFF. Java strings are sequences of UTF-16 code units**: a BMP character is one char, and a supplementary character is a surrogate pair of two chars (a high surrogate in U+D800 to U+DBFF followed by a low surrogate in U+DC00 to U+DFFF). That's why "hi \uD83D\uDE00".length() is 5, not 4: length() counts chars, not characters.

  3. 3

    For full-Unicode work use the code point APIs (Java 5): codePointAt(i), codePointCount(begin, end), codePoints() (an IntStream, Java 8), offsetByCodePoints, Character.charCount(cp), Character.toChars(cp), Character.toString(int) (Java 11) and the int overloads like Character.isLetter(int). StringBuilder.reverse() keeps surrogate pairs together; reversing a char[] by hand breaks them.

  4. 4

    An encoding (Java calls it a charset) turns code points into bytes and back. UTF-8 uses 1 byte for ASCII, 2 for most European letters, 3 for most Asian scripts and 4 for emoji; it's the web's standard. UTF-16 uses 2 or 4 bytes. ISO-8859-1 (Latin-1) uses exactly 1 byte but only covers U+0000 to U+00FF. Convert with s.getBytes(StandardCharsets.UTF_8) and new String(bytes, StandardCharsets.UTF_8); StandardCharsets arrived in Java 7.

  5. 5

    Mojibake (garbled text like café) happens when bytes are decoded with a different charset than they were encoded with. Before Java 18 the default charset came from the operating system (often windows-1252 on Windows), so code that called getBytes() or new FileReader(file) without a charset behaved differently per machine. JEP 400 made UTF-8 the default in Java 18, but always passing a charset explicitly is still the safe habit.

  6. 6

    One visible character can be several code points. é can be the single code point U+00E9 (precomposed) or e + U+0301 (a combining accent). They look identical but aren't equals. java.text.Normalizer.normalize(s, Normalizer.Form.NFC) converts both to the same form. A grapheme cluster (what a person calls "one character", including flag emoji and family emoji built from several code points) is found with BreakIterator.getCharacterInstance() or the regex \X (Java 9).

Explain it without notes

01

What is the difference between a code point, a char and a byte?

02

Why does "\uD83D\uDE00".length() return 2, and how do you count real characters?

03

What is mojibake, how does it happen, and how do you prevent it?

04

What did JEP 400 change in Java 18, and why should you still pass a charset?

05

Two Strings look identical on screen but equals returns false. What could cause this and how do you fix it?

Practice

01

Write static String codePoints(String s) that returns the code points of s as U+XXXX values separated by spaces. Print it for "Hi" and for "\u0915\u093E" (the Hindi syllable kaa).

02

Write static int utf8Length(String s) that returns how many bytes s takes in UTF-8, and print it for "chai", "caf\u00e9" and "\uD83D\uDE00".

03

Write static boolean sameText(String a, String b) that compares two Strings after NFC normalization. Test it with "caf\u00e9" and "cafe\u0301".

Trade-offs

  • ↔

    UTF-16 Strings give O(1) charAt and length, but those count code units, not characters. Code-point-correct processing is O(n) and more verbose; use it wherever text can contain emoji or rare scripts (which, for user input, is everywhere).

  • ↔

    UTF-8 is compact for ASCII-heavy text and universal on the web; UTF-16 can be smaller for mostly Asian-script text. In practice the interoperability of UTF-8 wins almost every time.

  • ↔

    Normalizing input makes comparison and search reliable, but changes the exact bytes users sent. Normalize for keys and comparisons; keep the original when you must reproduce it exactly (signatures, legal text).

Done when you can

  • Done when you can explain code point, char, code unit and byte, and convert between them.

  • Done when you count and iterate characters with the code point APIs.

  • Done when you always pass a charset when turning text into bytes or back.

  • Done when you can diagnose mojibake from its symptoms and fix the decoding side.

  • Done when you normalize text before comparing it and truncate on code-point boundaries.

  • Done when you know what changed in Java 18 (JEP 400) and what didn't.