Command Palette

Search for a command to run...

Back to the lesson: Topic 6.8 — Characters, Unicode and Encodings
Core Java · Example 2 of 3 Java 7+

Encodings: bytes, mojibake and the replacement character

UTF-16 output starts with FE FF, a byte-order mark. Unencodable characters become ? (byte 3F) when encoding, and malformed bytes become U+FFFD when decoding.

Unicode gives every character in every script a number called a code point; a Java char is a 16-bit UTF-16 code unit, so characters beyond U+FFFF (like most emoji) take two chars. When text becomes bytes (files, network), an encoding such as UTF-8 decides the bytes, and using the wrong one on either side garbles the text.

Change the code and press Run (Ctrl+Enter). Try to predict the output first, then break it on purpose and read the error. Your edits are saved and match the lesson page.

Practice questions

Write the code in the editor, run it, then open the model answer to compare.

01

Write static String codePoints(String s) that returns the code points of s as U+XXXX values separated by spaces. Print it for "Hi" and for "\u0915\u093E" (the Hindi syllable kaa).

02

Write static int utf8Length(String s) that returns how many bytes s takes in UTF-8, and print it for "chai", "caf\u00e9" and "\uD83D\uDE00".

03

Write static boolean sameText(String a, String b) that compares two Strings after NFC normalization. Test it with "caf\u00e9" and "cafe\u0301".

Explain it without notes

01

What is the difference between a code point, a char and a byte?

02

Why does "\uD83D\uDE00".length() return 2, and how do you count real characters?

03

What is mojibake, how does it happen, and how do you prevent it?

04

What did JEP 400 change in Java 18, and why should you still pass a charset?

05

Two Strings look identical on screen but equals returns false. What could cause this and how do you fix it?

Encodings: bytes, mojibake and the replacement character Java 7+
Sign in to run this example in your browser.

Expected output

UTF-8   (5 bytes): 63 61 66 C3 A9
Latin-1 (4 bytes): 63 61 66 E9
UTF-16  (10 bytes): FE FF 00 63 00 61 00 66 00 E9
emoji in UTF-8: F0 9F 98 80
Hindi ka in UTF-8: E0 A4 95
decoded as UTF-8:   caf\u00E9, equals original: true
decoded as Latin-1: caf\u00C3\u00A9, length 5
cut-off UTF-8 decodes to: c\uFFFD
emoji to Latin-1: 3F