length() vs code points, and surrogate pairs
The program prints only ASCII so the output looks the same on every console. Character.getName needs Java 7.
Unicode gives every character in every script a number called a code point; a Java char is a 16-bit UTF-16 code unit, so characters beyond U+FFFF (like most emoji) take two chars. When text becomes bytes (files, network), an encoding such as UTF-8 decides the bytes, and using the wrong one on either side garbles the text.
Change the code and press Run (Ctrl+Enter). Try to predict the output first, then break it on purpose and read the error. Your edits are saved and match the lesson page.
Practice questions
Write the code in the editor, run it, then open the model answer to compare.
Write static String codePoints(String s) that returns the code points of s as U+XXXX values separated by spaces. Print it for "Hi" and for "\u0915\u093E" (the Hindi syllable kaa).
Write static int utf8Length(String s) that returns how many bytes s takes in UTF-8, and print it for "chai", "caf\u00e9" and "\uD83D\uDE00".
Write static boolean sameText(String a, String b) that compares two Strings after NFC normalization. Test it with "caf\u00e9" and "cafe\u0301".
Explain it without notes
What is the difference between a code point, a char and a byte?
Why does "\uD83D\uDE00".length() return 2, and how do you count real characters?
What is mojibake, how does it happen, and how do you prevent it?
What did JEP 400 change in Java 18, and why should you still pass a charset?
Two Strings look identical on screen but equals returns false. What could cause this and how do you fix it?
Expected output
plain: length 4, code points 4
accented: length 4, code points 4
emoji: length 5, code points 4
codePointAt(3): U+1F600, chars needed: 2
charAt(3) is a high surrogate: true
charAt(4) is a low surrogate: true
code points: U+0068 U+0069 U+0020 U+1F600
name: GRINNING FACE
(char) cast gives U+F600
first 4 code points: 5 chars