Topic 6.4
Essential String Methods
In one line
The String class has a method for almost every everyday text job: measuring, searching, slicing, trimming, changing case, replacing, splitting, joining and repeating. Knowing their exact rules (end-exclusive indexes, regex arguments, trim vs strip, how split drops empty pieces) prevents most text bugs.
Think of it like this
A Swiss Army knife. One handle, many blades: a knife for cutting, scissors for trimming, a magnifier for finding small things. You don't build a new tool every time; you learn which blade does which job, and the one blade that's sharper than it looks. String is that knife, and this topic is a tour of its blades, including the two that cut you if you hold them wrong (split and replaceAll, which secretly take a regular expression).
Words you'll meet
New words in this topic, in plain English. Come back here whenever one feels fuzzy.
- Index
- A position number inside a String, starting at 0 for the first character.
- End-exclusive
- A range that includes the start position but stops just before the end position, written
[start, end). - Whitespace
- Characters you can't see but that take space: spaces, tabs, line breaks, and some special Unicode spaces.
- Regular expression (regex)
- A mini-language for describing text patterns, such as "one or more digits". Topic 6.7 teaches it.
- Delimiter
- The character or text that separates pieces, like the comma in
tea,samosa,cake. - Locale
- A setting that says which language and country rules to follow, such as how to upper-case letters or format numbers.
- Stream
- A sequence of values you can process step by step with operations like
mapandfilter(Phase 10).
Step by step
01Indexes and the end-exclusive rule
In "samosa", s is at 0 and the last a at 5. substring(2, 5) returns indexes 2, 3, 4: "mos". The end is excluded, which makes two things easy: the length is end - begin, and splitting at a point is s.substring(0, k) + s.substring(k) with no off-by-one.
begin may equal end (you get ""), and end may equal length(). Anything else outside 0..length() throws. Since Java 9 the message shows the range: Range [2, 10) out of bounds for length 4 on JDK 21.
02Searching: indexOf and friends
indexOf returns a position or -1. That -1 is a trap: s.substring(s.indexOf(":") + 1) silently returns the whole String when there's no colon (because -1 + 1 is 0). Check for -1 before using the result.
To find every occurrence, loop with the from argument: for (int i = s.indexOf(x); i >= 0; i = s.indexOf(x, i + 1)). Each search is O(n·m) in the worst case for a pattern of length m; for heavy searching, see string algorithms like KMP in the DSA course (/dsa).
String line = "key=value";
int eq = line.indexOf('=');
if (eq >= 0) {
String key = line.substring(0, eq); // "key"
String value = line.substring(eq + 1); // "value"
}03trim vs strip, isEmpty vs isBlank
trim() is from Java 1.0 and predates Unicode thinking: it removes any character with a code <= 32 from both ends. That includes control characters, but not Unicode spaces like the no-break-ish em space \u2003 or the ideographic space \u3000 used in Chinese and Japanese text.
strip() (Java 11) uses Character.isWhitespace, so it handles those spaces. (One catch: the no-break space \u00A0 is not "whitespace" in Java's definition, so neither method removes it. Replace it explicitly when cleaning web text.)
isBlank() is true when the String is empty or every character is whitespace. It's the right check for "did the user type anything?".
04replace vs replaceAll: literal vs regex
replace(".", "-") treats . as a dot. replaceAll(".", "-") treats it as the regex "any character" and replaces everything. Despite the names, both replace all occurrences; the difference is literal vs regex.
In replaceAll's replacement text, $ and \ are special too: $1 means "group 1". So replaceAll("X", "$5") throws IndexOutOfBoundsException: No group 5. Use Matcher.quoteReplacement or plain replace when the replacement is ordinary text.
05split: the method with the most surprises
split(regex) cuts the String around each match. Three rules catch people out. First, the argument is a regex, so split(".") and split("|") don't do what they look like; escape them as split("\\.") or use Pattern.quote(".").
Second, trailing empty strings are removed: "a,b,,".split(",") has length 2. Pass a limit of -1 to keep them, which matters for CSV-like data where empty last columns are real.
Third, a positive limit n caps the number of pieces: "a=b=c".split("=", 2) gives [a, b=c], perfect for key=value lines whose value may contain =. A leading empty string is kept when the input starts with a delimiter (",a".split(",") is [, a]), except that a zero-width match at the start never produces one.
06Changing case safely
toUpperCase() with no argument uses Locale.getDefault(). On a machine set to Turkish, "title".toUpperCase() gives TİTLE with a dotted capital İ, which breaks comparisons against "TITLE". This bug has hit real libraries and build tools.
For anything a machine reads (enum names, HTTP headers, file extensions, map keys) use toUpperCase(Locale.ROOT). For text shown to a person, use their locale. Upper-casing can also change the length: "ß".toUpperCase() is "SS".
07Java 11+ helpers: repeat, lines, strip, isBlank
Java 11 added repeat(n), lines(), strip/stripLeading/stripTrailing and isBlank(). Java 12 added indent(n) and transform(f); Java 15 added stripIndent(), translateEscapes() and formatted(...) (Topics 6.5 and 6.6).
lines() is lazy and handles all three line endings, so prefer it over split("\n"), which leaves a stray \r on every line of a Windows file.
String bar = "=".repeat(20); // a 20-character divider
long nonBlank = text.lines()
.filter(l -> !l.isBlank())
.count(); // count the non-empty linesTry it yourself
- 1
Parse a log line
Take
"2026-10-04 ERROR disk full on /data". Using onlyindexOfandsubstring, print the date, the level and the message on separate lines. Predict what happens if the line has no second space, then guard against it. - 2
Keep the empty columns
In the split example, change the CSV to
"name,age,,"(two empty last columns). Predict the lengths with no limit and with-1, then run. Which one would you use to read a spreadsheet export? - 3
Count words properly
Count the words in
" chai is hot ". First trysplit(" ")and print the array: what do the empty strings come from? Then usestrip()followed bysplit("\\s+")and compare (Topic 6.7 explains the regex\s+).
Code & diagrams
Expected output
length: 18
charAt(6): s
indexOf(chai): 0
indexOf(chai, 1): 14
lastIndexOf(chai): 14
indexOf(tea): -1
contains(samosa): true
startsWith(chai): true
endsWith(samosa): false
substring(6, 12): [samosa]
substring(14): [chai]
occurrences of chai: 2
key [city], value [Pune]\u2003 and \u00DF are Unicode escapes (Topic 6.8), used so the source file stays plain ASCII.
Expected output
trim: [hello]
strip: [hello]
stripLeading: [hello ]
stripTrailing: [ hello]
trim length: 4
strip length: 2
"".isEmpty(): true
" ".isEmpty(): false
" ".isBlank(): true
MASALA CHAI / masala chai
upper-case of strasse with sharp s: STRASSE
original unchanged: Masala Chaisplit(".") matches every character, so every piece is empty, and empty trailing pieces are all removed: length 0.
Expected output
replace(., -): 10-0-0-1
replaceAll(., -): --------
replaceFirst: tea chai
split(".") length: 0
split("\."): [10, 0, 0, 1]
split(,): [a, b, , c]
split(,, -1): [a, b, , c, , ]
split(,, 2): [a, b,,c,,]
leading: [, a]
pipe trap: [a, |, b]
key token, value abc=123Expected output
tea + samosa + cake
ababab -----
lines: 4
non-blank, upper: [ONE, TWO, THREE]
vowels: 5
valueOf: 42true null
sorted: eilnst, anagrams: trueBreak it on purpose
Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.
Break #1
substring past the end
Call "chai".substring(2, 10).
Break #2
split on a dot
Write String first = "10.0.0.1".split(".")[0];.
Break #3
A dollar sign in replaceAll
Write "cost: X".replaceAll("X", "$5") to insert a price.
Myth vs fact
Myth
replace replaces only the first match and replaceAll replaces all.
Fact
Both replace every occurrence. replace is literal; replaceAll uses a regex. For only the first match, use replaceFirst (also regex).
Myth
trim() and strip() are the same.
Fact
trim removes characters with codes up to 32; strip removes Unicode whitespace. They differ for characters like \u2003 and for control characters such as \u0000, which trim removes and strip keeps.
Myth
split returns one piece per delimiter plus one.
Fact
Not with the default limit: trailing empty pieces are dropped, so "a,,".split(",") has length 1. Use split(regex, -1) to keep every piece.
Myth
toUpperCase() behaves the same on every machine.
Fact
It uses the default locale. Pass Locale.ROOT for machine-readable text to avoid the Turkish dotted-I problem.
Pro corner
Extra depth for experienced readers. New to this? Skip it for now and come back later.
- ▸
String.splithas a fast path that skips the regex engine when the pattern is a single character that isn't a regex metacharacter (or an escaped one, like"\\."). Any longer pattern compiles a newPatternon every call; in a hot loop, precompile withPattern.compileand callpattern.split(s). - ▸
indexOfandcontainsare HotSpot intrinsics with SIMD implementations, but the algorithm is still naive in the worst case. For repeated searches of one pattern in huge text, a precompiled regex or a dedicated algorithm (KMP, Boyer-Moore) can win. - ▸
String.valueOf(char[])andnew String(char[])copy the array;String.copyValueOfis identical and exists for historical reasons.String.valueOf(Object)returns"null"for null, butString.valueOf((char[]) null)throwsNullPointerException, because overload resolution picks thechar[]version for a barenullliteral. - ▸
toLowerCase/toUpperCaseshort-circuit and returnthiswhen nothing changes (current JDKs), and handle special cases like Greek final sigma, Lithuanian dots and the German sharp s, which is why their result can be a different length from the input.
Remember this
- 1
Measuring and reading.
length()counts UTF-16chars (Topic 6.8).charAt(i)reads one, with indexes from0tolength() - 1.isEmpty()is true only for"";isBlank()(Java 11) is also true for text that's only whitespace. Every method that takes an index throwsStringIndexOutOfBoundsExceptionwhen it's out of range. - 2
Searching.
indexOf(x)returns the first position of acharor String, or-1;indexOf(x, from)starts looking atfrom;lastIndexOfsearches backwards.contains,startsWithandendsWithreturn booleans. All are case-sensitive and none uses regular expressions. To search ignoring case, useregionMatches(true, ...)or lower-case both sides withLocale.ROOT. - 3
Slicing.
substring(begin)takes frombeginto the end;substring(begin, end)takes[begin, end), so its length isend - begin. The result is a new String (a copy since Java 7u6).subSequencedoes the same but returns aCharSequence.toCharArray()gives you an editable copy of the characters. - 4
Cleaning up.
trim()removes leading and trailing characters with codes up to' '(space, tab, newline, other control characters).strip(),stripLeading()andstripTrailing()(Java 11) remove Unicode whitespace as defined byCharacter.isWhitespace, which includes characters like the em space (\u2003) thattrimmisses. Preferstripin new code.toUpperCase()/toLowerCase()depend on the default locale; passLocale.ROOTfor machine text like keys and protocol words. - 5
Replacing and splitting.
replace(char, char)andreplace(CharSequence, CharSequence)replace every literal occurrence.replaceAll,replaceFirst,splitandmatchestake a regular expression (Topic 6.7), so.means "any character" and|means "or".splitalso drops trailing empty strings unless you pass a negative limit:"a,b,,".split(",")gives[a, b], butsplit(",", -1)gives[a, b, , ]. - 6
Building.
String.join(sep, parts)(Java 8),repeat(n)(Java 11),String.valueOf(anything)(null-safe:String.valueOf((Object) null)is"null"),concat, andlines()(Java 11, aStream<String>of lines split on\n,\ror\r\n).chars()(Java 8) streams thecharvalues asints. Since Strings are immutable, every one of these returns a new String or stream; none changes the original.
Explain it without notes
Explain the end-exclusive rule of substring and why it's convenient.
What's the difference between trim() and strip(), and between isEmpty() and isBlank()?
Why does "1.2.3".split(".") return an empty array, and how do you fix it?
Explain the limit argument of split.
Why should you pass a Locale to toUpperCase for machine-readable text?
Practice
Write static String initials(String fullName) that returns the upper-case initials of each word, e.g. "amit kumar singh" gives "AKS". Words are separated by single spaces.
Write static String maskEmail(String email) that keeps the first letter of the name and the whole domain, e.g. "ravi@hectal.in" gives "r***@hectal.in". Use indexOf and repeat.
Count how many times the letter a appears in "masala chai" without a loop, using replace and length.
Trade-offs
- ↔
splitandreplaceAllare short and expressive but compile a regex each call (outside the single-character fast path). In hot code, precompile aPatternor useindexOf-based parsing. - ↔
striphandles Unicode whitespace correctly;trimalso removes control characters. Pickstripfor user text, and be explicit about other characters (like the no-break space) you want removed. - ↔
Chaining String methods (
s.strip().toLowerCase(Locale.ROOT).replace(...)) is readable but creates an intermediate String per step. That's fine for normal sizes; for megabytes of text processed in a loop, a single pass with aStringBuildercan be much cheaper.
Done when you can
Done when you can slice any String with
substringwithout off-by-one errors.Done when you always check
indexOffor-1before using it.Done when you know which methods take a regex (
split,replaceAll,replaceFirst,matches).Done when you can predict
splitresults with limits 0, negative and positive.Done when you use
strip/isBlankand passLocale.ROOTfor machine text.Done when you can use
join,repeat,linesandcharsto build and inspect text.