Command Palette

Search for a command to run...

Hectal
PHASE 6Advanced ~15 min· topic 8 of 8

Topic 6.8

Measuring Network Performance: curl Timings, mtr & iperf3

In one line

'The network is slow' becomes solvable once you measure the right thing. curl -w breaks a request into DNS, TCP connect, TLS, time to first byte, and total. mtr shows latency and loss for every hop along the path. iperf3 measures raw throughput between two machines you control. Measure from where users are, compare against a baseline, and fix the slowest phase.

0/8 · 0%

Think of it like this

A slow food delivery. Was it the time to find the address, the traffic on the road, waiting at the restaurant, or the walk up the stairs? Timing each step tells you who to talk to.

Words you'll meet

New words in this topic, in plain English. Come back here whenever one feels fuzzy.

Time to first byte (TTFB)
The time until the first byte of the response arrives.
mtr
A tool that continuously traces the route to a host, showing loss and latency per hop.
iperf3
A tool for measuring maximum throughput between two hosts.
Percentile (p95)
The value below which 95% of measurements fall, showing the slow tail.
Baseline
Normal measurements kept for comparison when something seems wrong.
Packet loss
Packets that never arrive, causing retransmissions and slowdowns.

Step by step

01Splitting one slow request

Meera says the Tiffin menu takes ages to load in Guwahati. Arjun reproduces it from a test machine in the region with a curl timing template, so each phase has a number.

curl-timing.txtwhole filetext
     dns: %{time_namelookup}s
 connect: %{time_connect}s
     tls: %{time_appconnect}s
    ttfb: %{time_starttransfer}s
   total: %{time_total}s
    size: %{size_download} bytes
terminal
$ curl -s -o /dev/null -w @curl-timing.txt https://api.tiffin.in/menu?city=guwahati
── expected output ──
dns: 0.004s
connect: 0.052s
tls: 0.108s
ttfb: 1.842s
total: 1.861s
size: 41822 bytes
DNS, TCP, and TLS took about 0.1 s together. The server took about 1.7 s before sending the first byte.
Splitting one slow requestdiagram
Rendering diagram…

02Not the network, the server

Time to first byte dominates, so the network isn't the problem: the menu query is slow for that city. The fix is an index on (city_id, available) in the database, not a faster link. After the fix, ttfb drops to 0.19 s.

03When it really is the network

Another day, downloads from the Mumbai servers to a partner in Kolkata are slow and connect times are erratic. mtr shows loss starting at one hop and continuing to the destination, which is real loss on that link. iperf3 between two test machines confirms throughput far below the link's capacity.

terminal
$ mtr -rwzc 100 partner-kol.example.net
iperf3 -c 198.51.100.24 -t 10 -P 4 | tail -3
── expected output ──
HOST Loss% Snt Last Avg Best Wrst StDev
1. AS16509 10.0.0.1 0.0% 100 0.3 0.3 0.2 0.9 0.1
2. AS16509 52.95.66.1 0.0% 100 1.1 1.2 0.9 3.2 0.3
3. AS9498 182.79.152.41 0.0% 100 2.0 2.3 1.8 6.1 0.6
4. AS9498 182.79.142.18 6.0% 100 28.4 31.9 27.6 88.2 9.8
5. AS9498 116.119.57.33 6.0% 100 29.9 33.4 28.8 91.0 10.1
6. ??? 198.51.100.24 7.0% 100 31.2 34.0 29.9 95.3 10.4
[SUM] 0.00-10.00 sec 112 MBytes 94.0 Mbits/sec 1824 sender
[SUM] 0.00-10.04 sec 110 MBytes 91.9 Mbits/sec receiver
Loss appears at hop 4 and continues to the end: real loss on that provider's link, and 1,824 retransmissions in the iperf run. Arjun sends the mtr report to the provider.

Break it on purpose

Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.

Break #1

Chasing a scary mtr hop

Priya sees 60% loss at one hop in an mtr report and opens an urgent ticket with the provider.

terminal
$ mtr -rwc 50 api.tiffin.in
── what you'll see ──
HOST Loss% Snt Avg
3. 182.79.152.41 0.0% 50 2.3
4. 182.79.142.18 60.0% 50 31.9
5. 116.119.57.33 0.0% 50 33.4
6. 203.0.113.40 0.0% 50 34.0

Myth vs fact

Myth

Slow requests are usually a network problem.

Fact

Timing the phases often shows the time is spent waiting for the server. Measure before blaming the network.

Pro corner

Extra depth for experienced readers. New to this? Skip it for now and come back later.

  • ▸

    Real user monitoring (the browser's Navigation and Resource Timing APIs, or mobile SDKs) reports the same phases from real users' devices, by region and network type. Synthetic checks (curl timings from several regions on a schedule) give a steady baseline. Use both.

Remember this

  1. 1

    curl timings (-w): time_namelookup (DNS), time_connect (TCP), time_appconnect (TLS), time_starttransfer (first byte = server think time + everything before), time_total. Each value is cumulative from the start.

  2. 2

    mtr combines ping and traceroute, repeating them to show per-hop loss and latency. Loss at one middle hop that doesn't continue to later hops is usually just that router rate-limiting ICMP, not real loss.

  3. 3

    iperf3 measures throughput between a server (iperf3 -s) and a client (iperf3 -c host). Use -P for parallel streams and -R to test the reverse direction.

  4. 4

    Latency vs throughput problems look different: high time to connect means distance or a slow network path, high time to first byte with fast connect means a slow server, and slow total with fast first byte means bandwidth or loss on large responses.

  5. 5

    Measure from the user's side (different regions, mobile networks) as well as from servers, using percentiles (p50, p95, p99), not averages.

  6. 6

    Keep a baseline: numbers from a normal day make it obvious what changed during an incident.

Explain it without notes

01

How do you tell a slow network from a slow server with curl?

02

Why can loss at a middle hop in mtr be ignored?

Practice

01

Measure the timing phases of a website you use from your machine.

02

Measure throughput between two machines you control.

Trade-offs

  • ↔

    Synthetic tests are consistent and easy to compare but don't match real users' networks. Real user data is realistic but noisy. iperf3 needs machines at both ends and uses bandwidth during the test, so avoid running it on busy production links.

Done when you can

  • I can split a request's time into phases with curl.

  • I read mtr correctly, ignoring ICMP rate limiting.

  • I can measure throughput with iperf3 and keep baselines.