Topic 16.7
Ride Sharing: Locations, Matching and Trips
In one line
Ride sharing separates fast-changing ephemeral data (driver locations every few seconds, kept in memory with geospatial indexes) from durable business data (riders, drivers, vehicles, trips, pricing, payments, ratings in a relational store). Matching queries nearby available drivers from the geo index; the trip record follows a state machine with an assignment that must be exclusive.
Think of it like this
A taxi dispatcher's board. Pins showing where every cab is right now are moved constantly and nobody archives them; the logbook of completed journeys, fares and complaints is kept for years.
Key ideas
- 01
Locations: 1M drivers × one update per 4 s = 250K writes/sec. Store the latest position in Redis GEO or an in-memory geo index partitioned by city (H3 or S2 cells), with a TTL so stale drivers disappear. Stream raw pings to Kafka for trip tracing and analytics.
- 02
Matching: query drivers within radius r of the rider in the same city cell (
GEOSEARCH ... BYRADIUS 2 km ASC COUNT 20), filter availability, rank by ETA, and offer to one driver at a time. - 03
Exclusive assignment:
UPDATE driver_state SET status = 'assigned', trip_id = ? WHERE driver_id = ? AND status = 'available'(or a Redis SET NX lease), so one driver never gets two trips. - 04
Trip:
trip(id, rider_id, driver_id, vehicle_id, status, requested_at, pickup geography, dropoff geography, fare_minor, surge_multiplier)with a state machine and a trip_event history. PostGIS geography columns with GiST indexes for durable geo queries. - 05
Pricing snapshots (surge, rate card version) are stored on the trip at request time; payments use the ledger design; ratings table with UNIQUE (trip_id, rater_role).
Code & diagrams
CREATE EXTENSION IF NOT EXISTS postgis;
CREATE TABLE trip (
id bigint PRIMARY KEY,
rider_id bigint NOT NULL, driver_id bigint, vehicle_id bigint,
status text NOT NULL CHECK (status IN ('requested','accepted','arrived','in_progress','completed','cancelled')),
pickup geography(Point, 4326) NOT NULL,
dropoff geography(Point, 4326),
rate_card_version int NOT NULL, surge numeric(4,2) NOT NULL DEFAULT 1.00,
fare_minor bigint, requested_at timestamptz NOT NULL DEFAULT now()
);
CREATE INDEX ON trip USING gist (pickup);
CREATE INDEX ON trip (rider_id, requested_at DESC);
CREATE INDEX ON trip (driver_id, requested_at DESC);
-- one active trip per driver
CREATE UNIQUE INDEX one_active_trip_per_driver ON trip (driver_id)
WHERE status IN ('accepted','arrived','in_progress');# driver location updates (per city key), plus a heartbeat with TTL
GEOADD drivers:blr 77.5946 12.9716 driver:881
SET driver:881:alive 1 EX 15
# nearby drivers for a rider
GEOSEARCH drivers:blr FROMLONLAT 77.6000 12.9750 BYRADIUS 2 km ASC COUNT 20 WITHDISTInterview problem
The problem
Design Uber's data layer
Design storage for 5M daily trips across 500 cities, 1M active drivers sending locations every 4 s, matching within 2 seconds, fare calculation, and trip history for riders and drivers.
When it breaks
Writing every location update to PostgreSQL
What you see
250K row updates/sec create massive WAL, bloat and replication lag; the trip database slows for everyone.
Fix & prevent
Keep current locations in memory; persist traces via a stream to an append-optimised store; write to the OLTP database only on trip state changes.
Explain it without notes
How do you prevent a driver from being assigned two trips at once?
Practice
Estimate location write rate for 300K drivers updating every 3 s.
Trade-offs
- ↔
Splitting ephemeral and durable data keeps each store efficient, at the cost of two systems and stream plumbing.
Done when you can
I can design geo location storage, matching, exclusive assignment and trip records.