Command Palette

Search for a command to run...

Hectal
PHASE 16Advanced ~9 min· topic 7 of 8

Topic 16.7

Ride Sharing: Locations, Matching and Trips

In one line

Ride sharing separates fast-changing ephemeral data (driver locations every few seconds, kept in memory with geospatial indexes) from durable business data (riders, drivers, vehicles, trips, pricing, payments, ratings in a relational store). Matching queries nearby available drivers from the geo index; the trip record follows a state machine with an assignment that must be exclusive.

0/8 · 0%

Think of it like this

A taxi dispatcher's board. Pins showing where every cab is right now are moved constantly and nobody archives them; the logbook of completed journeys, fares and complaints is kept for years.

Key ideas

  1. 01

    Locations: 1M drivers × one update per 4 s = 250K writes/sec. Store the latest position in Redis GEO or an in-memory geo index partitioned by city (H3 or S2 cells), with a TTL so stale drivers disappear. Stream raw pings to Kafka for trip tracing and analytics.

  2. 02

    Matching: query drivers within radius r of the rider in the same city cell (GEOSEARCH ... BYRADIUS 2 km ASC COUNT 20), filter availability, rank by ETA, and offer to one driver at a time.

  3. 03

    Exclusive assignment: UPDATE driver_state SET status = 'assigned', trip_id = ? WHERE driver_id = ? AND status = 'available' (or a Redis SET NX lease), so one driver never gets two trips.

  4. 04

    Trip: trip(id, rider_id, driver_id, vehicle_id, status, requested_at, pickup geography, dropoff geography, fare_minor, surge_multiplier) with a state machine and a trip_event history. PostGIS geography columns with GiST indexes for durable geo queries.

  5. 05

    Pricing snapshots (surge, rate card version) are stored on the trip at request time; payments use the ledger design; ratings table with UNIQUE (trip_id, rater_role).

Code & diagrams

ride.sqlsql
CREATE EXTENSION IF NOT EXISTS postgis;
CREATE TABLE trip (
  id bigint PRIMARY KEY,
  rider_id bigint NOT NULL, driver_id bigint, vehicle_id bigint,
  status text NOT NULL CHECK (status IN ('requested','accepted','arrived','in_progress','completed','cancelled')),
  pickup  geography(Point, 4326) NOT NULL,
  dropoff geography(Point, 4326),
  rate_card_version int NOT NULL, surge numeric(4,2) NOT NULL DEFAULT 1.00,
  fare_minor bigint, requested_at timestamptz NOT NULL DEFAULT now()
);
CREATE INDEX ON trip USING gist (pickup);
CREATE INDEX ON trip (rider_id, requested_at DESC);
CREATE INDEX ON trip (driver_id, requested_at DESC);
-- one active trip per driver
CREATE UNIQUE INDEX one_active_trip_per_driver ON trip (driver_id)
  WHERE status IN ('accepted','arrived','in_progress');
locations.redisbash
# driver location updates (per city key), plus a heartbeat with TTL
GEOADD drivers:blr 77.5946 12.9716 driver:881
SET driver:881:alive 1 EX 15
# nearby drivers for a rider
GEOSEARCH drivers:blr FROMLONLAT 77.6000 12.9750 BYRADIUS 2 km ASC COUNT 20 WITHDIST

Interview problem

The problem

Design Uber's data layer

Design storage for 5M daily trips across 500 cities, 1M active drivers sending locations every 4 s, matching within 2 seconds, fare calculation, and trip history for riders and drivers.

When it breaks

Writing every location update to PostgreSQL

What you see

250K row updates/sec create massive WAL, bloat and replication lag; the trip database slows for everyone.

Fix & prevent

Keep current locations in memory; persist traces via a stream to an append-optimised store; write to the OLTP database only on trip state changes.

Explain it without notes

01

How do you prevent a driver from being assigned two trips at once?

Practice

01

Estimate location write rate for 300K drivers updating every 3 s.

Trade-offs

  • ↔

    Splitting ephemeral and durable data keeps each store efficient, at the cost of two systems and stream plumbing.

Done when you can

  • I can design geo location storage, matching, exclusive assignment and trip records.