Command Palette

Search for a command to run...

Hectal
PHASE 8Advanced ~17 min· topic 5 of 5

Topic 8.5

Capstone: Tiffin's Production Reference Architecture

In one line

The capstone puts the whole course together: Tiffin's production architecture on AWS, from DNS to database, and from the account structure to the deploy pipeline. For every part, the question is why it's there, what it protects against, and what it costs. Being able to draw this, explain each choice, and name its trade-offs is what system design interviews and real architecture reviews ask for.

0/5 · 0%

Think of it like this

The final map of a city you've been exploring street by street. You already know each street (topics). Now you see how they connect, why the hospital is near the highway, and where the flood barriers are.

Words you'll meet

New words in this topic, in plain English. Come back here whenever one feels fuzzy.

Reference architecture
A documented, reusable design showing how components fit together.
Blast radius
How much of a system is affected when one part fails or is compromised.
Single point of failure
One component whose failure stops the whole system.
Defence in depth
Several independent layers of protection, so one failure doesn't expose everything.
Architecture decision record (ADR)
A short document recording a design decision, its reasons, and trade-offs.
Unit cost
Cost per business outcome, such as cost per 1,000 orders.

Step by step

01The whole picture

This is what happens when Meera in Guwahati opens the Tiffin app and orders lunch: each box links back to the topic that explained it.

request-path.txtwhole filetext
1. DNS          tiffin.in alias -> CloudFront (Topic 7.2)
2. Edge         WAF rules + Shield, TLS ends at edge (8.4)
3. Static       app bundle and thumbnails from S3 via OAC, cached (4.1, 7.2)
4. API          /api/orders -> ALB -> pod in private subnet (2.4, 5.2)
5. Auth         JWT checked by API, secrets from Secrets Manager (7.1)
6. Data         menu from ElastiCache, order written to Aurora (4.3, 4.4)
7. Event        OrderPlaced -> EventBridge -> SQS / Step Functions (6.2, 8.3)
8. Observe      trace across services, 5xx and p99 alarms (6.3, 6.4)
The whole picturediagram
Rendering diagram…

02Failure walk-through

A good architecture answers 'what happens when X fails?' for each part. The team keeps this table in the repo and tests rows during game days.

failure-modes.txtwhole filetext
failure                       what happens                                       recovery
one EC2 node dies             pods rescheduled, ALB stops routing to it          automatic, ~1 min
AZ ap-south-1b down           2 of 3 AZs serve, Karpenter adds nodes, Aurora     automatic, minutes
                              fails over if writer was in 1b
payment partner slow          timeouts + circuit breaker, orders queue as        automatic, degrade
                              "payment pending"
bad deploy                    Argo Rollouts canary fails analysis, rolls back    automatic
menus cache flushed           Aurora takes load briefly, readers absorb it       automatic
Aurora data deleted by bug    PITR to new cluster, copy rows back (4.5)          manual, < 1 h
region ap-south-1 down        promote Aurora Global secondary, scale ap-south-2, manual runbook, 30 min
                              Route 53 failover
prod credentials leaked       IR playbook: contain, scope, rotate (7.5)          manual, first hour

03Cost and trade-offs

The architecture costs about ₹7.5 lakh a month, roughly ₹2.1 per order at current volume. The biggest choices and their trade-offs are recorded as ADRs, so a new engineer can see why things are the way they are.

terminal
$ ls docs/adr/
── expected output ──
0001-eks-over-ecs.md many teams, Argo CD, operators; accepted higher ops cost
0002-aurora-postgresql.md relational orders + PITR; DynamoDB only for tracking
0003-single-region-plus-warm-standby.md active/active not worth 2x cost at our size
0004-graviton-and-spot.md ~30% compute saving; Spot only for stateless pods
0005-vpc-endpoints-over-nat.md cut NAT bill by 60%
0006-eventbridge-for-domain-events.md decoupling teams; accepted harder debugging, added tracing
Output annotated with each ADR's one-line summary.

Break it on purpose

Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.

Break #1

Find the single point of failure

An early version of the diagram had one NAT gateway, Aurora with a single instance, and the Redis cache as a single node holding user sessions.

terminal
$ # game day: stop the cache node
── what you'll see ──
every logged-in user was logged out at once, then 30,000 re-logins hit Aurora and the auth service together
# the cache had quietly become a database for sessions

Myth vs fact

Myth

A good architecture uses as many AWS services as possible.

Fact

A good architecture uses the fewest services that meet the requirements, each for a clear reason. Every extra service is more to learn, monitor, secure, and pay for.

Pro corner

Extra depth for experienced readers. New to this? Skip it for now and come back later.

  • ▸

    Keep the architecture diagram, failure-mode table, and ADRs in the same repo as the infrastructure code, and review them in the same pull requests that change the architecture. Diagrams kept elsewhere go out of date within months.

Remember this

  1. 1

    Edge: Route 53 alias → CloudFront with WAF and ACM certificate → private S3 (OAC) for static files, ALB for /api/* (Topics 7.2, 8.4).

  2. 2

    Compute: EKS in private app subnets across three AZs, Karpenter with Graviton and Spot for stateless pods, Pod Identity per service, AWS Load Balancer Controller (Topics 5.2, 2.4).

  3. 3

    Data: Aurora PostgreSQL (writer + reader, Multi-AZ storage), ElastiCache for menus, DynamoDB for tracking, S3 for invoices and photos, all encrypted with customer keys (Phase 4, Topic 7.1).

  4. 4

    Async: EventBridge for domain events, SQS with DLQs for work, Kinesis for rider pings, Step Functions for fulfilment, Lambda for thumbnails and webhooks (Phases 5, 6, Topic 8.3).

  5. 5

    Operations and security: multi-account landing zone, Identity Center, CloudTrail to a locked log archive, GuardDuty/Config/Security Hub, CloudWatch alarms on symptoms, traces, AWS Backup with locked copies, warm standby in Hyderabad (Phases 6, 7).

  6. 6

    Delivery: GitHub Actions with OIDC (no stored keys) builds images to ECR and plans Terraform. Argo CD deploys to EKS (CI/CD and GitOps courses).

Explain it without notes

01

Walk through what protects Tiffin if an Availability Zone fails.

02

Why did Tiffin choose warm standby instead of active/active across regions?

Practice

01

Draw your own reference architecture for a simple web app on AWS and label each component with the topic it comes from.

02

Write a failure-mode table with at least six rows for your architecture.

Trade-offs

  • ↔

    Every layer of resilience and security adds cost and complexity. The right architecture matches business needs (RPO, RTO, compliance, budget) and team size, and records the risks it accepts.

Done when you can

  • I can draw and explain a full production architecture on AWS.

  • I can say what happens when each component fails.

  • I record major decisions and their trade-offs as ADRs.

  • I know the cost of the architecture and its unit cost.