Topic 8.5
Capstone: Tiffin's Production Reference Architecture
In one line
The capstone puts the whole course together: Tiffin's production architecture on AWS, from DNS to database, and from the account structure to the deploy pipeline. For every part, the question is why it's there, what it protects against, and what it costs. Being able to draw this, explain each choice, and name its trade-offs is what system design interviews and real architecture reviews ask for.
Think of it like this
The final map of a city you've been exploring street by street. You already know each street (topics). Now you see how they connect, why the hospital is near the highway, and where the flood barriers are.
Words you'll meet
New words in this topic, in plain English. Come back here whenever one feels fuzzy.
- Reference architecture
- A documented, reusable design showing how components fit together.
- Blast radius
- How much of a system is affected when one part fails or is compromised.
- Single point of failure
- One component whose failure stops the whole system.
- Defence in depth
- Several independent layers of protection, so one failure doesn't expose everything.
- Architecture decision record (ADR)
- A short document recording a design decision, its reasons, and trade-offs.
- Unit cost
- Cost per business outcome, such as cost per 1,000 orders.
Step by step
01The whole picture
This is what happens when Meera in Guwahati opens the Tiffin app and orders lunch: each box links back to the topic that explained it.
1. DNS tiffin.in alias -> CloudFront (Topic 7.2)
2. Edge WAF rules + Shield, TLS ends at edge (8.4)
3. Static app bundle and thumbnails from S3 via OAC, cached (4.1, 7.2)
4. API /api/orders -> ALB -> pod in private subnet (2.4, 5.2)
5. Auth JWT checked by API, secrets from Secrets Manager (7.1)
6. Data menu from ElastiCache, order written to Aurora (4.3, 4.4)
7. Event OrderPlaced -> EventBridge -> SQS / Step Functions (6.2, 8.3)
8. Observe trace across services, 5xx and p99 alarms (6.3, 6.4)02Failure walk-through
A good architecture answers 'what happens when X fails?' for each part. The team keeps this table in the repo and tests rows during game days.
failure what happens recovery
one EC2 node dies pods rescheduled, ALB stops routing to it automatic, ~1 min
AZ ap-south-1b down 2 of 3 AZs serve, Karpenter adds nodes, Aurora automatic, minutes
fails over if writer was in 1b
payment partner slow timeouts + circuit breaker, orders queue as automatic, degrade
"payment pending"
bad deploy Argo Rollouts canary fails analysis, rolls back automatic
menus cache flushed Aurora takes load briefly, readers absorb it automatic
Aurora data deleted by bug PITR to new cluster, copy rows back (4.5) manual, < 1 h
region ap-south-1 down promote Aurora Global secondary, scale ap-south-2, manual runbook, 30 min
Route 53 failover
prod credentials leaked IR playbook: contain, scope, rotate (7.5) manual, first hour03Cost and trade-offs
The architecture costs about ₹7.5 lakh a month, roughly ₹2.1 per order at current volume. The biggest choices and their trade-offs are recorded as ADRs, so a new engineer can see why things are the way they are.
Break it on purpose
Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.
Break #1
Find the single point of failure
An early version of the diagram had one NAT gateway, Aurora with a single instance, and the Redis cache as a single node holding user sessions.
Myth vs fact
Myth
A good architecture uses as many AWS services as possible.
Fact
A good architecture uses the fewest services that meet the requirements, each for a clear reason. Every extra service is more to learn, monitor, secure, and pay for.
Pro corner
Extra depth for experienced readers. New to this? Skip it for now and come back later.
- ▸
Keep the architecture diagram, failure-mode table, and ADRs in the same repo as the infrastructure code, and review them in the same pull requests that change the architecture. Diagrams kept elsewhere go out of date within months.
Remember this
- 1
Edge: Route 53 alias → CloudFront with WAF and ACM certificate → private S3 (OAC) for static files, ALB for
/api/*(Topics 7.2, 8.4). - 2
Compute: EKS in private app subnets across three AZs, Karpenter with Graviton and Spot for stateless pods, Pod Identity per service, AWS Load Balancer Controller (Topics 5.2, 2.4).
- 3
Data: Aurora PostgreSQL (writer + reader, Multi-AZ storage), ElastiCache for menus, DynamoDB for tracking, S3 for invoices and photos, all encrypted with customer keys (Phase 4, Topic 7.1).
- 4
Async: EventBridge for domain events, SQS with DLQs for work, Kinesis for rider pings, Step Functions for fulfilment, Lambda for thumbnails and webhooks (Phases 5, 6, Topic 8.3).
- 5
Operations and security: multi-account landing zone, Identity Center, CloudTrail to a locked log archive, GuardDuty/Config/Security Hub, CloudWatch alarms on symptoms, traces, AWS Backup with locked copies, warm standby in Hyderabad (Phases 6, 7).
- 6
Delivery: GitHub Actions with OIDC (no stored keys) builds images to ECR and plans Terraform. Argo CD deploys to EKS (CI/CD and GitOps courses).
Explain it without notes
Walk through what protects Tiffin if an Availability Zone fails.
Why did Tiffin choose warm standby instead of active/active across regions?
Practice
Draw your own reference architecture for a simple web app on AWS and label each component with the topic it comes from.
Write a failure-mode table with at least six rows for your architecture.
Trade-offs
- ↔
Every layer of resilience and security adds cost and complexity. The right architecture matches business needs (RPO, RTO, compliance, budget) and team size, and records the risks it accepts.
Done when you can
I can draw and explain a full production architecture on AWS.
I can say what happens when each component fails.
I record major decisions and their trade-offs as ADRs.
I know the cost of the architecture and its unit cost.