All case studies
Case study

Scaling a high-traffic microservices backend on AWS

An AWS microservices backend serving thousands of requests per second, using ECS autoscaling, Redis caching and async processing to relieve PostgreSQL.

Context
High-traffic consumer platform (client under NDA)
Role
Tech Lead / Backend
Key point
Thousands of req/s, autoscaled
  • Go
  • Node.js
  • TypeScript
  • AWS ECS
  • ALB
  • Aurora PostgreSQL
  • Redis
  • SQS
  • CloudWatch
  • SNS

The problem

The system relied heavily on the primary database for high-volume reads and on synchronous processing.

As traffic grew, the database was becoming the likely bottleneck, and application capacity was fixed.

Architecture: before → after

Before
  1. Database-centric reads
  2. Synchronous processing
  3. Fixed application capacity
After
  1. Redis for hot data
  2. SQS for async work
  3. ECS autoscaling (~60% threshold)
  4. DLQ for failed jobs
  5. CloudWatch alarms → SNS

What I did

  • Split the backend into roughly 20 services running on ECS behind an Application Load Balancer, connected through private DNS.
  • Configured ECS autoscaling on CPU / memory at around a 60% threshold so capacity follows traffic.
  • Moved frequently read data into Redis so hot paths no longer hit PostgreSQL directly.
  • Pushed non-critical work onto SQS queues processed by workers, with a dead-letter queue for failed jobs.
  • Added CloudWatch metrics and alarms with SNS notifications for operational visibility.

How I led it

  • When the team had to prepare the backend for roughly 10x traffic growth, I broke the problem down into parts the team could own.
  • Aligned the team on one approach before implementation instead of letting each service solve scaling differently.

Outcome

  • Production traffic in the thousands of requests per second (peaks below 10K req/s), served by an autoscaled fleet.
  • Direct read pressure on PostgreSQL reduced by serving hot data from Redis.
  • Traffic spikes absorbed by autoscaling and queues instead of a fixed fleet.

Traffic figure is the real production range (peaks below 10K req/s). I'll add measured CloudWatch numbers (latency, DB CPU, error rate) when they can be published.