Scaling, Load Balancing & Rate Limiting

How to serve 10 users, then 10 million. Vertical vs horizontal scaling, load balancers, stateless servers, replication, sharding, rate limiting and resilience basics.

Intermediate⏱ 6 min readLesson 9 of 12#backend#scaling#load-balancing#rate-limiting#system-design

The big idea

A small café with one barista is fine for 20 customers a day. At 2,000 customers you can either:

  • Give the barista a faster espresso machine → vertical scaling (scale up).
  • Hire more baristas and add a host who sends each customer to a free one → horizontal scaling (scale out) with a load balancer.

Vertical scaling makes one server bigger; horizontal scaling adds more serversVertical scaling makes one server bigger; horizontal scaling adds more servers

Vertical (up)Horizontal (out)
HowBigger CPU / RAMMore servers
Simplicity✅ No code changesNeeds stateless design
LimitHardware ceiling, expensive at the topPractically unlimited
FailureOne server = single point of failureSurvives losing a server

The scaling journey

Drawing diagram…

Don't jump to step 7 on day one (remember KISS and YAGNI). Each step solves a real, measured bottleneck.

Load balancers

A load balancer spreads requests across healthy servers and stops sending traffic to broken ones.

Drawing diagram…

Algorithms

AlgorithmHow it picksGood when
Round robinA, B, C, A, B, C…Servers are identical
Weighted round robinBigger servers get moreMixed server sizes
Least connectionsServer with the fewest active requestsRequests vary in duration
IP hash / consistent hashingSame client → same serverNeed stickiness or cache locality

Layer 4 vs Layer 7

  • L4 (transport): routes by IP and port. Very fast, doesn't read HTTP.
  • L7 (application): reads HTTP, so it can route /api/* to API servers and /images/* to storage, terminate TLS, add headers. Examples: NGINX, HAProxy, AWS ALB, Envoy.

Stateless servers: the key to scaling out

If a server keeps user sessions in its own memory, the next request might land on a different server that doesn't know the user.

Drawing diagram…

Rule: any server should be able to handle any request. Keep sessions in Redis or in tokens, files in object storage (S3), and state in the database.

Scaling the database

The database is usually the hardest part to scale.

Read replicas

Most apps read far more than they write. Copy the data to replicas and send reads there.

Drawing diagram…

⚠️ Replication lag: a replica may be a few milliseconds behind. After a user updates their profile, read their profile from the primary ("read your own writes").

Sharding (partitioning)

When one primary can't handle the writes or the data size, split the data across several databases by a shard key.

Drawing diagram…
ChallengeWhy
Choosing the shard keyA bad key creates "hot" shards (all celebrities on one)
Cross-shard queries and JOINsMust query every shard and combine the results
Re-shardingAdding shards moves data. Consistent hashing minimises that
Transactions across shardsHard; avoid them by design

Rate limiting

Protect your API from abuse, bugs and traffic spikes by limiting requests per user, IP or API key. Over the limit → 429 Too Many Requests with a Retry-After header.

🪣 A bucket holds up to 10 tokens and refills at 1 token per second. Each request takes a token. Empty bucket → request rejected. Allows short bursts but enforces an average rate.

Drawing diagram…
AlgorithmIdeaTrade-off
Fixed windowCount per minute (resets at :00)Simple; allows 2× bursts at window edges
Sliding windowCount over the last 60 sSmoother, slightly more work
Token bucketTokens refill steadilyAllows bursts, smooth average ✅
Leaky bucketQueue drains at a fixed rateVery smooth output, adds latency

In a multi-server setup, keep counters in Redis so all servers share the same limits (see Caching & Redis for code).

Resilience basics

At scale, something is always failing. Design for it:

PatternPurpose
TimeoutsNever wait forever on a dependency
Retries with exponential backoff + jitterSurvive brief glitches without hammering the service
Circuit breakerStop calling a service that keeps failing; fail fast and give it time to recover
BulkheadSeparate resource pools, so one slow dependency can't consume every thread
Graceful degradationShow cached or partial results instead of an error page
Health checksLet load balancers route around sick servers

These are covered in depth in Microservices → Resilience Patterns.

Back-of-the-envelope numbers

MetricRough figure
1 million requests/day≈ 12 requests/second average (peak maybe 5–10×)
One Node.js / Go serverHundreds to thousands of simple requests/second
One Postgres primaryThousands of simple writes/second, many more reads
One Redis node~100,000 operations/second

Key takeaways

  • Scale vertically first (simple), horizontally when needed (resilient, unlimited).
  • Load balancers spread traffic and route around failures; L7 can route by path.
  • Stateless app servers are the foundation of horizontal scaling.
  • Databases scale with caching, read replicas, then sharding.
  • Rate limit with a token bucket in Redis; return 429 with Retry-After.
  • Assume failure: timeouts, retries with backoff, circuit breakers.