Stripe logo

How Stripe runs four different rate limiters, not just one

A single token bucket stops the easy attacks - real production traffic needed three more layers behind it

Most rate limiters answer one question: how many requests per second is a user allowed? Stripe's engineering team found that question alone doesn't protect payments infrastructure. A user could stay under any per-second limit while still keeping dozens of expensive requests running at once. A low-priority analytics query and a live charge-creation call look identical to a limiter that only counts requests. And when something inside Stripe's own infrastructure degrades, rejecting traffic evenly across every endpoint is exactly the wrong response. Their answer was to run four separate rate limiters in production, each one built to catch a specific failure mode the others miss.

Terms worth knowing before you read on

Token bucket

A rate-limiting algorithm where each user has a bucket of tokens that drains with each request and refills at a steady rate. An empty bucket means the next request gets rejected until it refills - simple to reason about, and it naturally tolerates short bursts as long as tokens are banked up.

Load shedding

Deliberately rejecting some requests during an overload, on purpose, so the system stays available for everyone else. The alternative - accepting everything until the whole fleet falls over - is worse for every user, not just the ones turned away.

Fail open

Designing a safety mechanism so that if it breaks, it lets traffic through instead of blocking everything. A rate limiter that depends on a cache being reachable needs a plan for what happens when that cache is down - and 'let requests through' is usually the safer failure mode than 'reject everything.'

Flapping

Rapidly oscillating between two states - shedding load, then restoring it, then shedding it again - because a system reacted too fast to a brief spike instead of a real sustained trend.

Interactive

Walk the pipeline

Step through each stage of how this actually works, in order.

Stage 1 of 5 · Request Rate Limiter

millions rejected/mo

First line of defense - each user gets N requests per second via a token bucket, with burst room for legitimate spikes.

The problem: one rate limiter isn't enough once you're payments infrastructure

Stripe's engineers point to four distinct scenarios a single rate limiter doesn't cover: an individual user's traffic spike threatening service for everyone else, a misbehaving script sending far more requests than intended, lower-priority requests like analytics queries crowding out critical transaction traffic, and internal failures where the system needs to selectively drop some traffic just to keep functioning at all.

The common thread is that a limiter which only asks 'how many requests per second' treats every one of those situations identically - and for a company processing real financial transactions, a query listing someone's past charges and a request creating a new charge genuinely aren't the same kind of traffic, even though a naive rate limiter can't tell them apart.

Layer one: a token bucket, tuned for bursts, not just averages

The Request Rate Limiter is Stripe's most heavily used layer, restricting each user to a set number of requests per second. It's built on the token bucket algorithm specifically rather than a stricter fixed-window count, because a token bucket naturally tolerates short bursts - a legitimate spike, like traffic during a real-time sale event - as long as a user has tokens banked up, instead of punishing them the instant they cross a strict per-second line.

It's applied identically in test mode and live mode, so a developer sees the same rate-limiting behavior during development that they'll hit in production - catching a badly-behaved integration before it ships, not after. This is the limiter doing the most work day to day: Stripe describes it rejecting millions of requests a month, the large majority of them in test mode.

Layer two: limiting concurrency, not just rate

Rate alone doesn't catch everything. A user could stay comfortably under a per-second limit while still keeping dozens of expensive requests in flight simultaneously, each one chewing through CPU on a heavy endpoint. The Concurrent Requests Limiter caps how many of one user's requests can be active at the same time instead of how many arrive per second.

That change in what's being measured changes how clients end up behaving, too: hammering the API and retrying immediately stops working once concurrency itself is capped, which nudges well-built clients toward queuing work and backing off instead. This limiter triggers far less often than the request-rate layer - Stripe cites roughly 12,000 rejections a month - but credits it with fixing persistent performance problems on their most CPU-intensive endpoints that per-second limiting alone never caught.

Layer three: reserving capacity before anyone even hits a wall

The Fleet Usage Load Shedder works proactively rather than reacting to any single user's behavior. Traffic gets split into critical methods (creating a charge) and non-critical methods (listing past charges), and a fixed share of total fleet capacity is reserved specifically for critical traffic - Stripe's own example reserves 20% of capacity for critical requests, leaving non-critical traffic an 80% share to work within.

Once non-critical traffic exceeds its allotted share, it starts getting shed with a 503 - even if critical traffic is nowhere near overloading the system on its own. The point isn't to wait for an actual overload before protecting what matters most; it's to make sure charge creation can never be crowded out by a flood of listing requests, structurally, before that crowding ever becomes a real incident.

Layer four: the last line of defense, tuned to avoid making things worse

The Worker Utilization Load Shedder is the layer that only fires during genuine incidents. It monitors real-time capacity across the fleet, and when things degrade, it sheds traffic progressively in priority order: test-mode traffic first, then GET requests, then POSTs, and only critical methods if the system is still struggling after shedding everything less important.

The subtler engineering problem here is flapping - reacting too fast to a brief spike can cause a system to oscillate between shedding and restoring load in a way that's worse than doing nothing at all. Stripe tuned this layer to shed and recover gradually, over a period of minutes, specifically to avoid that oscillation. It's the least-triggered limiter of the four - around 100 rejections a month - because it exists purely for the worst days, not the routine ones.

The quiet detail that makes any of this safe to run: failing open

Every one of these four limiters depends on infrastructure - a shared cache - that can itself fail. Stripe wrapped each limiter's logic in explicit exception handling so that if the rate limiter's own dependency breaks, requests are let through rather than blocked, backed by feature flags that can disable a misbehaving limiter immediately.

The lesson underneath that detail generalizes well beyond rate limiting: a safety system that can itself become the outage is worse than having no safety system at all. The correct default failure mode for something optional is 'get out of the way,' not 'block everything.'

Takeaway

Stripe's rate limiting isn't one clever algorithm - it's four separate, purpose-built layers, each catching a failure mode a single token bucket alone would miss: raw request rate, concurrency per user, proactive capacity reservation for what matters most, and a last-resort shedder specifically tuned to avoid making incidents worse through flapping. All of it is deliberately built to fail open, because a protection mechanism that can itself take down the whole API isn't actually protecting anything. The transferable lesson: 'add a rate limiter' is really four different questions, and most systems only ever answer the first one.

Source

“Scaling your API with rate limiters”

By Paul Tarjan, on Stripe’s engineering blog

This page explains, in plain language, the rate limiting architecture described in Stripe's own engineering blog post, written by Paul Tarjan. All credit for the original work, research, and writing belongs to him and Stripe - this is our own explanation of the same publicly documented system, not a copy of their text.