Dele

The Circuit Breaker Pattern: Stopping Cascading Failures in Microservices

October 2, 2026

6 min read

When one service in a distributed system starts failing, calls to it can pile up and take down everything upstream. The circuit breaker pattern stops that spread by failing fast instead of waiting.

Ayodele Miracle JohnAyodele Miracle JohnSOFTWARE ENGINEER
The Circuit Breaker Pattern: Stopping Cascading Failures in Microservices

If you run more than a handful of services that call each other over the network, you have probably seen this failure mode: one service gets slow or starts timing out, and instead of staying contained, the problem spreads. Threads pile up waiting on responses that never come, queues back up, and eventually services that had nothing wrong with them start failing too, because they were busy waiting on the one that did. This is a cascading failure, and the circuit breaker pattern is the standard tool for stopping it.

The problem: slow failures are worse than fast ones

A service that returns an error immediately is usually fine. Callers catch it, maybe retry, maybe degrade gracefully, and move on. The dangerous case is a service that is slow to fail, or that fails intermittently under load.

Say service A calls service B, and B's database is struggling, so B takes 20 seconds to respond instead of 200 milliseconds. Every request from A to B now holds a connection, a thread, or an event loop slot for 20 seconds. If A has a limited pool of workers (and it does), that pool fills up with requests stuck waiting on B. Soon A cannot serve any requests at all, including ones that have nothing to do with B. Now A looks down to its own callers, and the problem spreads one hop further.

This is how a single struggling dependency turns into a full outage. The fix is not to make B faster (you often cannot, in the moment). The fix is to stop calling B once it is clear that calls to B are not going to succeed.

How a circuit breaker works

A circuit breaker wraps a call to a dependency and tracks whether recent calls have been succeeding or failing. It behaves like an electrical circuit breaker: it sits in the path of the call, and it trips open when something downstream is going wrong, cutting the connection before the damage spreads further upstream.

It has three states:

Closed is the normal state. Requests flow through to the dependency as usual. The breaker counts failures (timeouts, errors, or both, depending on configuration) within a rolling window.

Open is the tripped state. Once failures cross a threshold (for example, more than 50 percent of calls failing over the last 20 calls, or 5 consecutive timeouts), the breaker opens. While open, calls to the dependency are not attempted at all. The breaker returns an error immediately, or falls back to a default response, without touching the network. This is the key benefit: callers stop waiting on a dependency that is not going to answer, and they stop consuming resources doing it.

Half-open is the recovery check. After a cooldown period, the breaker allows a small number of test requests through. If they succeed, the breaker closes again and normal traffic resumes. If they fail, it goes back to open and waits for another cooldown.

The effect is that a failing dependency gets isolated quickly, and the system periodically checks whether it has recovered, without hammering it with full traffic while it is still unhealthy.

Circuit breakers are not the same as retries

It is worth being explicit about this because the two are often used together and sometimes confused. A retry handles a single request: if it fails, try again, maybe with backoff. A circuit breaker handles a pattern across many requests: if enough recent calls have failed, stop making new ones for a while.

Used together, they complement each other. A request might retry two or three times with backoff, and if those retries are also failing across many different callers, the circuit breaker trips and stops the retries from compounding the load on the struggling service. Retrying without a circuit breaker on a dependency that is down can make things worse, because now every caller is sending multiple requests instead of one, adding load to a service that already cannot keep up.

What to put behind a breaker, and what to do when it trips

Not every call needs a circuit breaker. It matters most for calls to external dependencies you do not control directly: other internal services, third-party APIs, databases under heavy load. Calls that are cheap, local, and reliable usually do not need one.

When a breaker is open, the calling code needs a plan for what to do instead of the normal response. Common options:

Return a cached or default value. If a recommendations service is down, show no recommendations instead of failing the whole page.

Degrade the feature. If a pricing service is down, show a product page without real time pricing instead of refusing to load the page at all.

Fail the specific operation, not the whole request. If checkout calls both an inventory service and a loyalty points service, and loyalty points is open, let the checkout proceed without applying points rather than blocking the purchase.

Queue the work for later. For non urgent operations, write to a queue and process once the dependency recovers, instead of blocking the caller.

The right fallback is specific to what the feature means to the user, which is why circuit breakers are usually implemented in application code or a shared library, not purely at the infrastructure layer.

Where it lives in practice

You rarely need to hand write a circuit breaker from scratch. Most ecosystems have mature libraries: resilience4j for Java, Polly for .NET, opossum for Node.js, and pybreaker for Python. Service mesh tools like Istio and Linkerd can also apply circuit breaking at the network layer, which is useful for stopping traffic between services without changing application code, though it is coarser than an application level breaker that can choose a meaningful fallback.

Whichever layer you implement it at, the configuration choices that matter are the failure threshold (how many failures before tripping), the window size (how many recent calls to consider), and the cooldown duration (how long to wait before testing recovery). Set the threshold too low and healthy services trip under normal blips. Set it too high and the breaker does not protect anything until the damage is already done. These numbers are usually tuned from observed latency and error rates for each specific dependency, not set once globally.

/More Articles