---
title: "What Is Middleware? The Onion Model, and Why the Order Is the Configuration"
description: "Your handler is not the first code that runs. In front of it sits a chain of middleware — auth, logging, request IDs, rate limiting — and that chain is not a queue of gates. Each step wraps the next, so every step runs twice: once on the way in, once on the way back out. Here is what a middleware actually is, why a step can refuse to call the next one, why reordering two of them silently changes what your API does, and what the chain honestly costs — with measured numbers from a running service."
keywords: "middleware, what is middleware, middleware pipeline, onion model, request pipeline, ASGI middleware, intercepting filter, cross-cutting concerns, call_next, rate limiting order, request id, FastAPI middleware, Express middleware, backend design"
created_at: "2026-08-27T09:00:00"
post_type: "anonym_post"
content_type: "technical_article"
---

![Banner](./banner.webp)

<div style="display: flex; flex-wrap: wrap; gap: 0.75rem; align-items: center; padding: 1rem 1.25rem; margin: 1.5rem 0 2rem; background: linear-gradient(135deg, #f8fafc 0%, #f1f5f9 100%); border: 1px solid #e2e8f0; border-radius: 12px;">
  <span style="font-size: 0.95rem; font-weight: 600; color: #475569; margin-right: 0.25rem;">Prefer to watch?</span>
  <a href="https://youtu.be/_Vlj8mVj6nk" target="_blank" rel="noopener noreferrer" style="display: inline-flex; align-items: center; gap: 0.4rem; padding: 0.55rem 1rem; border-radius: 8px; background: #ff0000; color: #ffffff; font-size: 0.875rem; font-weight: 600; text-decoration: none;">▶ The full episode</a>
  <a href="https://youtube.com/shorts/AJ7XToKwSk0" target="_blank" rel="noopener noreferrer" style="display: inline-flex; align-items: center; gap: 0.4rem; padding: 0.55rem 1rem; border-radius: 8px; background: #0f172a; color: #ffffff; font-size: 0.875rem; font-weight: 600; text-decoration: none;">⚡ The 2-minute version</a>
  <a href="https://t.me/SoftwareEngineerBlog" target="_blank" rel="noopener noreferrer" style="display: inline-flex; align-items: center; gap: 0.4rem; padding: 0.55rem 1rem; border-radius: 8px; background: #229ed9; color: #ffffff; font-size: 0.875rem; font-weight: 600; text-decoration: none;">✈ Telegram</a>
</div>

One small service. One process, one port, six business endpoints plus a health check.

Three things had to happen on almost every one of them: check the caller's token, write a log line with a request id, start and stop a timer. Nine lines. They were written once, and then pasted by hand into endpoint after endpoint.

```python
# the SAME nine lines, pasted by hand into endpoint after endpoint
async def ep_account(scope, receive, send):
    user = user_for(header(scope, b"authorization"))        # pasted
    if user is None:                                        # pasted
        return await send_json(send, 401, {"error": "no"})  # pasted
    started = clock()                                       # pasted
    log.info("GET /account", rid=rid(scope))                # pasted
    result = handlers.account(user)                         # the actual work
    metrics.observe("GET /account", clock() - started)      # pasted
    await send_json(send, 200, result)
```

Then the sixth endpoint was added. It was copied from the one next to it, and in the copy the two auth lines did not come along.

```python
async def ep_create_review(scope, receive, send):
    started = clock()                                       # pasted
    log.info("POST /reviews", rid=rid(scope))               # pasted
    result = handlers.create_review(user_of(scope), body)   # the actual work
    metrics.observe("POST /reviews", clock() - started)     # pasted
    await send_json(send, 200, result)   # the auth lines are not here
```

Sweep the API with no `Authorization` header at all and you get this:

```
GET     /books            401  {"error": "unauthorized"}
GET     /books/1          401  {"error": "unauthorized"}
POST    /orders           401  {"error": "unauthorized"}
GET     /orders/1         401  {"error": "unauthorized"}
GET     /account          401  {"error": "unauthorized"}
POST    /reviews          200  {"review": {"user": "anonymous", "book": 1, "stars": 5}}
GET     /healthz          200  {"ok": true}
```

`POST /reviews` answered everyone. Nothing threw. No test failed. No error log line was ever written — the endpoint was working exactly as its code said. Counting the cross-cutting lines inside the endpoint functions gives **50** across the module: nine on each of five endpoints, and **five** on the sixth. The missing four are the entire incident.

That is the problem middleware solves. Not "less typing" — we will get to the honest bill later, and it is not smaller. **One place.**

---

## What a middleware actually is

A middleware is a function that is handed **two** things: the request, and *the next thing to call*.

That second argument is the entire idea.

```python
def timing(request, call_next):
    started = clock()                      # 1. before  — runs on the way IN
    response = call_next(request)          # 2. hand it on, and WAIT here
    took = clock() - started               # 3. after   — runs on the way OUT
    response.headers["x-took-ms"] = took
    return response
```

Look at line 2. It is not "finish my job and hand over." It **blocks** until everything deeper in the chain has run and come back. Which is precisely why line 3 is able to know how long the whole rest of the request took.

If you delete the `call_next` argument, you no longer have a middleware. You have a hook.

---

## The chain is an onion, not a queue

Almost everyone's first mental picture of a middleware chain is a row of turnstiles: the request passes through gate one, then gate two, then gate three, then reaches your handler. It is a tidy picture and it is wrong.

The steps do not stand in a row. **Each one wraps the next.** So every step runs *twice*: once on the way in, and once on the way back out.

Here is a real trace, printed by a chain of five:

```
--> ERRORS
--> REQUEST ID
--> LOG + TIMER
--> AUTH
--> RATE LIMIT
    [ YOUR HANDLER RUNS ]
<-- RATE LIMIT
<-- AUTH
<-- LOG + TIMER
<-- REQUEST ID
<-- ERRORS

entered 5 steps, exited 5 steps
exits are the entries REVERSED : True
```

Read the two halves. Going in, the order is the list order. Coming out, it is the list **reversed**. That reversal is the whole shape, and it is not a stylistic detail — it is what makes several of the most common middlewares possible at all:

- A **timer** on the outside can measure the whole request, because its "after" half runs last.
- An **error handler** on the outside can catch an exception thrown by anything beneath it.
- A **response-header** step can set a header only after the response exists.

Measured on a chain of four: the outermost layer's mean was **68.3 µs**, the innermost (the handler alone) **33.9 µs**. The handler was **49.7 %** of what the outermost layer measured, and the outer figure was greater than or equal to the inner one on **every one of 2,000** requests. That is the onion, in a number.

---

## A step is allowed to refuse

The second consequence of "you are handed the next thing to call" is that you may decline to call it.

```python
def auth(request, call_next):
    if not valid_token(request.headers.get("authorization")):
        return json({"detail": "unauthorized"}, status=401)   # no call_next
    request.state.user = user_for(request)
    return call_next(request)        # only a valid caller gets past this line
```

Fifty requests with no token, through a chain of `auth → rate limit → database lookup → handler`:

```
AUTH reached           50
RATE LIMIT reached      0   <- never ran
database lookup         0   <- never ran
your handler reached    0   <- never ran
responses              50 x 401
```

Your handler did not decide to refuse those. **Your handler was never asked.** A browser CORS preflight that never reaches your code is the same move, and so is a WAF rule, and so is an API gateway's quota check.

This is also the first hint of a real operational problem. If a request is rejected two layers above your handler, your handler's own metrics never see it. In one run, **160** requests arrived, the handler counted **50**, and the gap of **110** reconciled exactly to 80 rejections plus 30 throttles. That is **68.8 %** of all traffic invisible to the dashboard most teams actually look at.

---

## The order is the configuration

Here is the part that makes middleware a design decision rather than a convenience.

```python
# the whole ordering decision, in one list, outermost first.
CHAIN = [errors, request_id, logging, auth, rate_limit]   # A: auth, then limit
CHAIN = [errors, request_id, logging, rate_limit, auth]   # B: limit, then auth
```

Same two components. Same code inside them. One line moved. Now send the identical 90-request burst at both — 30 from one authenticated user, 30 from thirty different authenticated users, 30 with no token, limit 10 per key:

<table>
  <thead>
    <tr>
      <th align="left">On the same 90 requests</th>
      <th align="right">ORDER A<br /><small>auth → limit</small></th>
      <th align="right">ORDER B<br /><small>limit → auth</small></th>
    </tr>
  </thead>
  <tbody>
    <tr><td>reached the auth check</td><td align="right">90</td><td align="right">10</td></tr>
    <tr><td>reached the rate limiter</td><td align="right">60</td><td align="right">90</td></tr>
    <tr><td><strong>200 OK</strong></td><td align="right"><strong>40</strong></td><td align="right"><strong>7</strong></td></tr>
    <tr><td>401 unauthorized</td><td align="right">30</td><td align="right">3</td></tr>
    <tr><td>429 too many requests</td><td align="right">20</td><td align="right">80</td></tr>
    <tr><td>rate-limiter keys used</td><td align="right">31 <small>(per user)</small></td><td align="right">1 <small>(per IP)</small></td></tr>
    <tr><td>the one heavy user, throttled</td><td align="right">20</td><td align="right">26</td></tr>
    <tr><td>thirty different users, throttled</td><td align="right">0</td><td align="right">27</td></tr>
  </tbody>
</table>

**40 successful responses versus 7.** That is not a mild trade-off, it is two different products.

And neither order is *wrong*. Order A is a **per-user quota**; order B is a **per-source flood gate**. The limiter can only key by user if something above it has already identified one — placed first it has no user yet, so it falls back to keying by client address, and behind a shared address that means thirty innocent users share one bucket. That is not a bug. It is the consequence of the position.

The same principle bites a request-id step. Placed **first**, all four downstream log lines carry the id. Placed **last**, only **1 of 4** does:

```
request-id FIRST                      request-id LAST
[log:edge]    rid=ffc68931            [log:edge]    rid=-
[log:access]  rid=ffc68931            [log:access]  rid=-
[log:audit]   rid=ffc68931            [log:audit]   rid=-
[log:app]     rid=ffc68931            [log:app]     rid=fe9f57b1
```

Note that it is 3 of 4 lost, not all four — the step still tags whatever runs beneath it. And the error handler is the sharpest case of all. Outermost, with a failing layer below it: **HTTP 500**, body `{"error": "internal server error"}`, caught. Innermost, with the failing layer above it: **no HTTP response at all** — the exception escapes the application entirely.

The worst property of an ordering mistake is that it is **silent**. On that 90-request burst under the wrong order: exceptions escaped **0**, exceptions caught **0**, ERROR log lines **0**, 5xx responses **0**, log lines written **90** — all INFO, every status code documented and expected. The only visible symptom, and only if you already knew your intent, was that **27** requests from users who had sent *one request each* came back `429`.

---

## The honest bill

Middleware is usually sold as a cleanup. Measured, it is a trade, and it is worth being precise about what you are trading.

**It does not mean fewer lines.** The endpoint module shrank from **92 SLOC to 52**. But the steps themselves are another **137 SLOC**, plus an 8-line list to order them. At six endpoints, the repository is *bigger*. What you bought is not brevity — it is that the auth rule exists in **one place**, so the sixth endpoint cannot forget it. On the chain variant, the same unauthenticated sweep returns 401 on all six business endpoints, **including the newest one, which nobody had to remember**.

**It costs something on every request.** A chain of eight layers against a bare handler: bare **4.3 µs**, full chain **30.7 µs** — an overhead of **26.4 µs/request**, about **7.1×**, roughly **3.3 µs per layer**. Quoted alone, that number sounds alarming. So here is the other one: measured end to end over a real socket, a request took **640.1 µs**, and the chain was **4.1 %** of it. Quote either figure by itself and you are misleading someone. The chain is expensive relative to a function call and cheap relative to a request.

**Every request pays, including the ones that did not need to.** One layer doing a small database read added **+13.8 µs/request** to `/healthz` — **2.02×** the same chain without it — and performed **5,300 of 5,300** database reads for an endpoint that needed exactly none.

**It is invisible from the handler, and the blast radius is the whole API.** Changing `if scheme != "Bearer":` to `if scheme == "Bearer":` is a **one-character** diff in a shared layer. Result: **6 of 6** business endpoints reject a valid credential, **1** file changed, **0** endpoint files changed, and `handlers.py` byte-identical before and after. Nothing in the handler you are staring at explains the failure.

---

## The same shape, in front of a model

If you work on LLM serving, you have this chain whether you called it middleware or not — it is what an inference gateway *is*. And the vocabulary maps cleanly:

- **Auth** becomes API-key resolution to a tenant and a model allowlist.
- **Rate limiting** becomes a **token** budget rather than a request budget, and the ordering lesson lands twice as hard: a limiter placed before authentication cannot key by tenant, so one noisy customer and thirty quiet ones share a bucket — the exact 27-out-of-30 failure above, except now it is a paying customer's SLA.
- **Request id** becomes the trace id that has to survive a response lasting several seconds.
- **Guardrails** — prompt-injection screening on the way in, PII or policy filtering on the way out — are the onion's two halves, and the reason they belong in a wrapper rather than in the handler is the same reason auth did: the next endpoint you add must not be able to skip them.
- **Cost accounting** is a pure "way out" step: you cannot bill for tokens you have not generated yet.

But one thing genuinely does *not* transfer, and it catches people. The onion's "on the way back out" half assumes the response is a single object handed back up the stack. When you stream tokens over SSE, the response object returns almost immediately and the body arrives afterwards. A timing middleware written the ordinary way will therefore record **time to first token**, not total generation time — a number that can be twenty times smaller and looks perfectly healthy on a dashboard. The same applies to an output filter: by the time your "after" half runs, the first tokens are already on the client's screen. Streaming-aware guardrails have to wrap the *iterator*, not the response.

The shape holds. The assumption that a request is one atomic unit of work does not.

---

## Copy-paste versus chain, side by side

<table>
  <thead>
    <tr>
      <th align="left"></th>
      <th align="left">Pasted into every endpoint</th>
      <th align="left">One chain, registered once</th>
    </tr>
  </thead>
  <tbody>
    <tr><td>cross-cutting lines inside endpoints</td><td>50</td><td>0</td></tr>
    <tr><td>total code</td><td>92 SLOC</td><td>52 + 137 SLOC (bigger)</td></tr>
    <tr><td>the newest endpoint</td><td>can silently omit auth</td><td>covered without being asked</td></tr>
    <tr><td>changing the auth rule</td><td>edit 6 files, hope you got them all</td><td>edit 1 file</td></tr>
    <tr><td>blast radius of a typo</td><td>one endpoint</td><td>every endpoint</td></tr>
    <tr><td>where the behaviour is written</td><td>in front of you, in the handler</td><td>somewhere else, in a list</td></tr>
    <tr><td>cost</td><td>~0</td><td>~3.3 µs per layer, on every request</td></tr>
    <tr><td>ordering bugs</td><td>impossible</td><td>possible, and silent</td></tr>
  </tbody>
</table>

---

## The verdict

Use middleware for the things that are genuinely true of **almost every** request — authentication, request ids, error handling, logging, rate limiting, CORS. Those earn the wrapper, because their failure mode is *omission*, and a chain makes omission impossible.

Do not use it for things that are true of *some* requests. A layer that reads the database for every call so that three endpoints can avoid a lookup is a tax collected 5,300 times to be spent 3 times; that belongs in a dependency the three endpoints ask for.

And then treat the order as what it is. It is not registration boilerplate at the bottom of a file — it is a configuration file for your API's behaviour, written in the least obvious syntax imaginable, and it fails without raising anything. Put the list somewhere a reviewer will look at it, write down *why* each step sits where it does, and test the order the way you would test a feature: send a burst, count the status codes, and check you got the product you meant to build.

Your handler is not the first code that runs. Know what is in front of it.

## References and further reading

**On the problem — cross-cutting concerns and the pasted nine lines**

- Gregor Kiczales et al., *[Aspect-Oriented Programming](https://www.cs.ubc.ca/~gregor/papers/kiczales-ECOOP1997-AOP.pdf)* (ECOOP, 1997) — the paper that named cross-cutting concerns and the scattering/tangling problem; the missing auth block in `POST /reviews` is a textbook instance of scattering.
- Deepak Alur, John Crupi & Dan Malks, *Core J2EE Patterns*, 2nd ed. (Prentice Hall, 2003) — the **Intercepting Filter** pattern: a configurable, ordered chain of filters around a request handler. This is the pattern middleware implements, described before the word "middleware" was common in web frameworks.

**On the shape — why each step wraps the next**

- *[PEP 3333 — Python Web Server Gateway Interface v1.0.1](https://peps.python.org/pep-3333/)* — the primary source for the wrapping definition: middleware is itself an application that plays server to the application below it. The onion is written into the spec, not into a framework.
- *[The ASGI Specification](https://asgi.readthedocs.io/en/latest/specs/main.html)* — the async successor, and the place to read how the model changes when a response is a stream of events rather than one return value.
- *[Writing middleware for use in Express apps](https://expressjs.com/en/guide/writing-middleware.html)* (Express documentation) — the shortest statement of the `next()` contract, including what happens when you decline to call it.
- Roy T. Fielding, *[Architectural Styles and the Design of Network-based Software Architectures](https://ics.uci.edu/~fielding/pubs/dissertation/rest_arch_style.htm)*, ch. 5 (dissertation, 2000) — the **layered system** constraint: why intermediaries a client cannot see are a deliberate property of the web, not an accident.

**On ordering, refusal, and what it costs you operationally**

- Michael T. Nygard, *Release It!*, 2nd ed. (Pragmatic Bookshelf, 2018) — the case for guards that run before your code at all (bulkheads, circuit breakers, handshaking), and why a component that sheds load must sit where it can actually see the load.
- Charity Majors, Liz Fong-Jones & George Miranda, *Observability Engineering* (O'Reilly, 2022) — instrumenting at the edge versus inside the handler; the direct answer to the 68.8 % of traffic the handler's own counters could not see.
- Jeffrey Dean & Luiz André Barroso, *[The Tail at Scale](https://research.google/pubs/the-tail-at-scale/)* (Communications of the ACM, 56(2), 2013) — why a small fixed per-request cost is usually the wrong thing to worry about, and variance is the right one.

**On the model-serving section**

- Woosuk Kwon et al., *[Efficient Memory Management for Large Language Model Serving with PagedAttention](https://arxiv.org/abs/2309.06180)* (SOSP, 2023) — the vLLM paper; useful here for why a request to a model server is not one atomic unit of work, which is exactly the assumption a "way back out" middleware makes.

If a reference you'd expect is missing, say so in the comments and I'll add it.

---

**Watch the reel:** [Middleware — the code that runs before your handler](https://youtube.com/shorts/AJ7XToKwSk0) · or the [full episode](https://youtu.be/_Vlj8mVj6nk).
