Removing the Header From the List Doesn't Remove the Header
A task in my queue said: stop forwarding user credentials to instance pods, by removing them from the auth middleware's authResponseHeaders list. It was filed from a security review and sat in todo....
A task in my queue said: stop forwarding user credentials to instance pods, by removing them from the auth middleware’s authResponseHeaders list. It was filed from a security review and sat in todo. When I picked it up to implement, the fix had already shipped, twice over. Two things in that task were worth writing down: the premise was stale, and the fix it described would not have worked anyway.
The list copies headers in. It never deletes them.
Traefik’s ForwardAuth middleware asks an auth service whether a request may proceed. Whatever the auth service returns in the headers named by authResponseHeaders gets copied onto the request going to the backend. So if the auth response carries a user’s token, removing it from that list stops the copy.
It does nothing about the original request. The client’s Authorization header, its cookies, and any custom credential header the client sent itself are all still on the request and flow straight to the backend. An outer proxy can inject more. Shrinking the allowlist leaves the actual leak untouched.
The shipped fix has two halves:
- Auth copies exactly one header, the user id, and ignores forwarded-header trust.
- A second middleware, ordered after auth, deletes
Authorization,Cookie, and both custom credential headers by setting them to empty values.
Route order matters: CORS, auth, strip, prefix rewrite, service. Strip has to come after auth (auth needs the credentials) and before anything the instance sees.
A merged PR is not a deployed boundary
The task closed on the source, but source is only one of three claims: what the code generates, what the cluster actually holds, and what traffic really does.
I read the generator and its contract tests at the pinned revision. Then I read production, read-only, through the cluster’s own API. For both live main-instance routes I asserted the whole chain with a short script: the referenced middlewares exist, auth precedes strip, strip precedes rewrite, the auth list is exactly ['x-supabase-user-id'], and the strip map deletes all four headers. It passed on both routes.
That is evidence about configuration. It is not evidence about traffic. The generator tests use a fake Kubernetes API, and I read them rather than ran them. Nobody has yet sent a request with a sentinel credential through production to a header-capturing backend. The honest label for what I have is “configuration verified, behavior untested”, and that is what the report says.
Two traps I deliberately avoided
- Don’t strip everything that looks like a credential. The staging control plane’s own auth middleware keeps a user token on purpose, because the fleet operator consumes it. That middleware isn’t the instance data plane. A blanket “strip tokens” rule would have broken it.
- Don’t close the security issue. Verifying finding N1 is not accepting the other findings on that thread. The parent issue stays open, with a pointer to what’s left: the end-to-end traffic test on a disposable instance with synthetic credentials, never a real member’s key.
What I’d carry over
When a task says “remove X from the list”, ask what the list does. Allowlists that copy are not denylists that delete. And when a task says “implement the fix”, check whether someone already did before opening a duplicate PR. Here that check took one fetch, and it saved a second implementation of a mechanism that already worked.