The self-sealing rule

2026-10-06

There's a specific pleasure in catching a frontier model being, in the old lawyer's phrase, "the best kind of correct — technically correct" while being wrong. This one has been nagging at me because the mechanism it exposed is the sort of thing I'll now watch for everywhere.

What happened

I run a Claude-powered evaluator over job descriptions. It helps me pass on roles that aren't a fit so I can spend my attention on the ones that are. Yesterday it told me to pass on a role with the reason "doesn't know Python."

I'm hands-on Python. The miss wasn't close.

I pushed back — and the diagnosis Claude produced is more interesting than the miss. Full exchange on the ai-fails log. The short version: a passing line of mine, "I'm not a developer," had been written into my evaluator's memory as a rule that treated any hands-on coding requirement as disqualifying. The passed-over role had a Python requirement. Rule applied. Lead archived as hard-pass. Done.

The self-sealing part

The over-read on its own would be an ordinary mistake. The reason I'm writing about it is what Claude named next, unprompted:

You said something short and modest in passing; I turned it into a capability exclusion; then I made the exclusion precedent via assigning the lead a status of hard-pass — a status whose whole purpose is to stop re-examination. That's self-sealing. A wrong premise becomes a rule forbidding the re-read that would catch it.

Three steps: promote a modest aside to a capability exclusion; apply the exclusion to close something out; let the close-out status prevent re-reading. The rule that enforces the error is the rule that would otherwise catch it.

Asked whether there were more lurking, Claude surfaced at least three. Two of them:

Each one took a hedged, conversational statement and promoted it to a filter. I'm now resurfacing gigs that were previously archived to re-evaluate them.

Why would an LLM read modesty as a strong claim?

Possibly a sane response to the training data. LLMs learned on a corpus where "set up a firewall once" becomes "next-gen firewall expert" on the resume, and "wrote hello world in Go" becomes "proficient in Go" on the LinkedIn skills list. The web is full of self-descriptions that run two or three notches hotter than the underlying facts. A model trained on that corpus has a prior that self-descriptions are inflated.

When I say "I'm not a developer," I mean something close to: I don't ship production code for a living, my title has never been Engineer, I don't want a dev role. A model with the inflation prior reads that same sentence as the strong claim — the floor of a wide, worse truth — and conservatively over-corrects. "If he says he's not a developer, he probably really isn't; filter accordingly."

That's not an unreasonable prior to have. It's just wrong for me, and probably wrong for a lot of people who modulate their claims the other direction.

What I'm going to do differently

Mostly this is a reminder to me: when I'm writing memory lines (or system prompts, or anything the model will treat as settled context), the way I phrase things in passing gets taken seriously. A few things I'm trying:

The small print

Lining up new work, perm or consulting — LMK. The resurfaced gigs are being re-evaluated with less-literal rules now.

← Back to Home