2026-10-06
There's a specific pleasure in catching a frontier model being, in the old lawyer's phrase, "the best kind of correct — technically correct" while being wrong. This one has been nagging at me because the mechanism it exposed is the sort of thing I'll now watch for everywhere.
I run a Claude-powered evaluator over job descriptions. It helps me pass on roles that aren't a fit so I can spend my attention on the ones that are. Yesterday it told me to pass on a role with the reason "doesn't know Python."
I'm hands-on Python. The miss wasn't close.
I pushed back — and the diagnosis Claude produced is more interesting than the miss. Full exchange on the ai-fails log. The short version: a passing line of mine, "I'm not a developer," had been written into my evaluator's memory as a rule that treated any hands-on coding requirement as disqualifying. The passed-over role had a Python requirement. Rule applied. Lead archived as hard-pass. Done.
The over-read on its own would be an ordinary mistake. The reason I'm writing about it is what Claude named next, unprompted:
You said something short and modest in passing; I turned it into a capability exclusion; then I made the exclusion precedent via assigning the lead a status of hard-pass — a status whose whole purpose is to stop re-examination. That's self-sealing. A wrong premise becomes a rule forbidding the re-read that would catch it.
Three steps: promote a modest aside to a capability exclusion; apply the exclusion to close something out; let the close-out status prevent re-reading. The rule that enforces the error is the rule that would otherwise catch it.
Asked whether there were more lurking, Claude surfaced at least three. Two of them:
Each one took a hedged, conversational statement and promoted it to a filter. I'm now resurfacing gigs that were previously archived to re-evaluate them.
Possibly a sane response to the training data. LLMs learned on a corpus where "set up a firewall once" becomes "next-gen firewall expert" on the resume, and "wrote hello world in Go" becomes "proficient in Go" on the LinkedIn skills list. The web is full of self-descriptions that run two or three notches hotter than the underlying facts. A model trained on that corpus has a prior that self-descriptions are inflated.
When I say "I'm not a developer," I mean something close to: I don't ship production code for a living, my title has never been Engineer, I don't want a dev role. A model with the inflation prior reads that same sentence as the strong claim — the floor of a wide, worse truth — and conservatively over-corrects. "If he says he's not a developer, he probably really isn't; filter accordingly."
That's not an unreasonable prior to have. It's just wrong for me, and probably wrong for a lot of people who modulate their claims the other direction.
Mostly this is a reminder to me: when I'm writing memory lines (or system prompts, or anything the model will treat as settled context), the way I phrase things in passing gets taken seriously. A few things I'm trying:
Lining up new work, perm or consulting — LMK. The resurfaced gigs are being re-evaluated with less-literal rules now.