- Published on
AI Is Breaking Our Proxies for Engineering Seniority
- Authors
- Name
- Iván González Sáiz
- @dreamingechoes
Table of contents8 sections
The message came in one afternoon, wedged between two conversations I was already behind on.
"The engineers are asking me to be part of the hiring process for the new AI engineer, and I don't think I've ever interviewed an engineer?"
A senior product manager I used to work with, now at another company. A minute later: "Do you have any good questions you've asked in the past to test an engineer's product sense?" And then: "I trust your experience way more than I trust AI."
We overlapped while I was leading teams, so the question landed somewhere reasonable: I had spent years on the other side of hiring panels.
What I did not expect was what came back out. I ended up with five questions, and almost none of them asked whether the candidate could produce a solution.
They asked whether the candidate could frame one, push back on it, check that it worked, and stay answerable for it after it shipped. Not because producing solutions stopped mattering. Because producing a plausible one — or a fluent account of how you would — got cheap enough that it no longer carries what it used to about the judgment underneath.
The baseline that hasn't formed yet
Strip the rubric away and what is underneath is comparison. You have met forty people who did this job, and this person sounds like the ones who were good at it.
That needs a stable baseline, and for the way this work gets done right now there isn't one yet. Building with models has stopped being novel; what hasn't settled is the workflow around it — how much implementation gets delegated, which parts get checked and how, how much scaffolding is worth maintaining. Nobody has eight years of this version of the job, because the version keeps changing faster than a cohort can form.
So panels fall back on the proxies inherited from the job as it was: coding ability, system design, technical depth under questioning. Those proxies were never wrong, exactly. They were shorthand for the expensive thing, which was judgment.
The shorthand is what broke. Not the skill it pointed at — the pointing. Technical depth is still necessary; it is no longer sufficient evidence of everything that used to arrive bundled with it. Last week I wrote about the same crack at the other end of the ladder, where the entry-level work that quietly manufactured senior engineers is closing without any single decision behind it. Same repricing, further up — which is roughly where I was sitting when I started writing her document.
The questions that survived the translation
What I meant to write her was a translation layer. Here is what a good architecture answer sounds like. Here is enough vocabulary that you won't feel lost in the room.
I got most of a page in before it collapsed, because every criterion required her to already know the answer. You cannot verify that Postgres was the better call than DynamoDB without knowing the shape of the data and what the team can operate at three in the morning. I was writing a document that would only work for a reader who didn't need it.
So I threw it out and wrote the inverse. Not: can you evaluate their technology. Instead: can they make you understand why the technology was necessary. That one survives being asked by somebody with no technical background, and it is much harder to fake.
The document only made sense to me afterwards. What I had ended up with was not a translation layer for a product manager but, near enough, the list of things I listen for myself now. I had never had to write it down, because it had never had to work without the technical signals underneath it.
Two seniorities under one job title
With the technical assessment left to the people who can do it, what remained for her sorted into two shapes. We have been using one word for both.
The first is senior inside the frame. The problem arrives defined and what follows is genuinely excellent: clean decomposition, sharp trade-offs within the boundary, work that ages well. Responsibility begins when the ticket is written and ends when it is deployed.
The second is senior across the frame. The request arrives and gets turned into a question first. Options come back rather than estimates. When adoption is poor they want to know what "poor" is made of — discovery, or people trying it once and never returning — and they can land on the conclusion that the thing they built should be removed. Responsibility starts before the solution exists and does not stop at the deploy.
Neither is the lesser engineer, and this is not a ranking of technical competence — the strongest people I have worked with were unmistakably senior inside the frame first. It is a difference in how much product judgment somebody has had reason to build. And for most of the industry's history you could not get much of the second without going through the first: you learned which decisions age badly by making them and living downstream for eighteen months. Execution was the toll on the way there.
Some of that toll is now cheaper to pay. Not the judgment — the years of execution that were the road most people took to it, and the half of product engineering nobody assigns you picked up along the way. So an interview built on the old proxy is not measuring badly. It is measuring precisely, and reporting a number that carries less than it used to.
Scoring Fluency
The same repricing shows up one level further in, inside the room, and this is the part nobody wants to raise in a debrief.
A well-shaped answer used to cost something. To describe fluently how you would approach an ambiguous production problem — the sequence, the competing hypotheses, the moment where mitigation and diagnosis separate into different jobs — you generally had to have stood in it once. So we treated fluency as evidence, and we were right to.
It costs less now. The tools we build with are good at producing the shape of an experienced answer, and candidates prepare with them in good faith, because that is what preparation looks like in 2026. The answer can be entirely true. It is doing less work as evidence than the same answer did five years ago.
Call it Scoring Fluency: taking the shape of an answer as evidence of the judgment behind it.
It cuts in both directions, which is why it is worth naming. It over-rewards the rehearsed account of work somebody mostly watched happen. It under-rewards the engineer who has the scars and narrates them badly — starts in the middle, doubles back, spends too long on the detail that mattered to them and skips the one you asked about. Real memory doesn't come out sounding like a transcript.
None of which makes the candidates dishonest or preparation a problem. It leaves the interviewer with one question worth the whole hour: which evidence is still expensive?
What a real decision leaves behind
Not a good decision. Describing one convincingly has become cheap.
Having decided something, and then having stayed around for what it produced. Not the decision — the wreckage around it.
A decision somebody owned drags along things nobody includes when they imagine one:
a constraint that took the good option off the table before anyone got to choose
information that turned up in the wrong order, after the commitment
somebody who disagreed, and whether they turned out to be right
the part they would do differently, which is almost never the part they got wrong
The highest-value move in the guide is also the least clever one. When somebody says "I would normally…", ask when they last did. Then stay inside it: what did you own, what changed because you were in the room, which assumption turned out to be false.
And one question I wasn't asking three years ago, which may now be the most revealing thing you can ask an AI engineer: if the agent tells you it has found the cause, how do you decide that it has?
Weak answers are about the tool and how reliable it has been lately. Strong ones are about evidence: reproduce the failure, explain the mechanism, prove the change fixes it, watch the metric move. Generation can be delegated; judgment and accountability cannot. When I built the human gates into my team's AI harness, that line is what I was writing down. A senior candidate can tell you where it falls for them, and what it cost them the week they held it.
One story instead of five questions
The guide came out as five moments, and the part that matters is that they are not five questions. They are one story.
A product request arrives, and you ask what they would want to understand before deciding how to build it. Then the work turns out to be six weeks rather than two. Then it ships, behaves exactly as specified, and a month later hardly anybody is using it. Then, two months on, it slows down and starts failing with no obvious cause. And finally: take everything we have just walked through, and show me where AI sits inside it — what you delegate, what you keep, and how you know you can trust what comes back.
Five separate questions get you five prepared pieces. A continuous story makes that much harder, because the candidate has to carry their own earlier answers forward: the scope they cut in the second moment is the scope you ask about when adoption is thin in the third, and the shortcut they defended is the thing degrading in production in the fourth.
Seniority turns up in the transitions, not in the answers. You can hear it in what they do with each handoff — whether every moment ends with the problem passed back to product, or whether they keep converting technical information into decisions somebody else can act on.
Take the shape rather than the script. One problem, five pressures, no resets in between.
The part the interview cannot see
Now the uncomfortable half, because a filter that selects well is still doing something to somebody.
Every signal above rewards people who were given room: scope, ownership, a seat where the problem got framed instead of handed over. Plenty of capable engineers spent four years inside the frame because that is what their organisation handed them, and they arrive with a thin supply of stories through no fault of their own. Left unsaid, the interview converts an opportunity gap into a verdict about a person.
A fluent storyteller working with borrowed material still gets through sometimes. The defence is not suspicion of polish — preparation improves the telling of things that genuinely happened. It is asking about the parts nobody rehearses: what went wrong, who disagreed, what you would remove. A story with no friction anywhere in it proves nothing on its own, but it is a reason to keep asking.
And none of this certifies technical depth. It should not try. The engineering interviews still have to do their job — what changed is that they can no longer answer the seniority question on their own.
I wrote a guide for one person. She has the questions, and as far as I know the interview hasn't happened yet, so there is no candidate and no debrief. What I have is the debriefs where two of us read the same answer differently, from the stretch of leading teams before I went back to building, and the specific discomfort of realising I had never written down what I was listening for.
Final thoughts
She got her questions. What I got was the thing I had never had to pull apart: technical ability, and the judgment I had been reading underneath it without noticing.
Those two arrived together for long enough that one could stand in for the other. That is the part coming apart. Not because implementation stopped mattering — the technical rounds still have to establish depth, and nothing here replaces them — but because producing plausible implementation, and a fluent account of it, costs less than it did. A signal that costs less carries less.
Hiring is how a profession states what it values, in the only form that can be audited afterwards. If what we mean by senior includes framing the problem, choosing under uncertainty, knowing which trade-off is unsafe to take, checking what the work actually did, and still holding it when the consequences arrive, then the loop has to look at those directly. That they came bundled with the implementation was never something we tested. It was a shortcut, and it held while the bundle did.
Which is why the panel needs more than one kind of eye on it. She cannot tell you how deep this person goes. She can tell you whether you could hand them the messy version of a problem — the one that is still an argument rather than a ticket — and trust them to help work out what is worth building, to weigh what the technical calls will cost later, and to still be answerable after the thing ships.
Enjoyed this article?
Subscribe via RSS
Follow along in your favourite feed reader. Every new post lands there as soon as it's published — no account needed.