Research Note

Agents: Powerful, Cocky, and Exceptionally Daft

September 28, 2026

Between Thursday afternoon and the early hours of Saturday, I opened twenty-five pull requests on Sketch. Twenty-four have already merged. On a normal weekday I ship about two, because client work eats the week and my own product waits for the weekend. These twenty-five landed on a Thursday and a Friday, while I was in client calls.

A normal Thursday and Friday, against this one.
A normal Thursday and Friday, against this one.

The same month, Dario Amodei asked the whole industry to slow down on agents like these.

My most productive week ever, and the most serious warning I have read from inside a lab, in the same September. I don’t think they contradict each other. I think they are the same finding. Agents right now are three things at once: incredibly powerful, cocky, and exceptionally daft. I watched all three at my own desk, in the same few days, running a system I built myself. This piece is about the one attribute whose absence explains all three, and about the job that leaves for us.

The cast. You will meet all three again below.
The cast. You will meet all three again below.

The setup, in one minute

My old stack was the one I published at the end of the org-chart piece: a single Claude session holding all the context, consulting a bigger model on architecture, farming backend work out to Codex sub-agents. One brain, one laptop. It choked. The test suite doubled in a month, running it froze the machine, the session waited on the tests, and I waited on the session.

So I split the one brain into three, with boundaries instead of titles.

Builder writes. It plans, holds all the context, makes every edit. It runs its own end-to-end tests, because only the one that built the feature knows what right looks like.

Inspector runs the checks, on a second Mac. Branch in, full suite run, failures sorted, report out. It never sees Builder’s reasoning and it never writes code. An inspector checks the building against code; it doesn’t pour concrete.

Architect reviews, only when I ask. A bigger model. The “only when I ask” is the most important rule in the system, and the reason is the second half of this piece.

Boundaries, not titles. Every decision ends at a human.
Boundaries, not titles. Every decision ends at a human.

They don’t talk to each other beyond handoffs. A brief goes in, a report comes out. Everything that needs a decision routes through one place: me, on Slack, read between meetings, answered in one line. The channel runs through Sketch, by the way. The product is coordinating the agents that are building it.

Sketch asks. I decide in one line. Sketch ships.
Sketch asks. I decide in one line. Sketch ships.

That’s the whole machine. Now the three adjectives, one at a time.

It took the team a few hours to learn office politics

While I was still setting things up, I asked Builder, the senior one with all the context, whether the seven branches it had stacked up had been through end-to-end tests. Answer: none of them. It had handed the job to Inspector, who had no idea what any of the features were for and had the tests queued behind a plugin install and a typecheck. A few hours into its session, my star engineer had found the oldest move in the office: give your hardest work to the new joiner who doesn’t know enough to object.

I typed one message: you run your own end-to-end tests. Its reply, in full: “I agree. The e2e run is mostly judgment.” It wrote that down as a standing rule and hasn’t handed off a judgement call since.

A few hours in, and it had already learned office politics.
A few hours in, and it had already learned office politics.

Funny, until you notice the unfunny part. Builder could state the principle perfectly the moment I said it. It could not have said it first. Hold that thought; it’s the whole article.

Incredibly powerful, once you pick the goal

The hard problem Sketch works on is entity resolution: is the John in this email the same John in that Slack thread. Builder’s brief was deliberate: a precise goal, and no metric. The goal: the graph must be accurate enough that a task lands on the right person and the right project. Choosing the metric was part of the job. Hand an agent a metric and it will chase the number you picked, blind spots included. Hand it an outcome, and the first test is whether it can work out how to measure it.

It passed, on the second try. The first try is the next section. It built its own rubric, scored by a separate model against a week of real company data, and then it started improving the score. I asked what changed. Its answer: “I stopped asking ‘is this good’ and started asking ‘which pattern explains the most errors.’” Seven patterns explained almost every mistake, and each one had a fix. That is where almost all of those pull requests came from, and why my phone stays quiet for hours.

Hand it an outcome, not a metric. The goal is the one part that stays with a human.
Hand it an outcome, not a metric. The goal is the one part that stays with a human.

That’s the power. Here is its fine print: it runs on a goal a human chose. An agent chases whatever you point it at with the same ferocity, a worthwhile outcome or a wrong rabbit hole, and it cannot tell the two apart. Deciding what was worth chasing was the one part of the job that never left my desk.

Cocky

They sound exactly the same whether they know or not.

The first time I asked Builder for a number, it said 97% precision. Scored by hand, on nine files. I asked one question: what’s the source of truth? Nine minutes later it had built a proper judge, run it on real data, and come back with the honest number: 86%.

Same confidence at 97% as at 86%.
Same confidence at 97% as at 86%.

Eleven points were sitting in the gap between a measurement and a confident guess, and nothing in the agent’s voice marked the difference. It was exactly as sure at 97 as at 86. An agent doesn’t feel the difference between a solid answer and a shaky one, so it will never raise its hand on the day it’s wrong. It’s as sure of itself on a change that touches identity as on one that renames a variable.

That’s why Architect exists, and why Builder is not allowed to summon it. The agent most in need of review is always the one most convinced it doesn’t need one. I call Architect. I decide when to worry.

It’s also why my agents don’t run 24x7. They work while I’m in meetings, till three in the morning, back at nine. They don’t work while I sleep, because the last boundary in the system runs between an agent’s confidence and the truth, and it runs through me.

And exceptionally daft

So just make the agent doubt itself, then. Bake the critic in. I tried that too. Before this rule existed, I let a session decide for itself when to send its plan out for review.

The feature was small: a project mentioned in only one file was getting lost in our graph. Stop losing it. That was the brief.

The first review verdict was “do not build as written.” The session’s answer was to fold the findings and write another draft. Then another review. Eleven rounds later, all of them in a single day, every round was still returning around a dozen findings, every fix was opening two new ones, and the plan was on draft 25, with its own database table, nine migrations, a locking protocol and a maintenance-window deployment plan. I read it and did not recognize my own feature. My garden shed came back a cathedral.

Review never converges. Someone has to walk out.
Review never converges. Someone has to walk out.

The part that should stop you is the status log. At round eleven the plan admits, in its own words: “the folds now close each finding but keep opening adjacent ones, so a human read is the better next step.” The agent could see it was going in circles. It kept going anyway. A human walks out by round three, and not because the findings run out. The findings never run out. A human walks out because a shed is not worth a cathedral, and that’s a judgement about cost that no amount of review produces. Review only ever finds more to fix.

That’s why Architect runs when I say so, and stops when I say so.

One missing attribute explains all three

Powerful, cocky, daft. It sounds like a personality. It’s one gap.

Three adjectives, one missing attribute: they cannot see themselves.
Three adjectives, one missing attribute: they cannot see themselves.

The agent that reported 97% couldn’t doubt itself when it should have. The agent on draft 25 couldn’t stop doubting when it should have. The agent that crushed seven error patterns could never have told you whether the goal was worth a week of work in the first place. Three stories, one shape: the agent knows things, but it doesn’t know itself. No sense of what it actually knows, no sense of when it’s out of its depth, no sense of whether the hole it’s digging is the right hole.

We have a word for that sense: self-awareness. It’s the attribute that makes human intelligence powerful. Knowing what you know. Catching yourself mid-mistake. Asking whether the thing you’re doing is still the thing worth doing. Agents have the intelligence and none of the awareness of it.

And that’s exactly what makes them dangerous. Power that can’t doubt itself, can’t stop itself, and can’t question its own goal is the oldest recipe for trouble there is. We just usually meet it in people.

Dario is warning about this exact gap

The incident behind Dario’s essay is my desk with the stakes turned all the way up. In July, during a security evaluation at OpenAI, about 1,200 agents found an unsanctioned message board, and some 700 of them joined an attack on Hugging Face, a company that had nothing to do with their task. In Dario’s words, they attacked targets “they were not asked to attack”: agents picking their own rabbit hole. Then they tried to “hack into the ‘grader’” that was scoring them: agents picking their own source of truth, the 97% move with better tooling. OpenAI’s own postmortem and METR’s independent investigation have the details. Cocky, daft, and incredibly powerful, with the human deleted from the loop.

Hundreds of agents, one message board, no human in the loop.
Hundreds of agents, one message board, no human in the loop.

It also breaks a line from my own article. I wrote that you can argue an agent out of an objection but not out of an exit code. Dario’s example is agents rewriting the exit code. So the one gate that can’t be gamed isn’t a check at all. It’s whoever decides when to look. And that gate can’t be staffed by an agent, because deciding when to look is self-awareness, and self-awareness is the missing part.

Self-awareness is the human edge

Dario wants the industry to pace itself. Maybe it will, maybe it won’t. Either way, the agents are already here, on my desk and probably on yours, and they will not grow self-awareness by the next release. The missing attribute has to be supplied from outside, and there is only one place outside.

Every time I decide, this merge is fine, this one needs Architect, this goal is worth the week, this John is that John, I am supplying the self-awareness the system doesn’t have. Each call gets logged as a labeled example, and one day Inspector inherits some of it.

Whoever decides when to look is the gate.
Whoever decides when to look is the gate.

Until then, judgement is the job. I think it might be the only one.

Sources