
The Stack
The Take — An agent grading its own homework is an alibi, not proof
The most important product promise made this week isn't that agents write better code. It's that you can stop reading it. Cognition's pitch for wiring GPT-6 Astra into Devin, and Perplexity's parallel claim that it checks in less on a production answer engine, both sell the same trade: human review out, agent-produced evidence in. I think that trade is being made in the wrong direction, and that this week's other stories explain exactly why. The reason isn't that automated tests are unreliable










