Jay Kruer on why LLM autonomy still runs into specification costs
A sharp counterweight to frontier-lab autonomy narratives: the bottleneck may be domain specification, validation, and review capacity rather than raw model capability.
Links: Original source · Shared link
Logged at IST: 2026-09-17 16:01 IST
What it is: Mario Zechner points to Jay Kruer's essay arguing that recent headline LLM wins do not erase the structural limits on autonomous knowledge work.
Gist: Kruer's core claim is that LLMs can learn many specific tasks, but brittle generalization and reward hacking make fully autonomous deployment depend on rigorous specifications or expert human review. Both are expensive. Navier-Stokes-style theorem proving is the rosiest case because the problem statement, proof checker, and math library already provide unusually strong specification and validation machinery.
For most knowledge work, Kruer expects LLMs to remain more like a cracked intern: useful and fast under adult supervision, but not something you hand the whole workplace to. The likely winners are domains where failure is cheap, tasks are narrow and well-guarded, or rigorous validation is already part of the cost structure, such as chip design and drug discovery.
Newsletter angle: Pairs well with agentic-software-factory optimism: the hard question is not just whether models can generate work, but who pays for the specification and validation surface around that work.