Built for, and by, humans
Systems judged by the real action they take and the value delivered to a person, not by leaderboard position or parameter count.
The benchmark culture of the last decade optimised for something measurable but not especially useful: how a model scores on a fixed test, in isolation, with no account of what it then does in the world. That was a reasonable proxy while models were tools a person drove directly. It stops being reasonable the moment systems act on their own behalf over long horizons.
An agentic system’s value is the quality of the action it takes and the outcome a person actually gets. Those are harder to measure than a leaderboard position, which is precisely why the industry has been slow to build the measurement. But what we choose to measure is what the hardware underneath eventually gets optimised for — so getting the metric wrong is not a reporting problem, it’s an architecture problem that compounds for a decade.
Designing for people also means designing for the person operating the system, not only the person consuming its output. Observability, explainability and the ability to attest what informed a decision are part of being built for humans.