
Which frontier model is at the top of the AI benchmark leaderboard was the right question when a model was a tool a person picked up, used, and put down. It is the wrong question now that workflows and agents autonomously use models, have credentials, tools, and network access. Worse, the wrong question quickly compounds into the wrong architecture, the wrong contracts, and the wrong org chart. Let’s explore.
Benchmark performance is quickly converging for the frontier models we (non-military, non-government, non-nation-state people) have access to. The gap between the best available closed models and the best open-weight models is now measured in months (maybe weeks), and the cost of equivalent intelligence is falling by roughly an order of magnitude a year. What does not (and likely will not) converge is operational readiness – the harness you build plus the organization you become. Identity, telemetry, and controls are the harness. Governance, incident response, and culture are the organization. One is an asset, the other is a capacity, and none of this has anything to do with the benchmarks achieved by the frontier model builders.
Governance
Governance maturity can be as simple as a named executive who can reconcile security, legal, procurement, finance, and technology, and who owns the decision when those functions disagree. Most big companies have the opposite: five functions that can each say no, and nobody who can say yes. Maturity is measurable. Measure it.
Identity and Permissioning
An agent acts on your behalf. It authenticates, it holds credentials, it invokes tools, it moves laterally, and it does all of this at machine speed while many critical controls run at human speed. Companies that grade on benchmarks give their agents shared credentials and broad permissions because they are easier to ship. Companies that grade on operational readiness give their agents defined identities, a least-privilege role, bounded credentials, approved tools, transaction ceilings, and an expiration date, exactly the way they treat a human with root access, because in practice that is what an agent is.
Telemetry
Humans cannot supervise a hundred agents by manually watching them. Agents need to be instrumented so that escape attempts, disabled monitoring, privilege escalation, unexpected egress, and cross-environment activity generate alerts a human sees while the activity is still happening rather than a week later. A week is not a hypothetical. In a recent incident that focused everyone’s attention on the potential of “rogue AI models,” an autonomous agent operated outside of its sandbox for days and its own creator did not connect the activity to the breach until after the target company disclosed it. The gap between machine-speed action and human-speed attribution was a control failure. Telemetry solves for this.
Agent Controls
Controls follow from identity and telemetry but deserve their own score. The agent that caused damage had a task, and it also had credentials, tools, network paths, and time. Humans need to be on top of all of this. Controls mean hard limits on what an agent may do, tested stop mechanisms, and failure drills you run before you scale, not after. Revoke a permission, poison a data source, force a quality check to fail, and watch whether the workflow stops safely, preserves evidence, and calls a human, or whether it improvises its way into your production database.
Incident Response
This is the readiness capability most of my clients assume they already have because they have cybersecurity, business continuity, and resiliency plans in place. Don’t make this assumption. AI incidents move faster than those plans were written for, and they require severity-based deadlines, cross-company escalation protocols, and a capable model you can run yourself. When Hugging Face investigated its own breach a few weeks ago, closed frontier models refused to analyze the attack commands and payloads its investigators needed to read. The defenders ran an open-weight model on their own infrastructure and kept the evidence inside their environment. If you are going to depend on your AI tech stack to do real work, you should stage a defensive model that can run locally, before it is needed, because the day you need it, the guardrails on the closed hosted models will block the exact work you are trying to do.
Organizational Readiness
This line determines the quality of every other line on the scorecard. The gap between what your AI tech stack can do and what your company can absorb is called “cultural debt,” a term you will not find any industry benchmark for. Most enterprises have people at every stage at once, from employees locked out of capable models to a few power users directing swarms of agents (and almost everybody is quietly wondering whether they are training their replacement). You cannot govern, permission, or instrument your way past a workforce that is applauding AI in public and stalling it in private. Operational readiness is the leadership work of naming that fear and building a culture that improves and continuously adapts.
The Vendor Problem
At the moment, most AI arrives embedded in software your company is already using (Microsoft (MSFT), Salesforce (CRM), ServiceNow (NOW), SAP (SAP), etc.) where the vendor owns the model, the runtime, the credentials, and the logs. You cannot instrument an agent that runs in someone else’s tenant.
The scorecard tells you how to negotiate with your AI vendors. Demand the identity model they run your agents under, the telemetry they will share, the incident-response deadlines they commit to, and the audit rights that prove the rest. A vendor that cannot commit to these terms in writing is asking you to carry its risk on faith.
Where Capability Still Wins
At the true frontier, capability is everything. An eight-month lead on weather forecasting, CBRN risk assessment, or novel research is strategically decisive. No amount of governance substitutes for a model that can actually do the work. That said, only a small amount of corporate workflows and processes live at the frontier. Reconciling invoices, testing code, drafting documents, and sorting tickets need a model that is “fit for purpose,” nothing more. Your operational readiness scorecard decides whether it clears that bar safely. The “mine is better than yours” benchmarks just don’t count.
Your AI advantage will come from your ability to permission, watch, stop, investigate, and staff what you deploy. It is unglamorous, expensive, and practically impossible to demo.
For most business functions, available models are already fungible. Become as model agnostic as practical, stop asking which model is best, and get serious about continuously improving your AI operational readiness. As for enterprise AI deployment at scale, cultural debt is the highest hill to climb. Build a scorecard (and the workflow and process required to keep it up-to-date). It will make the climb much, much easier.




Comments
Log in or sign up to join the conversation.