Gating frontier-model access behind a cost-justification form is a telescope pointed at the ground. You will satisfy every control objective and see nothing.
A widely circulated account has Mitchell Hashimoto — one of the best systems engineers of his generation — spending two hours and about $40 of Fable to push a gnarly low-level performance optimization to a level he says he could not have reached himself. Sit with the second half of that sentence. Not "faster than he could have." Better than he could have.
Here is the part that should stop a security leader cold: no routing table on earth would have assigned that task. A routing table can only allocate work someone has already imagined and queued. This one was never on the board. It existed because a person with a decade of contact-hours instinct got to point an expensive model at a problem nobody had scoped, on a whim, for the price of a team lunch.
The loud question in every regulated shop right now is "how do we control frontier-model spend before it controls us." I run security and DevOps for a fintech that has to prove its controls to more than 1,500 financial institutions and their examiners, so I feel the pull of that instinct as much as anyone. And I want to argue it is the wrong question, or at least the wrong first one. Gate access to the frontier behind cost-justification forms and approval workflows and you do bound the spend — but you also quietly defund the one organ your company has for sensing what the model can newly do. Governance, done right, is not a chokepoint on that sensing. It is the capacity to absorb what the sensing turns up, safely and reversibly. Governance is an absorption capacity, not a spend approval.
A routing table can only allocate work you already imagined
The same account is careful to include the boring half, and the boring half matters. On routine feature work, the story goes, the cheap open-weight model came in under a dollar, GPT-5.5 ran about a dollar-fifty, and Fable ran around nine — and the outputs were indistinguishable. There, routing to the cheapest competent model is correct, and it is dull, and dull is the goal. That is minimum-effective-intelligence routing working exactly as designed: stop paying frontier prices for janitorial work.
The two-hour, $40 experiment is a different animal entirely. Fable's reported list price — on the order of $10 per million input tokens and $50 per million output — makes the spread against a cheap model real, but against the outcome the absolute number is a rounding error. The value was never in the price. It was in a human being allowed to pose a question the organization had never thought to ask, and getting back an answer that beat what the best available person could produce. You cannot put that on a routing table, because the routing table's whole job is to move already-defined work to the cheapest competent tier. Discovery is not on the menu.
Now watch the reflex. In a regulated fintech, the instinctive response to "frontier models are expensive and a little scary" is to delegate all frontier contact to the function whose mandate is reducing spend, and to gate access behind a cost-justification form. Done reflexively, that hands your single sensing organ for what the model newly makes possible to the team chartered to say no. It is a telescope pointed at the ground. Every control objective satisfied; nothing seen.
And this is landing in policy right now. Sam Altman reportedly told an enterprise audience that blowing through the annual AI budget by the first quarter has become nearly a meme; Uber's CTO reportedly confirmed it happened to them. That anxiety is being written into AI-usage policy as approval gates this quarter — before anyone stops to ask whether the gate guards the right thing. The froth does not help: one restaurant chain reportedly named AI twenty-two times in its IPO filing, which is exactly the kind of number that puts "get AI spend under control" on every board agenda in July. The pressure is real. The instinct it produces is wrong.
Because a $40 approval form is the change-advisory board's old reflex rebuilt around a token meter. It is heavyweight ceremony gating the cheap, safe, reversible thing — a read-only experiment on public data — while the variable that actually carries risk sails through ungoverned. If you have ever watched a CAB spend forty minutes on a config toggle and then wave through a schema migration, you already know this failure mode. You are optimizing the wrong number.
Gate the data boundary, not the dollar
Here is the claim I will defend to any examiner: every control objective that actually matters — data classification, non-human identity, a complete audit trail — can be met without a single spend-approval chokepoint. "Is this use safe and compliant" and "is this spend justified" are two different questions. Conflating them into one approval form is a category error, and an expensive one, because it starves exploration while barely touching risk.
Put the gate where the risk actually lives: the class of data allowed to flow to a given model, and what the model's output is permitted to touch. That is the act-or-interpret and data-classification boundary, and it has nothing to do with the dollar figure. A $40 experiment on public or synthetic data needs no human approval at all. A $0.50 call carrying regulated customer PII needs the full control stack — classification, a scoped credential, the whole audit trail. Key your approval to cost and you have inverted the risk precisely: heavy friction on the cheap safe thing, a green light on the expensive dangerous one.
It gets worse, because the number you are gating is collapsing under you. OpenAI is reportedly weighing steep token-price cuts heading into its fight with Anthropic — and has reportedly floated a stake to the US government along the way. Whatever the details, the direction of token pricing is down and to the right. Standing up a permanent bureaucratic chokepoint to guard a cost that is actively falling is spending your scarce governance capital on the wrong axis. You will still be running the approval workflow long after the thing it guards has become a rounding error.
The mechanics to replace the gate already exist, and they are the same ones I would use to meter any AI spend. A hard cap enforced inline at a single internal gateway bounds the dollar blast radius without a human approving anything — the request that would blow the budget is refused or downgraded at the wire, in seconds, not reconstructed from an invoice a month later. And per-request attribution finally makes the spend observable after the fact: Amazon Bedrock shipped request-level usage attribution on May 20, 2026, and Microsoft Foundry shipped project-level cost attribution at the end of that same month. Caps bound the cost; classification bounds the data. Between them, "who is allowed to pose a $40 question without asking" stops being a purchase order and becomes what it always should have been — an IAM entitlement, scoped by data tier.
The impressive number is "a day." The durable asset is what made the day safe.
The other story everyone is quoting this month is Stripe, which reportedly migrated a fifty-million-line Ruby codebase in a single day against a manual estimate measured in months. Take the specifics with a grain of salt; the pattern is what matters, and the seductive read is "10x execution." That is the wrong asset to admire.
The durable asset is the thing that had to already exist for a one-day migration to be trustworthy at all: the test coverage, the review infrastructure, the change-verification pipeline that could tell you the fifty million lines still did what they did before. Speed is a firehose. Verification is the funnel. Ten-x execution pointed at one-x verification is not a capability — it is a brand-new mechanism for burning money and shipping risk faster than you can catch it. The migration was safe because the funnel was already built. Nobody demos the funnel.
And here is the part that should reassure a security leader rather than alarm one: the controls that make AI-generated change defensible to an examiner are the exact same controls that let you absorb the speed safely. Provenance on every change. A decision log you can replay. A deterministic-checks pass and then an independent, hostile-reviewer pass — the sign-off has to come from something other than the model that wrote the change, because author and approver can never be the same process. Those are audit controls and absorption controls at once. The asset that keeps you defensible and the asset that lets you go fast are one asset.
Which is where "can our control plane verify what came back" becomes the number a board should actually be tracking. Absorption capacity is half technical — the verification pipeline — and half organizational: who is allowed to route work to the frontier, and who reviews what it returns. A board that only asks "what did we spend on tokens" is staring at the one variable that is collapsing and ignoring the one that compounds.
Volatility is the case for building the muscle, not gating the door
There is a structural reason the approval workflow cannot work here, and it has nothing to do with bureaucracy being slow. An approval workflow assumes a stable catalog you can vet in advance. The frontier is not a stable catalog. Fable 5 launched on a Tuesday — June 9, 2026 — and, as reported, was gone by that Friday under a US export-control directive, offline for weeks before it came back. Routing tables rerouted around the gap within hours. A committee that met to approve it on Monday would have been guarding a corpse by Friday. You cannot pre-clear a catalog that reshuffles itself on a regulator's timeline.
Worse for the central-committee model, the patterns worth having are discovered at the edge, by the people doing the work, not in a spend meeting. One practitioner, CJ Zafir, reportedly halved a weekly burn rate with a "plan expensive, execute cheap, review expensive" pattern — a frontier model to plan, a cheaper coding model to execute, the frontier model again to review. A justify-it-in-advance gate structurally cannot discover that, because discovering it required permissionless experimentation at the edge. The gate does not just slow the discovery down. It makes the discovery impossible.
And do not imagine you can shortcut this by having a central approver simply bless "the best model." Opus 4.8 was last month's benchmark winner, and a careful practitioner still would not reflexively default to it for every task. Leaderboard rank is not a routing decision. No central approver, however senior, can pre-know the right tool for a given task better than the contact-hours instinct of the person whose hands are on the problem.
So build the muscle instead of gating the door. The absorption muscle is concrete: a gateway that holds the model at arm's length behind one governed egress, pinned model IDs with a scheduled re-validation cadence, hard caps, per-run observability, and the verification pass. The harness is the durable asset you build once; the model behind it stays a swappable part. That is the posture that turns a model suspension into a config change instead of an incident, and turns a $40 experiment into a governed, reversible event instead of a thing you either forbid outright or pray about quietly.
Four moves to stop gating the wrong number
Four moves, and none of them require a new platform:
- Split the safety gate from the spend gate, in writing. One mandatory gate keyed to data class and the action boundary. One spend question that is a cap, not an approval. Stop letting a dollar figure stand in for a risk assessment — they are different questions and they deserve different controls.
- Name, per data tier, who may pose a $40 question without asking. Grant permissionless frontier access on non-regulated, public, or synthetic data behind a metered gateway with a hard cap. Reserve human approval for the data boundary, never the dollar. Write it down next to your architecture decision records so it is a standing policy, not a favor someone grants.
- Build the verification pipeline before you scale the execution. Test coverage, an independent hostile-reviewer pass, per-run provenance, and a hard rule that whatever wrote a change is never what signs it off. That funnel is the durable, board-reportable asset — it is what lets you absorb 10x execution instead of drowning in it.
- Re-cut the board metric. Retire "what did we spend on tokens." Replace it with two sentences: who is allowed to pose a $40 question without asking, and can our control plane verify what came back. The first measures your sensing capacity; the second measures your absorption capacity. Those are the two numbers that compound.
Governance is an absorption capacity, not a spend approval. The company that wins the next two years is not the one that spent the least on tokens. It is the one that could safely say yes to the $40 question a person on the edge thought to ask — and then prove, afterward, exactly what came back.
So I will put it to you the way I would put it to a board: where have you drawn the line between a spend cap and a spend approval — and has anyone actually measured what their approval gate cost them in the things they never got to see? Tell me in the comments.
