Walk the floor of a science fair and every project is alive. Each has a poster, a champion who can explain it, and a demo that worked at least once. Nothing is ranked. Nothing is killed. The whole point is that everything gets to exist side by side, and the ribbons are mostly for showing up.
That is the exact shape of most enterprise AI programs right now. Forty pilots, or sixty, or a dozen — the number is a distraction. Every one has a sponsor, a slide, and a screenshot of the day it worked. None has a graduation gate, a kill date, or a number that says how it ranks against the thirty-nine others competing for the same engineers, data-access decisions, and scarce reviewer attention. That is not a portfolio. It is a science fair with a cloud bill.
A venture investor runs the opposite thing. A VC runs a book: a ranked set of bets, each sized to its expected value, most expected to return nothing, a few funded hard because the math says so, and the losers cut early so the capital flows back to the winners. The enterprise AI program that survives the next two years is the one that stops running a science fair and starts running a book.
Let me be precise about where I stand, because this is a chief-information-officer's problem and I am not sitting in that chair yet. I run security and DevOps for a fintech that has to prove its controls to more than 1,500 financial institutions and their examiners, and we built AgentOS, a governed internal agent platform with real users. That platform is where the demand side lands on my desk: every week, new use cases queue up asking to be pushed to production. I do not own the enterprise IT P&L. I do own the gate that decides which agent use cases graduate — so I have thought hard about ranking a portfolio you cannot afford to run whole. This is that method.
A portfolio is ranked. A science fair is just alive.
The disease is not that companies run too many pilots. It is that they refuse to rank the ones they run. Ranking forces a comparison, a comparison forces a judgment, and a judgment means telling a sponsor their thing lost. So the ranking never happens, and forty pilots sit in what people call pilot purgatory — permanently mid-stage, never scaled, never stopped, quietly consuming the two resources that actually constrain an AI program.
Those resources are not tokens. A widely cited MIT report on the state of AI in business last year found that the large majority of enterprise generative-AI pilots — the figure quoted everywhere was around 95 percent — produced no measurable P&L impact. The lesson is not that most pilots fail. Most bets in any real portfolio do not pan out, and that is fine when winners are funded and losers cut. The failure is running all of them at once, forever, as if that were free.
It is not free, and the cost is not the cloud bill. The scarce resource is attention — senior engineering time, data-access decisions, the few people who can review a model's output without rubber-stamping it. Forty unranked pilots spread that too thin to make any single bet succeed. Five ranked ones concentrate it, and concentration is what turns a promising pilot into production.
Score every use case, not the person who pitched it
The antidote is a rubric applied to every use case before it competes for a dollar of engineering time — the same three axes for the executive's pet idea and the intern's side project. This is how a disciplined investor reads a book. You do not fund the best pitch. You fund the best risk-adjusted expected value, and you make yourself write the number down.
- Expected value, not headline value. The probability the use case actually works in production, times the annual value if it does. A one-in-five shot at a big number and a near-certain shot at a small one can score the same, and both beat a beautiful demo with no path to a real workflow. Forcing the probability term onto the page kills "this could be huge," because "could" is where the whole disagreement lives.
- Feasibility, measured where projects actually die. Not "can the model do it" — models can do a startling amount. Is the data reachable and clean enough. Does it integrate with a system someone owns. Will the humans whose workflow it changes actually adopt it. Most pilots die on feasibility, not capability, and feasibility is the axis a good demo is built to hide.
- Risk, priced as reversibility. What data class it touches, what its output may act on, and how hard it is to unwind. In a regulated shop this is not a footnote — a use case that touches regulated customer data and acts autonomously carries a risk weight that can sink an otherwise strong expected value, and it should. Risk is the discount rate on the whole thing.
The output is not a precise number — anyone who says their pilot-scoring model is precise is selling something. It is a defensible ordering. When two use cases land a row apart, argue which is higher; that argument is cheap and useful. When one sits twenty rows above another, the conversation is over, and no one's charisma decided it.
Fund stages, not projects
The second discipline a VC has that most AI programs lack is that they do not write the whole check up front. They fund a stage, watch what it returns, and decide whether the next stage earns a bigger check. A pilot is not a promise to scale. It is money spent to buy information about whether scaling is worth it — an option, not a commitment. Treat every pilot as bought information and the weight of "killing" it drops away, because you got what you paid for: an answer.
The mechanism is old and boring and it works. The stage-gate model — Robert Cooper's product-development framework, older than any of this — puts a decision gate between stages, where a use case graduates, iterates, or dies. Adapted to an AI portfolio, it is four stages and three gates:
- Idea. Scored on the rubric, ranked against the book, cheap to enter and cheap to reject. Most ideas should stop here, and that is the stage working, not failing.
- Funded pilot. A small, time-boxed check to answer one question: does this clear its feasibility and risk on real data. The graduation criterion is written before the work starts — a specific, measurable result, not "it looked promising."
- Limited production. Real users, real data, real controls, a bounded blast radius. The gate is adoption and reliability under real load, and this is where the governance and control-plane work finally earns its keep — the third gate, not the first. You do not build the full control stack for a use case that has not earned one.
- Scaled. The winner. The one you concentrate the freed-up attention on. A real portfolio has few of these, and that is the point.
The checks get bigger as the evidence accumulates, never before. The most expensive mistake in enterprise AI is writing a production-sized check at the idea stage because a demo was compelling — money poured into a use case that has not survived a single gate.
Kill criteria are a feature, not a failure
Here is the line that separates a portfolio from a science fair, and it is written at funding, not at the funeral: every pilot gets a kill condition and a kill date the day it is funded, before anyone's identity is wrapped up in it. If by that date it has not cleared a specific bar, it stops, and that is not a mark against anyone. Up front — while the sponsor is still unattached — is the only time you can write that honestly, because once three months of effort have gone in, the sunk-cost reflex makes every flat line look about to turn up.
A healthy portfolio has a kill rate, and a meaningful one. If nothing in your AI program has been stopped this quarter, you are not running a disciplined book — you are hoarding, and the tell is that your oldest pilots are your least defensible. Gartner's TIME model — tolerate, invest, migrate, eliminate — exists precisely because application portfolios rot when nobody is allowed to say "eliminate," and an AI portfolio rots the same way, faster. Killing a pilot is not a loss; it returns capital to the pool — the engineers and the reviewer attention that can now go to a bet with a real chance. A VC does not mourn the write-off. It was priced in from the first check.
This is where I have to separate the argument from two adjacent ones, because they get conflated constantly. Governing a single agent well and metering a single experiment honestly are real problems, worth solving on their own terms. But they sit downstream. Whether a specific use case has earned a governed agent, a metered budget, or a slot on the reviewer's calendar at all — that is the portfolio question, upstream of every per-agent control. You can run an immaculate control plane and still be running a science fair, if the thing it governs is forty bets nobody ranked.
The portfolio is an operating model, not a spreadsheet
None of this works as a one-time triage. A scoring exercise you run once and file is a science fair with a spreadsheet stapled to it. The portfolio is an operating model: an owner, a cadence, and a cost you can actually see.
The owner is a single accountable person who runs the book — the intake gate every new use case passes through, and the one who convenes the rebalance. The cadence is quarterly at minimum: re-score the live pilots against the new ideas, re-rank, graduate what cleared a gate, kill what hit its kill date, and move the freed attention to the top. And the cost has to be legible. Technology Business Management and the FinOps practices around it exist to attach a real, defensible number to each line, so expected value is weighed against what a pilot actually consumes, not what its sponsor wishes it cost. A portfolio you cannot cost is a portfolio you cannot rank, because half the ranking is the denominator.
This is also the version of an AI program a board can actually govern. Forty green demo screenshots let a board decide nothing. The book — the ranked bets, the graduation rate, the kill rate, the concentration of spend behind the top few — lets it do its job: ask whether the allocation matches the strategy. I report to boards today, and the gap between those two artifacts is the gap between a meeting that produces a decision and one that produces a nod. Give the board the book.
Run the book, not the science fair
The method is four moves. Score every use case on expected value, feasibility, and risk before it competes for an engineer, and rank the book out loud. Fund in stages with graduation criteria written up front, and keep the checks small until the evidence is big. Write a kill condition and a kill date the day you fund a pilot, and be proud of your kill rate. Put a single owner on the book with a quarterly rebalance, so the ranking is a living operating model instead of a slide you presented once.
So I will ask you the question I ask myself every time a new use case shows up at the platform gate: what is the last AI pilot your organization actually killed, and did you write its kill condition the day you funded it, or the day you finally admitted it was dead? Tell me in the comments — I read every reply.
