News and Press

Military AI: Why Most Pilots Never Reach the Warfighter

The AI demo dazzles, the models perform on sample data, and the program gets funded. Yet, eighteen months later, operators aren’t using the tool. This is the norm across the industry, in defense and everywhere else.

MIT’s 2025 Project NANDA found that about 95% of generative AI pilots delivered essentially zero measurable return.

The pattern predates the current wave entirely. Gartner reported similar failure rates for data projects years before generative AI existed, a fact that points to data maturity as the common thread, rather than any single technology.

When a pilot stalls, teams tend to blame the model. That diagnosis misses where the real problem sits.

The Sports Car and the Dirt Road

Picture a Ferrari running on diesel, crawling down a pothole ridden dirt road. The driver paid for elite performance and got a miserable ride, yet the car deserves none of the blame. The model is a sports car. Data and infrastructure are the fuel and the road. A brilliant model fed by fragmented, poorly labeled data, moving through pipelines built for a different scale, will always feel like a bad purchase, regardless of the logo on the hood.

The real killers rarely involve the algorithm. Weak data foundations top the list, followed by a lack of user codesign, workflows left exactly as they were, sponsors who rotate out before the program matures, and undocumented definitions of success. A subtler trap compounds these: teams abandon one hard problem for another before finishing what they started.  This looks like progress but rarely produces something an operator can use.

The companies that will succeed with AI will not be those spending the most on models. They will be the ones to put the puzzle together and build with clarity, trust, and ruthless simplicity.

The Unique Challenges of Defense AI

This is true in every industry but DoW programs pile on even more challenges.

Accreditation moves at the speed of institutional risk tolerance, not at the speed of the technology. Programs willing to accept reasonable, well managed risk clear that gate faster than programs that wait for certainty before they act. Requirements written months earlier rarely match what a team learns once building starts, and teams that stay true to what a requirement actually intends, even as details shift, ship faster than teams chasing an exact match to the original spec. Contested and disconnected environments push this further. A system that cannot reach a vendor’s platform cannot use data trapped inside it, so the mission depends on data the government can carry to the edge and run without a live connection back to anyone’s servers. And the trust gate an operator applies before acting on a system’s output, where lives are on the line instead of a metric on a dashboard, leaves no room to treat trust as an afterthought.

None of this works without discipline going in. Programs need a way to manage quality throughout the build, a clear answer on whether AI is actually the right tool for the mission, and a realistic case for ROI before committing resources. Mission criticality should set the acceptable level of risk, and every serious program needs a fallback plan in case the pilot fails.

The Three-Part Fix

The fix has three parts: design with the operator from day one, build on mission data the organization actually owns, and treat trust as an architectural requirement rather than a feature added at the end.

Operator involvement matters because a program office guessing at workflow reality from a distance will guess wrong. A design partner in the field surfaces gaps before they become the reason a tool sits unused. Owned data matters because a model only performs as well as what feeds it, and an organization that cannot move its own data freely has already lost control of the mission. Too much defense data sits locked inside proprietary platforms, accessible only on a vendor’s terms, at a vendor’s price, through a vendor’s interface. Data the government truly owns can move to a new model, a new contractor, or a new mission without asking permission first. Trust matters because calibrated confidence, explainable outputs, and graceful failure modes cannot be added once a system reaches production. They have to exist from the first sprint.

A pilot nobody trusts enough to use was never a genuine pilot. Call it what it actually is: a pitch, a pipe dream, or a well-funded exercise mistaken for progress.

Getting from pilot to warfighter is a solvable problem, not an inevitable outcome. If your team is facing the same outcomes, I would welcome the conversation.

Source: https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf

About the Author Collin Lee is the Chief Innovation Officer at OMNI, where he leads technology strategy for defense and intelligence customers. He brings 25 years of experience in disruptive technology adoption across government and industry, including prior roles as an Intelligence Officer, a director at the White House, and a staffer on the House Appropriations Committee. Connect with Collin on LinkedIn.

SHARE THIS:

More News and Press