What It Costs to Prove Autonomy Works
Every claim here carries its source. Feedback wanted, especially corrections.
Building autonomy has gotten dramatically cheaper. Proving it works has not. That inversion is recent, quiet, and specific to autonomy, and it now decides who can field what.
It lands hardest on the systems designed to be the future: cheap, numerous, and built to keep learning after they are fielded. Their economics only close if the cost of proving falls with the cost of building. Whether it does is the question this essay prices, ending at the one number the public record has never contained.
Here is what that record actually says.
The rulebook is empty, and it says so
On 29 May 2026, the Air Force issued MIL-HDBK-516D, the current airworthiness certification criteria for military air systems, superseding a 2014 edition. It contains a new Section 1.8, written specifically to address AI. Its key sentence:
"As part of this update, it was determined that the development of AI AW standards was nascent and therefore were not included." (MIL-HDBK-516D §1.8.1)
The document that defines what airworthiness means for military aircraft, revised two months ago, contains no criteria for AI or autonomy, and says so in its own text.
Section 1.8.5 offers a program two paths. One: provide an analysis showing the AI element "does not have the ability, under all expected operating conditions, to adversely impact the integrity of SCF operation." That is an absolute claim over a behavior space that §1.8.3 concedes "cannot be assessed with the traditional development and verification standards." Two: consult the airworthiness authority for "additional AW considerations," an unbounded, program-by-program negotiation.
The Army says the same thing in plainer words. DEVCOM engineers, writing on qualifying UAS swarms: "no standard criteria for airworthiness qualification exists."
An empty rulebook does not mean nothing is required. It means the requirement is discovered program by program, at negotiation, by whoever holds the signature. That is the most expensive kind of requirement there is.
Autonomy starts at the top of the risk scale, by definition
MIL-STD-882E, the DoD system-safety standard, assigns software a control category. The highest, Software Control Category 1, is defined as:
"Software exercises autonomous control without predetermined safe intervention capability." (MIL-STD-882E, Table IV)
Read that carefully: autonomy without a verified independent intervention path is top criticality by definition, not by assessment. And undemonstrable claims at top criticality classify as high residual risk, which moves the acceptance signature up the chain. Per DoDI 5000.88, that runs from program manager to program executive officer to component acquisition executive, with user-representative concurrence required at the top two levels.
The commonly assumed escape, wrapping the autonomy in a runtime monitor, relocates the burden rather than removing it. NASA's runtime assurance guidance is unambiguous: "the use of RTA does not replace or eliminate SUO assurance obligations", and "assurance of RTA cannot be a lower level than that of the SUO." ASTM F3269, the runtime assurance standard itself, agrees: "RTA components are required to meet the design assurance level dictated by a safety assessment process." A monitor changes the shape of the evidence problem; it does not shrink the total. Unless the monitor's case is reusable. Hold that thought. The record turns out to be silent on it, and it returns at the end.
What proving costs, where anyone has measured it
The best measured numbers in the public record, all from contractor cost data:
- System test and evaluation runs about 21% of fixed-wing aircraft development cost, measured across 16 programs from contractor cost data reports, stable across three decades (RAND MG-109).
- Systems engineering and program management add about 12%, from 26 programs in the same data source (RAND MG-413).
So roughly a third of aircraft development is neither designing nor building the thing.
Now look at what happens to development cost as a share of a program, using four unmanned aircraft programs' own Selected Acquisition Reports:
| Program | Quantity | RDT&E share of acquisition |
|---|---|---|
| MQ-9 Reaper | 436 | 12% |
| MQ-25 Stingray | 76 | 25% |
| RQ-4 Global Hawk | 45 | 43% |
| MQ-4C Triton | 27 | 52% |
(Sources: each program's SAR, public in the WHS FOIA reading room; base years differ across rows, so compare shares, not dollars.)
Monotonic: the fewer units you buy, the more of your money goes to development. Triton spent more than half its acquisition dollars on development for 27 airframes. The per-unit non-recurring burden is published too, as PAUC minus APUC: $3.2M per aircraft for MQ-9 at 436 units; $175.4M per aircraft for Triton at 27.
Now read the column a second time, as economics instead of history. When the vehicle gets cheaper, every major term shrinks with it: materials, engines, production hours. The proving does not. It is priced by the behaviors the system might exhibit and by the chain of signatures required to accept them, and neither of those reads the sticker price. A fleet of many cheap units is therefore a bet that proving cost follows hardware cost down. Nothing in the record shows that it ever has.
One more datum worth sitting with: MQ-25's development contract went from an initial target of $805.3M to an estimate at completion of $3,213.9M, four times the target, on a fixed-price-incentive vehicle, per its own December 2023 SAR.
Nobody can separate the proving from the building
You would think, given the numbers above, that someone could state what certification costs as a fraction of a program. Nobody can. Not the programs, not the FAA, not the analysts.
An FAA-funded study of advanced air mobility certification costs, written by Wichita State's National Institute for Aviation Research in 2022, says it directly:
"None of the models separated the costs of certifying an aircraft from the cost of developing it." (FAA/ASSURE A36)
The reason is structural. Cost accounting tracks what money bought (airframes, engines, test hours), never why. A test sortie flown to find a bug and a test sortie flown to generate evidence for a signer are the same line item. RAND's 21% contains both, undifferentiated. So that 21% is a ceiling on the cost of convincing, not a measurement of it, and no existing ledger, public or otherwise, can do better, because the distinction was never a reportable attribute.
It gets worse. The software-certification cost multipliers everyone cites ("DO-178 adds 25–40%," "certified code costs $100 a line") trace back, through a 2009 EUROCONTROL report, to a single consultancy's experience figures and a middleware vendor quoted in a trade magazine. There is no dataset under any of them. The FAA's own research report on software assurance costs (DOT/FAA/TC-15/57) declines to give a number at all.
And for autonomy specifically, there is no published cost estimate for certifying a learning-enabled aviation system. Anywhere. NASA's certification-considerations reports don't have one. The academic surveys don't have one. If someone quotes you a number, ask for the source.
Congress noticed the gap, for what it's worth: FY2021 NDAA §147 directed a study on "measures to assess cost-per-effect." I found no evidence the resulting measure exists.
The number that decides whether autonomy stays fielded
Everything above is the entry cost, paid once to field the first configuration. For autonomy, it is the wrong number to worry about.
A learned component changes. It retrains on new data, adapts to new threats, improves. Each change reopens some fraction of the assurance case: some of the analysis, some of the test, some of the approval. Call that fraction the reopening ratio.
Multiply it by how often the system changes and you have the recurring cost of keeping autonomy fielded. This is where the two worlds separate. A program that changes its fielded software every few years can survive almost any reopening ratio; the cadence absorbs it. A system built to retrain and redeploy on software timelines cannot. For it, ratio times change rate is not a maintenance line. It is the difference between a fleet that improves in the field and one whose velocity stops at the authorization boundary, aging while the threat moves.
As far as the public record shows, the reopening ratio has never been published for any program. Not in the SARs, not in the T&E literature, not in the certification research. The one number that decides whether fielded autonomy can keep pace with its own development is not in the record.
Which means the people who know it are the people who have lived it: the test leads and chief engineers who shipped a model change and then discovered what it re-triggered.
So here is my ask, and the reason this essay exists: if you have taken an autonomy software change through re-verification, on any platform, in any regime, I want to know what fraction of the case actually reopened, what it cost, and who required it. Even a rough answer. Even "we just pushed it and nothing reopened," which would be its own kind of finding.
I'll compile what I learn, anonymized and aggregated, and publish it back. The field is currently pricing its most important cost at "unknown," and that seems fixable.
Sources for every figure are linked inline. Corrections welcome; that's the point.
Answer the reopening-ratio question