NSPM-11: Turn AI Assurance Into Evidence the Mission Owner Can Use
Back to Signal
AIDefenseGovernmentCompliance

NSPM-11: Turn AI Assurance Into Evidence the Mission Owner Can Use

July 3, 2026Jess Loban

Read the assurance pillar alongside the adoption language

The memorandum sets four pillars: adoption, adaptation, assurance and accountability. It calls for security and functionality testing, evaluation, validation and verification; it also addresses government control over mission-dependent AI and protection against unauthorized disablement or material modification.

Its deadlines include a 90-day update to DoD Directive 3000.09 with annual review, and a 120-day review and update of procurement processes for rapid onboarding from multiple vendors. Those are directed actions, not proof that all resulting guidance has been issued. Accountability remains with commanders, directors and agency heads, including obligations involving lawful use, privacy and civil liberties.

This broad direction builds on existing requirements. The January 2023 version of DoDD 3000.09 already required appropriate human judgment and verification, validation and testing for covered autonomous weapon systems. Programs should map their actual obligations instead of assuming assurance began with this memorandum or that one identical approval gate applies everywhere.

The useful deliverable is an evidence package

An assurance claim needs a defined system and a defined use. Saying a model is accurate says little until the reader knows the task, data, failure consequence and tested operating conditions. A release can improve average performance while becoming worse on a small set of consequential cases.

Build a package that answers:

  • What was tested? Identify the model, supporting software, retrieval data, tools and permissions.
  • Where did it work? Describe representative tasks, measured results and the conditions under which those results apply.
  • Where did it fail? Include misleading confidence, missing information, hostile inputs and cases requiring human intervention.
  • What can change? Document supplier updates, administrative access, dependency changes and the government's approval controls.
  • Who can act on the evidence? Name the owner of release, restriction, incident response and reevaluation decisions.

The supplier-control question deserves engineering attention. A contractual commitment has to be supported by the architecture: access controls, update mechanisms, dependency inventory and recoverable configurations. A system that cannot identify what changed is difficult to assure after a vendor refresh.

Generate evidence while the system is being built

CDAO's evaluation frameworks cover models, human interaction, integration and operational performance. That breadth matters: a model can pass an isolated test while users misinterpret its outputs or its surrounding application gives it excessive authority.

  1. Define important use cases and unacceptable outcomes before integration.
  2. Preserve versioned test inputs, expected behavior and observed results.
  3. Run regression checks when a model, data source or connected tool changes.
  4. Test how users recognize uncertainty and recover from an incorrect recommendation.
  5. Keep the responsible official's decision, restrictions and monitoring plan with the release record.

This approach makes TEVV an ongoing engineering activity. It also gives the program a way to explain why evidence transfers to a new configuration—or why additional testing is needed. Logging should support reproducibility and oversight while respecting classification, access and retention requirements.

Accountability works best when officials receive a clear decision rather than an undifferentiated stack of reports. Show the mission benefit, remaining uncertainty and available restrictions. Vendors that deliver that clarity make their capability easier to evaluate, integrate and sustain.

Sources and further reading

Spartan X’s AI and cybersecurity practices connect these policy obligations to system design, testing and release decisions, giving mission owners a clearer basis for adopting capability with confidence.

Share this article
LinkedIn

BUILD WITH US

Ready to Solve Hard Problems?

Spartan X builds AI systems, autonomous platforms, and cybersecurity solutions for defense and national security.