What the May announcement establishes
The Department's classified-networks announcement, dated May 1, now lists eight providers: SpaceX, OpenAI, Google, NVIDIA, Reflection, Microsoft, Amazon Web Services and Oracle. It describes agreements to deploy advanced AI in IL6 and IL7 environments. An agreement is not evidence that every provider, model and intended workflow was already operating across those networks on the announcement date.
The distinction matters for planning. A program needs to know the specific service available, the authorized environment, the permitted data and the approved mission use. A provider's inclusion in an enterprise announcement does not answer those questions for a particular office. Nor should the behavior of a model in an unclassified demonstration be assumed to carry unchanged into a classified application with different sources and users.
The opportunity remains substantial: faster document review, more effective retrieval and assistance with complex analytical work. The engineering obligation is to show where those improvements hold, what errors remain and how users recognize the limits.
The assessment schedule comes from statute
Section 1533 of the FY2026 NDAA establishes a CDAO-led assessment effort. Its deadlines include a cross-functional team by June 1, 2026, functional leads by January 1, 2027, a standardized framework by June 1, 2027, and assessment of major AI systems by January 1, 2028. These are statutory milestones, not evidence that each milestone has been completed.
That work does not mean classified AI operates without existing controls until 2027. Cybersecurity authorization, classification rules and mission-specific approvals already apply. The practical gap to examine is whether those controls are accompanied by adequate evidence about the AI's behavior in the intended task.
The DoD Risk Management Framework includes ongoing monitoring, assessment and authorization activities. Calling every Authority to Operate a static, one-time check misstates that framework. A better question is whether a program's monitoring captures relevant model, retrieval and workflow changes as well as conventional security controls.
Test the system the analyst will actually use
A model version alone is an incomplete assessment boundary. Prompts, retrieval permissions, document collections, tool access and user behavior can change the result without changing the model's weights. A source update may improve currency while introducing conflicting definitions. An expanded search permission may improve recall while exposing material the task should not use.
CDAO's AI test-and-evaluation frameworks cover model behavior, human interaction, systems integration and operational performance. JATIC-related tools support reproducibility and traceability; they do not replace a mission owner's acceptance decision.
For a classified analytical workflow, useful evidence includes:
- Source fidelity: whether conclusions follow from the retrieved material, including caveats and conflicting reporting.
- Permission behavior: whether retrieval and outputs respect the user's actual access and dissemination restrictions.
- Change effects: whether a new model, prompt or document collection changes previously acceptable results.
- Human performance: whether analysts notice uncertainty, reject unsupported answers and retain enough time to review consequential outputs.
- Recovery: whether the team can withdraw a defective configuration and reconstruct affected work.
Evaluation data and logs need protection too. An audit package can itself contain classified prompts, sensitive metadata or revealing patterns of use. Design evidence access for authorized reviewers rather than assume an explanation can always be exported outside the operational environment.
Use model comparison as one check
Two models can disagree in ways that reveal an issue worth reviewing. They can also agree on the same false statement because they share sources, similar training material or an incorrect premise. Consensus is therefore a useful signal to test, not a substitute for independent evidence.
For important claims, connect model comparison to source verification and a reviewer who can resolve disagreement. Measure whether the added model actually catches errors on representative tasks, how many false alarms it creates and what it costs in latency and analyst effort. Where both models inherit the same retrieved document, a separate check of that document may add more value than a third generated answer.
Build a reviewable deployment package
- Define the intended tasks, prohibited uses and accountable mission owner.
- Record the model, retrieval sources, permissions and configuration used in testing.
- Test representative successes, ambiguous cases and known failure conditions with operational users.
- Establish change triggers, monitoring measures and a workable rollback procedure.
- Retain evidence in a form that authorized operational, security and oversight reviewers can inspect.
Suppliers that make these steps repeatable offer value beyond model access. Provenance, drift evaluation and usable audit records become part of operating the capability, rather than a separate exercise after deployment.
Sources and further reading
- Department classified-networks AI agreements
- FY2026 NDAA, section 1533
- DoDI 8510.01: Risk Management Framework
- CDAO AI test-and-evaluation frameworks
Spartan X's AI consulting, cybersecurity and engineering practices connect model evaluation to the surrounding mission system: the data it can reach, the actions it can support and the evidence needed to defend its use.



