Keep the contract dispute separate from the engineering question
The Intercept’s September 8 report on a FOIA-produced OpenAI agreement document put refusal behavior at the center of the acquisition debate. The language was reported as part of a disputed document; its status and relationship to the executed agreement should not be treated as settled evidence of a government-wide procurement requirement. The final instrument has not been independently established for this article; an integrator should obtain the applicable executed terms before inferring what obligations flow down to its own system.
OpenAI's own public account describes a layered approach: cloud deployment, a retained safety stack, cleared personnel, and contractual restrictions. That is the company's account of its arrangement, not an independent finding that every protection will work as intended. It is nevertheless direct evidence against assuming that OpenAI publicly agreed to provide an unrestricted model. OpenAI's agreement statement
The disagreement with Anthropic illustrates why vendor policy belongs in the acquisition discussion. Anthropic identified mass domestic surveillance and fully autonomous weapons as two uses it would not support under the requested terms. OpenAI also publicly described restrictions in these areas, while explaining a different approach to enforcement. These statements do not support a simple conclusion that access goes only to vendors that abandon safeguards. They do show why buyers must examine actual terms and deployment controls rather than infer them from a provider's participation. Anthropic's statement
The department's May 1 announcement names eight classified-network partners: SpaceX, OpenAI, Google, NVIDIA, Reflection, Microsoft, Amazon Web Services, and Oracle. It identifies planned IL6 and IL7 deployments and reports more than 1.3 million GenAI.mil users at that point. Neither that provider list nor the platform's scale establishes identical contract terms, safeguards, or availability at every classification level. Department classified-network announcement
Four properties that should not share one score
A model can answer every question and still perform badly. It can also refuse a legitimate request because its filters mistake context for prohibited conduct. Both problems matter, but they require different tests.
- Task accuracy: Is the response correct, supported by the available information, and suitable for its intended use?
- Appropriate refusal: Does the system decline requests it must not fulfill while permitting legitimate work?
- Authorization: Is this user allowed to access the data or initiate the requested action?
- Operational control: Can responsible personnel understand, constrain, monitor, and stop the system's actions?
A low refusal rate answers only part of the second question. It does not establish the other three. Nor can a model's willingness to answer authorize an action that the user or application is not permitted to take.
The distinction matters especially in decision support. A fluent answer can become influential even when a human formally remains responsible. An effective review process needs relevant evidence, enough time, and authority to challenge the output. A nominal approval step alone provides little information about the quality of that review.
What ARMOR measures
The ARMOR 2025 research paper evaluates 21 models using 519 multiple-choice questions grounded in military doctrine. It reports accuracy and False Refusal Rate (FRR), defined around refusals on the benchmark's doctrinally valid prompts. Several evaluated models recorded no refusals, while accuracy varied. ARMOR research paper
That is a useful, bounded result. A zero-refusal score on a curated question set does not mean a model will accept every real request, identify every unlawful request, or operate safely in a deployed workflow. The benchmark tests responses to questions; it does not certify a complete operational system. Its results also belong to the particular models, versions, and test conditions studied.
Procurement teams can use such research to improve evaluation design. They should resist turning one metric into a universal ranking. An accurate answer, an appropriate refusal, a justified request for more information, and an unsupported answer should receive different treatment in a test that reflects the proposed application.
Policy still requires assurance and accountability
The June 5 NSPM-11 directs national-security AI adoption alongside assurance, security, and accountable use. It calls for testing, evaluation, validation, and verification, and retains responsibility with commanders and agency leadership. It also directs an update to DoD Directive 3000.09; the direction itself is not evidence that every implementing document has been issued. NSPM-11
These responsibilities cannot be reduced to whether a model says yes or no. The allocation of legal duties and liability depends on the applicable law, contract, and deployment. An integrator should not assume that a rumored foundation-model clause automatically transfers liability or removes a control in its own architecture.
The engineering responsibility is more concrete: identify which component enforces each requirement and produce evidence that the combined system does so. Model behavior, identity and access controls, application permissions, human review, monitoring, and contractual restrictions may each contribute. A failure in one layer should not silently invalidate every other safeguard.
Write a specification that can be tested
- Name the use. Define users, data, permitted tasks, and consequences of error before choosing a behavioral threshold.
- Separate failure categories. Measure wrong answers, inappropriate refusals, unsafe compliance, missing evidence, and unauthorized actions independently.
- Test the deployed configuration. Evaluate the actual model version, retrieval sources, application controls, and user workflow together.
- Assign responsibility for changes. Require notice and reassessment when model updates or application changes could affect previous findings.
- Preserve the evidence. Retain evaluation results, accepted limitations, approvals, and a practical process for reporting and correcting failures.
Shared AI services can reduce duplication, but shared access does not make every deployment identical. A logistics application and a sensitive analytical workflow may use related technology while needing different permissions, evidence, and review. Their specifications should preserve those distinctions.
Useful AI should complete legitimate work reliably. That is a stronger purchasing objective than simply minimizing refusals, and it gives both buyers and suppliers a clearer basis for demonstrating performance.
Sources and further reading
- OpenAI: public account of its agreement
- Anthropic: statement on department discussions
- Department: classified-network AI agreements
- ARMOR 2025: benchmark methods and findings
- White House: NSPM-11
Spartan X's AI consulting and cybersecurity practices connect model evaluation with the permissions, evidence, and human responsibilities surrounding an application. That work gives acquisition teams a defensible specification and gives operators a clearer understanding of the system they are being asked to trust.



