Before a SNAP Fraud Flag Becomes a Decision: Calibration, Review and Resident Protection
Back to Signal
State & LocalBenefits AdministrationAIGovernmentCompliance

Before a SNAP Fraud Flag Becomes a Decision: Calibration, Review and Resident Protection

August 28, 2026Jess Loban

What Michigan teaches about automated decisions

Michigan's MiDAS controversy involved automated unemployment-insurance fraud determinations. The state's Bauserman settlement notice describes inaccurate auto-adjudication and collection of money, and the Attorney General announced final approval of a $20 million settlement in January 2024. The notice describes the October 2013–August 2015 period and reports that Michigan auditors found auto-adjudicated fraud allegations were wrong more than 90 percent of the time. That is a historical unemployment finding, with its own system and denominator.

Correction: The earlier version attributed historical Michigan unemployment figures to a recent SNAP audit and incorrectly suggested that no adverse actions occurred. MiDAS should not be presented as evidence about the error rate of a modern SNAP AI model. Its relevance is the governance risk of allowing automated signals to become consequential decisions without sufficient validation and recourse.

That distinction matters because a bad analogy can lead to a bad procurement requirement. An unemployment-system failure cannot provide an estimated error rate for a SNAP model. It can motivate rigorous testing of the decision process around that model.

Understand precision, recall and the denominator

Precision asks how many flagged cases are confirmed as the condition being investigated. Recall asks how many actual cases the system finds. Both depend on the definitions, data and reference outcomes used to evaluate them; neither can be inferred from the number of alerts alone.

Consider a hypothetical tool that flags 100 cases and, after appropriate review, identifies 10 confirmed cases of fraud. Its precision on those reviewed flags is 10 percent. That does not tell us recall: we still need to know how many fraud cases existed among unflagged records. It also does not justify treating the other 90 households as having committed fraud.

Low precision consumes investigative capacity and can burden innocent households. High precision achieved by flagging only obvious cases may miss substantial wrongdoing. The operating choice must account for both, along with review resources, legal requirements and the consequences of an error. There is no universal threshold that makes every benefits use case acceptable.

Keep fraud, payment error and eligibility separate

USDA's June 15, 2026 grant announcement made up to $5 million available for SNAP Fraud Framework implementation, including analytics, investigative coordination, oversight and training. It does not establish that an AI model is required or that receiving a grant demonstrates acceptable performance.

An improper payment can arise from an administrative mistake or an eligibility discrepancy without intentional fraud. A mismatch in wage data may reflect reporting periods, a corrected record or the wrong person. A fraud investigation therefore needs evidence beyond a discrepancy score. Do not optimize an AI procurement against a payment-error metric and then describe every improvement as fraud prevention.

The award conditions and solicitation should state how performance will be demonstrated. Where the grant leaves a delivery choice to the state, the agency still needs an acceptance test and someone responsible for it.

Make the system support a real investigation

In a useful triage design, the system presents the underlying discrepancy, data source, observation date and relevant rule so a trained reviewer can decide what to do next. A model-generated explanation is not evidence unless it is traceable to the record. Review time depends on complexity; a universal thirty-second case review would be an unsafe design assumption.

The reviewer must be able to clear an alert, request clarification or escalate it without being penalized for disagreeing with the tool. The process should distinguish an alert from an investigation, a confirmed finding and an adverse action. Each has different evidentiary and procedural requirements. Applicable notice, hearing and correction obligations remain part of the program's workflow.

The stakes are real even when a tool only routes work. Repeated requests for documents, long holds and inaccessible notices can prevent an eligible household from obtaining food assistance. Monitor those burdens alongside investigator productivity. Commercial false positives can also cause serious harm; the case for careful benefits design does not depend on dismissing them as mere inconvenience.

Revalidate when data or policy changes

Eligibility rules, source feeds and household circumstances change. A historical pattern can become a poor guide when the current program differs from the period represented in the training or evaluation data. A change in alert volume may signal increased fraud, an upstream formatting change or a broken linkage; investigate before interpreting it.

Review performance by relevant case type and population where lawful and statistically meaningful. Use independent adjudication or another defensible reference process, and account for the fact that reviewed cases may not represent all cases. Keep model versions, rule versions and decisions linked so an error can be traced and corrected.

A deployment sequence agencies can use

  1. Define the decision boundary. Specify what the tool can flag, what requires human judgment and what it must never determine on its own.
  2. Establish a representative baseline. Test on current program data with known limitations. Report precision, recall where estimable, review burden and errors—not a single accuracy score.
  3. Run a bounded evaluation before consequential use. Compare recommendations with independent review and examine cases the tool misses as well as those it flags.
  4. Test notices and correction. Ensure staff and residents can identify the reason for an action, correct data and use the applicable review or appeal process.
  5. Schedule and trigger revalidation. Initial reviews at 60 and 180 days can be useful planning checkpoints, not claimed federal requirements. Reassess sooner after policy, model or source-system changes.
  6. Define stop conditions. Set thresholds for unexplained changes, harmful errors and unmanageable queues. Give a named official authority to pause the tool and repair affected cases.

Sources and further reading

Spartan X's AI consulting and engineering disciplines are directly relevant to this calibration problem: defining what the model may do, testing its evidence and designing the review process around the consequences of an error.

Share this article
LinkedIn

BUILD WITH US

Ready to Solve Hard Problems?

Spartan X builds AI systems, autonomous platforms, and cybersecurity solutions for defense and national security.