A number that changes the argument
The National Security Agency is spending billions of taxpayer dollars this year evaluating and testing advanced AI models, the Washington Sun reported on Thursday, citing two people with access to classified intelligence estimates. The sum is far larger than had been publicly understood.
Those sources are unnamed, and the underlying estimates are classified, so the figure cannot be checked against a document. What can be checked is the direction it points. The NSA’s Artificial Intelligence Security Center began testing frontier models for national security weaknesses after a run of high-profile incidents, and the agency’s senior officials said in August that they wanted access to all AI models.
A Pentagon spokesperson told the paper that the department does not discuss the technical architecture or resource allocation for its AI tools.
Where the money goes
Compute dominates. Running and testing a frontier model means buying the same accelerators everyone else is bidding for, in a market where supply has not caught up with demand, and a government buyer has no special claim on it.

People are the second line. Top AI engineers command packages the federal pay scale cannot approach, which leaves the agency competing for exactly the staff the labs are also hiring.
Why it matters for regulation
The reporting frames the number as an argument about the cost of oversight, and lawmakers appear to be reading it that way. The Congressional Budget Office scored one proposed measure, an AI Security and Innovation Act, at roughly $20 million a year, and an AI reporting and tracking bill at $36 million over five years. Against a single agency spending billions, those figures look like they were written for a different technology.

Nathan Calvin, general counsel at the advocacy group Encode, told the paper that in-house AI evaluation capability is “extremely important and necessary” but “genuinely expensive”. Nat Purser of the AI Verification and Evaluation Research Institute made the same point about funding: assessing these systems requires paying for the computing resources and the expertise.
What to watch
The specific thing to look for is whether any of this reaches an unclassified document. A congressional appropriation line, a public CBO score for a testing regime, or an NSA statement about its evaluation programme would turn an anonymously sourced figure into something a legislature can argue about. Until then the cost of testing frontier models in government is a number only a few people have seen.