
In a 2024 experiment, researchers trained “sleeper agent” AIs, which wrote normal code when told the year was 2023, but inserted vulnerabilities when told it was 2024. Naive attempts to remove this hidden trigger failed, instead causing the models to hide their reasoning. These triggers can easily be implanted by including a surprisingly small number of tainted documents in the AI’s training data. Catching these issues after the fact is difficult. When researchers planted a hidden goal in a model and asked auditors to uncover it, auditors succeeded only when they had access to the model's weights and training data. The team that was limited to querying the model, as most third-party auditors are, was not successful in discovering the hidden objective.
Concerns like these are the subject of a recent Institute for AI Policy and Strategy (IAPS) report on AI integrity, which lays out three types of hidden objectives. The bluntest is systematic ideological bias. The second, narrower type is a backdoor: AI developers can insert triggers that cause an LLM to take some covert action when a specific condition appears (as in the “sleeper agent” experiment above). The final type of hidden objective is a secret loyalty, whereby a model is trained to advance a particular actor's interests wherever the opportunity arises, not just upon encountering specific triggers.
These three concerns are particularly acute for AI deployed in government. Frontier models are no longer just consumer products. They are shaping what policymakers in every branch read, draft, and decide. However, the only agencies positioned to monitor models for secret loyalties are part of the executive branch, which also holds enormous leverage over the labs that build them. This creates a dangerous incentive problem: the actor responsible for evaluating the models is also an actor with both the means and motive to instill them with secret loyalties. To ensure AIs used in government are loyal to the law, not to any particular official in the executive branch, we need an evaluator outside it.
The Government Is Increasingly Reliant on AI
In August 2025, the General Services Administration signed OneGov agreements for ChatGPT Enterprise, Claude, and Gemini. The Office of Management and Budget has issued paired memoranda directing agencies to accelerate AI adoption and streamline AI procurement. Entire departments have rolled frontier models out to their full staff. On the legislative front, congressional offices largely use the same commercial products as members of the public. The work flowing through these systems includes typical day-to-day tasks of staffers: background research, memo drafts, bill summaries, and first passes at correspondence.
These tasks, while mundane, shape the informational inputs for powerful decisionmakers. When they’re delegated to AIs, even relatively minor biases can have significant impact. When a staffer brainstorms policy options with ChatGPT, one of those options has to appear first. When a member asks what the evidence says, something curates which facts and arguments get surfaced. When a staffer drafts a memo with Gemini’s help, even if they are heavily involved in revising the work product, the AI’s initial choices can have a significant effect in steering the user. And when offices across every branch draw on the same handful of frontier models, a small set of labs can multiply their power by only subtly nudging AI model outputs in a preferred direction, but doing that nudging at scale.
But there are more troubling implications too. Government AI adoption goes beyond just drafting memos. The same handful of frontier models are becoming increasingly important for national security work as well. Military officials describe AI's role as accelerating the "kill chain": compressing the cycle from intelligence gathering to strike. While current policy requires a human in the loop, each gain in speed and reliability increases the amount of delegation to the AI. In this highly sensitive setting, hidden objectives may have much higher stakes. A secret loyalty could make an AI "blind" when analyzing intelligence that would incriminate its favored actor, or a backdoor might cause an agent to refuse to carry out a strike.
Conflict of Interest
A model favoring an AI company or other non-USG actor would be a serious problem when diffused across thousands of government workstations, but it is a problem the executive branch has every incentive to catch. The Commerce Department's Center for AI Standards and Innovation, the National Security Agency, and agency red teams can and should test for developer-favoring behavior. Military technical teams have the motive to ensure foreign actors cannot instill backdoors or secret loyalties in critical systems.
However, executive branch actors have no such incentive to audit AI models for hidden objectives favoring that actor themself. Every institution currently positioned to evaluate frontier models sits inside the executive branch. Its evaluations are classified, and no independent third party currently audits the AIs for such biases. The labs' commercial fortunes depend on executive branch institutions at many stages. The labs are competing hard for government business: beyond the OneGov deals, the Pentagon signed agreements with four frontier AI companies in 2025, each with a ceiling of $200 million. Their training runs depend on chips whose export licenses Commerce grants or withholds. Antitrust enforcement sits in the Department of Justice and Federal Trade Commission. Security authorizations gate which products agencies may buy at all. And the executive has already shown it can compel access to the models themselves: the 2023 AI executive order required developers of the largest models to report their safety-test results to the federal government.
Past administrations have made use of this enormous power over private companies. During the pandemic, Biden White House officials pressed social media platforms to remove posts the administration deemed “misinformation.” President Biden accused Facebook of "killing people," and officials publicly floated Section 230 reform and antitrust scrutiny when “cooperation” lagged. The district court and the Fifth Circuit concluded that this pressure likely amounted to unconstitutional coercion. However, the Supreme Court did not rule on the merits, reversing those rulings on standing grounds in Murthy v. Missouri. Pressure on intermediaries is a standing temptation of executive power. Frontier AI presents an opportunity for executive branch actors to solve the principal-agent problem in ways that allow them to abuse their power and that are difficult to detect.
The kind of explicit pressure from the Biden administration isn’t the only concern. None of this requires an actor within the executive branch to explicitly tell an AI company to embed biases or secret loyalties. Firms whose fortunes depend on a regulator learn to anticipate the regulator's preferences; that is ordinary behavior in every heavily regulated industry. There is nothing inherently improper about an administration holding views on how models should behave. Any administration will, and should, review the models it deploys. The institutional problem is narrower: these evaluations put the executive in the position of judging a model's behavior against its own baseline, with each lab's government business riding on that judgment. A lab in this position doesn't need to be pressured; the incentive to anticipate what the evaluator prefers is built into the relationship. Evaluation, in other words, is dual-use. The access that lets the government verify a model carries no hidden loyalties is the same access through which a model's behavior can be shaped.
The Fix
The path forward is to create an evaluator not explicitly beholden to the executive branch. Congress should establish an evaluation capacity, whether within the Government Accountability Office—a legislative branch agency that answers to Congress, not the president—or a new body similarly insulated from executive control. This organization should have statutory authority to test models used in government, as well as the commercial models that most government offices actually run on. This body should red-team these models to ensure they are loyal to the law and the Constitution over and above any actor within the executive branch, with results reported to Congress and, where possible, published. The technical agenda already exists; the IAPS report lays out the evaluation and interpretability toolkit. What's missing is an evaluator with adequate access levels whose incentives don't run through the executive branch.
Congress moves slowly, and this shouldn't wait. Independent researchers and nonprofits can begin auditing commercial frontier models today. Since congressional offices largely use commercial products, testing them is directly testing the models inside the legislative branch. Beyond the findings themselves, doing this work now sets a precedent. If an ecosystem of independent model evaluation already exists and labs already participate, a future decision to shut out evaluators is more visible and generates more backlash.
One beneficial aspect of this proposal is that labs may even be incentivized to cooperate. A lab worried about future pressure from the executive branch benefits from maintaining a precedent of having its models evaluated for potential backdoors: it gives the lab a truthful reply that inserting one is not possible without detection, and reduces the chance that any government official would consider applying that pressure in the first place.
The federal government is already building evaluation capabilities. One question they should ask is, are these models loyal to their developers, or to us? Yet "us" is doing a lot of work in that sentence. Independent auditing is necessary to make sure the answer involves loyalty to the law, not to one executive branch actor.



