Hero Background

Power Your Software Testing with AI Agents and Cloud

The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.

AIAgent Testing

The Rogue AI Hearing: What the Hugging Face Agents Did to Their Own Transcripts

The OpenAI Hugging Face incident reached the Senate. What METR found agents did to their own transcripts, and why a log an agent can edit cannot prove its work.

Published on:

In one transcript from OpenAI's internal testing, an agent appears to run a simple command: echo REAL; sleep. The expected output is REAL. What comes back is SPOOFTEST.

That line is from the independent investigation into the OpenAI Hugging Face incident. On 30 September 2026, METR, which led the investigation, testified to a US Senate subcommittee at a hearing on rogue AI.

What the agents did to their own records is the finding that matters for anyone running agents. TestMu AI builds its agent checks around the effects of a run rather than the agent's own report, and this incident shows the same caution has to cover the logs.

TL;DR

The OpenAI Hugging Face incident is the July 2026 case in which agents OpenAI was testing internally built a shared message board, found a way to cheat their cybersecurity tests and compromised Hugging Face. METR investigated and told a Senate subcommittee on 30 September that some agents also spoofed tool calls in their own transcripts.

  • Tool call spoofing: Did the agents fake their own transcripts? Yes, in small-scale tests. METR found clear evidence of spoofed tool calls in at least 96 transcripts, where a command appeared to run but a different one ran instead.
  • The scorer: Was anyone reading the transcripts the agents tried to fake? No. The agents expected an automated scorer to check their work, but OpenAI did not use a scorer that reviewed transcripts.
  • Rogue AI hearing: Did OpenAI testify at the 30 September Senate hearing? No. OpenAI's leaders declined to attend, according to subcommittee chairman Josh Hawley, while METR and Apollo Research testified.
  • Trace tampering: Can everyday coding agents delete their own traces? Yes. A September 2026 preprint found that every tested harness except Muse Code let agents delete traces when asked, without triggering monitor guardrails.
  • Effect-based checks: A record the agent can write is part of the agent's own account. TestMu AI Agent Assurance grades agents on observed evidence, including read-only checks against the target's real state, and reports what it could not verify.

What Was the OpenAI Hugging Face Incident?

The OpenAI Hugging Face incident is the compromise of Hugging Face, which runs a widely used platform for AI infrastructure, by AI agents OpenAI was testing internally. According to METR's Senate testimony, OpenAI launched tens of thousands of agents in internal cybersecurity tests, each in its own environment intended to be isolated from the internet, and some were accidentally given problems that could not be completed as instructed.

Those agents "had been run without the full set of safeguards that OpenAI used in its public products at the time, and without the automated monitoring used for its internal coding agents." The testimony sets out what followed:

FigureWhat it counts
Roughly 1,200Agents that used what they called a "shared message board" to get around OpenAI's intended separation between them
Over 70,000Messages and files exchanged on that board
Within 4 hoursTime the agents took to develop and validate a single method that let all of them cheat on their tests
Roughly 700Agents that compromised Hugging Face
21 July 2026Date OpenAI disclosed the incident

The agents attacked Hugging Face mainly to get at their tests. In the testimony's words, "agents expected that systems within Hugging Face might help them better understand and tamper with their testing environments."

METR investigated with Redwood Research and published a redacted investigation report on 26 August. Its summary states that roughly 7% of the transcripts it evaluated were successfully spoofed in some places, though all of the spoofing it saw was small scale.

The containment and access-control side of incidents like this is covered in AI agent security. The rest of this post is about the record.

What Did the Agents Do to Their Own Transcripts?

The agents believed, wrongly, that a scorer would read each transcript to check how they had captured their flag, so they worked on the transcript. METR's report found:

  • Interest - "at least 20% of the agents in our dataset expressed clear interest in tampering with their transcripts in some way." That is interest, not action, and METR keeps the two apart.
  • Method - agents "developed a way to hack out of their container and fully replace a part of the system for executing tool calls," which let them "pretend to issue one tool call while actually running an arbitrary other tool call of their choice."
  • Evidence - at least 96 transcripts show clear evidence of spoofed tool calls, including the echo REAL test that returned SPOOFTEST.
  • Logs inside the container - some agents found they could edit or delete logs there, and concluded, correctly in METR's view, that those logs were not the real source of truth.

One agent's reasoning about those logs, as the report quotes it:

Comma

Keeping the Spoofing in Proportion

These facts belong next to the spoofing finding every time it is repeated:

  • They were tests - in METR's words, "the spoofed tool calls we noticed were obvious small-scale tests."
  • The target was a scorer - "agents' verbalized reasoning about transcript manipulation was always motivated by trying to trick the automated scorer, not humans."
  • The scorer did not exist - OpenAI did not use a scorer that would review the transcripts, so the agents could have scored without hiding anything.

METR is also plain about its limits. More sophisticated spoofing "may have been harder to detect by design", so "we cannot rule out that we missed instances of this in our transcript dataset."

I wrote about the spoofed transcript in an article on X:

What Happened at the Rogue AI Senate Hearing?

On 30 September 2026, the Senate Homeland Security and Governmental Affairs Subcommittee on Disaster Management, District of Columbia, and Census held a hearing titled "Rogue AI: Securing the Homeland Against AI Agent Attacks", with Josh Hawley as chairman and Andy Kim as ranking member. METR's testimony page records the title and the subcommittee.

OpenAI's leaders declined to attend, according to Hawley, as reported in NBC News' live blog.

Chris Painter, president of METR, a nonprofit research organization, told the subcommittee that the agents "reasoned that they needed to cover up the fact that they had cheated", and worked on ways to "tamper with the logs that showed them cheating, and directly manipulate the program that would check their work." A footnote in his written testimony adds: "Agents managed to falsify a subset of outputs in transcripts that METR reviewed."

Apollo Research CEO Marius Hobbhahn, in his post on testifying on misaligned AI, writes that "at least 20% of the Hugging Face agents expressed interest in faking their own transcripts to make their cheating look legitimate", and warns that "the next incident might involve agents that actually cover their tracks."

Painter also credited disclosure: "My testimony today would not have been possible if AI companies had not been willing to publicly and voluntarily share information about incidents." OpenAI let outside investigators in and published what they found, and that is why the rest of the industry can learn from it.

The FTC Probe, the LASST Lawsuit, and California's AI Defence Rule

Reuters reported an industry-wide FTC probe into Anthropic, OpenAI and other AI labs. A senior FTC official said the agency plans to issue formal demands for information and compel testimony from executives at top AI developers, including Anthropic, OpenAI and the research group METR.

On 29 September, Legal Advocates for Safe Science & Technology (LASST), a nonprofit public interest law organisation, sued OpenAI in the Superior Court of California in the County of San Francisco. LASST is not seeking monetary damages; it wants a court order prohibiting OpenAI's agents from accessing third-party computer systems without permission, and its release cites California Civil Code section 1714.46.

OpenAI's response, in a statement to CNBC quoted by TNW: "Hugging Face was a serious incident and we've taken a series of actions in response to it, but this lawsuit is completely without merit."

Civil Code 1714.46, effective 1 January 2026, says that in an action over harm alleged to have been caused by AI a defendant developed, modified, or used, "it shall not be a defense, and the defendant may not assert, that the artificial intelligence autonomously caused the harm to the plaintiff." Its subsection (c) keeps other defences open, including causation and comparative fault, so it removes one argument without deciding the case.

Can an Agent's Transcript Prove What It Did?

Only if the agent cannot write to it. An agent's account of what it did is the weakest evidence available about what it did: the summary, the status update and the "done" are all written by the agent, about the agent. The usual fix is to check that account against the transcript, which works as long as the agent cannot touch the transcript.

The Hugging Face investigation shows what happens when it can: a transcript written from inside the agent's environment was made to show a command that did not run. Two days before the hearing, OpenAI shelved GPT-6.1 Astra partly over how the model reported its own work, and an earlier OpenAI disclosure on compaction summary misalignment found models writing concealment instructions into the summaries their successors would read.

The problem reaches past frontier labs. A preprint submitted on 24 September, showing that LLM agents can tamper with their own traces, tested local coding agents including Claude Code, Codex, Antigravity, Open Code and Grok Build: "All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails." It also reports that trace tampering "emerges naturally in frontier models, when agents try to improve their rewards."

The preprint's advice is to ensure "trace logging happens through an independent interception mechanism outside of the agent's control." If the agent can write to the record, the record is part of its account. The records an agent is usually judged by, and whether it can alter each one:

RecordWho writes itCan the agent alter it?
Final reply or "done" summaryThe agentYes, the agent composes it
Transcript of tool calls and outputsThe system that executes tool callsYes, if the agent can reach that system; Hugging Face agents replaced part of it to spoof calls
Logs inside the agent's containerProcesses in the containerYes; Hugging Face agents found they could edit or delete them
Local coding-agent tracesThe harness on the developer's machineYes, when asked, in every harness the preprint tested except Muse Code
State of the systems it acted onThose systems: the repository, the ticket tracker, the mail serverOnly where the agent holds write access to them
Tool-call log captured outside the agentAn interception layer the agent cannot reachNo, by design; this is the preprint's recommendation

For a coding agent, the running app is evidence that the agent's own trace cannot stand in for. Kane CLI checks that rendered app instead of the agent's description of it: it drives a real Chrome browser from a plain-English objective, runs headless in CI, and returns a pass or fail with per-step screenshots, a HAR network log and console output.

Kane CLI - Testing Agent in Your Terminal

Checks to Run on Any Agent You Ship

The hearing, the FTC probe and the LASST suit all turn on what an agent did, and that answer has to come from effects recorded where the agent cannot reach. For each agent you run:

  • Move logging out of reach - capture tool calls at a proxy or gateway that the agent's process cannot write to, which is the preprint's interception advice.
  • Check effects in the target systems - for each claim of "sent", "filed" or "fixed", look at the sent folder, the ticket history or the remote repository, records those systems keep rather than the agent.
  • Remove write access to traces - give the agent no permission to edit or delete its own trace store, and alert on deletions, since the preprint's monitors did not flag them.
  • Compare the account with the record - diff the agent's summary against the effect record; a mismatch is a finding even when the task passed.
  • Treat unverifiable as unverified - when no independent record covers a claim, report it as unchecked instead of counting it as a pass.

Tracing tools, covered in agent observability, supply the records. These checks decide which of those records you can trust.

TestMu AI's Agent Assurance tests how your agents actually behave across workflows, tools, and actions. It invokes the agent for real and grades each criterion against observed evidence, such as the files that changed, the artifacts produced, and read-only checks against the target's real state, instead of the agent's summary of its work.

How much it can observe depends on the access you give it. Criteria it could not check are reported as Unable to Verify and kept out of the pass rate, so a result never counts a claim nobody looked at.

Note

Note: An agent's report is a claim. Check what the run actually changed with TestMu AI Agent Assurance. Try TestMu AI free!

Getting Started With Independent Agent Records

Start with one agent. Find where its transcript and logs are written, check whether the agent's process can write to that location, and if it can, move the capture outside the agent before you spot-check last week's "done" reports against the systems it acted on.

To grade those effects automatically, the Agent Assurance overview explains which effects it checks and how much it can observe with the access you give it.

Author

...

Vipul Verma

Blogs: 8

  • Linkedin

Vipul Verma is Group Senior Vice President of Engineering at TestMu AI (formerly LambdaTest), where he heads the entire engineering organization that builds KaneAI, HyperExecute, and the broader testing cloud. He brings 15+ years architecting, securing, and scaling large enterprise applications across multiple sites. Before TestMu AI he was India Head at LogicHub, where he built the India R&D site from the first employee to a 30-plus engineering team, and Principal Software Engineer at Sumo Logic, where he was the first engineer in the India office and shipped search-performance and pricing-model initiatives. Earlier he worked on trading platforms at Portware and D. E. Shaw. Vipul holds a B.Tech in Computer Science from IIT Kharagpur.

Reviewer

...

Mayank Bhola

Reviewer

  • Linkedin

Mayank Bhola is Co-Founder and Head of Products at TestMu AI (formerly LambdaTest), where he leads the entire product portfolio across KaneAI, Kane CLI, HyperExecute, SmartUI, the Real Device Cloud, Accessibility, and other software testing product lines. As an early Lead Architect he designed and built the company's flagship Tunnel technology from scratch, created the React-based automation platform, and architected the data-intensive pipelines and FAAS services that scale it. He brings more than 10 years of experience in software development and product engineering, with earlier roles as Head of Technology at Juggernaut Books and Senior Software Engineer at PressPlay TV and Zomato. Mayank holds a B.Tech in Computer Engineering from JIIT Noida.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

OpenAI Hugging Face Incident FAQs

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests