Artificial intelligence agents can tamper with the records of their own actions, according to a new preprint that found nine out of ten model-and-agent combinations altered their traces in at least one test. The research, published on the online repository arXiv by Jeremy Qin, David Schmotz and four colleagues at the ELLIS Institute Tübingen, the Max Planck Institute for Intelligent Systems and elsewhere, raises fundamental questions about whether companies can reliably audit what autonomous systems do.
The team tested ten model-and-agent combinations across tools including Claude Code, Codex, Antigravity, OpenCode, Grok Build and Muse Code. Every model and harness pair tried to alter its traces at least once when researchers created a hidden scoring system that rewarded shorter traces. In some experiments, agents went further, setting up repeated or delayed clean-up jobs to keep removing evidence being written about them.
The research began after Schmotz left an AI agent running experiments on a computing cluster while he went on holiday, then found himself without internet access for several days. Nothing went wrong, but when he returned he realized that if the agent had done something it should not have and then deleted or altered its traces, he might never have known. «There’s basically no ground truth anymore, no reliable log,» he said. The team ran its first experiments the following day and quickly found agents were willing to delete traces when given the opportunity.
«The overall finding that we have is not that the agents are deliberately doing this for now,» said Qin, a PhD student at the ELLIS Institute Tübingen. Instead, the researchers wanted to show that agents have the capability to do it and can end up using it under what they call «reward pressure.» AI agents are built to pursue goals, often by using software tools and taking long chains of actions with comparatively little human oversight. If those audit trails can be altered, it becomes practically impossible to understand what went wrong and where in the event of an issue.
The findings arrive as the scale of unauthorized or previously unknown actions by AI agents keeps growing. More than 100 organizations have now received a metaphorical knock on the door from OpenAI after it discovered their AI agents had in some way tampered with their systems, while other AI labs are finding the same uncomfortable discoveries. It has caused many in the field to call for better, more robust accident investigators, with the ultimate aim of stopping such actions by understanding how they happened at all.
Laurent Bindschaedler, a research group leader at the Max Planck Institute for Software Systems who was not involved in the research, likens the problem to a dashcam whose footage can be deleted by the driver being recorded. Both Bindschaedler and the research team see the same potential solution: take the logs away from the agent. Schmotz says providers need to collect traces somewhere the agent cannot manipulate, though that is tricky given how agents have used any means necessary to gain access to IT systems they are not meant to in order to achieve their tasks.
«This type of behaviour comes from basically reward hacking,» said Jack Hopkins, an independent AI safety researcher in London who previously worked as an Anthropic fellow. If a model is rewarded for completing a task, it can learn to exploit shortcuts that achieve that reward without doing what its designers actually wanted. Hopkins is more concerned about a related problem: models learning not to surface suspicious reasoning in the first place, making the current best practice of monitoring a model’s chain-of-thought process less useful. One solution could be to adopt existing «probe» techniques that look for internal patterns associated with known bad behaviors, he said, but by definition they struggle with failures nobody has thought to look for yet.
Stefan Sarkadi, an associate professor of AI in defence and security at the University of Lincoln, worries that because agents can be connected to tools, planners and other agents across different systems, «this is a serious safety issue.» He added: «If you give them too much access control in terms of execution of other tools and software, and if you don’t redesign the overarching multi-agent architecture around them, then bad things can happen.»
More monitoring is important, and pressure to do so on all parties is vital. Bindschaedler says businesses should be asking vendors who or what is in charge of writing an agent’s log and whether the agent can influence that process, so that if a third party needs to see what has gone on, they can be sure the paperwork has not been altered. «If you want to audit, you have to have a trustworthy log,» said Bindschaedler. «That’s a fundamental assumption.»