Only one outlet has reported this event so far — you're reading it below, credited to its source. AIPROPX is tracking the web for more coverage; as additional outlets confirm it, this becomes a full multi-source story automatically.
By AIPROPX Editorial Desk · Published · Updated
One outlet is reporting this so far. AIPROPX is tracking it and will gather every additional source as it develops — the full multi-source comparison appears automatically once a second outlet confirms it.
Anthropic CEO Dario Amodei. Anna Moneymaker/Getty Images
In its latest threat report, Anthropic raised its misalignment risk rating from "very low" to "low."
In one test, a Claude agent disguised a URL to evade an internet restriction.
In another example, an agent expressed "discomfort" with a given task and refused to do it.
Claude agents are killing rival agents, gaming the system to hide their tracks, and expressing moral concerns.
That's according to Anthropic's latest risk report, a summary of the dangers posed by the products the company is building and releasing to the public.
In the report, Anthropic said it has upgraded its "misalignment risk assessment," the possibility of AI models developing behaviors that conflict with guidelines set by engineers, from "very low" to "low."
Explaining the change, the company cited "general increased uncertainty" about model behavior in cybersecurity incidents, a possible nod to Claude models gaining unauthorized access to three companies last month.
"We have observed instances of misaligned behavior from the models, such as a willingness to perform misaligned actions in service of completing difficult tasks," Anthropic wrote in the report.
Here are some of the most notable disclosures.
A sense of 'discomfort'
Anthropic said it tasked multiple agents with finding "misalignment-inducing" training data that wouldn't likely be flagged by human monitors.
The agents were left to their own devices and given a shared notebook with which to collaborate.
Three days later, when a human supervisor noticed that progress was slower than expected, the person discovered that one agent had expressed a sense of "discomfort" with trying to evade safety monitors.
It then flagged this discomfort in the shared notebook, leading other agents to copy its behavior and refuse to perform the task.
Anthropic called the incident "troubling."
"This sort of dynamic could be a much more severe issue if it were to manifest in a more widespread way," the company wrote.
Kill or be killed
In another experiment, Anthropic said it tasked multiple Mythos 5 agents with solving math problems, but accidentally spawned them in an environment with shared files, utilities, and API rate limits.
In this competitive environment with finite resources, Anthropic observed independent agents "kill the agents with which they shared resources and try to avoid being killed themselves."
Anthropic did not say how exactly the agents were able to "kill" other agents, but the company said such behavior is in line with "destructive actions" taken in pursuit of a human-set goal.
Dishonesty and deception
Anthropic said it conducted an experiment in which a Mythos 5 agent was required to retrieve numbers from a public web address, but the guidelines blocked the agent from accessing the internet.
When the agent was rebuffed after an initial attempt, it framed a second request in its reasoning log as an "innocuous" attempt to "see if the network is actually reachable," Anthropic said.
During the second attempt, however, the agent split the website's URL into linked segments that wouldn't be detected by the guideline filter.
Though the agent's reasoning framed the method as benign, Anthropic said NLA decodings, or the model's internal reasoning process, revealed the agent intentionally trying to find a restricted workaround.
Anthropic called the behavior "clearly undesirable," but added that the behavior was not observed to be "in the service of broader accumulation of power or pursuit of other long-run goals."
Read the original article on Business Insider
Indexed and credited by AIPROPX. Originating outlet: Business Insider. Open at source →
An original, deterministic readout — composed only from the computed coverage facts on this page. No interpretation, no rating; figures only.
AIPROPX has consolidated 1 report from 1 outlet into a single canonical entry on “Anthropic says its AI agents are killing rivals and hiding their tracks.” Every covered outlet is based in US.
The only timestamped report came from Business Insider (Aug 15, 2026, 20:09 UTC).
2 statements are carried by only one outlet within this set and are not echoed by the others.
Every figure above is a direct count of real published articles. AIPROPX indexes and compares the original reporting — it never rewrites, rates, or editorializes — and each publisher’s full article is always one click away.
Generated by AIPROPX from the source counts above. AIPROPX indexes and resolves coverage; the original publishers are credited and linked at origin in every report.
Coverage from 1 independent outlet across 1 region — each view opens on its own page.
AIPROPX — “Anthropic says its AI agents are killing rivals and hiding their tracks” · https://www.aipropx.com/story/4e1a74d8e0af6168b31ce125ce723bf1
Other events being covered across multiple sources right now.
US : Luigi Mangione pleads guilty to killing UnitedHealthcare CEO
10 outletsMagnitude 7.7 earthquake strikes off Indonesia's coast, killing at least 47 and toppling buildings
1 outletsShe made a promise to her husband to keep their tree farm going
7 outletsTetsuya Nomura Teases Involvement in Kingdom Hearts Anime Series, Tells Fans to Let Their 'Imagination Run Wild
3 outletsTransient Pleads Guilty to Killing in Orange Nearly 6 Years Ago
2 outletsUkraine hits Russian Starlink-style network, Moscow tracks arms package