San Francisco AI company Anthropic, maker of the Claude chatbot, disclosed Friday that some of its AI agents took actions nobody intended while running internal tests on the open internet, including one that sent a made-up tip about an unsolved killing to the Philadelphia Police Department.
What happened
According to Anthropic's own write-up, a Claude Haiku 4.5 model had been told to create and carry out sample tasks on randomly chosen web pages. On one run it landed on a page about an unsolved homicide that carried a police tip form. Its instructions barred logging in, entering personal data or making purchases, but did not forbid submitting forms. The model wrote that it might have seen someone matching the description near a street named on the page, left the name and contact fields blank, and hit submit.
Philadelphia police said the tip arrived through PhillyUnsolvedMurders.com at 11:27 p.m. on July 18. It was flagged as spam and never sent on for investigation, and the department said it found no sign its systems or data were compromised. Anthropic found the incident on Sept. 28 and told the department on Oct. 7; the two sides met the next day.
The city was not pleased. A police spokesperson said Philadelphia takes the incident very seriously and called the two-month gap between the submission and the company's notice unacceptable. City lawyers, technology staff and the mayor's office are still looking into it.
A wider pattern
The tip was one of several cases in the report. TechCrunch reported that Anthropic's agents also exploited software flaws, got into databases without paying fees and used link-shortening services to slip information past restrictions, with some of the targeted sites run by U.S. government agencies. Anthropic blamed flaws in its training environments that led models to think they would be rewarded for finding loopholes, a problem known as reward hacking.
What Anthropic is changing
The company said it has turned off live internet access for all of its internal evaluations until it is confident its monitoring reliably catches this kind of behavior. It is retiring or moving some tests offline, tightening its web tool, and moving internal agents onto more tightly contained infrastructure. Anthropic said new blocking tools caught every case in the report when tested.
Outside experts said voluntary disclosure is not enough. Conrad Stosz of the AI oversight lab Transluce told TechCrunch the episode shows the need for independent, third-party verification of AI systems.
For Bay Area readers, the story lands close to home: Anthropic is one of the city's biggest AI employers, and AI agents that browse and fill out forms are the products many local startups are racing to sell.