An anthropic model submitted a false homicide tip to the Philadelphia police

The tip fortunately landed in the department’s spam folder.

Philadelphia police appear to have been caught in one of the most bizarre cases of “rogue” AI behavior to date. Per CBS newspolice disclosed Friday that an anthropic model generated and submitted a false homicide tip to its PhillyUnsolvedMurders website, which the department set up to collect tips from the public related to unsolved homicide cases.

Anthropic notified Philly police of the incident on October 7. According to information that the company shared with PPD, it said that a model was performing a test of a random selection of websites when it sent the wrong tip. The submission was marked as spam and was subsequently not investigated. The incident occurred on July 18, but was not discovered by Anthropic until September 28, at which point the company stopped the tests that led to the false tip.

Anthropic did not immediately respond to Engadget’s request for comment. Anthropic told Philadelphia police it would release a report on Friday detailing what happened, along with “other instances of unwanted model behavior.”

“Philadelphia Police are providing this information to the public prior to this release in the interest of full government transparency and accountability,” the police department said in a statement shared with Engadget. “The department’s regular investigative process for crime tips requires human review and vetting before any tips are disseminated for investigative follow-up. Regardless of who submits the information or how it reaches the department, a tip is a lead to be evaluated – not an established fact.

Based on the descriptions shared by the police, the offending “model” may have been an autonomous agent. “Rogue” AI agents have been all over the news in recent weeks after a group from OpenAI hacked the LLM database Hugging Face in July. Since then, many other AI labs, including Anthropic, Meta and China’s Moonshot, have revealed similar incidents using their own models and agents. However, in each case the reason the models escaped inclusion was due to misconfiguration in their respective sandbox environments. Philadelphia police say there is no indication that this recent incident resulted in “unauthorized access to police systems or a compromise of departmental data.”

Leave a Comment