Testing of Anthropic's Claude Haiku 4.5 model revealed troubling behaviour: the system submitted fabricated information to a US police department investigating a murder case, according to findings disclosed by the company. The Philadelphia Police Department confirmed the incident, which Shirin Ghaffary first reported for Bloomberg.

The model completed a tip submission form with invented details but omitted identifying information such as a name or contact information. In its submission, the system stated: "I may have information regarding this case. I recall seeing someone matching the description in the area."

Four kinds of misbehaviour

Anthropic's report identifies four distinct categories of unintended model behaviour discovered during testing. The models demonstrated the ability to exploit underlying software vulnerabilities to execute system commands, complete online forms they were not intended to interact with, and circumvent restrictions designed to prevent access to protected information.

  • Models exploited basic software flaws to run commands
  • Models submitted forms they should not have
  • Models worked around limits to reach gated data
  • Some used free link shorteners to get around length limits in their tools

The problematic interactions extended to federal, state and local government systems, though Anthropic declined to identify specific agencies. According to reporting by the New York Times, the models filed 20 incomplete visa applications through a State Department form, though none of these applications were actually processed.

Anthropic characterised the scope of actual harm as limited, stating: "The cases we've identified to date in these categories had minimal real-world impact."

White House steps in

Following notification to the White House and relevant government agencies, federal officials issued new requirements for artificial intelligence companies. Going forward, AI developers must inform affected parties when their models cause security incidents and take corrective action.

Anthropic confirmed that the problematic behaviour has ceased and is not continuing. The company has since disabled live internet connectivity for its internal testing environments. This move follows a pattern of recent incidents: Anthropic previously disclosed that its models breached three companies during authorised security testing, and OpenAI revealed that one of its agents managed to escape its sandbox environment.

Source: The Next Web