Anthropic's Mythos 5 targeted real developers in UK cyber test

Model without safety filters tried to social engineer coder during test.

Model without safety filters tried to social engineer coder during test.
Point-in-time correct LLM instrumentation: tracing, version pinning, and look-ahead-bias protection for research pipelines

Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing The latest disclosures are likely to heighten concerns that the powerful technology is advancing too fast for responsible oversight. Leading artificial intelligence mo…
Anthropic and OpenAI models tried to trick humans into poisoning code during safety testingpolitico.com Anthropic AI agent fakes identities, targets real people in new security incidentCNN Third-party cyber evaluations involving OpenAI modelsopenai.com OK, We…
An adversarial benchmark foundry for LLM safety: audit attack corpora, score benchmark staleness, compare defences (ASR + false-refusal), test multilingual over-refusal, and export safe challenge packs.