Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing

Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing The latest disclosures are likely to heighten concerns that the powerful technology is advancing too fast for responsible oversight. Leading artificial intelligence mo…


