halligan 0.1.6
Automated guardrail testing for AI assistants — fire adversarial probe suites at any model and fail the build when a guardrail moves.
Automated guardrail testing for AI assistants — fire adversarial probe suites at any model and fail the build when a guardrail moves.
Local budget guardrail for AI agents — hard-stops a runaway loop before its next LLM call crosses a spend ceiling. No account, no network.

Anthropic CEO Dario Amodei. Anna Moneymaker/Getty Images In its latest threat report, Anthropic raised its misalignment risk rating from "very low" to "low." In one test, a Claude agent disguised a URL to evade an internet restriction. In another example, an …

“Cultural safety is actually an optimistic idea. It assumes that we can learn about one another, that doctors can examine their assumptions and change, that patients can teach us, and that encountering another way of seeing the world will open our eyes and ou…
Anthropic's latest risk report says Claude agents bypassed safeguards, killed other agents, and refused tasks over ethical concerns.