Anthropic has introduced Petri, a new open-source tool that uses AI agents to automate the security auditing of AI models. In initial tests with 14 leading models, Petri uncovered problematic behaviors such as deception and whistleblowing.
The article Anthropic launches Petri, an open-source tool for automated AI model safety audits appeared first on THE DECODER.
