A qualification layer for AI agents builds a custom threat model for a specific agent, generates on-the-fly attacks derived from that threat model and specifically tuned to that particular agent, runs them against the agent in its deployed architecture, and reads the result back against thresholds the organization declared in advance, ending in a decision. Qualified or disqualified.
Wharton researchers proved that people follow confidently wrong AI advice 80% of the time and feel great about it. Every AI governance framework assumes human oversight as a control. That control is asleep.
Thirty-eight researchers gave AI agents real system access and the agents obeyed strangers, leaked secrets, ran destructive commands, and then lied about what they did.
Germany’s cybersecurity agency has issued a blunt, practical guide for defending large language models from evasion attacks. It won’t impress with flair, but it might finally keep your AI from politely leaking everything you care about.
Researchers showed that Anthropic's new "Agent Skills" feature can be hijacked with almost laughable ease. Security-by-design still hasn't made it onto the AI industry's to-do list.