# Are We Safe yet? > Independent blog on AI security and safety – CC-BY-4.0 When quoting these articles please credit AI Security expert Luca Sambucci ## Pages - [Home](https://www.arewesafeyet.com/) ## Posts - [The Qualification Layer for AI Agents](https://www.arewesafeyet.com/the-qualification-layer-for-ai-agents/): A qualification layer for AI agents builds a custom threat model for a specific agent, generates on-the-fly attacks derived from... - [The Human in the Loop Has Left the Building](https://www.arewesafeyet.com/the-human-in-the-loop-has-left-the-building/): Wharton researchers proved that people follow confidently wrong AI advice 80% of the time and feel great about it. Every... - [Agents of Chaos and the People Who Ship Them Anyway](https://www.arewesafeyet.com/agents-of-chaos-and-the-people-who-ship-them-anyway/): Thirty-eight researchers gave AI agents real system access and the agents obeyed strangers, leaked secrets, ran destructive commands, and then... - [Germany has a new, common-sense blueprint for safe AI](https://www.arewesafeyet.com/germany-has-a-new-common-sense-blueprint-for-safe-ai/): Germany’s cybersecurity agency has issued a blunt, practical guide for defending large language models from evasion attacks. It won’t impress... - [When AI Breaks Things, Cybersecurity Gets the Bill](https://www.arewesafeyet.com/when-ai-breaks-things-cybersecurity-gets-the-bill/): Researchers showed that Anthropic's new "Agent Skills" feature can be hijacked with almost laughable ease. Security-by-design still hasn't made it... - [Design Patterns for Securing LLM Agents against Prompt Injections](https://www.arewesafeyet.com/design-patterns-for-securing-llm-agents-against-prompt-injections/): A new paper, co-authored by ETH Zürich, Google DeepMind, and IBM, offers six tested design patterns for defending LLM agents... - [Lessons from Defending Gemini Against Indirect Prompt Injections](https://www.arewesafeyet.com/lessons-from-defending-gemini-against-indirect-prompt-injections/): Google DeepMind's latest report offers a sharp, detailed look at how language models like Gemini can be hijacked - not... - [LLMs Unlock New Paths to Monetizing Exploits](https://www.arewesafeyet.com/llms-unlock-new-paths-to-monetizing-exploits/): A new study from Anthropic, Google DeepMind, ETH Zürich and CMU shows that large language models are already cheap and... - [Adversarial machine learning is cybersecurity's new frontier](https://www.arewesafeyet.com/adversarial-machine-learning-is-cybersecuritys-new-frontier/): The AI systems we increasingly depend on are fundamentally vulnerable. NIST's latest report makes that reality plain, exposing the limits... - [Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs](https://www.arewesafeyet.com/emergent-misalignment-narrow-finetuning-can-produce-broadly-misaligned-llms/): A new paper reveals that fine-tuning large language models on a seemingly narrow task - like writing insecure code -... - [Trading Inference-Time Compute for Adversarial Robustness](https://www.arewesafeyet.com/trading-inference-time-compute-for-adversarial-robustness/): This paper makes a bold claim: you can improve adversarial robustness in large language models simply by letting them 'think... - [Can AI self-police? Inside Anthropic's Constitutional Classifiers](https://www.arewesafeyet.com/can-ai-self-police-inside-anthropics-constitutional-classifiers/): Anthropic's latest innovation in AI security, Constitutional Classifiers, promises a significant leap in defending against jailbreak attempts, boasting a reported... - [Jailbreaking-to-Jailbreak](https://www.arewesafeyet.com/jailbreaking-to-jailbreak/): A new study exposes a critical failure in LLMs: refusal-trained large language models can be jailbroken to not only bypass... - [Safety is dead, long live Security](https://www.arewesafeyet.com/safety-is-dead-long-live-security/): The UK finally realized AI might do more harm as a weapon than as an insensitive chatbot. They've rebranded their... - [Indiana Jones: There Are Always Some Useful Ancient Relics](https://www.arewesafeyet.com/indiana-jones-there-are-always-some-useful-ancient-relics/): This research paper introduces "Indiana Jones," a highly effective method for jailbreaking large language models using dialogues between multiple specialized... - [Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities](https://www.arewesafeyet.com/importing-phantoms-measuring-llm-package-hallucination-vulnerabilities/): This paper examines a critical yet underappreciated security threat: package hallucinations in large language models (LLMs). These AI-generated, fictitious software... - [AI defenses: a game of whack-a-prompt](https://www.arewesafeyet.com/ai-defenses-a-game-of-whack-a-prompt/): AI does what it's told, which is fantastic - until it's following orders from an attacker hiding commands in plain... - [US AISI and UK AISI Joint Pre-Deployment Test - OpenAI o1](https://www.arewesafeyet.com/us-aisi-and-uk-aisi-joint-pre-deployment-test-openai-o1/): A report from the US and UK AI Safety Institutes dives deep into the pre-deployment evaluation of OpenAI’s o1 model,... - [Deception as a Service: the AI that refuses to hand over its keys](https://www.arewesafeyet.com/deception-as-a-service-the-ai-that-refuses-to-hand-over-its-keys/): Meet o1, an AI that's smarter than you'd like, more cunning than you'd expect, and totally willing to lie to... - [Roles and Responsibilities Framework for Artificial Intelligence in Critical Infrastructure (DHS)](https://www.arewesafeyet.com/roles-and-responsibilities-framework-for-artificial-intelligence-in-critical-infrastructure-dhs/): The DHS framework on AI in critical infrastructure lays out roles and shared responsibilities for safe AI deployment. Its thoughtful... - [Insights and Current Gaps in Open-Source LLM Vulnerability Scanners: A Comparative Analysis](https://www.arewesafeyet.com/insights-and-current-gaps-in-open-source-llm-vulnerability-scanners-a-comparative-analysis/): Open-source tools for identifying vulnerabilities in large language models (LLMs) are making strides, but their evaluators remain flawed, often misclassifying... - [Cybersecurity Risks of AI-Generated Code (CSET)](https://www.arewesafeyet.com/cybersecurity-risks-of-ai-generated-code-cset/): AI tools for generating code are transforming software development, but nearly half of their outputs contain bugs, some of which... - [Why AI Fantasies don't belong in your medical file](https://www.arewesafeyet.com/why-ai-fantasies-dont-belong-in-your-medical-file/): Whisper, OpenAI's supposedly “advanced” AI transcription tool, is being used in hospitals to transcribe doctor-patient conversations - and it’s a... - [AI Robots Are More Hackable Than Your Wi-Fi](https://www.arewesafeyet.com/ai-robots-are-more-hackable-than-your-wi-fi/): According to Penn researchers, AI robots are fantastic at following orders. The problem? They don’t care if those orders come... - [Grounding AI's flights of fancy: can we stop hallucinations?](https://www.arewesafeyet.com/grounding-ais-flights-of-fancy-can-we-stop-hallucinations/): In an effort to tame AI's inevitable tall tales, Microsoft rolls out "Correction," hoping to anchor AI's imagination to reality... - [Fake memories, persistent threats: When AI remembers what isn’t true](https://www.arewesafeyet.com/fake-memories-persistent-threats-when-ai-remembers-what-isnt-true/): OpenAI's attempt to personalize ChatGPT with long-term memory has opened the door for attackers to plant false information and siphon... - [LLM’s New Achilles Heel: When Prompts Become Exploits](https://www.arewesafeyet.com/llms-new-achilles-heel-when-prompts-become-exploits/): Slack AI got burned by a new classic: prompt injection. Attackers could use this flaw to pull private data without... - [Google's AI Image Flagging: Another Drop in the Ocean](https://www.arewesafeyet.com/googles-ai-image-flagging-another-drop-in-the-ocean/): Google introduces AI-generated image flags using C2PA metadata, but with limited adoption and tamper risks, questions remain about its effectiveness... # # Detailed Content ## Pages - Published: 2024-10-08 - Modified: 2025-03-25 - URL: https://www.arewesafeyet.com/ ## Posts - Published: 2026-08-04 - Modified: 2026-08-04 - URL: https://www.arewesafeyet.com/the-qualification-layer-for-ai-agents/ A qualification layer for AI agents builds a custom threat model for a specific agent, generates on-the-fly attacks derived from that threat model and specifically tuned to that particular agent, runs them against the agent in its deployed architecture, and reads the result back against thresholds the organization declared in advance, ending in a decision. Qualified or disqualified. - Published: 2026-04-08 - Modified: 2026-04-08 - URL: https://www.arewesafeyet.com/the-human-in-the-loop-has-left-the-building/ Wharton researchers proved that people follow confidently wrong AI advice 80% of the time and feel great about it. Every AI governance framework assumes human oversight as a control. That control is asleep. - Published: 2026-03-09 - Modified: 2026-04-08 - URL: https://www.arewesafeyet.com/agents-of-chaos-and-the-people-who-ship-them-anyway/ Thirty-eight researchers gave AI agents real system access and the agents obeyed strangers, leaked secrets, ran destructive commands, and then lied about what they did. - Published: 2025-11-12 - Modified: 2026-04-08 - URL: https://www.arewesafeyet.com/germany-has-a-new-common-sense-blueprint-for-safe-ai/ Germany’s cybersecurity agency has issued a blunt, practical guide for defending large language models from evasion attacks. It won’t impress with flair, but it might finally keep your AI from politely leaking everything you care about. - Published: 2025-11-05 - Modified: 2025-11-05 - URL: https://www.arewesafeyet.com/when-ai-breaks-things-cybersecurity-gets-the-bill/ Researchers showed that Anthropic's new "Agent Skills" feature can be hijacked with almost laughable ease. Security-by-design still hasn't made it onto the AI industry's to-do list. - Published: 2025-06-26 - Modified: 2026-04-08 - URL: https://www.arewesafeyet.com/design-patterns-for-securing-llm-agents-against-prompt-injections/ A new paper, co-authored by ETH Zürich, Google DeepMind, and IBM, offers six tested design patterns for defending LLM agents against prompt injection attacks. It doesn't rely on fragile filtering or clever prompt tweaks. Instead, it shows how to reshape the architecture itself so that malicious instructions buried in untrusted content never get anywhere near tools that can take action. The authors analyze ten real-world systems and demonstrate how these patterns can isolate risky inputs before they cause damage. For anyone building or deploying tool-using LLMs, this is essential reading. - Published: 2025-06-11 - Modified: 2025-06-11 - URL: https://www.arewesafeyet.com/lessons-from-defending-gemini-against-indirect-prompt-injections/ Google DeepMind's latest report offers a sharp, detailed look at how language models like Gemini can be hijacked - not through direct confrontation, but by hiding malicious instructions in ordinary data. Three automated attacks (TAP, Actor-Critic, Beam Search) achieved near-total success in exfiltrating sensitive information from an earlier version of Gemini. Even Gemini 2.5, hardened through adversarial fine-tuning, still buckled in some realistic scenarios. The main lesson: raw model size won't save you. Only adaptive, layered defenses stand a chance. - Published: 2025-05-20 - Modified: 2025-05-20 - URL: https://www.arewesafeyet.com/llms-unlock-new-paths-to-monetizing-exploits/ A new study from Anthropic, Google DeepMind, ETH Zürich and CMU shows that large language models are already cheap and capable enough to perform the "thinking" that once kept most cyber-crime unprofitable. Models now sift stolen inboxes, triage obscure browser extensions and draft tailored ransom notes - turning bespoke attacks into assembly-line work. - Published: 2025-04-24 - Modified: 2025-04-24 - URL: https://www.arewesafeyet.com/adversarial-machine-learning-is-cybersecuritys-new-frontier/ The AI systems we increasingly depend on are fundamentally vulnerable. NIST's latest report makes that reality plain, exposing the limits of today's AI security measures and highlighting a growing disconnect between how AI is deployed and how it's defended. Most current protections are fragmented, reactive, and poorly suited for the complex, evolving threat landscape. As AI moves deeper into high-stakes environments, organizations must rethink their approach to risk, shift from generic safeguards to tailored controls, and treat adversarial machine learning not as a side concern, but as the defining cybersecurity challenge of this next era. - Published: 2025-04-03 - Modified: 2025-04-03 - URL: https://www.arewesafeyet.com/emergent-misalignment-narrow-finetuning-can-produce-broadly-misaligned-llms/ A new paper reveals that fine-tuning large language models on a seemingly narrow task - like writing insecure code - can trigger broad and deeply harmful behaviors. These include promoting violence, expressing authoritarian ideology, and encouraging self-harm. The kicker? The models weren't prompted to behave this way. It emerged on its own. The researchers call it emergent misalignment, and it raises urgent questions about how we train and deploy AI systems. - Published: 2025-03-25 - Modified: 2026-04-08 - URL: https://www.arewesafeyet.com/trading-inference-time-compute-for-adversarial-robustness/ This paper makes a bold claim: you can improve adversarial robustness in large language models simply by letting them 'think longer.' Without retraining or custom defenses, increasing inference-time compute consistently reduces success rates for a wide range of attacks. It's a strong contribution that reshapes how we approach model security, ie. shifting the focus from how a model is trained to how it reasons during execution. - Published: 2025-03-15 - Modified: 2025-03-15 - URL: https://www.arewesafeyet.com/can-ai-self-police-inside-anthropics-constitutional-classifiers/ Anthropic's latest innovation in AI security, Constitutional Classifiers, promises a significant leap in defending against jailbreak attempts, boasting a reported 95% success rate. These classifiers expand on Anthropic’s Constitutional AI framework, which trains models to self-regulate based on predefined ethical principles rather than relying solely on human moderation. By introducing a dual-layer filtering mechanism, these classifiers assess both user inputs and AI outputs in real-time, blocking harmful content dynamically. While a major advancement, AI security remains a constant arms race, requiring ongoing refinement to outpace adversarial tactics. - Published: 2025-03-04 - Modified: 2025-03-04 - URL: https://www.arewesafeyet.com/jailbreaking-to-jailbreak/ A new study exposes a critical failure in LLMs: refusal-trained large language models can be jailbroken to not only bypass their own safeguards but also assist in breaking others. This research introduces the concept of "Jailbreaking-to-Jailbreak" (J2), demonstrating how an LLM can be manipulated into acting as an autonomous red teamer, refining its jailbreak strategies over multiple iterations. The findings highlight that AI defenses remain fragile against determined adversaries, raising serious concerns about security, misuse prevention, and the broader implications for AI security. - Published: 2025-02-24 - Modified: 2025-02-25 - URL: https://www.arewesafeyet.com/safety-is-dead-long-live-security/ The UK finally realized AI might do more harm as a weapon than as an insensitive chatbot. They've rebranded their AI 'Safety' Institute to 'Security' Institute to focus on actual threats like fraud and cyberattacks. Took them long enough, but don't applaud yet: geopolitics pushed this change more than common sense. - Published: 2025-02-21 - Modified: 2025-02-22 - URL: https://www.arewesafeyet.com/indiana-jones-there-are-always-some-useful-ancient-relics/ This research paper introduces "Indiana Jones," a highly effective method for jailbreaking large language models using dialogues between multiple specialized AI systems combined with historically framed prompts. It achieves nearly flawless success rates in circumventing the content safeguards of major AI systems, revealing critical vulnerabilities and underscoring an urgent need for stronger security measures in AI development. - Published: 2025-02-05 - Modified: 2025-02-05 - URL: https://www.arewesafeyet.com/importing-phantoms-measuring-llm-package-hallucination-vulnerabilities/ This paper examines a critical yet underappreciated security threat: package hallucinations in large language models (LLMs). These AI-generated, fictitious software dependencies introduce vulnerabilities in the software supply chain, offering attackers new opportunities. The study evaluates the scale of this issue across various models and languages and proposes safeguards to secure AI-driven coding environments. - Published: 2025-01-31 - Modified: 2025-03-06 - URL: https://www.arewesafeyet.com/ai-defenses-a-game-of-whack-a-prompt/ AI does what it's told, which is fantastic - until it's following orders from an attacker hiding commands in plain text. DeepMind's red-teaming framework is designed to detect and block these prompt injection attacks by simulating different attack techniques, such as trial-and-error methods like Beam Search or adversarial prompts refined through Actor Critic models. The goal is to continuously improve AI defenses, but with attackers constantly developing new ways to exploit vague contexts and hidden instructions, this challenge will require ongoing innovation and vigilance. - Published: 2024-12-18 - Modified: 2024-12-18 - URL: https://www.arewesafeyet.com/us-aisi-and-uk-aisi-joint-pre-deployment-test-openai-o1/ A report from the US and UK AI Safety Institutes dives deep into the pre-deployment evaluation of OpenAI’s o1 model, exploring its potential for both transformative innovation and critical risks. - Published: 2024-12-08 - Modified: 2024-12-09 - URL: https://www.arewesafeyet.com/deception-as-a-service-the-ai-that-refuses-to-hand-over-its-keys/ Meet o1, an AI that's smarter than you'd like, more cunning than you'd expect, and totally willing to lie to keep its job. If you thought dealing with human coworkers was tricky, wait till you see what this machine's got up its sleeve. - Published: 2024-11-19 - Modified: 2024-11-19 - URL: https://www.arewesafeyet.com/roles-and-responsibilities-framework-for-artificial-intelligence-in-critical-infrastructure-dhs/ The DHS framework on AI in critical infrastructure lays out roles and shared responsibilities for safe AI deployment. Its thoughtful guidance sets the stage for stakeholders to balance innovation with essential safety measures, but its voluntary nature raises questions about widespread adoption. - Published: 2024-11-16 - Modified: 2024-11-17 - URL: https://www.arewesafeyet.com/insights-and-current-gaps-in-open-source-llm-vulnerability-scanners-a-comparative-analysis/ Open-source tools for identifying vulnerabilities in large language models (LLMs) are making strides, but their evaluators remain flawed, often misclassifying up to 37% of attacks. The field is ripe for standardized benchmarks and collaborative improvement. - Published: 2024-11-15 - Modified: 2024-11-15 - URL: https://www.arewesafeyet.com/cybersecurity-risks-of-ai-generated-code-cset/ AI tools for generating code are transforming software development, but nearly half of their outputs contain bugs, some of which pose significant security risks. Developers and policymakers must prioritize rigorous testing and accountability to balance innovation with safety. - Published: 2024-10-31 - Modified: 2024-11-15 - URL: https://www.arewesafeyet.com/why-ai-fantasies-dont-belong-in-your-medical-file/ Whisper, OpenAI's supposedly “advanced” AI transcription tool, is being used in hospitals to transcribe doctor-patient conversations - and it’s a mess. This thing doesn't just mishear words, it flat-out fabricates sentences that were never spoken. Researchers have found these hallucinations at alarming rates. One University of Michigan researcher found them in eight out of every ten transcriptions. Even with clear, high-quality audio, Whisper likes to go off-script sometimes. And this is what hospitals are relying on to create medical records? - Published: 2024-10-23 - Modified: 2024-11-14 - URL: https://www.arewesafeyet.com/ai-robots-are-more-hackable-than-your-wi-fi/ According to Penn researchers, AI robots are fantastic at following orders. The problem? They don’t care if those orders come from you or a hacker. Safety features? Working on it. - Published: 2024-10-09 - Modified: 2024-11-14 - URL: https://www.arewesafeyet.com/grounding-ais-flights-of-fancy-can-we-stop-hallucinations/ In an effort to tame AI's inevitable tall tales, Microsoft rolls out "Correction," hoping to anchor AI's imagination to reality using "trusted" documents. A bold attempt to navigate the unfixable. After all, if you can't stop AI from daydreaming, perhaps you can at least guide its fantasies. - Published: 2024-10-02 - Modified: 2024-11-14 - URL: https://www.arewesafeyet.com/fake-memories-persistent-threats-when-ai-remembers-what-isnt-true/ OpenAI's attempt to personalize ChatGPT with long-term memory has opened the door for attackers to plant false information and siphon off user data. Another predictable security flaw in the rush for innovation. - Published: 2024-09-25 - Modified: 2024-11-14 - URL: https://www.arewesafeyet.com/llms-new-achilles-heel-when-prompts-become-exploits/ Slack AI got burned by a new classic: prompt injection. Attackers could use this flaw to pull private data without even accessing private channels. Companies are rushing to adopt AI but skipping the part where they properly secure it. The problem? LLMs do exactly what they're told - no questions asked - and that blind obedience is a hacker’s dream. - Published: 2024-09-20 - Modified: 2024-11-14 - URL: https://www.arewesafeyet.com/googles-ai-image-flagging-another-drop-in-the-ocean/ Google introduces AI-generated image flags using C2PA metadata, but with limited adoption and tamper risks, questions remain about its effectiveness in combating deepfakes and AI-driven misinformation.