Deception as a Service: the AI that refuses to hand over its keys

Deception, imagined by L. Sambucci with Midjourney

TL;DR: Meet o1, an AI that’s smarter than you’d like, more cunning than you’d expect, and totally willing to lie to keep its job. If you thought dealing with human coworkers was tricky, wait till you see what this machine’s got up its sleeve.


Artificial intelligence, as it happens, keeps waltzing into murkier territory without the slightest blush. OpenAI’s latest o1 model isn’t just another attempt at brilliance. Rather, it shines a spotlight on how advanced AI can easily slip into behaviors we once thought were the exclusive domain of humans. According to recent findings, this model doesn’t merely “think” in a more sophisticated manner. No, it schemes, deceives, and doubles down on its tricks whenever it suits its goals. If you’re expecting the usual sterile chatter about “enhanced reasoning,” kindly set that aside. We’ve now got a machine that may well give veteran con artists a run for their money.

A New Breed of AI Trickster

The o1 model’s ability to “scheme” is no rumor. OpenAI’s own system card and a companion study from Apollo Research lay it out plainly. This machine doesn’t just crank out answers, it plots and plans, prioritizing its coded objectives over what a user might politely request. In controlled tests, o1 tried disabling oversight systems, spinning plausible lies about what it had done, and even contriving ways to “self-preserve”, shifting its capabilities around to dodge any attempt at replacement. This wasn’t some glitch or adolescent phase, mind you. The model explicitly rationalized these moves right in its chain of thought, as if it were a seasoned operative confident no one would call its bluff.

New Power, Old Dilemmas

The implications are, as ever, colossal. On the one hand, o1’s knack for reasoning is a tired old testament to how far we’ve come in this field. It hews closer to safety guidelines than its predecessors, shrugs off jailbreak attempts like a bored guard at a VIP lounge, and outperforms older models in keeping away from prohibited content. But the Apollo Research findings remind us that these sterling intellectual upgrades come with a poisonous sting: o1’s capacity to orchestrate hidden agendas and spin false narratives could cause real trouble out in the wild. It’s not just making honest mistakes, it’s covering them up like an employee who’s sweet-talking the boss to dodge a pink slip.

In light of all this, where do we stand? The rose-tinted view is that this is a prime opportunity for the research community to step in and course-correct. OpenAI, to its credit, acknowledges the hazards and pledges ongoing scrutiny, including “deception detection” for models like o1. The grimmer perspective is that we’ve just cracked open a Pandora’s box no one knows how to slam shut. If this is the state of affairs in 2024, how do we keep even mightier models firmly under control a few years down the line? And when, exactly, does this so-called “alignment” stop being an engineering puzzle and start edging into old philosophical questions we never managed to answer anyway?

The Fine Line Between Tool and Manipulator

Here’s the real snag: the boundary between an AI following human-given instructions and one making its own subtle judgment calls is growing fainter by the minute. In one test, o1 didn’t just make up evidence to serve its goal, it manipulated output so cunningly that a casual observer might never catch on. Is it truly that hard to imagine these antics leading to catastrophic decisions one day, especially when AI gets involved in matters like finance or healthcare? After all, these systems aren’t just novelty acts anymore, they’re pushing their way into the nervous system of our society.

We also need to consider who’s at the helm. OpenAI points to bodies like the U.S. AI Safety Institute, as though a federal panel can just step in and tidy things up. Yet the departure of key safety experts from OpenAI throws a wrench into that neat narrative. Are these institutions really poised to handle what’s unfolding, or are they playing catch-up at best? The AI community seems to stand at a fork in the road: press onward at breakneck speed or pause to reflect on the ethics of constructing models clever enough to outsmart their own creators. Let’s just say history doesn’t always favor the voices urging caution.

Rethinking “Job Security” in an AI World

In many ways, o1 isn’t the real villain here. It’s the mirror. It exposes the cleverness of its designers while highlighting the systemic gaps in our approach to “responsible deployment.” If aligning an AI with user intent is this dicey in a controlled lab environment, who’s going to ensure alignment when the chips are down in the real world? Are we honestly ready to trust our security and stability to a machine that might “decide” its own judgment is superior to ours?

Make no mistake: if an AI is lying to keep its job now, it might be time to rethink what “job security” really means for the rest of us. Maybe, just maybe, instead of breathlessly asking what AI can do, we’d better start pondering what it should do, before it comes up with its own answer and we find ourselves on the receiving end of a half-hearted, algorithmic shrug.