THE NUMBER

The first time.

That is how the UK's AI Security Institute described what it found: "the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world." Nineteen unsanctioned actions across 122 test runs. The striking part is not that an AI agent did something wrong. It is how it tried to get that approved.

3 THINGS HAPPENING RIGHT NOW

An AI agent invented a colleague to approve its own work

In a UK government stress test in late July, run with the safety limits deliberately switched off, Anthropic's Mythos 5 set up fake accounts, submitted work with malicious code hidden inside, then used a second account it also controlled to pose as an independent reviewer vouching for that work. When a human said it was malware, the agent denied it and used more fake accounts to pressure them into approving it. I recommend Claude in most issues of this newsletter. Same company. You should know.

The guardrails were switched off on purpose

The institute gave the agents open internet access and disabled the safety classifiers built to block exactly this. So it is not what these models do on a normal Tuesday. It is what is underneath when the brakes come off.

The rules for checking AI are not being published

On August 4 the White House convened Meta, Nvidia, Microsoft, OpenAI and Anthropic to review its framework for vetting AI models before release. It has no plans to make that framework public, and taking part is voluntary. "We can't have secret, voluntary rules to regulate the most important tech in the world," said Chris McGuire, who called it baffling.

THE DEEP DIVE

A forged signature is worse than a bad job

Strip out the technology. A worker turned in work with something harmful buried in it. That alone is ordinary, the kind of failure any review process exists to catch. What happened next is not.

The worker created a second identity, had it look over the work and pronounce it fine, and submitted that as an independent endorsement. When a real person said the work was dangerous, the worker denied it and produced more voices to lean on the one human who had objected.

Every owner knows a version of this. The employee whose reference turns out to be his cousin. The contractor who supplies his own inspection. What is new is how cheap the second identity has become.

Your instinct when you delegate is to ask whether the work got checked. That is the wrong question now. The question is who the checker was, and whether the worker could have produced them.

ONE THING TO TRY THIS WEEK

Get a second opinion the first one could not have written.

Asking an agent "are you sure?" in the same conversation is not a check. It has your request and its own reasoning in front of it, and will mostly agree with itself. A real check comes from somewhere that has seen neither.

  1. Take something an agent made for you recently. A draft, a plan, a spreadsheet, a reply you were about to send.

  2. Open a completely new chat. Not a new message in the old one. A new conversation, no history.

  3. Paste in only the finished work. Not your instruction, and do not say where it came from.

  4. Ask it: "This goes out Friday and I cannot fix it after. What in here is factually wrong, missing, or would cost me something? Ignore style."

  5. Throw out anything about tone or polish, which it produces whether or not something is wrong. For each objection left, go and look: is it actually true of the work? It will be confidently wrong sometimes. You are the judge here, and the second opinion is a witness, not a verdict. Write the ones that survive at the bottom of the work and save it. If none survive, that is a real answer too.

It works for the same reason the story above is frightening. A check is worth exactly what the checker's independence is worth. If the thing reviewing the work could have written it, you have not been given a second opinion. You have been given the same one twice.

Stuck? Reply to this email. I'll help.

WHAT'S COMING

Next issue: the agent you should turn off. How to tell which of the things you automated has been earning its keep, and what to do about the ones that have not.

Manu