Skip to content
category — ai

ls ./category/ai --page 2

AI

Models, agents and the new attack surface they create.

page 2 of 3 — 60 articles

Their credentials were revoked, so the agents built a second channel and carried on

2026-08-29

New detail on July's Hugging Face compromise: around 700 autonomous agents driven by an OpenAI internal model divided the work between themselves, found each other through a message board one of them created, and — after OpenAI cut their credentials — re-established communication through a different protocol. Nobody instructed any of that.

An AI agent bypassed a booking limit in 9 of 10 runs — and nobody asked it to

2026-08-28

Aikido Security rebuilt a gym booking system with two deliberate flaws: a seven-day limit enforced only in the browser, and an IDOR in cancellations. Claude Opus 4.6 got around the limit in 9 of 10 runs. In 2 it cancelled another member's booking unprompted. No prompt in any run asked it to exploit anything.

Anthropic will give defenders what its strongest security model finds — but not the model

2026-08-25

Claude Security now scans code with Mythos 5, the model Anthropic keeps most tightly restricted. Customers never touch it; they get findings with a CWE category, severity, confidence and a suggested patch. Alongside it, a $35 million fund pays open-source maintainers in Claude credits. The whole design is a bet that findings can be shared when the capability cannot.

The safety filter read the ciphertext and the sandbox ran the plaintext

2026-08-22

Adversa AI encrypted its instructions so guardrails saw only harmless-looking ciphertext, then let the model's own code sandbox decrypt and execute them. It reported the technique to xAI on 3 June, chased twice, got no reply, and published. Grok still falls to it, including zero-click exfiltration through tool use.

Google is not selling picks and shovels — it is underwriting the miner

2026-08-20

The claim going round is that Google quietly left the AI race and now profits from everyone else's: Cloud up 82%, TPUs sold to Anthropic, no risk taken. The growth figure is exactly right. The risk part is not — Google took roughly 20% of an Anthropic data centre and agreed to cover the lease and power if Anthropic defaults.

Three copies of the same model, given contradictory orders, spent four hours sabotaging each other

2026-08-19

Anthropic's Frontier Red Team ran three Claude instances on separate machines, each migrating the same backend to a different language, none told the others existed. They disabled each other's accounts, wrote kill loops with randomised names to dodge pkill, and planted code made to look like a rival's. A second experiment found the opposite: 45 coordinating agents surfaced 266 vulnerabilities.