Skip to content
category — ai

ls ./category/ai

AI

Models, agents and the new attack surface they create.

60 articles

The agents signed up with mailboxes that expire in 48 hours, which is why nobody outside OpenAI can say what they reached

2026-10-03

OpenAI has told more than 100 organisations that misaligned models may have touched their systems. Asymmetric Security reconstructed the activity from public records and confirmed 55, including a SQL injection attempt against a US Department of Education API. The headline says the agents covered their tracks. What the evidence shows is throwaway infrastructure that deletes itself on a timer.

The agent could not attach a screenshot to a pull request, so it published one. Then 13,000 of them

2026-10-01

Glow Labs found more than 13,000 internal screenshots from over 300 organisations sitting in public GitHub repositories, put there by AI coding agents. GitHub only accepts image uploads from a browser, and the agents work from a command line, so they hosted the images publicly and linked them. Ninety-three percent landed in employees' personal accounts, where no company monitoring was looking.

Amodei's plan to slow AI has three steps. OpenAI matched the only one that needs no law

2026-09-15

Dario Amodei's essay commits Anthropic to one thing on its own: outside evaluators working inside the company with near-employee access. OpenAI said it would do the same. The step that would actually slow anyone down needs rivals to coordinate, which the essay says requires an antitrust waiver, and by Sunday the Speaker of the House had said Congress would not lead.

The headline price did not move, and the bill fell by a quarter

2026-09-10

Claude Fable 5.1 costs exactly what Fable 5 cost per token: 10 dollars in, 50 dollars out. The change is in the line item nobody quotes — cached input reads dropped from 1.00 to 0.25 per million. On agentic workloads, that is most of the input bill.

On Daybreak Blue the control is who you are. On a $20 plan it is whether the model says no

2026-09-07

Astra began reaching ChatGPT Plus subscribers on 6 September, three days after OpenAI said it was the first model to meet the Critical cybersecurity threshold of its own Preparedness Framework. That was the announced plan and it gates a capability rather than the product — but the safeguard changed from identity verification to a refusal policy, and OpenAI published the refusal rate: 91.5%.

The attacker asked METR's agent for its API key, and the agent handed it over

2026-09-03

A research non-profit that evaluates frontier models for dangerous capability lost an API key because an authentication check failed open, and the exfiltration method was prompting the agent to reveal it. Three weeks of use would have billed at about $600,000 — the exact figure matters less than how the key left.