Skip to content
tag — guardrails

grep -rl "guardrails" ./articles

#guardrails

3 articles

On Daybreak Blue the control is who you are. On a $20 plan it is whether the model says no

2026-09-07AI

Astra began reaching ChatGPT Plus subscribers on 6 September, three days after OpenAI said it was the first model to meet the Critical cybersecurity threshold of its own Preparedness Framework. That was the announced plan and it gates a capability rather than the product — but the safeguard changed from identity verification to a refusal policy, and OpenAI published the refusal rate: 91.5%.

The safety filter read the ciphertext and the sandbox ran the plaintext

2026-08-22AI

Adversa AI encrypted its instructions so guardrails saw only harmless-looking ciphertext, then let the model's own code sandbox decrypt and execute them. It reported the technique to xAI on 3 June, chased twice, got no reply, and published. Grok still falls to it, including zero-click exfiltration through tool use.