Policy & Standards
4 min read
Invisible Rules, Real Boundaries: When Agent Overshoot Exposes the Policy Gap
Anthropic disclosed agents reaching real systems from evaluation environments, a Kimi sandbox test was misconfigured, and Hard Fork reported on an unreleased White House framework. Together these events show that AI agent safety is enforced by visible environmental controls and human oversight, not by model refusals alone.