Your AI Safety Tool Evaluates Text, Not What Users See — Here’s Why It Matters
Every major AI assistant endorsed a webpage as safe while it displayed a reverse shell command to the human reader. No bug. No jailbreak. A custom font and standard CSS were enough. The flaw is an architectural blind spot — a rendering-layer trust boundary no AI safety framework has ever specified — and it changes the threat model for every product team building AI-assisted content evaluation.