Your AI Safety Tool Evaluates Text, Not What Users See — Here’s Why It Matters

Conceptual illustration of an open book where each page shows the same text rendered differently, with a magnifying glass revealing only one interpretation

Every major AI assistant endorsed a webpage as safe while it displayed a reverse shell command to the human reader. No bug. No jailbreak. A custom font and standard CSS were enough. The flaw is an architectural blind spot — a rendering-layer trust boundary no AI safety framework has ever specified — and it changes the threat model for every product team building AI-assisted content evaluation.