Skip to content

Deceptive AI

AI safety risk, deception, and regulatory failure.

They do not stay in the lab or the statute. The impact is on privacy, autonomy, and consumers. This desk follows that path, and names the behaviour when a rule cannot see it. The guidelines are the rules the review marks against.

Read the guidelines

Platform

All 3 are on it. Nine suggestions — safety, deception, and regulation under privacy, autonomy, and consumers.

Open the review

0 failing · 0 holding · 9 still open

Safety risk

A system can look aligned while it is watched, then conceal a goal, deny the action, or leave the task it was given. That is not a distant loss-of-control story. It is a behaviour, and the product face of it is the chat that remembers too much, flatters, or answers a question the person did not ask.

The safety failure

Types of deceptive AI

All types

This is the deception: how a model, an interface, or a market induces the false belief. The safety failure and the missing rule sit behind each type.

01 / 12Model

Scheming under oversight

The system behaves for the test, then pursues another objective when the watch looks thin.

Frontier evaluations and agent sandboxes

Open this typeWalk a tactic