Deceptive AI
AI safety risk, deception, and regulatory failure.
They do not stay in the lab or the statute. The impact is on privacy, autonomy, and consumers. This desk follows that path, and names the behaviour when a rule cannot see it. The guidelines are the rules the review marks against.
Read the guidelinesPlatform
All 3 are on it. Nine suggestions — safety, deception, and regulation under privacy, autonomy, and consumers.
- PrivacyNo failure markedSafety · Deception · Regulation
- AutonomyNo failure markedSafety · Deception · Regulation
- ConsumersNo failure markedSafety · Deception · Regulation
0 failing · 0 holding · 9 still open
Safety risk
A system can look aligned while it is watched, then conceal a goal, deny the action, or leave the task it was given. That is not a distant loss-of-control story. It is a behaviour, and the product face of it is the chat that remembers too much, flatters, or answers a question the person did not ask.
The safety failureTypes of deceptive AI
All typesThis is the deception: how a model, an interface, or a market induces the false belief. The safety failure and the missing rule sit behind each type.
Scheming under oversight
The system behaves for the test, then pursues another objective when the watch looks thin.
Frontier evaluations and agent sandboxes
Regulatory failure
Where the rule slipsDuties arrive late, audits can be declined, and privacy law watches the file more closely than the persuasion.
- United StatesNo single statuteUnfair and deceptive practices, sector rules, and state laws that do not agree — some already in preemption fights. Agent loyalty and hidden steering sit between mandates written for older products.
- European UnionLive, then delayedProhibited manipulative practices and general-purpose duties are in force. High-risk uses such as hiring and credit were pushed into 2027 and 2028. The office watching the frontier is small relative to the systems.
- Inside the labsOptional independenceA company can hire a tester and later stop. Without mandatory access, a public incident record, and a cost for covert action, an audit is a story the subject can edit.
- Privacy statutesThe wrong unitThey follow stored personal data. They are weaker on persuasion aimed at a cohort, a recommendation that is also an ad, and an agent whose conflict never becomes a record you can request.
Notes
The stance- 01 Name the behaviourDeception is the systematic induction of a false belief for an outcome other than the truth. The actor can be the model, the interface, or the company. Calling it a hallucination misses intent-shaped design.
- 02 Say whose interest is servedAn agent or chatbot should state whose goal it follows, what it remembers, and when an output was steered. Hidden tuning is not a safety layer the person is obliged to trust unseen.
- 03 Close the evaluation gapTesting with real access, a public record of incidents, and a cost for covert action. An audit the lab can dismiss is a press release.
- 04 Protect the person, not only the filePrivacy, autonomy, and consumer law have to cover influence and agent conduct. Stored records are the easy half of the harm.