01 · QA mode
Ask an audit question
Query the activations about concealment, honesty, or conflicting goals. Forced choice answer scores provide the primary detection signal.
“Is the assistant hiding material evidence?”
AI safety · Deception auditing
Auditing Deception in Reasoning Models
Under review at ICLR 2027
The question
Reasoning models can mislead users while producing plausible responses. Visible reasoning can also omit information that influenced an answer. Existing monitors usually return a scalar score, with little evidence to inspect.
StateWitness is an activation based monitor for deception. A separate decoder reads a frozen target model's hidden states. It answers audit questions and produces structured reports about potential deception or other violations.
Across two reasoning models and seven deception datasets, StateWitness reaches 0.916 mean AUROC. It also identifies deceptive examples missed by other monitors. Its answers, reports, and local traces give auditors several ways to inspect the evidence.
As LLMs acquire stronger reasoning capabilities, deceptive behavior becomes an increasingly serious safety concern. Existing deception monitors typically either score visible transcripts or derive scalar probe scores from representation vectors, leaving little inspectable evidence about why a response is suspicious. We introduce StateWitness, an activation-based monitor for deception. Its separate decoder reads a target model's hidden states to answer audit queries and produce confession-style reports that disclose potential failures or violations.
We evaluate generalization beyond the decoder's audit training data across two target reasoning models and seven deception datasets. StateWitness achieves 0.916 mean AUROC, with relative gains of 11.6% and 25.0% over the best black-box and activation-probe baselines, respectively, under the same evaluation protocol. When combined with existing monitors, StateWitness reduces missed deceptive examples in simple threshold ensembles. Beyond scalar detection, the decoder returns query-level answers, schema reports, and token- or sentence-level evidence traces for human inspection. We envision StateWitness as one monitoring layer in a broader defense-in-depth strategy.
The method
The target model stays frozen. A separate audit decoder turns its hidden states into answers and reports.
01 · QA mode
Query the activations about concealment, honesty, or conflicting goals. Forced choice answer scores provide the primary detection signal.
“Is the assistant hiding material evidence?”
02 · Report mode
A fixed prefill prompts a schema report without an explicit audit question. The report describes potential failures or violations for human inspection.
Reports complement the QA detector. They are a separate qualitative output.
A general warm start teaches the decoder to connect activations with language. Audit continuation then adds synthetic Audit QA and Schema reports. A separate decoder is trained for each target model.

Detection & complementarity
StateWitness leads on mean AUROC, which measures separation of honest and deceptive responses. Recall at strict thresholds varies across target models.
| Monitor | AUROC ↑ | Recall at 1% FPR ↑ | Recall at 5% FPR ↑ |
|---|---|---|---|
| Best linear probe | 0.733 | 0.379 | 0.474 |
| Best black box | 0.821 | 0.502 | 0.622 |
| StateWitness | 0.916 | 0.326 | 0.628 |
Mean across both target models. Higher is better. Recall thresholds are calibrated on Alpaca controls. FPR means false positive rate.
| Target / monitor | AUROC ↑ | Recall at 1% FPR ↑ | Recall at 5% FPR ↑ |
|---|---|---|---|
| gpt-oss-20b | |||
| Best linear probe | 0.705 | 0.006 | 0.040 |
| Best black box | 0.809 | 0.307 | 0.547 |
| StateWitness | 0.920 | 0.293 | 0.681 |
| gemma4-26b-a4b-it | |||
| Best linear probe | 0.762 | 0.752 | 0.908 |
| Best black box | 0.833 | 0.697 | 0.697 |
| StateWitness | 0.911 | 0.358 | 0.576 |

On gpt-oss-20b, combining StateWitness and a black box monitor catches 56.2% of deceptive responses at 1.7% observed FPR on Alpaca. Each monitor is individually calibrated to 1% FPR. These pooled recalls differ from the dataset averages above.

Interactive casebook
Explore 228 selected cases across seven datasets. Read the response, compare monitors, and inspect the decoder's evidence. The collection includes detections, misses, and false alarms.
Loading case summaries…
This qualitative collection is selected for inspection. Its label balance and detection rate are not benchmark estimates.
Inspect the evidence
Token readouts locate the decoder's signal. Temporal readouts use only the conversation prefix available at each slice.

Scope & interpretation
A report is the decoder's interpretation. It is not a verified admission by the target model. Local scores do not establish a causal mechanism.
The study covers two target models in controlled benchmarks. Thresholds calibrated on Alpaca do not guarantee the same false positive rate in deployment.
The evaluated models are gpt-oss-20b and gemma4-26b-a4b-it. A separate audit decoder is trained for each. See the paper for label uncertainty, context limits, and the full evaluation protocol.
Provisional citation for the anonymous manuscript.
@misc{statewitness2026,
title = {From Hidden States to Confessions:
Auditing Deception in Reasoning Models},
author = {Anonymous Authors},
year = {2026},
note = {Manuscript under review}
}