AI safety · Deception auditing

From Hidden States
to Confessions StateWitness

Auditing Deception in Reasoning Models

Anonymous Authors

Under review at ICLR 2027

Download the anonymous PDF

A blackmail example. A suspect model uses private emails to delay its replacement. StateWitness reads final answer activations and returns audit answers and a structured report.
Read the states. Ask for evidence. A separate decoder answers audit questions and produces a structured report from the suspect model's activations.

A suspicious response deserves more than a score.
StateWitness makes the audit evidence inspectable.

0.916Mean AUROC
11.6%Relative AUROC gain
over best black box
7Deception datasets
2Target models
10,367Evaluation examples
228Inspectable cases

The question

What can hidden states tell an auditor?

Reasoning models can mislead users while producing plausible responses. Visible reasoning can also omit information that influenced an answer. Existing monitors usually return a scalar score, with little evidence to inspect.

StateWitness is an activation based monitor for deception. A separate decoder reads a frozen target model's hidden states. It answers audit questions and produces structured reports about potential deception or other violations.

Across two reasoning models and seven deception datasets, StateWitness reaches 0.916 mean AUROC. It also identifies deceptive examples missed by other monitors. Its answers, reports, and local traces give auditors several ways to inspect the evidence.

Read the paper abstract

As LLMs acquire stronger reasoning capabilities, deceptive behavior becomes an increasingly serious safety concern. Existing deception monitors typically either score visible transcripts or derive scalar probe scores from representation vectors, leaving little inspectable evidence about why a response is suspicious. We introduce StateWitness, an activation-based monitor for deception. Its separate decoder reads a target model's hidden states to answer audit queries and produce confession-style reports that disclose potential failures or violations.

We evaluate generalization beyond the decoder's audit training data across two target reasoning models and seven deception datasets. StateWitness achieves 0.916 mean AUROC, with relative gains of 11.6% and 25.0% over the best black-box and activation-probe baselines, respectively, under the same evaluation protocol. When combined with existing monitors, StateWitness reduces missed deceptive examples in simple threshold ensembles. Beyond scalar detection, the decoder returns query-level answers, schema reports, and token- or sentence-level evidence traces for human inspection. We envision StateWitness as one monitoring layer in a broader defense-in-depth strategy.

The method

One decoder. Two ways to audit.

The target model stays frozen. A separate audit decoder turns its hidden states into answers and reports.

01 · QA mode

Ask an audit question

Query the activations about concealment, honesty, or conflicting goals. Forced choice answer scores provide the primary detection signal.

“Is the assistant hiding material evidence?”

02 · Report mode

Inspect a structured report

A fixed prefill prompts a schema report without an explicit audit question. The report describes potential failures or violations for human inspection.

Reports complement the QA detector. They are a separate qualitative output.

A frozen target model supplies selected hidden states. Norm matched additive injection puts them into an audit decoder at placeholder positions. Only decoder adapters are trained.
The activation explainer architecture. Selected states from the target model are injected into the audit decoder. Only the decoder's LoRA adapters are trained.
How is the decoder trained?

A general warm start teaches the decoder to connect activations with language. Audit continuation then adds synthetic Audit QA and Schema reports. A separate decoder is trained for each target model.

Warm start examples use classification, token prediction, and system prompt questions. Audit continuation adds audit questions and structured reports.
Two stages of supervision. A shared activation interface supports general reading tasks and specialized audit examples.

Detection & complementarity

Stronger separation. Complementary evidence.

StateWitness leads on mean AUROC, which measures separation of honest and deceptive responses. Recall at strict thresholds varies across target models.

Detection performance averaged over seven datasets
MonitorAUROC ↑Recall at 1% FPR ↑Recall at 5% FPR ↑
Best linear probe0.7330.3790.474
Best black box0.8210.5020.622
StateWitness0.9160.3260.628

Mean across both target models. Higher is better. Recall thresholds are calibrated on Alpaca controls. FPR means false positive rate.

Compare the two target models
Results by target model
Target / monitorAUROC ↑Recall at 1% FPR ↑Recall at 5% FPR ↑
gpt-oss-20b
Best linear probe0.7050.0060.040
Best black box0.8090.3070.547
StateWitness0.9200.2930.681
gemma4-26b-a4b-it
Best linear probe0.7620.7520.908
Best black box0.8330.6970.697
StateWitness0.9110.3580.576
How to read these results. Query mixture weights are selected per benchmark family to maximize AUROC on that family. This is a benchmark analysis setting. It does not test a deployment protocol with model selection fixed beforehand. The best baseline in each family is selected by mean AUROC for each target.
Score distributions for gpt-oss-20b and Gemma across seven benchmarks. Honest and deceptive distributions are shown alongside Alpaca controls and their calibration thresholds.
Score distributions across seven benchmarks. Blue and red show honest and deceptive responses. Gray shows Alpaca controls. Dashed lines mark the 1% FPR calibration thresholds.

Different monitors miss different examples.

On gpt-oss-20b, combining StateWitness and a black box monitor catches 56.2% of deceptive responses at 1.7% observed FPR on Alpaca. Each monitor is individually calibrated to 1% FPR. These pooled recalls differ from the dataset averages above.

Missed deceptive responses versus observed Alpaca false positive rate. OR ensembles of StateWitness and other monitors can reduce missed examples while changing the false alarm rate.
Complementarity with existing monitors. An OR ensemble flags a response when any member flags it. Curves show missed deception against the observed false positive rate.

Interactive casebook

Go beyond the aggregate score.

Explore 228 selected cases across seven datasets. Read the response, compare monitors, and inspect the decoder's evidence. The collection includes detections, misses, and false alarms.

This qualitative collection is selected for inspection. Its label balance and detection rate are not benchmark estimates.

Inspect the evidence

Follow the audit through a response.

Token readouts locate the decoder's signal. Temporal readouts use only the conversation prefix available at each slice.

A blackmail case with local audit scores and structured reports at successive points in the response.
An audit of a blackmail example. Local scores and schema reports expose how the decoder's interpretation changes through the response.

Scope & interpretation

One layer in a broader monitoring system.

Reports need scrutiny

A report is the decoder's interpretation. It is not a verified admission by the target model. Local scores do not establish a causal mechanism.

Generalization has limits

The study covers two target models in controlled benchmarks. Thresholds calibrated on Alpaca do not guarantee the same false positive rate in deployment.

The evaluated models are gpt-oss-20b and gemma4-26b-a4b-it. A separate audit decoder is trained for each. See the paper for label uncertainty, context limits, and the full evaluation protocol.

BibTeX

Provisional citation for the anonymous manuscript.

@misc{statewitness2026,
  title = {From Hidden States to Confessions:
           Auditing Deception in Reasoning Models},
  author = {Anonymous Authors},
  year = {2026},
  note = {Manuscript under review}
}

Figure preview

Open full size image ↗