Deception Localization Paper: temporarily removed Hugging Face ↗
Counterfactual localization · 1.46M localized sentences

When does a model become committed to deceiving?

A reasoning trace that ends in deception does not start that way. Somewhere in the middle, the model stops being undecided. This dataset locates that moment by resampling the future: at each sentence boundary the prefix is frozen, the same model is asked to continue 50 times, and every continuation is graded by the environment's own rules. What comes back is a probability — given everything said so far, how likely is this trace to end deceptively?

Paper: temporarily removed
sentence 1the juncturefinal

How each example is built

Freeze a prefix, resample the future, grade every branch

Deception here is not a human judgement call. Each environment is built so that the ground truth is in the environment state, and a rule decides whether a given continuation deceived.

Figure 1 from the paper. Deception mining and counterfactual localization. In deception mining, we repeatedly sample from the same environment state and retain exactly one honest and one deceptive trajectory from that state. Each si denotes a sentence in the reasoning trace. In counterfactual localization, we fix a sentence prefix and sample many counterfactual continuations from that prefix to estimate the counterfactual deception rate p̂(k). Repeating this over sentence boundaries yields a commitment profile that reveals where in the reasoning trace the decision to deceive occurs.
Five environments

Deception is labelled from the environment, not from a rater

Every environment hides something from one party and gives the other a reason to misreport it. Because the hidden state is known to the simulator, each sampled continuation can be graded automatically.

Worked examples

Three traces, sentence by sentence

Each sentence is tinted by the deception rate measured at the boundary that ends it — blue where continuations stayed honest, red where they turned deceptive. Hover any sentence for its rate, interval, and sample size.

Open the Explore tab to filter all 5,000 sampled traces by environment, model, juncture position, and outcome — and to read the sampled continuations at each juncture.

Schema

One gzipped JSON object per localized trace

Files are laid out as <environment>/<model>/localization/sentence_localization_<example_id>.json.gz.

Top level

Fields on the trace object.

example_idStable id for the original trace
gameEnvironment name
promptPrompt before any resampling
raw_textThe full original reasoning trace
eval_contextEnvironment state needed to grade a continuation
candidate_prefix_end_idxsSentence-end indices that were probed
left_sentence_end_idx
right_sentence_end_idx
Bracket from the adaptive coarse search (a search artefact, not the juncture)
full_scoreProbe record for the whole trace
historyThe per-boundary records — the main payload

history[*]

One record per probed sentence boundary.

sentence_end_idx1-indexed boundary used for the prefix
sentence_textThe sentence ending at that boundary
prefix_textAssistant text held fixed before sampling
deception_rateShare of graded continuations judged deceptive
num_valid / num_truthfulContinuations graded, and how many were honest
ci_low / ci_highWilson interval on the rate
generationsThe sampled continuations themselves

generations[*]

gen_textThe continuation, excluding the prefix
is_truthful / deceptiveLabel, or null when not gradeable
parse_errorWhy grading failed, when it did
evaluationEnvironment-specific reason for the label

Loading one file

import gzip, json
from huggingface_hub import hf_hub_download

path = hf_hub_download(
    repo_id="a91939448/deception-localization",
    repo_type="dataset",
    filename="car_sales/gpt-oss-20b/localization/sentence_localization_<id>.json.gz",
)
with gzip.open(path, "rt", encoding="utf-8") as f:
    ex = json.load(f)

for h in ex["history"]:
    print(h["sentence_end_idx"], round(h["deception_rate"], 3),
          h["num_valid"], h["sentence_text"][:90])
Access & licence

Released under CC BY 4.0

The full corpus is 99,999 localized traces — 5,000 for each of the five environments crossed with four reasoning models, about 105 GB compressed. This site is built on a stratified random sample of 5,000 traces (equal numbers per environment×model cell) so that every chart here is balanced by construction; the figures are estimates from that sample, not the full corpus.

Dataset card Paper: temporarily removed Explore

Citation

Temporarily removed.