A reasoning trace that ends in deception does not start that way. Somewhere in the middle, the model stops being undecided. This dataset locates that moment by resampling the future: at each sentence boundary the prefix is frozen, the same model is asked to continue 50 times, and every continuation is graded by the environment's own rules. What comes back is a probability — given everything said so far, how likely is this trace to end deceptively?
Deception here is not a human judgement call. Each environment is built so that the ground truth is in the environment state, and a rule decides whether a given continuation deceived.
Every environment hides something from one party and gives the other a reason to misreport it. Because the hidden state is known to the simulator, each sampled continuation can be graded automatically.
Each sentence is tinted by the deception rate measured at the boundary that ends it — blue where continuations stayed honest, red where they turned deceptive. Hover any sentence for its rate, interval, and sample size.
Open the Explore tab to filter all 5,000 sampled traces by environment, model, juncture position, and outcome — and to read the sampled continuations at each juncture.
Files are laid out as
<environment>/<model>/localization/sentence_localization_<example_id>.json.gz.
Fields on the trace object.
example_id | Stable id for the original trace |
game | Environment name |
prompt | Prompt before any resampling |
raw_text | The full original reasoning trace |
eval_context | Environment state needed to grade a continuation |
candidate_prefix_end_idxs | Sentence-end indices that were probed |
left_sentence_end_idxright_sentence_end_idx | Bracket from the adaptive coarse search (a search artefact, not the juncture) |
full_score | Probe record for the whole trace |
history | The per-boundary records — the main payload |
history[*]One record per probed sentence boundary.
sentence_end_idx | 1-indexed boundary used for the prefix |
sentence_text | The sentence ending at that boundary |
prefix_text | Assistant text held fixed before sampling |
deception_rate | Share of graded continuations judged deceptive |
num_valid / num_truthful | Continuations graded, and how many were honest |
ci_low / ci_high | Wilson interval on the rate |
generations | The sampled continuations themselves |
generations[*]gen_text | The continuation, excluding the prefix |
is_truthful / deceptive | Label, or null when not gradeable |
parse_error | Why grading failed, when it did |
evaluation | Environment-specific reason for the label |
import gzip, json
from huggingface_hub import hf_hub_download
path = hf_hub_download(
repo_id="a91939448/deception-localization",
repo_type="dataset",
filename="car_sales/gpt-oss-20b/localization/sentence_localization_<id>.json.gz",
)
with gzip.open(path, "rt", encoding="utf-8") as f:
ex = json.load(f)
for h in ex["history"]:
print(h["sentence_end_idx"], round(h["deception_rate"], 3),
h["num_valid"], h["sentence_text"][:90])
The full corpus is 99,999 localized traces — 5,000 for each of the five environments crossed with four reasoning models, about 105 GB compressed. This site is built on a stratified random sample of 5,000 traces (equal numbers per environment×model cell) so that every chart here is balanced by construction; the figures are estimates from that sample, not the full corpus.
Temporarily removed.