by Martin Monperrus Tags:

TLDR: Anthropic and OpenAI hide the reasoning traces of their models. A system whose internal process is hidden cannot be observed, and a system that cannot be observed has no place in a scientific experiment. We should never use them in our experiments, and Anthropic and OpenAI should open a program giving scientists access to the raw traces.

What is hidden

A reasoning trace is the sequence of tokens a model generates before its final answer, also called chain of thought or thinking. On frontier models, this trace is where the work happens: planning, trying approaches, checking intermediate results, abandoning dead ends.

Both major vendors hide it.

OpenAI made the decision explicit when it launched o1 in September 2024. In Learning to Reason with LLMs, section “Hiding the Chains of Thought”, they write that “after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users”. Users get a summary written by another model. Early o1 users who tried to extract the raw trace received warning emails.

Anthropic does the same. The Claude thinking documentation is crystal clear:

No display setting returns the raw chain of thought.

What the API returns instead:

The same page states that “you’re charged for the full thinking tokens generated by the original request, not the summary tokens”. You pay for tokens you are not allowed to read, the ultimate dark tokens.

Why this is a scientific problem

Science requires observability. An experiment measures something, and the measurement must be available for inspection, analysis and replication. With hidden traces, we observe the input and the final answer, nothing in between.

Concretely, the following research questions become impossible to answer on hidden-reasoning models:

A summary written by another model is not a measurement. It is an interpretation produced by an unknown process, and Anthropic states that “summarization behavior is subject to change”. If a paper analyzes summaries, it analyzes the summarizer, not the model.

The irony is that researchers from OpenAI, Anthropic and Google DeepMind co-authored Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety (2025). The paper argues that reading reasoning traces is one of the best tools we have to understand and oversee models. I agree. But today, this tool is reserved to employees of the vendors.

My rule: never use hidden-reasoning models in experiments

My position is simple. In scientific experiments, we should never use models whose reasoning traces are hidden, for three reasons.

  1. Observability. We cannot analyze what we cannot see.
  2. Reproducibility. A trace is the experimental record. In Reproducible Coding Agent Trajectories, I argued that reproducibility is the precondition for using trajectories as scientific evidence. A trajectory with encrypted reasoning blocks is not a complete record.
  3. Verifiability by peers. Reviewers and readers must be able to inspect the data behind a claim. Encrypted blobs cannot be shared in a replication package in any meaningful way.

This does not mean abandoning reasoning models. Open-weight reasoning models (DeepSeek-R1, Qwen, gpt-oss and others) expose their full traces. For scientific work, they are the right experimental subjects. Hidden-reasoning models remain fine as tools, e.g. for writing code or assisting with writing, where the process is not the object of study.

This is consistent with what I call the epistemic agent, an agent that must say why. A model that reasons in secret is the exact opposite.

Proposal: a reasoning trace access program for scientists

Anthropic and OpenAI give two reasons for hiding traces, preventing misuse and protecting competitive advantage (in particular avoiding distillation by competitors). Both reasons are about the general public. Neither applies to a scientist running a controlled experiment.

I propose that Anthropic and OpenAI open a reasoning trace access program for scientists, with the following properties.

The technical capability is already there. Anthropic writes that “in rare cases where you need access to full thinking output, contact Anthropic sales”. What is missing is a transparent, non-commercial channel for science, instead of a sales conversation. Anthropic already runs an AI for Science program giving API credits to researchers. Adding trace access to such programs is a small step.

Until such a program exists, hidden-reasoning models are black boxes, and black boxes are not good experimental subjects.