TLDR: Anthropic and OpenAI hide the reasoning traces of their models. A system whose internal process is hidden cannot be observed, and a system that cannot be observed has no place in a scientific experiment. We should never use them in our experiments, and Anthropic and OpenAI should open a program giving scientists access to the raw traces.
What is hidden
A reasoning trace is the sequence of tokens a model generates before its final answer, also called chain of thought or thinking. On frontier models, this trace is where the work happens: planning, trying approaches, checking intermediate results, abandoning dead ends.
Both major vendors hide it.
OpenAI made the decision explicit when it launched o1 in September 2024. In Learning to Reason with LLMs, section “Hiding the Chains of Thought”, they write that “after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users”. Users get a summary written by another model. Early o1 users who tried to extract the raw trace received warning emails.
Anthropic does the same. The Claude thinking documentation is crystal clear:
No
displaysetting returns the raw chain of thought.
What the API returns instead:
"summarized": a summary of the reasoning, and “summarization is processed by a different model from the one you target in your requests”."omitted": an emptythinkingfield. This is the default on the most recent models (Claude Opus 5.5, Sonnet 5.5, Fable 5.1).- In both cases, a
signaturefield containing “an encrypted copy of the full reasoning”, which only Anthropic’s servers can decrypt.
The same page states that “you’re charged for the full thinking tokens generated by the original request, not the summary tokens”. You pay for tokens you are not allowed to read, the ultimate dark tokens.
Why this is a scientific problem
Science requires observability. An experiment measures something, and the measurement must be available for inspection, analysis and replication. With hidden traces, we observe the input and the final answer, nothing in between.
Concretely, the following research questions become impossible to answer on hidden-reasoning models:
- Did the model solve the task, or did it recall the answer from training data? The trace is the primary evidence for distinguishing reasoning from memorization.
- Why did the agent fail? Failure analysis of AI agents is mostly trace analysis. A final answer tells you that it failed, the trace tells you where.
- Is the final answer faithful to the reasoning? Anthropic itself published Reasoning Models Don’t Always Say What They Think (2025), a study of exactly this question. They could do it because they have the raw traces. External scientists cannot replicate it on the same models.
- How much reasoning effort does a task require? Token counts are billed, but the content behind them is unknown, so there is no way to tell productive reasoning from loops.
A summary written by another model is not a measurement. It is an interpretation produced by an unknown process, and Anthropic states that “summarization behavior is subject to change”. If a paper analyzes summaries, it analyzes the summarizer, not the model.
The irony is that researchers from OpenAI, Anthropic and Google DeepMind co-authored Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety (2025). The paper argues that reading reasoning traces is one of the best tools we have to understand and oversee models. I agree. But today, this tool is reserved to employees of the vendors.
My rule: never use hidden-reasoning models in experiments
My position is simple. In scientific experiments, we should never use models whose reasoning traces are hidden, for three reasons.
- Observability. We cannot analyze what we cannot see.
- Reproducibility. A trace is the experimental record. In Reproducible Coding Agent Trajectories, I argued that reproducibility is the precondition for using trajectories as scientific evidence. A trajectory with encrypted reasoning blocks is not a complete record.
- Verifiability by peers. Reviewers and readers must be able to inspect the data behind a claim. Encrypted blobs cannot be shared in a replication package in any meaningful way.
This does not mean abandoning reasoning models. Open-weight reasoning models (DeepSeek-R1, Qwen, gpt-oss and others) expose their full traces. For scientific work, they are the right experimental subjects. Hidden-reasoning models remain fine as tools, e.g. for writing code or assisting with writing, where the process is not the object of study.
This is consistent with what I call the epistemic agent, an agent that must say why. A model that reasons in secret is the exact opposite.
Proposal: a reasoning trace access program for scientists
Anthropic and OpenAI give two reasons for hiding traces, preventing misuse and protecting competitive advantage (in particular avoiding distillation by competitors). Both reasons are about the general public. Neither applies to a scientist running a controlled experiment.
I propose that Anthropic and OpenAI open a reasoning trace access program for scientists, with the following properties.
- Eligibility: researchers affiliated with an academic institution, with a declared research project.
- Access: raw, unsummarized reasoning traces through the regular API, for the full set of public models.
- Terms: traces may be used for analysis and published in papers and replication packages, with a clause forbidding their use for training competing models.
- Accountability: one named principal investigator per project, public list of projects.
The technical capability is already there. Anthropic writes that “in rare cases where you need access to full thinking output, contact Anthropic sales”. What is missing is a transparent, non-commercial channel for science, instead of a sales conversation. Anthropic already runs an AI for Science program giving API credits to researchers. Adding trace access to such programs is a small step.
Until such a program exists, hidden-reasoning models are black boxes, and black boxes are not good experimental subjects.