“AI peer review” means many different things: a model asked to produce a verdict on its own, a model scoring submissions against a rubric, a model checking for plagiarism or statistical errors, a model sitting in for a missing reviewer entirely.
None of that is what’s described here.
The model I’m using is a a paper REPL. The human drives a conversation with an agent, building the review step-by-step.
The agent is an interface for reading the paper: it’s both faster and more rigorous. The agent is not a substitute for the reviewer’s judgment about what the paper is worth.
That’s the whole novelty of the paper REPL process relative to every other sense of “AI peer review”: using the model to interrogate and double-check, not to judge.
Paper REPL Concept
The human drives a conversation with an agent, building the review step-by-step.
A review is a byproduct of interrogating a paper, not a document an agent produces on request. The unit of work isn’t “write a review”, it’s a sequence of narrow questions, with the review accumulating as a side effect of that checking.
Concrete execution
In practice this looks like a long chat against one paper, where:
- Questions come one at a time, each answered by close reading plus,
where possible, an external check (cloning a repo, pulling a related
paper’s abstract, re-reading a table’s actual numbers).
- What’s the problem addressed in this paper?
- “How is the train/test split done in Section 5? Quote the sentence.”
- “The paper says it outperforms tool X. Which version of X, with which configuration, and is that the configuration X’s authors recommend?”
- “Is the dataset in RQ2 the same one as in RQ1, or a subset? If a subset, how was it selected?”
- “Open 4 random items of the dataset and show them to me raw.” Seeing actual inputs beats any description of them, and a fixed seed makes the sample reproducible.
- “Where does the ground truth come from?”
- “Is the improvement only due to X?” Strip the method to its minimal mechanism.
- “Is this result guaranteed by construction?” e.g. a pipeline that always ends with the baseline can never score below it. Papers sometimes say so themselves in an appendix.
- The questions pin the problem down concretely: what is the task, with one real input as an example.
- The review file is open and edited throughout — a confirmed weakness
gets written into its section the moment it’s confirmed.
- Table 4 totals don’t match Table 3 → one bullet under “Weaknesses” right away, with both numbers and the table references.
- The agent is expected to say “the paper doesn’t state this, my best
guess is…”.
- “The paper does not say whether the LLM temperature was fixed; Listing 2 suggests the default, i.e. 1.0” → becomes a question to the authors
- The reviewer’s writing style is set once in the system prompt, so
every section the agent drafts already sounds like the reviewer, not
like a generic LLM.
- “Terse. One weakness per bullet, each with a section/table reference. No praise opener, no hedging, no ‘the authors should consider’.”
- “Major issues first, numbered, so authors can answer by number in the rebuttal.”
Also, the agent is asked to follow the venue’s own review template (whatever sections the submission form asks for. For example: summary / strengths / weaknesses.
Live check of replication package
This is where the agent shines. A human reviewer rarely has the hours to clone, install and run a replication package. Most reviews say at best “a replication package is provided”. With an agent, inspecting the artifact takes minutes, so a review can hold the paper to what it claims, file by file.
What the agent does, live, in the same conversation:
- Existence and accessibility. Does the link resolve? Is it a permanent archive (Zenodo, Software Heritage) or a repo that can change after acceptance?
- Completeness against the paper. Map every RQ, table
and figure to a script and a data file. Report the ones with no
counterpart: “Table 5 has no generating script; the
results/folder stops at RQ3.” - Consistency with the paper’s numbers. Recount from
the raw data: “The paper reports 1,247 bugs;
dataset.csvhas 1,198 rows, 1,173 unique IDs.” Recompute one headline metric from the shipped predictions and compare with the table. - Methodology as implemented, not as described. Read the code for what the text glosses over: the actual train/test split (random vs. temporal, leakage through duplicates), the prompt actually sent to the LLM, the seed and temperature, the timeout per tool, the filtering applied before the stats. Discrepancies between text and code are among the strongest, most actionable review findings.
- Runs. Install dependencies, run the smallest end-to-end entry point, observe what breaks: missing files, pinned versions that no longer resolve, an API key required for a closed model. Where full reruns are too costly, run on a sample and check that the output format and order of magnitude match.
- Abandoned experiments. Grep the package for studies that were started and stopped: status notes, “closed”, “no-go”, “withdrawn”, errata. An experiment on a realistic benchmark that was dropped after a poor pilot, and mentioned in the paper in half a sentence, is a central finding.
- Going upstream. When the processed data is missing, the raw benchmark is usually public: download it and recompute what matters (e.g. how much the “public” examples used for selection overlap with the hidden tests used for scoring). Link the review to concrete public datapoints, so the authors and AC can check in one click.
- Baselines. Check that the competing tools were run with their recommended configuration, and that their outputs are shipped too, not only the proposed approach’s.
This raises the scientific bar: the paper is no longer judged on its narrative, but on the evidence it ships. Claims that cannot be traced to an artifact become explicit weaknesses; claims that can be traced become verified strengths. That is the standard of scientific reproducibility we need.
Confidentiality applies: the package is run locally, and never uploaded to a third-party service.
Limitations
- Ceiling on artifact-based checks. The highest-value step (cloning and reading real artifacts) only fires when something is actually shared. A paper with no code, no data, and no repo link degrades this back toward close reading plus domain skepticism — still better than a direct “review this,” but without its strongest source of leverage.
- Time cost. This is a long conversation, not a single prompt.
- Domain pushback is only as good as the reviewer’s own judgment. The agent can hold a claim up against general practitioner expectations, but it has no privileged access to what’s normal in a subfield.
- The write-up itself may leak identity. Notes, examples, or shared process docs drawn from a real review may allow someone reverse-engineer the reviewer.