Rhetorical Influence AI Review: Same Paper, Different Score in 4,200 Tests

4 hours ago 7
rhetorical influence AI review

A growing body of scientific writing is now judged, at least in part, by artificial intelligence, and a new study suggests that the words a researcher chooses can shift an AI reviewer’s verdict even when the underlying data never change. Researchers examining rhetorical influence on AI review found that framing, not evidence, can quietly tilt scores up or down in large language model-based peer review systems, raising fresh questions about how reliable machine judges really are.

The concern isn’t hypothetical. Peer review, the decades-old system where researchers vet each other’s work before publication, is already buckling under a flood of submissions. A recent Frontiers journals survey cited by Ars Technica found that over half of peer reviewers already lean on AI tools somewhere in the review process, often just to keep up with the volume. That reliance is exactly what makes the question of rhetorical bias in AI review so pressing: if machines are already grading papers, and word choice alone can move their scores, the stakes for scientific fairness go up fast.

That’s the gap a team led by Ming Li set out to measure. Their paper, submitted in August 2026 and built around ICLR 2026 submissions, isolates rhetoric from substance to see exactly how much presentation can sway an AI-based judgment, and the results paint a picture of bias that is structured, uneven, and in some cases surprisingly predictable.

Key takeaways

  • Rhetorical framing alone measurably shifted AI review scores even though the underlying scientific content stayed identical.
  • Researchers built a corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions to test the effect.
  • Evidence framing and novelty stance drove the biggest swings in overall assessment scores, more than any other rhetorical dimension.
  • Papers with lower original AI scores tended to rise after rhetorical rewriting, while higher-scoring papers tended to fall.
  • A stricter review protocol cut average scores by 1.36 points but did not reliably reduce the underlying rhetorical sensitivity.

Study Overview and Methodology

The research team designed a controlled experiment specifically to strip away every variable except how a paper is worded, isolating rhetorical sensitivity as the one thing that changes between versions of the same manuscript. That design choice is what gives the findings their weight: any score difference between manuscript versions can be traced to presentation rather than to genuine differences in the science.

Controlled Manuscript Corpus from ICLR 2026 Submissions

To test this cleanly, the researchers built a corpus of 4,200 full-paper manuscripts derived from 120 anonymized submissions to ICLR 2026, one of the field’s major machine learning conferences. Each original paper was rewritten into multiple rhetorical variants, creating a large enough dataset to spot patterns that a handful of papers couldn’t reveal.

Rhetorical Rewriting and AI Review Protocols

Two separate large language model rewriters altered six distinct rhetorical dimensions of each manuscript, pushing each one in opposite directions, more confident versus more cautious framing, for instance, while leaving the reported scientific content untouched. Five different LLM reviewers then evaluated all the resulting manuscripts, once under a standard review protocol and once under a stricter one, so the team could compare how rhetoric played out across different evaluation conditions. The study also tested joint, recursive, and reviewer-guided rewriting passes to see whether combining strategies amplified the effect.

Impact of Rhetorical Choices on AI Review Scores

The headline finding is that rhetorical sensitivity in AI-based peer review is not random noise, it follows a clear hierarchy, with some stylistic choices mattering far more than others. That structure is what turns this from an anecdotal quirk into something evaluators need to actively design around.

Key Rhetorical Dimensions: Evidence Framing and Novelty Stance

Among the six rhetorical dimensions tested, evidence framing and novelty stance produced by far the largest positive-negative contrasts in overall assessment scores. Scope framing formed a weaker second tier, while the remaining dimensions showed smaller and less consistent effects. In other words, how confidently a paper claims to be novel or how assertively it presents its evidence can matter more to an AI reviewer than several other stylistic factors combined.

Score Movement Patterns Based on Original Review Scores

The direction of the shift depended heavily on where a manuscript started. Papers that originally received lower AI review scores tended to climb after rhetorical rewriting, while papers that already scored highly tended to drop. The clearest directional contrasts, meaning the sharpest gap between a positively and negatively reframed version of the same paper, showed up in the middle range of scores, where a manuscript’s fate is arguably most contested.

Complexity and Dependence in Rewriting Workflows

One might assume that stacking rewriting techniques would produce bigger score swings, but the data didn’t support that. More elaborate workflows, including joint, recursive, and reviewer-guided rewriting, did not reliably yield larger gains than simpler approaches. Joint rewriting effects turned out to be strongly dependent on which rewriter model was used, and giving reviewers explicit guidance did not consistently outperform a simple, unguided second look at the manuscript. Repeated rewriting passes brought diminishing and configuration-dependent returns rather than steadily compounding advantages. Across every condition tested, the rewriter mainly determined how far apart the opposing rhetorical variants ended up, while the reviewer determined the size and direction of the resulting score change.

Evaluation Protocol Effects and Implications

Tightening the review rules lowered scores across the board, but it didn’t fix the underlying rhetorical bias, which is arguably the more important takeaway for anyone building or relying on AI review systems.

Effect of Strict Review Protocol and the Case for Robust Evaluation

Switching to a strict review protocol lowered the mean overall assessment score by 1.36 points compared to the standard protocol. That’s a meaningful drop in absolute terms, yet it did not consistently reduce how sensitive the AI reviewers remained to rhetorical framing. Stricter rules made reviewers harsher overall, but not necessarily fairer in the sense of resisting stylistic manipulation.

That distinction matters for anyone tracking the rise of AI-based peer review. If tightening the rubric doesn’t neutralize rhetorical sensitivity, then simply asking AI reviewers to be more critical isn’t a fix on its own. The findings point toward a different kind of solution: evaluation systems that are explicitly built to be robust against content-preserving rhetorical variation, meaning systems that judge a paper the same way regardless of whether its evidence is framed confidently or cautiously, as long as the underlying science hasn’t changed.

For a field already worried about volume, burnout, and inconsistent human reviewers, as reporting from Ars Technica on the broader peer review crisis has documented, the temptation to lean further on AI graders is only going to grow. This study is a reminder that automating the judgment doesn’t automatically remove the bias, it just changes where the bias comes from.

FAQ

How do rhetorical choices influence AI-based peer review scores?

Rhetorical choices such as evidence framing and novelty stance significantly affect AI review judgments while the scientific content remains unchanged.

What dataset was used to analyze rhetorical sensitivity in AI review?

A controlled corpus of 4,200 manuscripts derived from 120 anonymized ICLR 2026 submissions was used for the analysis.

Do more complex rewriting strategies lead to better AI review scores?

No, more elaborate rewriting workflows did not reliably yield larger gains in AI review scores.

What is the effect of using a strict review protocol on AI review scores?

The strict review protocol lowered the mean overall assessment by 1.36 points but did not consistently change rhetorical sensitivity.

Article produced with the assistance of artificial intelligence and reviewed by the editorial team.

Read Entire Article