
Banning researchers from using AI to assess the merits of scientific papers doesn’t make much difference to the conclusions the reviewers reach – in large part because many of them ignore the prohibition and use AI anyway.
A team led by computer scientists at Microsoft Research ran a large randomised experiment during the 2026 International Conference on Machine Learning (ICML), one of the world’s biggest AI conferences, which took place in Seoul, South Korea in July.
Some 24,661 papers were submitted to the conference, requiring oversight from 17,886 reviewers. When researchers submitted their work, they could choose how it would be reviewed. The first option was to have it assessed by reviewers using a conservative policy that entirely banned the use of large language models (LLMs) to help with reviewing. The other option was to have the assessment performed by reviewers using a permissive policy that allowed the use of LLMs to help understand papers, check related work and polish reviewers’ own writing, but not to judge the merits of a paper or draft a review of the work.
The team led by Microsoft Research scientists discovered that papers reviewed under the two policies had almost identical acceptance rates – 27 per cent under the stricter policy and 26.5 per cent under the permissive one – while average review scores were 3.31 and 3.32 out of 6, respectively. Reviewer confidence was also virtually unchanged. Reviews written under the permissive policy were around 5.5 to 7 per cent longer, and were rated as slightly higher quality by experts, although a reviewer-by-reviewer comparison found no meaningful difference.
The reason for that lack of difference became clear after an anonymous survey of 1486 of the reviewers who took part in the exercise. There, 22.5 per cent of reviewers who were instructed not to use AI admitted to using an LLM anyway. “Self-reported non-compliance of 23 per cent is definitely higher than what I was anticipating,” says Miro Dudík at Microsoft Research, who was involved in the study and was also one of the conference organisers.
Those who used AI when told not to did so to brainstorm feedback, draft review text, read the papers and summarise their strengths and weaknesses. “Our post-conference survey reveals some reasons why reviewers might end up violating policy, like large reviewing workloads and insufficiently clear rules,” says Dudík.
“This study shows that banning AI in peer review is very difficult to enforce,” says Kayvan Kousha at the University of Wolverhampton, UK. “Because of the workload of academics, there is a temptation” to use AI, says Kousha.
In a separate part of the study, the researchers analysed the reviews using Pangram, an AI text detector. They found that only 52.2 per cent of reviews under the conservative policy were classified as fully human-written, compared with 37.0 per cent under the permissive policy – though the researchers warn AI text detectors aren’t perfect.
“I believe our work has relevant insights for other computer science conferences – and potentially venues in other fields as well – that are facing similar challenges,” says Sunnie S. Y. Kim at Microsoft Research.