Over 30% of new papers on arXiv now read as AI-written, according to a study of 12,750 papers from 2021 to 2026. The rise is sharp, reaching 39% in early 2026, but the numbers are not without caveats. A detector calibrated to flag only 0.4% of pre-ChatGPT papers as machine-written sets the baseline, revealing a stark increase in AI-generated content across fields like computer science, where 65% of recent papers show signs of AI involvement.
This article breaks down how the study was conducted, what the data shows, and why some fields still show low AI writing rates. You’ll get a clear picture of what works in measuring AI-generated academic text, and where the methods fall short.

The Hidden Reality: AI-Written Papers Are Already a Major Presence on arXiv
A third of new arXiv papers now appear to be AI-written, but the numbers tell only part of the story. The study used a detector calibrated to flag just 0.4% of pre-ChatGPT papers as machine-written, setting a baseline that highlights the real surge in AI-generated content. This method avoids overestimating the impact, but it still reveals a dramatic shift, especially in computer science, where 65% of recent papers show AI involvement. The challenge lies in interpreting these results accurately, as some fields, like mathematics, show negligible AI use, raising questions about detection limitations and the nature of the content itself.

How the AI Detection Methodology Was Built and Calibrated
Calibrating the AI detector for academic writing
The AI detection tool was specifically trained on academic writing to ensure it could distinguish between human and machine-generated content accurately. It was tested extensively on a range of disciplines, ensuring it could adapt to the unique language patterns found in different fields. This training allowed the tool to focus on the structure and style of academic papers, rather than generic text, improving its precision in identifying AI-generated content.
Setting a 0.4% false-positive rate as the baseline
To establish a reliable baseline, the detector was calibrated to flag only 0.4% of pre-ChatGPT papers as machine-written. This threshold was chosen based on a control group of papers from 2021 and 2022, ensuring that the tool did not overestimate the presence of AI-generated content in pre-LLM literature. This approach avoids inflating the results and provides a clearer picture of the actual rise in AI-written papers after the introduction of large language models.
The Data Sample and Scope of the Study
Sampling 12,750 papers across ten field groups
The study sampled 12,750 papers from ten distinct field groups, with roughly 25 papers per field per month, spanning from January 2023 to July 2026. This approach ensured broad representation across disciplines, from computer science to mathematics. Control data from 2021 and 2022 provided a baseline to measure changes over time. The selection process avoided over-reliance on any single area, ensuring the results reflect trends across multiple domains.
Using version-1 PDFs to avoid contamination
To prevent newer text from influencing older analyses, the study used version-1 PDFs of each paper. This means that even if a paper was revised in 2026, its 2023 version was analyzed using the text from that year. This method avoids contamination and ensures that the data reflects the language and structure of the paper as it was originally published. The full body text, rather than abstracts, was analyzed to capture a more accurate signal of AI involvement.

Key Results: AI-Written Papers by Field and Over Time
Computer science leads with 65% AI-written papers
The study shows that computer science is the most affected field, with 65% of recent papers flagged as AI-written. This is a dramatic shift from the pre-LLM control level of just 0.2%. The rise is steep and consistent, reflecting the field’s heavy use of AI tools for writing and research. Other fields like quantitative biology and electrical engineering also show significant increases, but none match the scale seen in computer science.
Mathematics shows a stark contrast with only 0.7% flagged
By contrast, mathematics remains the outlier, with just 0.7% of papers flagged as AI-written. This is far below the pre-LLM control level of 0.0% and suggests that AI-generated content is either rare or difficult to detect in this field. The limitations section of the study explains that the low value may be due to the nature of mathematical writing, which relies more on symbols and structured logic than natural language patterns.
Where the Measurement Breaks: Limitations and Misinterpretations
The false-positive floor and its implications
The study’s calibration sets a 0.4% false-positive rate for pre-ChatGPT papers, but this doesn’t eliminate all misclassification. If the detector flags 30% of new papers as AI-written, but also marks 0.4% of pre-LLM papers, the real change is 29.6%, a nuance that gets lost in headlines. This floor prevents overestimating AI’s role, but it also means the numbers are not absolute. The method avoids false positives in historical data, but it can’t account for all variations in writing style or tool evolution.
The control years of 2021 and 2022 anchor the study, but they don’t capture all pre-LLM writing nuances. Some human papers might still be misclassified, especially in fields where writing styles are more varied. This means the 30% figure is a best guess, not a definitive measure of AI’s impact on academic writing.
Why some fields show low AI detection rates
Fields like mathematics and high-energy physics show low AI detection rates, but this doesn’t necessarily mean AI isn’t being used. The structure of mathematical writing, with its dense formulas and minimal prose, may be harder for current detectors to flag. The study acknowledges this, noting that the low value in mathematics is hard to interpret due to the nature of the field.
Other fields, like computer science, have high AI detection rates, but this could reflect both increased use of AI tools and the field’s writing style, which is more text-heavy and easier for detectors to analyze. These differences highlight the limitations of a one-size-fits-all detection approach.

Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.
What This Means for Researchers and Institutions
Understanding AI’s role in research workflows
AI is no longer a novelty in academic writing, it is a tool being used at scale. The study’s calibration shows that 65% of new computer science papers on arXiv now read as AI-written, meaning researchers are integrating AI into drafting, editing, and even generating content. This shift requires a new mindset: AI is augmenting, not replacing, human effort. Institutions must train researchers to recognize AI-generated content and ensure proper attribution.
Tools that flag AI writing are improving, but they are not perfect. The 0.4% false-positive rate for pre-ChatGPT papers shows that even the best detectors can misclassify. Researchers should treat AI detection as a guide, not a definitive answer. Transparency in how AI is used, whether in drafting, data analysis, or ideation, is critical for maintaining trust and academic integrity.
What the future might look like for AI in academic writing
The rise of AI in academic publishing is not a passing trend. With 32% of new papers flagged as AI-written in the most recent quarter, institutions must prepare for a future where AI is a standard part of the research process. This includes updating policies on authorship, data handling, and peer review to account for AI’s growing role.
Fields like computer science are leading the charge, but others are catching up. As AI tools improve, they will likely become more integrated into workflows, raising questions about originality, authorship, and quality. Researchers and institutions must act now to define clear guidelines and ensure that AI enhances, not undermines, the integrity of academic work.
Looking Ahead: The Road to More Accurate AI Detection
Improving detection accuracy across disciplines
The current study shows that AI detection tools are still field-specific. For example, mathematics papers show only 0.7% AI involvement, but this may be due to the nature of the field rather than a lack of AI use. Future tools must be trained on more diverse writing styles and disciplines to avoid blind spots. The calibration method used in the study, setting a 0.4% false-positive rate for pre-ChatGPT papers, provides a solid foundation, but it needs to be refined for every academic domain.
The need for transparency and control
Researchers and institutions must demand transparency in AI detection tools. If a paper is flagged as AI-written, it should be clear why and how confidently. This is especially important as AI writing becomes more common. Tools that allow users to adjust detection thresholds or understand the reasoning behind flags will be more useful. Without this, the value of AI detection remains limited, even with high accuracy.
Source: unslop.run