How Often Is a Master's Thesis Flagged as AI? A 2026 Study of 1,163 Theses

A research team at Vrije Universiteit Brussel ran one detector over every master's thesis their faculty received in 2024–2025. It flagged nearly half, and for about three in four of those it estimated half the text or less. The same team tested four detectors on papers whose authorship they knew. Here is what they found about human-written work, what the thesis figures do and do not mean, and what to do if your own thesis is flagged.

HumanPen Team

· 6 min read

How often is a master's thesis flagged as AI?

In one faculty's 2024–2025 master's theses, Pangram flagged 529 of 1,163, or 45.5%. For the flagged theses, the median estimate was that 30% of the text was AI-generated, and only 1.9% were estimated above 80%. The researchers had no way to check those figures against how the theses were actually written. In a separate controlled test with known authorship, Turnitin, Pangram, GPTZero and Copyleaks all scored every human-written paper between 0 and 20%.

The study is "Who wrote this? Evaluating the reliability of AI detection tools in higher education", by Marijke Van Vlasselaer, Filip Van Droogenbroeck and Bram Spruyt of Vrije Universiteit Brussel, published open access in the International Journal for Educational Integrity on 29 June 2026. The authors declare no financial or institutional connection with any of the tools they tested.

The controlled test: 160 papers with known authorship

The researchers built four sets of 40 English academic papers, each at least 4,000 words long:

  • Human-written. Master's theses by non-native English-speaking students, submitted between 2017 and 2019 and drawn from an institutional archive.
  • Fully AI-generated. Complete papers produced by a ChatGPT model.
  • Hybrid. The human papers with their literature review and conclusion replaced by AI-generated text.
  • Rewritten hybrid. The hybrid papers after the AI sections had been rewritten by a single ChatGPT prompt asking for text that "sounds more like it was written by a real person".

They then compared each tool's score with the known share of AI text in each paper.

On the human-written papers, all four tools put every paper between 0 and 20%. Pangram and Copyleaks returned 0%. Turnitin showed an asterisk rather than a number for three of them; Turnitin's own documentation uses the asterisk for scores between 1% and 19%, which we explain in what the asterisk on a Turnitin AI score means. GPTZero gave some of them a small score. The authors set this against a 2023 study of non-native writers: "Unlike Liang et al. (2023), who reported substantial false-positive bias against non-native writers using seven early-generation GPT detectors, we found no false positives for Pangram, Copyleaks, and Turnitin". That earlier research is covered in why non-native English writers get flagged more often. The authors write that false positives were "almost absent", which surprised them given earlier research, and add a caution of their own: "some tools may use conservative thresholds to avoid falsely labelling human-written text as AI-generated, which may increase the number of false negatives."

On the papers that did contain AI text, the tools disagreed sharply. Pangram's estimates came closest to the known share across the three AI categories. The other three tools, Turnitin among them, badly underestimated the fully AI-generated papers, all of which came from the one model used for that set; the authors link the gap to its being the newest model available at the time. On the hybrid papers, Pangram placed 37 of 40 within the correct range, Turnitin 24, Copyleaks 12 and GPTZero none.

One Turnitin detail matters for reading their figures: the researchers counted papers with a Turnitin asterisk in the lowest band. Our comparison of Pangram and Turnitin also discusses this study.

The real-world run: a year of master's theses

Because Pangram had tracked the known AI content most closely, the researchers used it to scan all 1,163 master's theses submitted in the academic year 2024–2025 across the programmes of their faculty of Social Sciences and Economics.

ResultFigure
Theses scanned1,163
Theses flagged with any AI-generated content529 (45.5%)
Median estimated AI share, flagged theses30%
Middle half of flagged theses17% to 49%
Range, flagged theses7% to 100%
Flagged theses estimated above 80%1.9%

The authors read the distribution this way: "the majority showed low-to-moderate AI usage, with very high AI usage (> 80%) being rare."

What the thesis figures do not tell you

The authors are explicit about the limits:

  • No ground truth. Nobody knew how the 1,163 theses were written, so "the flagging percentages cannot be interpreted as confirmed prevalence rates of AI use." The authors still call the analysis "a valuable first indication", but the figures are Pangram's estimates, not confirmed rates.
  • One faculty, one tool. The theses came from social sciences and economics at one university, and only Pangram was run on them.
  • A snapshot. The AI text in the test was generated in May 2025, and the authors say their results "are valid only for a specific snapshot in time."
  • One rewriting method. The "rewritten hybrid" set used a single ChatGPT prompt. The authors note that students may combine several methods, so the results for that set may not carry over.

The human-written papers were also a particular group: master's theses from one faculty, submitted between 2017 and 2019. A clean result on those papers is a good sign, not a guarantee for every thesis.

If your thesis is flagged

  1. Check which detector the score came from. This study found large differences between tools on the same papers. A Pangram estimate, a Turnitin percentage and a GPTZero score are not interchangeable.
  2. Check whether the tool can read your whole thesis. Turnitin, for example, does not produce an AI writing report for submissions over 30,000 words; see can Turnitin check a whole thesis.
  3. Look at which sections were flagged. Where the report highlights passages, those are the ones to be ready to explain, with your drafts and version history.
  4. Ask how the score will be used. The authors' own position is that a flag "should trigger a more thorough review of the student's paper content, rather than relying on an AI detection tool to draw a final conclusion."

For a second 2026 study that measured how often AI-polished human writing is flagged, see does AI polishing get your writing flagged.

Where HumanPen fits

When the flagged part of a thesis is a set of passages you need to rework, HumanPen's humanize option is built for that job. Import the Turnitin or iThenticate report and rewrite only the passages it flags, returned as the same editable file. Your university's rules on AI assistance come first, and every returned chapter is worth a full read before submission.

Frequently asked questions

Did any detector wrongly flag the human-written papers? Not above 20%. All four tools put every human-written paper between 0 and 20%. Turnitin showed an asterisk for three of them and GPTZero gave some a small score; Pangram and Copyleaks returned 0% for all. The authors caution that conservative thresholds can reduce false positives at the cost of missing more AI text.

Does 45.5% mean nearly half of master's students used AI? Not as a confirmed figure. It means Pangram estimated some AI-generated content in 45.5% of the theses. The authors read it as "a valuable first indication" that AI-assisted writing is already common, but with no ground truth they say the percentages "cannot be interpreted as confirmed prevalence rates of AI use."

Which detector was most accurate? In this test, Pangram's estimates came closest to the known AI share. That is one study, with one set of papers generated in May 2025; detectors and AI models have changed since.

Was Turnitin tested on the real theses? No. Only Pangram was run on the 1,163 theses.

Sources

KEEP READING