Can a High AI Score Lower Your Grade? An Experiment With 214 Teachers
Being accused is one worry. A quieter one is that a high AI score sits next to your paper while it's being graded and nobody ever mentions it. A study published in July 2026 tested that situation directly, with one paper, four versions of an invented detection report and 214 teachers. The gap it found is big. So is the list of reasons not to read it as a forecast of your own grade.
HumanPen Team
· 9 min read
Can a high AI score lower your grade even if no one accuses you?
In one experiment it did, by a wide margin. 214 university teachers graded the same course paper after seeing a detection report. When the report said 7% AI, called the text low risk and highlighted nothing, the paper averaged 73.30 out of 100. When it said 87%, called the text high risk and showed four of the paper's sentences in red, it averaged 51.61. The paper itself never changed.
That doesn't mean your grade will drop by 22 points, or at all. The report was invented by the researchers, the paper was written in Chinese, the teachers graded it in an online survey rather than a real course, and only 16 of the 214 (7.5%) had ever used an AI detection tool. What the study does show is that a score and the verdict printed beside it can change how the writing itself gets judged, not only whether the reader suspects AI.
The study is "Automation bias in teachers' evaluation of student writing: effects of algorithmic warnings and visual risk cues in AI detection reports", by Peitao Du and Tingting Liu of Al-Farabi Kazakh National University and Xujin Xian of The Catholic University of Korea. It appeared in Frontiers in Psychology on 7 July 2026 and is open access. The authors report no funding and no commercial or financial conflicts of interest.
What the teachers saw
The researchers wrote one paper, "The Effects and Challenges of Online Education", pitched as a middling undergraduate essay: fluent, with a basic structure, conventional ideas and not much depth. Into it they placed four formulaic sentences. They had no obvious factual errors or logical breaks. They were just the kind of stock phrasing that can read as generic.
Then they built four versions of a report from a detector that doesn't exist, called "AcademicCheck":
| Group | Report said | The four sentences in red? |
|---|---|---|
| A | 7% AI | No |
| B | 7% AI | Yes |
| C | 87% AI | No |
| D | 87% AI | Yes |
Each teacher was randomly given one version, with 52 to 55 teachers per group. They looked at the report summary first and read the paper second. Every report also carried a short verdict along with the number. The 7% versions said "Low risk" and described "predominantly human-authored writing characteristics"; the 87% versions said "High risk" and "a high likelihood of AI-assisted writing." In the red versions the report also pointed to the highlighted passages as needing manual review, so "red" here means color plus a prompt, not color alone. When they were done, teachers were told the report was simulated.
What changed when the number went up
| Group | Report | Average paper score (0–100) | Likely to step in (1–7) |
|---|---|---|---|
| A | 7%, no red | 73.30 | 3.40 |
| B | 7%, red | 71.87 | 3.72 |
| C | 87%, no red | 60.64 | 4.89 |
| D | 87%, red | 51.61 | 5.66 |
The 87% report did most of the work. Swapping the 7% "Low risk" report for the 87% "High risk" one, with nothing highlighted, cost the paper about 13 points on average. And teachers didn't just rate it as more likely to be AI-written. They also gave lower ratings for its originality, its language and its logical structure, even though the text in front of them was word for word the same.
The red sentences mattered most when the report already said high risk. At 7%, highlighting four sentences didn't move the 0–100 score by a statistically significant amount. It did pull down how teachers rated the language, which fits: those were the sentences they had just been pointed at. At 87%, the same highlighting took roughly another 9 points off.
Teachers said they would do more. The last column asked how likely they were to require revisions, remind the student about writing independently or following AI-use rules, or ask the student to explain how the paper was written and sourced. It went from 3.40 to 5.66 on a 7-point scale. That's what teachers said they would do, not something anyone was observed doing.
Why this isn't a forecast of your grade
The authors call their own results "preliminary and context-specific". Here is why that label fits.
- The report was fictional and deliberately extreme. "AcademicCheck" was invented so that teachers' opinions of real tools wouldn't get in the way. The two numbers were picked to be far apart. The authors say the contrast "was intentionally strong" so they could see whether teachers react to a clearly low versus a clearly high warning, and that applying the results to real-world scores in between calls for caution.
- One paper, written in Chinese. Every teacher read the same Chinese-language essay. Its "medium-quality" level was the researchers' own call, and they say it wasn't checked with a pilot study or outside raters. A strong paper or a weak one might not behave the same way.
- Most of the teachers had never used a detector. They were recruited online through the survey platform Wenjuanxing and teacher communities, for a small payment, and about two-thirds taught social sciences. Only 16 (7.5%) had used AI detection or AI plagiarism-checking tools. Neither the article nor its supplementary file breaks the results out for those 16.
- It was a survey, not a gradebook. Teachers had the report and the paper and very little else, and their scores counted for no one. The authors' own conclusion is that the findings "should not be generalized directly to all real grading contexts."
- The split is unusually clean. In every group, most teachers' scores sat within about 5 points of their group's average, yet the averages for A and D ended up almost 22 points apart. The authors themselves say an effect this large "should be interpreted cautiously." They put it partly down to the stark 7%/87% contrast and to how little context teachers had: "In real teaching, teachers may see intermediate rates and have richer information about students and writing processes, so actual bias may be smaller."
What it does say about real grading
Set the specifics aside and one finding still stands: the teachers saw the report before they read the paper, and the report changed how they read it. The authors' explanation is that a number like "AI content: 87%" turns the job from judging the paper into checking whether the system got it right. One of their suggestions for universities is a two-stage process in which teachers "complete their quality evaluation before reviewing the detection report".
Most of our readers deal with Turnitin rather than a made-up tool, so three things Turnitin documents are worth putting next to the study.
- You usually can't see what your teacher sees. Turnitin's FAQ says "only instructors and administrators are able to see the indicator", though an instructor can download the AI report as a PDF and share it. Can Students See Their Turnitin AI Score? Student vs Instructor Views goes through who sees what.
- Low scores come without highlights. Turnitin says "no score or highlights are attributed for AI detection scores in the 1% to 19% range." The report shows an asterisk instead. So group B's combination, a low number with sentences marked, isn't something a current Turnitin report would show you. Under 20%, there are no AI-marked sentences to steer anyone's eye. More in What does the asterisk (*%) mean on a Turnitin AI score?
- Turnitin says the number isn't a grade. The same FAQ says the percentage "should not be used as the sole basis for action or a definitive grading measure by instructors." Nobody in this experiment was told to use the score as a grading measure, though. They were asked to grade a paper, and the report came along with it. Is a Turnitin Score a Grade? covers where a grade is supposed to come from.
So the study can't tell you whether your own teacher was swayed, or by how much. What it does is move "the report might color how my paper gets read" from paranoia to a reasonable thing to ask about.
What you can do about it
None of these steps needs anyone to have accused you of anything.
- If a grade looks off, ask about the criteria, not the score. "Which rubric criteria lost marks, and which passages show it?" is a normal question after any grade. Highlighting also changed what teachers wrote back. When asked what they would tell the student to revise, 65.4% of teachers who saw the 7% report with red sentences suggested language edits, against 7.5% who saw it without red, and suggestions about developing the content fell from about half to almost none. (The authors treat this coding of open answers as exploratory.) If your feedback is all about wording and says nothing about your argument, ask about the argument too.
- Ask whether an AI report was used, and if so, ask for the PDF. A percentage alone tells you very little. The report shows which passages were marked.
- Have your process ready to show. The follow-ups this study asked about weren't penalties but requests: revise this, explain how you wrote it, show your sources. Teachers shown the 87% report with red sentences leaned clearly toward making them. Drafts, notes and version history let you answer that kind of request quickly.
- If you're still revising and have a report, read the marked passages first. These were the four sentences marked in red in the experiment, in the authors' English translation: "With the rapid development of information technology, online education has gradually become an important component of higher education." "This paper analyzes this issue from three aspects: learning effects, interaction modes, and technical conditions." "In conclusion, online education brings convenience but also faces many practical challenges." "Therefore, this issue deserves continued attention and further research from the education community." Once marked, lines like these dragged down how the language was rated. Openings, roadmaps and conclusions that could sit in anyone's essay are worth putting in your own words anyway.
Where HumanPen fits
If you have a Turnitin or iThenticate report, you're still revising and your course allows this kind of help, HumanPen's humanize option works on the flagged passages. Import the Turnitin or iThenticate report and rewrite only the passages it flags, returned as the same editable file. Go through the returned file yourself before you hand it in.
Frequently asked questions
Did the teachers know the report was fake? The researchers didn't tell them while they were grading. Afterwards, they were told it was simulated and wasn't a real academic integrity judgment.
Was this a study of Turnitin? No. The report came from an invented tool, "AcademicCheck", so that opinions about real detectors wouldn't skew the results.
Did the red highlighting matter at 7%? Not for the 0–100 score, where the difference wasn't statistically significant. It did lower teachers' ratings of the paper's language.
Turnitin shows an asterisk below 20%. Could that have the same effect? This study didn't test it. Its low condition was the number 7% plus a "Low risk" verdict, not an asterisk. An asterisk does tell the teacher that something was detected, so how teachers read it is something this study cannot answer.
Sources
- Peitao Du, Tingting Liu and Xujin Xian, Automation bias in teachers' evaluation of student writing: effects of algorithmic warnings and visual risk cues in AI detection reports, Frontiers in Psychology 17:1889402, published 7 July 2026; read 2 October 2026 (sections 3–6, Tables 1 and 3; supplementary file, Appendix A, Tables B1 and F4).
- Turnitin, Turnitin's AI writing detection capabilities FAQs.
KEEP READING