Falsely flagged for AI: how to appeal, and what evidence your university will accept

You wrote it. A percentage says otherwise, and now there is a meeting in your calendar. This page covers what to do in the first hour, which evidence has actually persuaded adjudicators in published cases, the arguments that hand the other side a win, and a letter you can copy.

HumanPen Team

· 20 min read

The short answer

This page assumes one thing: you wrote the work. If you did not, most of what follows will make your position worse rather than better, and there is a short section near the end about that case instead.

Preserve your writing record before you reply to anything, then ask in writing for three things: the exact rule you are said to have broken, every piece of evidence the decision-maker will see, and the procedure and deadline. A detection score is not proof on its own. What actually decides these cases is a dated record of how the document came to exist, matched to the specific passages that were flagged.

Two organisations say the useful part of that out loud. Turnitin's own guide states that its AI writing detection model "may not always be accurate (it may misidentify human-written, AI-generated, and AI-paraphrased text), so it should not be used as the sole basis for adverse actions against a student", and adds that determining misconduct "takes further scrutiny and human judgment in conjunction with an organization's application of its specific academic policies". The Office of the Independent Adjudicator, the independent student complaints scheme for England and Wales, puts the burden where it belongs: "The responsibility is on the provider to prove that the student has done what they are accused of doing, not on the student to disprove it."

That does not mean you sit back and let them prove it. It means the evidence you bring is there to give a decision-maker something concrete to weigh, not to discharge a burden you never had.

The first hour

Five things, in this order. The first three are time-sensitive in a way the rest are not.

  1. Do not touch anything. Not the submitted file, not the folder it lives in, not the drafts, not the file names. Re-saving a document changes its modification date, and a folder you tidied on the day you were accused is the single easiest thing for the other side to raise. Make a copy if you want to annotate. Work in the copy.
  2. Find out whether you have version history at all. This is the question that decides what kind of case you can make, and most people assume the answer is yes when it is no. Microsoft's support documentation is blunt about it: "Version history in Microsoft 365 only works for files stored in OneDrive or SharePoint in Microsoft 365." If you wrote your essay in Word on your laptop's hard drive, there is no version history to produce. Google Docs keeps one, but you need the right access: "To browse earlier versions of a file, you need permission to edit that file." If it lives on a course account or a shared drive you could lose access to, check today.
  3. Capture what you have while you still have access. Course accounts get closed, shared drives get revoked, and a version history is not something you can download as a file. Take dated screenshots or exports in whatever form your institution will accept, and note where the original still lives so it can be inspected directly if that is asked for.
  4. Write your own timeline today, from memory. When you picked the topic, where you were working, which sources you read first, what you changed after feedback and why. This is not evidence. It is the index you will need in a week, when you are trying to work out which of forty files matters. Memory for this decays fast, and it can decay before anyone asks you: in one of the published OIA cases the student argued that so much time had passed between submission and their viva that it affected what they could recall about their own work.
  5. Get the procedure and the deadline, and find your adviser. The academic misconduct procedure is a document your institution publishes. Read it before you write a word of reply, because it tells you what stage you are at, what the decision standard is, and how long you have. Students' union advisers do this work every week and it costs you nothing.

Work out what you are being accused of

"The AI detector flagged your essay" is not a charge. It is a reason someone opened a file. Before you defend anything, get the allegation into a form you can answer.

Ask for these four things in writing, and keep the request neutral:

  • The specific provision of the academic misconduct policy you are said to have breached. In one case the OIA upheld, part of its recommendation was that the provider review its regulations so AI-related offences sat in the core policy, and so that notifications told students which offence they faced and why.
  • All the evidence the decision-maker will see, including the detection report itself. The OIA's casework note says students "should also be given sufficient notice of meetings and be provided with all relevant evidence, including detection software reports, to allow them to respond effectively to allegations." The same note tells decision-makers to "understand the strengths and limitations of detection software, and weigh this evidence carefully against other available information."
  • Which tool produced the report, which version, and when it was run. Vendors change models. A report from last year and a report from this morning are not the same measurement of the same thing.
  • Which part of the work is in question. The OIA calls this good practice: providers should "consider whether the whole or part of a submission is thought to be AI-generated, and clearly set out what aspect of the assessed work led to the suspicion that AI had been used inappropriately."

One practical note about the report. Turnitin's FAQ states that the AI writing detection indicator and report are not visible to students, and in the same breath that "with the PDF download feature, instructors can download and share the AI report with students". So it exists in a form that can be handed to you. Ask for it instead of reasoning from someone's description of what was highlighted.

Two things about the number will stop you arguing about the wrong thing. Turnitin says it analyses only what it calls qualifying text, meaning prose written in standard grammatical sentences, and that the resulting percentage "is not necessarily the percentage of the entire submission". Separately, for AI detection scores in the 1% to 19% range it displays no number at all, only an asterisk, which it attributes to avoiding potential incidence of false positives in that band. If your report shows an asterisk where the percentage should be, that is what you are looking at.

The evidence, ranked by what it can carry

The ordering here is ours. It is based on what the OIA's published AI cases turned on, not on any official scale, and your institution's procedure may weigh things differently.

EvidenceWhat it can showWeightHow to get it, and what stops you
Version history of the submitted documentThe text accumulated over time rather than arriving wholeStrongest, when it existsGoogle Docs keeps one but you need edit permission on the file. Word only keeps one for files stored in OneDrive or SharePoint. Local files have none
Dated drafts and outlines as separate filesThe same thing, in a form you controlStrongYour own storage, email attachments, LMS submission points. Do not rename or reorganise them
Notes and annotated readingsThat the claims in the essay came from sources you handledStrongYour own files. The OIA note tells providers "It can be helpful to explicitly ask students to supply copies of any notes or drafts", so offering them is not overreach
Feedback exchanged before submissionA third party saw the work in progress on a date you did not controlStrongSupervisor emails, LMS comments, seminar notes, a peer's markup
Your own explanation of the workWhy the question was framed that way, how sources were chosen, what changed and whenStrong, and often decisiveA meeting or viva. Prepare it as an account of your decisions, not a memory test
Earlier work of yours, for style comparisonThat this is how you writeMixedEasy to obtain, but it cuts both ways: task, deadline and language support differ between assignments, and so does the writing
Published research on detector error ratesContext about the toolWeak on its ownFreely available, and worth almost nothing detached from your case. It only does work when attached to your report and your writing record
A clean score from a second detectorVery littleNegativeDo not. See below

If rows one to four all come back empty, say so plainly rather than padding. A short honest answer ("I wrote this in Word on my own machine and did not keep drafts") plus a strong row five is a real position. An invented row one is not.

What the published cases actually turned on

Most advice on this subject is written from theory. The OIA publishes case summaries, so for providers in England and Wales there is a small set of decided cases showing what fair and unfair handling looked like. Three of them are about AI and academic misconduct.

[CS072501](https://www.oiahe.org.uk/resources-and-publications/case-summaries/ai-and-academic-misconduct-cs072501/), justified. The student gave the panel details of their essay preparations, planning and notes, and pointed out that detection software had previously flagged another of their essays incorrectly. The OIA was satisfied the panel had engaged with the substance of the essay and had not relied on the detection report alone. Its finding turned on what came next: "it didn't properly consider all the points the student raised or the evidence they presented, including their essay planning and preparation." Some of the evidence the panel relied on had not been shared with the student, so they had not had a fair chance to comment on how this essay compared with their others, and the decision gave neither sufficient reasons nor an explanation of why that penalty was the right one. The provider was asked to reconsider. The student later told the OIA that on reconsideration the provider decided there had not been any academic misconduct.

[CS072502](https://www.oiahe.org.uk/resources-and-publications/case-summaries/ai-and-academic-misconduct-cs072502/), justified. This one is about the viva. "None of the evidence relied upon by the provider was shared with the student ahead of the viva", the student was not invited to put in written mitigation or discuss the allegations in person although the procedure required it, and the viva itself tested subject knowledge but "did not give the student an opportunity to also explain how they had worked on the assignment". The student had also said the long gap between submission and viva affected their recall. The appeal was then dismissed without meaningful consideration of any of it.

[CS072503](https://www.oiahe.org.uk/resources-and-publications/case-summaries/ai-and-academic-misconduct-cs072503/), not justified. Read this one too, because it is the useful counterweight. Here the provider considered an appropriate range of evidence, including the student's disclosure at the viva, their draft notes and their written responses, followed its own procedures, and applied a penalty the OIA considered proportionate. The student's argument was essentially that honest disclosure at the viva and the effort they had put in should have counted for more. The complaint was not upheld.

The pattern across the three is not "the detector was wrong". It is procedural: was the evidence shared, was the student able to answer it, was a range of evidence considered, were reasons given. Those are the things you can ask for, and asking for them is not adversarial.

The one passage from Turnitin you should quote in full

If your writing is dense, repetitive in places, or heavily paraphrased from sources, this is the strongest thing you can put in front of a decision-maker, because it is the vendor describing your situation. It comes from Turnitin's detection FAQ:

"Sometimes false positives (incorrectly flagging human-written text as AI-generated), can include content without a lot of structural variation, text that literally repeats itself, or text that has been paraphrased without developing new ideas. If our indicator shows a higher amount of AI writing in such text, we advise you to take that into consideration when looking at the percentage indicated."

The second sentence does the work, and it is the one that gets dropped everywhere it is quoted. Turnitin is not conceding a flaw there. It is telling the person reading the report to discount the number for text of that kind. That is advice addressed to your marker, from the company that sold them the tool.

Only use it if it genuinely fits. A literature review, a methods section, a systematic description of a standard procedure, a paper written by someone working in their second or third language: those are the shapes it describes. If your essay is none of them, quoting it makes you look like you are reaching.

One more, shorter, that applies to everyone. Turnitin says elsewhere in its guidance that the score "is not meant to provide definitive answers in isolation", and that when educators use it "as a single data point rather than a definitive response, then it is being used as intended."

Four arguments that will backfire

"AI detectors are unreliable, therefore this one is wrong." This is the most common opening and the weakest. It invites the obvious reply: Turnitin publishes a target of keeping its false positive rate "under 1% for documents with over 20% of AI writing". Whether that figure holds up is a real question, but arguing about it puts you in a debate you cannot win in a twenty-minute meeting and drags attention away from the only ground where you are strong, which is your own writing record. Argue about your document, not about the category of tool.

Burstiness and perplexity. Half the internet will tell you that AI detectors measure sentence-length variation and word predictability, and that academic prose scores badly on both. Turnitin's FAQ says its model "is not explicitly programmed to evaluate specific signals such as 'burstiness,' 'perplexity,' or other individual metrics sometimes referenced in public discussions", and that it "learns statistical patterns from our training data" instead.

Be careful how far you take that, because it is narrower than it looks. The same page also says its classifiers are trained on differences in word probability, and perplexity is a measure of word probability. So the defensible version is: Turnitin denies computing those two named scores as explicit rules. It does not deny that word statistics matter. If you walk into a meeting and say the vendor has debunked the whole theory, somebody with the FAQ open can correct you in one line. The safer move is to say nothing about mechanism at all and spend the time on your drafts, where there is nothing to correct.

The 2023 release note about introductions and conclusions. This one circulates widely in appeal templates and it is a trap. Turnitin did publish a note saying it had observed a higher incidence of false positive detection in the first few or last few sentences of a document, often generic introduction or conclusion content. The very next sentence of that note says: "As a result, we have changed our detection logic to help reduce these false positives." It sits on the model release notes page under the heading "Updates to address our customers' false positive concerns", dated May 2023. It is an announcement of a fix, not a description of a current weakness. Anyone who opens the page you cited reads the next line, and you lose the room.

A second detector, or a rewrite. Running the essay through another checker to get a clean score proves nothing (detectors disagree with each other constantly), and a second flag hands the panel another number to point at. Rewriting the submitted essay after the allegation is worse. It looks like tampering, and it changes the one artefact everyone is trying to reason about.

A fifth, unofficially: do not speculate in writing about a marker's motives. It never helps and it stays on file.

The letter

Keep it under two pages. Attach an index of evidence rather than a folder. Send it in a way that gives you proof of the date. Everything in square brackets is an instruction to you, not text to send.

Subject: Response to academic misconduct allegation, [module code] [assignment title], student [number]
Dear [name or Academic Misconduct Panel],
I am writing in response to the notice dated [date] concerning [assignment title], submitted on [date] for [module code].
I wrote this assignment myself. I did not use a generative AI tool to produce any part of the text.
Tools I did use, with the specific feature in each case: [reference manager, spell checker, grammar checker, translation tool, statistical software, transcription. If none beyond the word processor, say so.]
Attached is a record of how the assignment was produced:
  1. [item] / [date or date range] / [what it shows]
  2. [item] / [date or date range] / [what it shows]
  3. [item] / [date or date range] / [what it shows]
In summary: I chose the topic on [date] after [reason]. My reading and notes date from [range]. The outline in item 2 was written on [date]. The first full draft (item 3) was written between [dates]. After [feedback / supervision / seminar] on [date] I changed [what] because [why]. The submitted version dates from [date].
On the specific passages that were flagged: [take them one at a time. For each, say why it reads the way it does: a standard formulation in the field, a close paraphrase of a source attached at item N, a methods description with limited room for variation, a passage I revised repeatedly. Do not answer with a general denial only.]
I would be grateful if you could confirm:
  1. the specific provision of [policy name] I am said to have breached;
  2. all evidence the decision-maker will consider, including the full detection report and any comparison material;
  3. which detection tool and version produced the report, and the date it was generated;
  4. the standard of proof that applies, and the route and deadline for appeal.
I would also ask that any decision set out the specific evidence it rests on beyond the detection score. [Only if the report came from Turnitin:] Turnitin's own guidance states that its AI writing detection model may misidentify human-written text and should not be used as the sole basis for adverse actions against a student.
I am happy to attend a meeting and to talk through how the assignment was written, and I can bring the underlying files.
Yours sincerely, [name] / [student number] / [programme] / [date]

Four things to check before you send it. Every attachment you mention is actually attached. Every date in it is right, because one wrong date undermines the rest. There is no speculation about anyone's motives. And you have read your own institution's procedure, because that is where the deadline and the response route come from, not from this page.

The meeting

Prepare it as an explanation of your decisions, not a memory test. Re-read the submitted version and the assignment brief. Know why the question was framed the way it was, which sources you leaned on and why, what a specific table or figure was built from, and what changed between drafts.

If a long delay, a disability, a communication difference or language needs affect how you recall or explain things, say so early and in writing, and link the request to whatever the procedure says about adjustments. The OIA's casework note tells providers to consider "whether assumptions about AI use could be biased against a student's writing style, for example if the student is disabled or has a communication difference, or if English is not the student's first language". That is a live consideration for the panel, not a favour you are asking for.

If you are asked something you do not remember, say you do not remember. Guessing wrong on a small detail costs more than the detail was worth.

How far this can go

Every institution's ladder is its own, and the only reliable description of yours is in your own academic misconduct procedure. The general shape:

  1. Informal or investigatory stage. A conversation, a request for an explanation, sometimes a viva. Often the whole thing ends here, and it is the cheapest place to produce your evidence.
  2. A formal panel or committee. A finding, and if there is one, a penalty. This is the stage that should give you reasons.
  3. Internal appeal. Usually on narrow grounds, most often procedural irregularity or new evidence, and always on a deadline.
  4. External review. Only after the internal route is exhausted.

For providers in England and Wales, step four is the Office of the Independent Adjudicator. Two mechanics matter. You will normally need a Completion of Procedures Letter, "which the provider will send you once you have completed its complaints or appeals procedures", and the complaint has to reach the OIA "within 12 months of the date of the provider's final decision (usually the date of the Completion of Procedures Letter)". The OIA describes itself as the independent student complaints scheme for England and Wales, so if you are studying elsewhere it is not your route.

If you are outside England and Wales, do not assume an equivalent exists, and do not assume it does not. Your institution's procedure has to state what happens after the final internal stage. That sentence is the one to find.

If you did use AI for part of it

Then the framing above is the wrong one for you, and borrowing it anyway is how people make their position worse. Two things from the record still matter.

Disclosure is weighed, but it does not erase a finding. In CS072503 the student disclosed their AI use at the viva and argued that should have counted in their favour. The OIA found the provider had considered that alongside the draft notes and written responses, had followed its procedures, and had applied a proportionate penalty. The complaint was not upheld. Being straight about it is still the better move, and it is the only move that survives contact with a viva, but it is not a defence.

The line between a permitted tool and a prohibited one is narrower than most people assume. Turnitin says its detector "is not tuned to target Grammarly-generated spelling, grammar, and punctuation modifications" and that in tests, "in most cases", changes made by Grammarly and similar grammar checkers were not flagged. The same answer then says this "excludes content generated by Grammarly's generative AI-powered features, including draft generation, paraphrasing, summarizing, and other features", and that content produced with those features "will likely be flagged as AI-generated by our detector". So "I only used Grammarly" is not one statement. Name the feature.

Frequently asked questions

Can I see the AI writing report myself? Not directly. Turnitin states that the AI writing detection indicator and report are not visible to students. It also states that instructors can download the report as a PDF and share it, so ask for that rather than working from someone's description of what was highlighted.

My report shows an asterisk instead of a percentage. What does that mean? Turnitin does not display a number for AI detection scores between 1% and 19%, and shows an asterisk instead. It attributes this to a higher incidence of potential false positives in that range.

If the report says 40%, does that mean 40% of my paper was AI-written? No. Turnitin says the model analyses only qualifying text, meaning prose in standard grammatical sentences, and that the percentage "is not necessarily the percentage of the entire submission". It also says a document containing several different writing types will show a disparity between the percentage and the highlighting, so the two not adding up is expected rather than a sign of a bug.

I used Grammarly. Am I in trouble? It depends entirely on which part of Grammarly and on what your assignment permitted. Turnitin says its detector is not tuned to target spelling, grammar and punctuation changes and that in most cases such changes were not flagged in its tests, but that content from Grammarly's generative features, including draft generation, paraphrasing and summarizing, will likely be flagged. Name the specific feature you used, and check the assignment brief as well as the general policy, since a brief can narrow what is allowed.

My essay was only 700 words. Does length matter? Turnitin says it does. In shorter documents of only a few hundred words the prediction is "mostly 'all or nothing'", because there is a single segment with no opportunity to overlap, and it adds that this means "some text that is a mix of AI-generated and original content could be flagged as entirely AI-generated." If your document is short, that sentence is worth putting in front of whoever is reading the score.

Should I run it through a different detector to prove my case? No. A clean score elsewhere does not establish anything, and a second flag gives the panel a second number. Spend the effort on your writing record instead.

I have no version history and no drafts. Is it hopeless? No, but your case now rests on being able to explain the work. That is the ground the published cases suggest is strongest anyway. Reconstruct what you can from things you did not create for this purpose: library borrowing records, emails, LMS access logs, seminar notes, messages to a classmate about the assignment.

Who actually makes the decision? People, applying your institution's policy. Turnitin says its score "is not meant to provide definitive answers in isolation" and should not be the sole basis for adverse action, and that determining whether misconduct occurred "takes further scrutiny and human judgment in conjunction with an organization's application of its specific academic policies". The number opens the file. It does not close it.

This page is general information, not legal advice. Academic misconduct procedures, evidence rules and appeal rights differ by country and by institution. The OIA material is directly relevant to higher education providers in England and Wales; elsewhere it is a useful example of what fair process looks like, not a rule that binds anyone.

KEEP READING