What Evidence Is Used in AI Misconduct Cases? A Study of 1,162 Case Files

If your work has been flagged for AI and you are waiting to hear what happens next, the useful question is what will actually be put in front of the person deciding. A study published in September 2026 went through three years of case files at one Australian university and counted. Here is what was in them, how much weight the researchers gave each kind of evidence, and what that suggests you prepare.

HumanPen Team

· 12 min read

What evidence actually shows up in an AI misconduct case?

At one Australian university, detector scores were a small and shrinking part of it. Across 1,162 AI-related misconduct cases from 2023 to 2025, outputs from AI-detection tools such as GPTZero made up 5.8% of the evidence logged in 2023, 8.8% in 2024 and 0.5% in 2025, which was five items in a year with 618 such cases. Over the three years, detector output appeared in 67 of the 1,162 cases. Far more common were admissions made in the student interview (in 34.3% of cases), references that could not be found (30.6%) and a marker's judgement that the writing read like AI (20.6%). The type and strength of the evidence only modestly explained how cases ended.

The study is "How strong is the evidence in generative AI-related academic misconduct allegations? A mixed-methods analysis", by Albert Munoz, Mercedez Hinchcliff, Cameron Langfield and Ann Rogerson of the University of Wollongong. It appeared in the International Journal for Educational Integrity on 2 September 2026 and is open access. The case files come from their own university and the authors work there. Rogerson is on the journal's editorial board, which the paper declares as a competing interest.

What the researchers looked at

The university logged 5,127 academic misconduct cases between January 2023 and December 2025. The team pulled out the 1,162 that involved suspected use of generative AI, 22.7% of the total. That share went up every year:

202320242025
AI-related cases214330618
Share of all misconduct cases that year11.9%20.0%36.9%

Being in this set does not mean a student was found to have done anything. The paper is direct about it: "This study does not adjudicate whether individual students engaged in the misuse of generative artificial intelligence."

Inside the files they found 1,855 separate pieces of evidence and sorted them into 15 types. Each piece was rated on three scales taken from legal scholarship: relevance, credibility, and inferential force, which is how strongly an item supports the allegation on its own. Those ratings are the researchers' judgement, applied across the corpus by a locally run language model that they calibrated against their own coding. They are not a test of whether any individual piece of evidence was correct.

What was in the files

All 15 types, by how many cases they appeared in. One case can hold several kinds of evidence, so the percentages add up to more than 100.

Evidence typeCases (of 1,162)Researchers' rating of inferential force
Admission made in an interview399 (34.3%)Strong for 97.2% of items
Cited references that could not be found356 (30.6%)Moderate for 98.0%
Text the marker thought read like AI239 (20.6%)Weak 70.6%, moderate 29.4%
Real sources cited for things they do not say196 (16.9%)Weak 53.6%, moderate 27.1%, strong 19.3%
Turnitin or other similarity report71 (6.1%)Weak for 100%
AI-detection tool output (examples given: GPTZero, ZeroGPT)67 (5.8%)Weak 92.5%, moderate 7.5%
Writing unlike the student's level or earlier work65 (5.6%)Moderate 51.4%, weak 48.6%
Student could not explain the work in an interview63 (5.4%)Moderate for 100%
Allegation with no evidence described at all58 (5.0%)Not rated
Prohibited behaviour recorded during an online exam51 (4.4%)Strong 68.4%, moderate 29.8%
Near-identical submissions from several students51 (4.4%)Weak 50.9%, moderate 41.8%, strong 7.3%
Exam proctoring records (logs, screen captures)37 (3.2%)Strong 86.1%
File metadata or system timestamps26 (2.2%)Moderate 67.9%, weak 32.1%
Written admission10 (0.9%)Strong for 100%
A named third party confirming sources were fake1 (0.1%)Moderate

References that could not be found were the fastest-growing type, from 10.4% of the evidence logged in 2023 to 30.0% in 2025.

Detector scores: rare, then almost gone

Share of that year's evidence items202320242025
AI-detection tool outputs5.8%8.8%0.5% (5 items)
Turnitin or other similarity reports4.0%5.5%3.1%

The percentages are of evidence items, not of cases, and the paper prints a count only for 2025. For context only, not as a denominator: the number of AI-related cases was 214, 330 and 618 in those three years. Over the three years, AI-detection tools account for 69 items in 67 of the 1,162 cases.

Before reading "detectors don't count" into this, two things.

The ratings were set, not measured. Every AI-detection item was rated low credibility, and that was a rule the researchers applied going in: the paper says these outputs "were treated in the present analysis as low-credibility evidence", and its credibility scale was built to account for "documented limitations such as AI detector false positive rates". The study did not check whether any score was wrong. If you want to see what a small error rate turns into across a whole university, what a 1% false positive rate means when a university submits 75,000 papers does that arithmetic.

The drop also lines up with a local rule. From 2025 the university's assessment policy has barred staff from uploading work to outside detectors. In its current wording, staff "are not permitted to upload student work to third party tools, including generative artificial intelligence (GenAI) or misconduct detection software". The same clause still allows "approved detection software" that the university supports, and similarity reports were still 3.1% of the evidence in 2025. The university's current misconduct procedure also lists software the university may use, including text-matching (Turnitin) and "language analysis software (e.g. Turnitin Authorship)". The authors read the decline as staff accepting that detector outputs cannot meet the balance of probabilities standard. Your university will have its own rule, and some are collected in 19 universities that limit AI detector use or evidence.

One thing the paper does not spell out is where Turnitin's AI writing score was filed. The similarity-report type is defined as text-matching evidence, and its example is a report showing "a similarity score of 1%, which is unusually low" for that kind of assignment. The AI-detection type names GPTZero and ZeroGPT.

The common one: "it reads like AI"

Among evidence drawn from the prose itself, the most frequent type was a marker's sense that it read like AI, in 239 cases. The paper's example of how staff put it: "generic transitional phrases, unnaturally uniform paragraph structure, and unsolicited definitions of basic terms not required by the task."

The researchers rated most of these items weak, because those features are "not exclusive to GenAI misuse." Their share of the evidence went from 19.5% of items in 2023 to 14.5% in 2025.

A smaller type compares the work with what the student normally hands in (65 cases), and its ratings split almost evenly between moderate and weak. The paper's background section names the problem with it: students' writing changes over time, "which can make legitimate improvement difficult to distinguish from anomalies attributable to misuse." If that is the claim you are facing, old assignments as a writing baseline goes through where a comparison with old assignments holds up and where it gets pushed back on.

Much of the case is made in the interview

The type that appeared in more cases than any other was an admission made during an interview (399 cases), rated strong almost every time. A second interview type, a student not being able to explain their own work, appeared in 63 cases. It was rated moderate across the board, because it is "consistent with GenAI use but equally consistent with alternative explanations such as poor comprehension or test anxiety".

The authors draw the conclusion themselves: "in a substantial proportion of cases, the strongest evidence available emerged only after the institution had already determined the allegation was sufficient to proceed."

It also moves fast. The median time from allegation to final outcome was 12 days (the middle half of cases took 6 to 20). Under the university's current coursework misconduct procedure, you are offered an interview, and before it the subject coordinator has to make sure you are emailed "the summary of the allegation and evidence that has been collected", and you can bring a support person, though they cannot represent you. The university's student guidance warns that "you may be asked to explain how you developed an assessment item".

None of this is advice to say less. If you did use AI in a way the task did not allow, falsely flagged for AI: how to appeal, and what evidence your university will accept has a section on that situation, and it comes down on the side of being straight about it. If you did not, the interview is where you get to show how the work was made.

Did stronger evidence decide the outcome?

Not clearly, in this data. The researchers tested whether the type, number and rating of evidence items lined up with how far a case escalated and how it ended. They found some associations, but "the explanatory power of evidence characteristics across all model specifications is modest".

Their explanation is the procedure itself. Escalation follows a staff member's finding, not the evidence ratings. Whether a case ends in a misconduct finding or in "poor academic practice", which carries an educational rather than a disciplinary response, tends to depend on things like whether it was a first offence. Dismissals at the committee stage can come from a student's account, a procedural problem or welfare considerations, and the case files do not capture what was argued at the hearing. In the authors' words there is "no minimum evidentiary threshold at any stage of the pipeline".

That last point showed up in the numbers. Fifty-eight cases (5.0%) had no evidence described beyond the allegation itself. Those empty allegations did become rarer, from 7.6% of evidence items in 2023 to about 2.5% in 2025.

Where this study stops

  • One university, three years. The University of Wollongong, 2023 to 2025. The authors say the patterns may not hold elsewhere.
  • Files, not hearings. "The evidence records capture what was documented in the case file, not what was argued, tested, or weighed during the hearing."
  • No breakdown by student. Undergraduate or postgraduate, domestic or international, and discipline are not in the data, so it cannot say who is accused more often.
  • The ratings are a framework. They come from the authors' own scheme, applied by a language model calibrated on 12 cases and then checked on 54 cases it had not seen, where its agreement with two human coders was statistically indistinguishable from the two humans' agreement with each other.
  • Insiders. The authors are employees of the university whose files they analysed, which they say "introduces potential bias, especially during the calibration exercise."
  • A rule changed midway. The 2025 restriction on uploading work to third-party detectors sits inside the period studied.

If you have been flagged or called to a meeting

  1. Ask what each piece of evidence is. A score, a passage someone thinks reads like AI, a comparison with your earlier work, a problem with a source: get them listed one by one. In 5.0% of the cases here, there was nothing beyond the allegation. Your procedure should say what you are entitled to see before the meeting. The rest of the preparation is in academic integrity meeting after an AI accusation.
  2. If a detector score is in the file, find out which tool it was and whether your university allows staff to use it.
  3. If the claim is about how the text reads, bring the record of how you wrote it. Drafts, notes, and the file's version history. How to keep version history in Word, Google Docs and Overleaf explains where each of those lives.
  4. Be ready to explain the work. Why the argument is structured the way it is, what your main sources actually say, what changed between drafts. In 63 cases here, not being able to explain was itself written down as evidence.
  5. Put your evidence in at the first stage. At this university, an appeal against a finding is only possible for a lack of due process or "new and substantial evidence that has not previously been considered." Check your own procedure for the equivalent line.

Where HumanPen fits

Nothing here helps with a case that is already open. That case is about the version you submitted, and rewriting it afterwards does not change what is in the file. This section is for an earlier point: you are still revising a draft, you have a Turnitin or iThenticate report for it, and you have checked that your subject outline allows an outside AI tool to reword your text, not only one named tool, and how it wants that help acknowledged. Wollongong's student guidance, for one, tells students to "use only the permitted gen AI technology", and counts use beyond the assessment instructions as misuse. If all of that holds, this is the case HumanPen was built for. Import the Turnitin or iThenticate report and rewrite only the passages it flags, returned as the same editable file. Check the result against your own draft before you submit, because you should be able to explain every sentence you hand in.

Frequently asked questions

Do universities still use AI detector scores as evidence? At this one, rarely by 2025: five items, 0.5% of the evidence logged that year, after a policy barred staff from uploading work to third-party detectors. Other universities set their own rules, so this does not tell you what yours does.

Is "it reads like AI" enough on its own? The researchers rated most of those items weak, because the features involved are not exclusive to AI use. That is their framework, not the university's rule. At this university the decision is made on the balance of probabilities, across all the evidence in the case. The paper spells out what that means: it has to be more likely than not that misconduct occurred, "with the burden of proof resting with the institution and the academic raising the allegation rather than the student".

In this study, what counted as the strongest evidence? Admissions, written or made in an interview, the records kept by exam proctoring software, and prohibited behaviour those records show. Each of these was rated strong for most of its items.

Does the study show whether the accused students had used AI? No. It counts and rates what was in the files. Inclusion in the 1,162 cases means AI misuse was suspected, not that it was found.

Sources

KEEP READING