AI grade inflation editorial card reading Your Rubric Cannot See It, kicker News Day Faculty, subtitle Two studies. The score went up. The substance did not.
| |

AI Grade Inflation Is Not a Cheating Story. It Is a Measurement Story.

At one large research university, A grades rose thirteen percentage points. The evidence of more learning never showed up. That was the smaller of the two studies.

AI grade inflation is the story most faculty have already heard, and it is the one people argue about in committee. The bigger study landed this week, and it was not about students at all. It was about professional scientists writing grant proposals for federal money, which is about as high stakes as writing gets. Same pattern. The score went up. The substance did not.

Put those two findings next to each other and you get a problem that has almost nothing to do with cheating, and almost everything to do with what our instruments are built to notice.

What the AI grade inflation numbers actually say

Start with the classroom, where AI grade inflation shows up first. Igor Chirikov at Berkeley’s Center for Studies in Higher Education analyzed more than 500,000 grades at a large research university from 2018 to 2025, using a difference in differences design. In courses with more AI exposed tasks, meaning writing and coding, the share of A grades rose by thirteen percentage points, about thirty percent relative to the 2022 baseline.

That is AI grade inflation, measured rather than asserted, and the paper is titled exactly that. The detail that matters most is not the size. It is where the increase concentrated. Grade increases were larger in courses where homework carried greater weight. Chirikov reads that as AI substituting for student work rather than broad learning gains from AI.

Now the grants. On August 12 a team from Northwestern’s Center for Science of Science and Innovation published in the Proceedings of the National Academy of Sciences. Dashun Wang, Yifan Qian and colleagues looked at about 5,700 confidential grant proposal submissions and 131,000 publicly released awards from the National Science Foundation and the National Institutes of Health, covering 2021 to 2025.

Proposals showing stronger signs of AI assisted writing were four percentage points more likely to receive NIH funding. They were also measurably less distinctive. Moving from low to high model involvement corresponded to roughly a five point drop in distinctiveness percentile at NSF and about four points at NIH. And the funding gains concentrated in lower impact publications rather than highly cited research. At NSF, no significant association appeared at all.

So AI grade inflation at the undergraduate level means better grades without better learning. At the federal level, better odds without better ideas.

Nobody in the second study was cheating

This is the part that should stop a faculty meeting cold.

Using a language model to draft a grant proposal is legal. It is normal. In many labs it is encouraged, because the alternative is a postdoc losing a weekend to formatting. Not one person in that PNAS dataset broke a rule.

And the outcome was still a measurable loss of distinctiveness across the national research portfolio.

That tells you the problem is not sitting in anyone’s ethics. It is sitting in the instrument. Both the rubric and the review panel were designed for a world where producing clear, complete, conventionally structured prose was itself weak evidence that someone had thought carefully. For a very long time that inference held, because fluency was expensive. It is not expensive anymore.

The working session

Do it once, on one assignment, by hand.

The Distinctiveness Criterion is a free twenty minute working session. You bring one real assignment prompt and its current top rubric row. You leave holding a rewritten prompt and one criterion that measures something a model cannot compress its way into. Nothing is saved and nothing is sent.

Start the working session

The mechanism, stated plainly

Here is the cleanest way I have heard the underlying capability described, and it is not from an education researcher. It is from someone who writes for a living: AI is very good at compressing a large amount of material down, and very good at combining things, and genuinely bad at generating new thought. Compressing, combining, but not ideating.

Hold that against your rubric.

Organization. Clarity. Completeness. Coverage of the required elements. Correct use of the source material. Appropriate academic register. Every one of those is a compress and combine task. Every one of them used to cost a student four hours, and cost them four hours of thinking as a side effect. That side effect is what we were actually grading. We just never had to say so, because you could not get the artifact without the thinking.

Now you can. So every criterion on your rubric that measures a compress and combine skill has quietly become free, and the points attached to it no longer carry information about the person.

Which leaves an obvious question, and it is the only question worth taking into next week: what is left on your rubric that could not be produced by compressing and combining?

For most assignments the honest answer is very little. Sometimes nothing.

The one thing that cannot be borrowed

There is a category of thing a model cannot compress or combine its way into, because it does not exist in the training data and never will. It is the specific.

The reading your student did of the data set your class collected on Tuesday. The argument that only holds if you sat through the third week of your course. The case from the clinical site your program actually places into. The counterexample from your own city, your own term, your own lab notebook. The claim that contradicts the assigned reading and has to defend itself.

None of that is available to compress. It is only available to a person who was there.

That is not a detection strategy and it is not a ban. It is a design move, and it is a small one. You are not rebuilding a course. You are changing what one criterion rewards, so that the score starts carrying information again.

The Berkeley finding tells you where to aim it, too. The inflation concentrated where homework carried the most weight. That is the exact place a single criterion change buys the most back.

Why the grant study is the one that should worry you

AI grade inflation is uncomfortable. The funding finding is structural.

Grades are an internal signal. If they drift, employers discount them, students notice within a couple of years, and the market corrects roughly. The federal research portfolio does not correct roughly. It compounds. Proposals that read as more fundable get funded, those projects produce the next round of proposals, and if the selection pressure is running slightly toward the conventional, that bias compounds across the entire national research agenda for a decade before anyone can see it in the outcomes.

Dashun Wang put the stake in one sentence. “Science advances by exploring ideas that don’t yet look obvious.”

An instrument that rewards polish will systematically fund the obvious. Not because anyone chose that. Because polish is what it can see.

What this actually asks of you

Not a policy. Not a tool ban. Not a detection subscription.

It asks you to look at one assignment, the one whose grades you already half distrust, and find the single place where a strong answer could only come from someone who was in your room. Then make that the criterion with the most points on it.

That is one assignment. You teach three more sections of it, plus two other preps, and the syllabus for the spring course landed in your inbox last week.

Which is the actual reason this stays undone in most departments. The move is not hard. The move is small. It is just that nobody has an afternoon to do it forty times.

The thing the studies did not measure

Both papers measured what happened to the score. Neither one could measure what happened to the student who wrote a genuinely strange, genuinely first hand paragraph and got marked down for organization.

That student is still in your class this fall. Right now your rubric cannot see them either.

Where this goes next

Rewriting one criterion is a twenty minute job. Rewriting forty of them across three preps is the reason it never happens. The Assessment Builder inside the Harvest Kernel Faculty Toolkit drafts assignments, rubrics and tagged question banks against your Course Record, so it already knows your course, your outcomes and your voice before it writes a word. You review and export. It is one stage of the Course Creation Pathway and it is the stage that does this exact job.

See the Assessment Builder in the Faculty Toolkit

If you would rather compare notes with other instructors doing the same rewrite this term, that conversation is happening in the free Harvest Kernel community.

Sources

  • Igor Chirikov, Artificial Intelligence and Grade Inflation, CSHE Higher Education Working Paper Series Vol 26.3, University of California Berkeley, May 13, 2026. cshe.berkeley.edu
  • Yifan Qian, Zhe Wen, Alexander C. Furnas, Yue Bai, Erzhuo Shao and Dashun Wang, The Rise of Large Language Models and the Direction and Impact of US Federal Research Funding, Proceedings of the National Academy of Sciences, August 2026. Preprint and Northwestern Now
  • Large language models reduce originality of research proposals, Times Higher Education, August 12, 2026. timeshighereducation.com

Similar Posts