Bold editorial type card on cream. Kicker reads NEWS DAY . FACULTY in sage small caps. Headline reads The Rubric Could Not See It in charcoal serif over three lines. Subtitle in sage italic reads Write the row a polished answer cannot earn. Gold and sage circles at right, Harvest Kernel sage seedling mark at lower left.
|

Your Rubric Rewards Polish. The Students Who Thought Better Scored the Same.

Causal reasoning training produced more original ideas. The rubric gave them no higher score.

That sentence is the finding, and it comes from OpenAI’s own research arm.

On August 27, 2026, OpenAI published results from a randomized experiment run with researchers at Bocconi University. More than 1,000 first year undergraduates worked a real business case, building marketing recommendations for the university’s own merchandise store. Students were randomly assigned by class period into one of four groups: access to ChatGPT, training in causal reasoning, both, or neither.

Trained human graders scored every submission on a five point rubric. Separately, the researchers ran automated text analysis on the same submissions, measuring how many ideas each one contained, how varied those ideas were, how much causal reasoning showed up in the writing, and how closely the work resembled recommendations written by three experts.

Two instruments, pointed at the same stack of student work. That is the part worth your attention, and it is not the part the headlines picked up.

What the rubric saw

The rubric saw ChatGPT clearly.

Students with access scored almost a full point higher on the five point scale. Their answers carried more ideas, followed clearer logic, and looked more like what an expert would have written. OpenAI’s own framing is that AI helped novices produce work that looked more professional.

Note the verb. Looked.

The researchers are careful to say students were not simply handing the assignment over. They still had to decide what to ask, judge the responses, and pick what went into the final submission. That is real work and it is the work most of us actually want students doing. But the rubric could not tell the difference between a student who did that work well and a student whose draft simply arrived better organized. It recorded the polish either way.

What the rubric missed

Now the other arm of the experiment.

The critical thinking group received training in causal reasoning. Not AI training. An exercise involving a game, examples, questions and feedback, teaching students to link cause to effect and to explain why a proposed solution might work or might fail.

Those students explained more clearly why their ideas might work and when they might fail. And they did not score higher on the rubric.

The rubric, as the researchers describe it, measured how well the recommendations addressed two standard marketing goals: increasing awareness and increasing use of the university store. A well organized answer that hit both goals scored well. Nothing in the instrument asked whether the idea was one nobody else had.

The text analysis found what the rubric could not. Across that group, students produced a wider range of ideas that were more distinct from what their peers submitted. The gain was real, it was measurable, and it was invisible to the thing the grade came from.

Read that once more with your own course in mind. A documented improvement in student thinking produced no change in the grade, because the grade was never looking there.

Do this before your next stack lands. Open one assignment rubric and audit it row by row for the criterion a polished answer cannot earn. We built the manual method as a free working session: The Missing Row: a one assignment rubric audit. You finish holding one new criterion, written in your own language, ready to paste into the rubric you already use.

The Berkeley study said the same thing from the other side

Two weeks ago we wrote that AI grade inflation is a measurement story, not a cheating story. The source was Igor Chirikov’s Artificial Intelligence and Grade Inflation, UC Berkeley Center for Studies in Higher Education Working Paper 26-3, May 2026. It analyzed more than 500,000 grades at one large research university from 2018 to 2025 using a difference in differences design.

In courses with more AI exposed tasks, writing and coding heavy work, the share of A grades rose by 13 percentage points after ChatGPT’s release, about 30 percent relative to the 2022 baseline. The increases were larger where homework carried more weight, which the paper reads as consistent with AI substituting for student work rather than broad learning gains.

Berkeley looked at half a million real grades and found the scores going up. Bocconi ran the controlled version and found out why: the instrument rewards the thing AI is best at. One study is observational and one is randomized, they were run by different people with different incentives, and they agree.

The uncomfortable half is that the vendor’s study is the one that names the limitation of the vendor’s product. OpenAI’s own write up says that if AI can help students produce polished, expert like work, then looking only at the final answer tells us less about what a student actually understands.

Your rubric is a hypothesis and it just got tested

Every rubric you have ever written is a claim about what matters. Each row says: this is a thing I can see from the outside, and seeing it tells me something is happening on the inside.

That claim held for a long time because producing a clear, well structured, logically coherent answer used to require the understanding underneath it. The correlation was strong enough to grade on. It was never the thing itself, and now it has come apart.

This is not a story about student ethics and it does not get fixed by policing. Both studies point at the instrument. The Bocconi students with ChatGPT access were doing legitimate assigned work. Their scores still went up in a way that told the grader less than it used to.

The fix is not a longer rubric. It is one honest row.

The one row your rubric is missing

Here is the manual method, and it takes about twenty minutes on one assignment.

Take the rubric for the next major assignment you will actually grade this term. Read each row and ask one question of it: could a well organized, entirely conventional answer earn full marks on this row? For most rows the answer is yes, and that is not a failure. Organization and coverage are real criteria and they should stay.

Then count the rows where the answer is no. Rows that require a specific choice the student made and had to defend. Rows that ask why this approach and not the obvious one. Rows that reward an idea the reader has not already seen twenty times in the same stack.

If that count is zero, you have found it. Write one row. Keep it short, keep it observable, and anchor it to something the student has to supply themselves: the alternative they rejected and why, the assumption their recommendation depends on, the condition under which their answer would be wrong.

That last one is the causal reasoning move the Bocconi training taught, and it is the one their rubric had no place to record.

Then, and this is the part that makes it stick, tell the students the row exists before they start. A criterion nobody knows about is a trap. A criterion announced in week one is a curriculum.

That is one assignment

You did the audit, you wrote the row, and the rubric for that assignment is now measuring something a polished answer cannot fake.

Then you look at the rest of the term. Four sections, six assignments each, a program review that wants alignment evidence for every one of them, and every rubric you have inherited from the last person who taught the course.

Doing that by hand once is a good afternoon. Doing it across a whole course is the part that never happens, which is why most rubrics are still the ones somebody wrote in 2019.

The Rubric Builder in the Harvest Kernel Faculty Toolkit does that specific job. It reads your Course Record, which is the five minute profile of what you teach and how you sound, and drafts precise, measurable criteria in your language rather than generic ones. Review first, always. Nothing goes to a student until you approve it. Every criterion it proposes is a draft you edit, not a decision it makes.

Do the first one by hand anyway. You will write a better row than any tool will, and you will know exactly what to change in the twenty that come after it.

Sources

Open the rubric for your next major assignment. Which row rewards an idea nobody else had?

Similar Posts