Editorial card reading Not Worse. Just Average. with the subtitle What LLMs did to federal grant proposals, beside a diagram of scattered dots converging tightly into a single dense center, in the Harvest Kernel cream, charcoal, sage and gold palette.
| |

AI Did Not Make Grant Proposals Worse. It Made Them Average.

Researchers handed their federal grant proposals to a model. The proposals did not come back weaker. They came back closer to what the agency had already funded.

That is the finding in a paper published in Proceedings of the National Academy of Sciences on August 11, 2026, by Yifan Qian, Zhe Wen, Alexander C. Furnas, Yue Bai, Erzhuo Shao and Dashun Wang of the Center for Science of Science and Innovation at Northwestern University’s Kellogg School of Management. The team examined confidential NSF and NIH proposal submissions from two large US R1 universities, funded, unfunded and pending, alongside the full population of publicly released NSF and NIH awards.

Every AI writing study you have read this year measured quality. This one measured distance. And distance turns out to be the number that decides what a national research portfolio looks like in ten years.

The word the study uses is distinctiveness, and it is not a synonym for good

Here is the sentence that matters, from the authors’ own abstract: “Across both private submissions and public awards, higher LLM involvement is consistently associated with lower semantic distinctiveness, positioning projects closer to recently funded work within the same agency.”

Read that twice. The proposals were not sloppier. They were not less rigorous. They were not full of fabricated citations. They were closer to the middle.

A grant proposal is not graded like an essay. It is graded against a field. Its job is to occupy a position nobody else is standing on, and to convince a panel that the position is worth funding. Semantic distinctiveness is the measure of how far your idea sits from the ideas the agency already bought. Move toward the center and you have not written a worse proposal. You have written a proposal that is harder to distinguish from three others in the same pile.

That is a different failure than the one everyone was watching for. The academic conversation about AI and research integrity has been almost entirely about detection: fabricated references, invented data, text nobody wrote. Those are real and they are catchable. This one is not catchable, because nothing about it is wrong.

The split nobody is talking about, and it is not about quality

The second finding in that abstract is the one I keep coming back to. LLM use rises sharply beginning in 2023, and it “exhibits a bimodal distribution, indicating a clear split between minimal and substantive use.”

Bimodal means there is no middle. There is not a bell curve of faculty gently increasing their AI use. There are two camps. One barely touches it. The other hands over real portions of the writing. And the two camps are sitting in the same department, submitting to the same program officer, and telling themselves the same story about how careful they are being.

If you are in the first camp, you are probably reading this feeling vindicated. Do not. The study did not find that abstinence wins. At NIH it found close to the opposite.

Before your next proposal, abstract or course description goes into a chat box: we built a free twenty minute working session called Twenty Minutes Before You Prompt. You do it by hand, on one real piece of your own writing, and you finish holding a Distinctiveness Brief you paste above your next prompt. No account, nothing saved, nothing sent anywhere.

The asymmetry is where it gets uncomfortable

The consequences of heavier LLM use were not the same at both agencies. At NIH, higher LLM involvement was positively associated with proposal success and with higher early stage publication output. At NSF, no comparable association was observed.

So at one of the two largest funders of American science, the proposals that drifted toward the center got funded more often. The system rewarded the drift.

Then comes the line that should stop a department chair mid sentence. Those NIH productivity gains, in the authors’ words, “are concentrated in nonhit papers rather than the most highly cited work.”

More proposals funded. More papers published. Fewer of them landing anywhere. The individual researcher’s numbers go up. The field’s ceiling comes down. The authors put it plainly: AI “can expand individual scientists’ productivity and impact while simultaneously contracting the collective focus of science.”

That is not a story about cheating. That is a story about a measurement system quietly paying people to converge.

The agencies already noticed, and they wrote it down

On July 17, 2025, NIH published Guide notice NOT-OD-25-132, “Supporting Fairness and Originality in NIH Research Applications.” The policy language is not ambiguous: “NIH will not consider applications that are either substantially developed by AI, or contain sections substantially developed by AI, to be original ideas of applicants.”

The background section of that same notice explains what triggered it, and the detail is worth carrying. NIH recorded evidence that AI tools had enabled individual Principal Investigators to submit more than 40 distinct applications in a single application cycle.

Forty. From one investigator. That is what a blank box plus a topic plus a deadline produces at scale, and it is why NIH also capped applications per PI per calendar year in the same notice.

Why AI grant proposals drift toward the center

Here is the part that gets treated as a mystery and is not one.

A language model is a machine for returning the center of a distribution. Ask it about a topic with no other input and it gives you the most probable version of writing on that topic, which is the average of everything already written. Point that machine at a funding proposal and it returns the average of what funding proposals look like. In a system where the training signal is public award abstracts, the average of what proposals look like is, quite literally, what already got funded.

The model is not failing. It is doing the one thing it was built to do, on an input that contained almost nothing but the topic.

Dean has been making this argument on the Harvest Kernel side for a while, and he is not the only one. In a content masterclass he keeps in his research vault, the speaker draws the boundary as cleanly as anyone has: AI is “very bad at generating new thought. It’s very good at compressing a ton of stuff down.” Compressing, combining, but not ideating. Another card in the same vault puts the prescription next to the diagnosis: “if you want unique answers, you have to feed AI unique perspectives and constraints.” And Jeff Su, asked where he draws his own line, says he never lets AI touch ideation, “because that’s where I can inject my personality and point of view.”

Three people from three different worlds, all landing on the same rule, roughly a year before a PNAS paper measured it across a national funding pipeline.

What actually changes on Monday

The conclusion here is not stop using it. Nobody is going back, and the NSF and NIH data show the second camp is already large.

The conclusion is that the distinctiveness of your output is set almost entirely by the distinctiveness of your input, and most people are putting in a topic.

Prompt and pray is the habit: open the box, type the subject, accept the shape that comes back. It feels like using AI. It is closer to asking the field to write for you.

The alternative takes twenty minutes and it is entirely manual. Before you prompt, write down what only you have. The observation from your own lab that has never been published. The reason this problem is yours and not the field’s. The constraint your institution puts on the work that nobody else in the pile is under. The failed version you ran in 2023 and what it ruled out. None of that is in the training data, because none of it has been written down yet.

Then prompt. The model still compresses and combines, which is what it is good at. But now it is compressing your material instead of the field’s, and the output moves away from the center instead of toward it.

That is the whole idea behind the Course Record on the teaching side of Harvest Kernel: a short profile of what you actually teach and how you actually sound, written once, read by every tool afterward, so the draft comes back in your voice rather than in the average voice of higher education. Same principle, different artifact. The research version is a paragraph of what only your lab knows.

The question the study leaves open

The authors are careful and they stay descriptive. They report associations, they name the limits, they do not claim the model caused the funding outcomes.

But there is a question sitting underneath their data that nobody has an answer to yet. If heavier LLM use pulls proposals toward what an agency already funded, and the agency keeps funding them, then next year’s model trains on this year’s awards. The center moves toward itself.

Nobody has measured that loop. It has only been running for three years.

Go and find the last thing you wrote that a model helped with. Not to check it for errors. Read it and ask a different question: is there a single sentence in here that only you could have written?

For faculty: the Course Record is the five minute profile that makes every tool in the Faculty Toolkit draft from your material instead of the field’s average. It is the first thing to build and it takes one sitting.

For everyone else: the conversation about where the line goes is happening in the Harvest Kernel Learning Community, free to join.

Sources

  • Qian, Y., Wen, Z., Furnas, A. C., Bai, Y., Shao, E., and Wang, D. “The rise of large language models and the direction and impact of US federal research funding.” Proceedings of the National Academy of Sciences, August 11, 2026. doi.org/10.1073/pnas.2601439123
  • NIH Guide Notice NOT-OD-25-132, “Supporting Fairness and Originality in NIH Research Applications,” released July 17, 2025. grants.nih.gov
  • Skillen, R. “Large language models reduce originality of research proposals.” Times Higher Education, August 12, 2026. timeshighereducation.com

Similar Posts