Was this forwarded to you by a friend? Sign up, and get your own copy of the news that matters sent to your inbox every week. Sign up for the On EdTech newsletter. Interested in additional analysis? Upgrade to the On EdTech+ newsletter.

Sometimes a bad academic study is just a bad academic study, particularly when its title is misleading. We should be able to admit that.

In Saturday’s post for premium subscribers, Morgan described just such a study that unfortunately is going to cloud an important topic—whether the use of course-specific AI tutors help or hurt learning, course grades, and course participation.

The study’s title sounds promising: “The Effects of Course-Integrated AI Tutoring on Student Performance and Engagement: A Randomized University Trial.” And the report’s introduction touts its first-of-a-kind nature.

❝

The current study presents findings from, to our knowledge, the first large-scale randomized experiment evaluating the effects of AI tutoring on student learning at a large U.S. public university.

The problem is that the report does not produce what it promises. It does not directly measure course-integrated AI tutoring, and while it may have included randomization, the treatment group did not take the same drug (i.e., AI tutoring). There may have been some useful limited findings included, but those will be buried beneath the headlines.

Much of today’s post draws on Morgan’s post, but we want to share the core analysis with a broader audience since the research itself is getting disseminated through media coverage. We also draw on a LinkedIn post from USC professor Stephen Aguilar (partially for the entertainment value).

The Four Flaws

There are three claims worth unpacking here: effects, AI tutoring, and student performance and engagement.

But first, what did they actually do? The paper explores the use of an AI-powered course-based chatbot, the Virtual Study Assistant (VSA), in a randomized controlled trial among 2,379 undergraduate students and 30 instructors across multiple disciplines at the University of Maryland. Instructors were randomly assigned to offer the tool or serve as controls within groups of identical or similar courses. The researchers also analyzed a smaller “exact match” sample restricted to sections of the same course.

Effects: The treatment group largely did not use the AI tool

Let’s start with the effects part, as described in the paper’s abstract.

❝

Among sections of the same course, tutor access reduced final grades by 0.37 standard deviations and learning management system participation by 0.90 standard deviations. Participation declines were similarly large in the full sample [snip]

That sounds consequential, but what, exactly, produced those effects? Only about 15% of students offered access ever used the tool—and only 13% in the same-course sample.

Aguilar described this problem quite well [emphasis added].

❝

Think of a drug trail, there's a placebo group who doesn't get the drug, and a treatment group who does.

In medical research the ENTIRE treatment group gets the drug.

In this experiment only 15% of the treatment group used AI.

Yet, EVERYONE in the treatment arm got lower grades. Not just the 15% who used AI. That's means the effect CANNOT be attributed to AI-tool use. You can't blame a thing that wasn't used for lower grades. That's like blaming a drug I didn't take for giving me a stomach ache.

To be clear, the average of the entire treatment group showed lowered grades, which is different than saying that everyone did. But the key point on such low adoption causing interpretation problems remains.

Nor did the typical user engage with the tool intensively throughout the semester. Among students who tried it, the average was approximately four sessions. About 40% had only one session, while a few heavy users pulled up the average. A session could contain multiple questions and responses, but this was hardly widespread, sustained use.

So how did an intervention that relatively few students used produce a substantial difference in average scores and recorded LMS participation? The authors offer possible explanations, rather than demonstrated mechanisms. Offering the tool may have encouraged broader AI use. Its availability could have signaled that AI use was acceptable, changing students’ use of other tools that the researchers did not track. The intervention may have changed instructor practices. Treatment instructors received guides, tutorials, and weekly summaries of tool use. Those resources could have changed how they taught, supported students, or discussed AI.

These possibilities matter because the intervention involved more than putting a chatbot in front of students. It also introduced information and signals that could affect instructors and students who never touched the tool.

The study therefore leaves us with a consequential finding and an unresolved explanation. Offering this particular intervention reduced average course scores and recorded LMS participation in the same-course sample. Whether using the tutor caused those declines remains unanswered. That distinction becomes important when readers interpret the result as evidence that using an AI tutor harms learning, as many people seem to be doing.

Tutoring: Not the same thing as providing direct answers

Second, what exactly was the “AI tutoring” being tested? The tool was a custom chatbot built on OpenAI technology that retrieved information from course materials in the LMS. It had two modes: direct instruction and tutoring. The first supplied answers; the second could guide students step by step toward an answer. The tool default was configured for direct instruction, and only two instructors chose tutoring mode.

Answering a course question is not, by itself, tutoring. Tutoring should attempt to find out what a learner understands, diagnose misunderstandings, and adapt explanations, hints, and practice accordingly. Some AI tutoring systems also maintain learner profiles across sessions, allowing that guidance to build on previous work. Putting course documents behind a chatbot does not automatically supply those capabilities.

Of the requests made to the tool, 73.8% sought information, explanations, or solutions; 11.3% involved practice or test preparation; and 0.7% sought feedback on students’ own attempts. The study therefore largely examined access to a course-integrated answer-giving assistant. And that “integration” really just meant technical LTI integration with the LMS, not integration into teaching methods. Yet “AI tutoring” gets top billing in the title. This is a shaky basis for presenting the results as evidence about tutoring. The findings deserve to be read as evidence about the assistant as configured, rather than generalized to more scaffolded AI tutoring.

The report called out this challenge.

❝

Overall, these findings suggest that instructors engaged with the VSA to some extent but did little to integrate it into their instruction, which may have contributed to the low student engagement described above.

Aguilar provided a useful reaction [emphasis added].

❝

Translation: We gave instructors limited support; to be honest we just turned the AI thing on and let instructors do whatever. Most of them stuck with business as usual.

EdTech isn't like some learning grenade you toss into ignorance and when it explodes people learn. EdTech was never a passive change-maker. EdTech is a tool, and how that tool is/isn't used is EVERYTHING.

Performance and Engagement: Indirect at best

Third, what did the researchers actually measure when they examined “student performance and engagement”?

For performance, the researchers looked at final course scores. Students offered the tool scored about three points lower out of 100 across the full sample, though the statistical evidence for that difference was weaker. When the comparison was limited to sections of the same course, the gap was about four points, with stronger statistical evidence that it reflected more than chance.

A four-point drop deserves attention. But the study does not establish that using the AI tutor caused students to learn less. It shows that students in sections offered the tool earned lower course scores, while leaving unresolved what caused that difference.

The engagement claim needs similar unpacking. The researchers measured recorded LMS participation, including assignment submissions, discussion responses, and quizzes. These are substantive course activities, not merely clicks. The count tells us that recorded participation fell. It does not tell us why, or whether, actual learning was affected.

What students do in an LMS depends on what their courses ask them to do there. A discussion-heavy course, a course built around online quizzes, and a course where most work happens elsewhere will generate different activity counts. Comparing sections of the same course helps, but does not make those counts a complete measure of engagement.

The study found fewer recorded LMS activities. That matters. What remains unclear is whether students disengaged from learning, shifted their work elsewhere, or responded to changes in teaching. The LMS can count activities. It cannot supply the missing course context, or explain what those activities meant.

A Bad Study Clouding Important Questions

Our disappointment with this research is with the distance between what the title and some of the paper’s language promise about AI tutoring and what the study actually establishes. The authors acknowledge important limitations, but readers still have considerable translation to do. And we are already seeing the reaction, including posts titled “The Hidden Downside: AI Tutors Tanked Student Grades – Here’s Why.”

There is nevertheless useful research here, and we wish the researchers had focused on that instead. Making a course-specific AI assistant available inside the LMS did not produce widespread or sustained use, and offering it did not improve the measured outcomes. In the same-course sample, average scores and recorded LMS participation were lower.

That is a finding EdTech needs: availability is not adoption, and adoption is not evidence of better learning. What students do with a tool—and how instructors integrate it—requires attention. We wish that useful contribution had received clearer billing, rather than leaving readers to disentangle it from broader claims about the effects of AI tutoring wrapped up in dense statistical and methodological jargon.

The main On EdTech newsletter is free to share in part or in whole. All we ask is attribution.

Thanks for being a subscriber.