Should an AI tutor teach how to prompt or how to reflect?

Cristina Agrigoroae © 2026 EPFL
A conversation with Jérôme Brender, whose research at the intersection of artificial intelligence and the learning sciences, focuses on how students engage with AI tools while learning to program, and what that engagement means for their learning in the long run.
More than 8 in 10 students are now using large language models (LLMs) for their coursework. In classrooms, the question is no longer whether students will use generative AI, but how. A student who doesn’t touch these tools risks falling behind peers who use them to move faster through assignments, debug code in seconds, or get instant explanations of a concept they missed in class.
But not all AI assistance is the same: a chatbot that hands over answers and one that asks the right questions can produce very different students, even when built on the exact same underlying model. Understanding which kind of AI support actually builds lasting skill, rather than just short-term output, is one of the more urgent open questions in education today.
That question is at the center of Jérôme Brender's doctoral thesis. His paper, "Reflective Dialogue or Prompt Refinement? Effects of Tutor Scaffolding on Students' Independent LLM Use for Programming," co-authored with Laila El-Hamamsy, Kim Uittenhove, Aitor Perez, Patrick Jermann, Francesco Mondada, and Engin Bumbacher, received the Best Paper Award at AIED 2026 conference in Seoul in July 2026.
The study, conducted in a graduate-level mobile robotics course at EPFL, compared two AI tutoring designs over a six-week intervention with 66 students: a Socratic-Guidance (SG) tutor, which scaffolds learning through reflective questioning, and a Prompt-Refinement (PR) tutor, which coaches students to write clearer, more explicit prompts. In the next stage, the study followed 52 of those students into a three-week course project, where everyone used an unconstrained version of the same chatbot, with no scaffolding at all, to see whether anything from the guided sessions actually stuck.
We sat down with Jérôme to unpack the findings of the study, what surprised him, and what an ideal AI tutor might look like.
Your paper compares two different tutoring approaches. Why compare two approaches rather than comparing one approach against no AI tutor at all?
This is a follow-up to previous studies where we looked at how to help students reflect better on their prompting practices, helping them articulate, decompose, and refine prompts so they could better understand the theoretical concepts of a robotics course. We built a prompt refinement tutor and compared it to the current approach in AI tutoring, Socratic guidance, a tutor that guides students through questioning rather than direct feedback, encouraging reflection. So, the goal was not simply to ask whether an AI tutor is better than no tutor. Instead, we wanted to compare two different mechanisms for supporting reflective LLM use: one acting on the student’s input (prompt of the student), and one acting on the tutor’s output (response of the LLM). Importantly, we also examined what happened when these scaffolds (AI tutor) were removed, to an unconstrained, general-purpose LLM.
What did you find most surprising when comparing the two?
There were two key findings. First, we ran three sessions where students used their assigned AI tutor. By the final session, students using Socratic guidance showed a higher learning gain than those using prompt refinement.
The more surprising finding came next, when we removed the AI tutor, or the scaffolding, entirely. Students who had used Socratic guidance still prompted better than those who'd used prompt refinement, even though Socratic guidance never explicitly taught them how to improve prompts. It seems the questioning approach pushed students toward more critical thinking and reflection, and that carried over even without the tool present.
When comparing understanding or mastery of the actual course content, we saw the same pattern. By the final session, the Socratic guidance group showed higher learning gains for the theoretical concepts, though the statistical effect had some limitations.
A key limitation is that we had no control group of students using no AI tutor at all, so we can't say whether either condition is high or low in absolute terms, only relative to each other.
Your conclusions note that students preferred the prompt-refinement tutor over Socratic guidance. What does that tell you?
It's a classic mismatch in learning sciences between perceived value and actual learning benefit. Students often prefer what's less cognitively demanding, even when it isn't what benefits them most. We noticed that students using prompt refinement were mostly just copy-pasting the suggested prompts rather than putting in the effort to rewrite them. It may have required less cognitive effort, so they rated it highly, but they weren't aware of the actual learning outcomes. Students using Socratic guidance rated the experience lower, even though they seem to learn more.
Given that Socratic guidance is perceived as less efficient, is there a risk that it gets abandoned or under-adopted in practice?
We should be careful here. Students still rated both tutors positively overall; Socratic guidance wasn't perceived negatively, just less favorably than prompt refinement. The difference in perceived usefulness was statistically significant, but still positive even for the Socratic guidance, and not a sign of outright rejection. It's more a reminder that if we want students to actually use a tool like this outside a study, we can't ignore how it feels to use it, even when we know it works.
The Socratic tutor initially hurt task performance. How do you interpret it?
Looking at on-task performance, via the Jupyter notebook tasks students had to solve, students using Socratic guidance performed worse than the prompt-refinement group in the first session. But by the second and third sessions, that gap had closed. My hypothesis is that Socratic guidance requires more upfront cognitive effort, so there's a learning curve. Once students adapted to the tool, they became more efficient and better learners, a slow start with a strong payoff.
Do you think the growing pressure to deliver coursework quickly, now that AI can generate answers instantly, plays into why students preferred the more "efficient" tool?
Yes, that is possible in general. But in this specific course, we try to push back against that pressure. Students can use any AI tools they want, but what matters is that they understand the underlying algorithm and why they chose a particular approach, not just fast implementation. We intentionally don't add more time pressure when students use LLMs. The emphasis stays on conceptual understanding, precisely so that using AI doesn't become just about speed.
If you were to design an ideal AI tutor combining insights from both approaches, what would it look like?
The first thing you do when you interact with an LLM is write a prompt, so the more you enrich that prompt, the more clear and explicit it is, and the more you question yourself and reason about what you're doing, the better the learning outcomes are likely to be.
So, an ideal AI tutor would really be a mix of both. It would have the Socratic side questioning you, but at the same time giving you feedback on how to prompt better: why a prompt isn't specific enough, why it needs to be more explicit, how to put a question into the prompt itself to add real value. The key is that these two forms of support would work together: the tutor would help students improve how they interact with the LLM, while still engaging them with the disciplinary content.
What do you think made this paper stand out enough to win the award?
I think it's the focus on what happens after the scaffolding is removed. A lot of research looks at how well an AI tutor performs while students are actively using it, but few studies examine learning to learn with the tool and how students transfer that experience back to using an unscaffolded LLM on their own. Our finding, that the type of scaffolding affects behavior even after it's taken away, speaks to a more meta-level question: not just "does the AI tutor help," but "does it teach students how to learn to use LLMs effectively."
What's the added value of this study for the learning sciences more broadly?
This was a large collaboration. Multiple entities were involved, including researchers, educators, designers of technology, digital education centers, and robotics course teams. The findings show that AI tutors can meaningfully shape how students learn at EPFL, and support the case for continuing to develop AI tutors in collaboration with stakeholders across the school, improving how we support learning in the era of AI overall. Since the tutor students liked least was the one that taught them the most, the real design challenge ahead is building one that does both, without losing students along the way.