Last updated: 2026-09-28

U
Undergraduate level

Teaching with GenAI: Require It, Then Ask What It Did

A ban on AI in assessed work is hard to sustain when students are already using it. In the HEPI/Kortext survey, 92% of undergraduates used generative AI and 88% had used it on assessed work [1]. That leaves two questions for anyone teaching: what to do about work that AI can produce end to end, and how to teach students to use it well. This page takes one lecturer's answer, which is to require AI use in assessment and ask students to reflect on it, and sets it against the evidence. The relationship that makes the approach work is covered in Supporting Students with GenAI.

The Position, Stated Up Front Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The approach begins with a position that students hear at the start.

"I tell students that I'm in favour of using AI, but of using it well." — Pat Parslow

Everything else depends on what "well" means, and two recent experiments show that it concerns the pattern of use far more than the presence of the tool. Bastani and colleagues ran a field experiment with nearly 1,000 high-school mathematics students in Turkey. Students using a plain ChatGPT-style tutor improved their practice grades by 48% over the control group, and students using a tutor prompted with teacher-designed hints improved by 127%. When access was removed for the exam, the plain-tutor group scored 17% below the control group, and the hints group was indistinguishable from it [2]. The interaction logs show how the two groups worked. In the first session, 67% of first messages to the plain tutor repeated the question or asked for the answer, against 37% for the hints tutor, and students sent significantly more messages to the hints tutor [2]. More engaged time went with no exam penalty, though the hints tutor differed from the plain one in more than the effort it drew out, so the study can't isolate effort as the cause.

Shen and Tamkin randomly assigned 52 developers to learn an unfamiliar asynchronous programming library with or without an AI assistant. On a quiz taken without AI, the AI group scored 17% lower (4.15 points out of 27), and the speed advantage was not statistically significant [3]. Watching the screen recordings, the authors identified six patterns of use. The three that kept learning intact were generating code and then asking questions to understand it, mixing code requests with requests for explanation, and asking only conceptual questions; those clusters averaged 65–86% on the quiz. The three that averaged below 40% were handing the whole task to the AI, delegating progressively more, and repeatedly asking the AI to debug [3]. The clusters are small, two to seven participants each, so the pattern is suggestive.

Time on task did not separate the groups. The slowest pattern, repeated AI-assisted debugging at 31 minutes on average, scored lowest, at 24%, while generation followed by comprehension took 24 minutes and scored 86% [3]. The authors' own summary is that cognitive effort may matter more than raw time on the task [3]. Across both studies the same tool produced different learning depending on how it was used or configured, and "using it well" therefore has a testable meaning: keeping the learner thinking, which is a different thing from keeping them busy.

Set the Task So That AI Is Required Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The assessments then carry that position into the work itself.

"I set assessments that require the use of AI." — Pat Parslow

Perkins, Roe and Furze's AI Assessment Scale (AIAS) gives five levels of AI involvement that an assessment can be designed around: No AI, AI Planning, AI Collaboration, Full AI and AI Exploration [4]. In the published scale, an AI Collaboration task is designed so that AI alone will not reach the required standard, and a Full AI task expects AI involvement and can't be completed either by AI or by a person working alone. A task pitched at those levels makes AI use part of what is assessed, and a submission that outsources the thinking to a model fails on its own terms.

Corbin, Dawson and Liu's distinction explains why this design matters. Discursive changes rely on communicating rules to students and leave the mechanics of the task unchanged, while structural changes alter how the task works [5]. A policy telling students to use AI sensibly is discursive. Mapping the AIAS levels onto that distinction is our reading, and the mapping is direct: a task that can't be finished without AI, and can't be finished by AI alone, is structural.

graph TD T["Task designed so AI alone
will not reach the standard"] --> W["Student works with AI
and keeps evidence of the process"] W --> S["Submission: the work plus a reflection
on learning and on productivity"] S --> F["Early, formative feedback
on the process"] F --> C{"Conversation
needed?"} C -->|yes| V["Talk it through:
what did you do,
what can you explain?"] V --> W C -->|no| N["Next task"] N --> T F -.->|"egregious over-use"| X["Formal academic-integrity route"] style X fill:#FFC857

Requiring AI carries an equity obligation. If the task can't be done without the tool, every student needs access to a suitable one, and institutional provision has been uneven; What Higher Education Needs to Do About GenAI sets out the sector picture and the Russell Group principle on equal access.uneven access is a real barrier to this approach

The Reflection Section: Learning and Productivity, Asked Separately Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

Each assessment ends with a reflection that students have to write.

"The assessments include a reflection section asking about the impact of AI use on learning and on productivity. Normally, those two are at odds with one another." — Pat Parslow

The evidence supports that observation, and adds a warning about self-report. In Bastani's experiment, the plain-tutor group did better in practice and worse in the exam [2]. In Shen and Tamkin's, AI use didn't deliver significant efficiency gains on average and still impaired learning [3]. Becker and colleagues' randomised trial of sixteen experienced open-source developers, working through 246 tasks in their own repositories, found that before starting the developers forecast a 24% reduction in completion time, afterwards estimated that AI had cut it by 20%, and were measured at 19% slower [6]. That trial measured productivity in experts and not learning in students, but it shows that perceived and measured effects can diverge in the same direction as a student's confident summary of how much AI helped.

A reflection section that asks about learning and productivity as separate questions makes the gap visible to the student, and gives the marker something specific to discuss. Read the answers as reports of perception: valuable material for conversation and unreliable as measurement. The table below is our own framing, offered as a way of naming the four combinations a student might describe.the gap is a feature, not a bug to close

Learned moreLearned less
FasterAI used to check and explain after an unaided attemptThe thinking delegated: a strong practice score and a weak unassisted one
SlowerStruggle first, AI as critic afterwardsTime spent, little retained

Prompts that pull the two questions apart work better than a single "reflect on your AI use":

  • What did you hand to the AI, and what did you keep for yourself?
  • What did you check, and how did you check it?
  • What can you now explain without the tool that you couldn't before?
  • Where did it save you time, and where did it cost you time?
  • What would you do differently next time?

Feedback in Place of Penalty Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

Feedback is where the position meets a student's actual work.

"If early work looks over-reliant on AI, I point that out in the feedback and don't penalise it directly, unless the abuse is egregious." — Pat Parslow

The case for that policy rests on how unreliable noticing is. Fleckenstein and colleagues found that 89 pre-service and 200 experienced teachers could not identify ChatGPT-generated essays among student-written ones, and both groups were overconfident [7]. Scarfe and colleagues found that 94% of entirely AI-written exam submissions in a UK undergraduate psychology programme were not flagged through the standard misconduct procedure, though they note that the measure can't separate a marker who noticed nothing from one who noticed and didn't report [8]. The second case is common in practice: the time a formal case takes and the evidence it needs mean that relatively mild over-use often goes unreported, a point developed in Supporting Students with GenAI. Suspicion can be right and can be wrong, and formal reporting captures only part of it. A process that answers suspicion with a conversation loses little when the suspicion is mistaken, and it reaches the mild cases that formal procedures leave alone. One that answers with a penalty loses a great deal when it is wrong.

Feedback of this kind needs early work to respond to, which means staged assessment with something to look at before the final submission. When a conversation does need a firmer basis, a viva is the established tool, and its cost in staff time is covered in Academic Misconduct and GenAI. Egregious cases go to the formal route as usual.vivas cost staff time, but catch real issues

Modelling Use in the Room Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

Collins, Brown and Newman's cognitive apprenticeship describes six teaching methods: modelling, coaching, scaffolding, articulation, reflection and exploration [9]. Mapped onto teaching with GenAI, they cover most of what this page has described. The mapping is ours.

MethodWhat it looks like with GenAI
ModellingWork an AI-assisted problem live, including a wrong turn and the check that caught it
CoachingGive feedback on students' process as well as their product
ScaffoldingProvide prompts and guides that are withdrawn as students gain judgement
ArticulationAsk students to explain what the AI produced and why they kept or changed it
ReflectionThe reflection section, comparing one's own process with others'
ExplorationOpen tasks where students choose how to use AI, as at the top of the AIAS

Scaffolding carries a condition that is easy to forget. Wood, Bruner and Ross's account is explicit that a scaffold is meant to be withdrawn as competence develops [10], and Scaffolding and GenAI gives a practical test for whether a given AI use is fading or replacing the student's own effort. The learning evidence for why that matters is in Avoiding the AI Short-Circuit.scaffolding must fade else it becomes crutches

A First Semester's Worth of Change Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

None of this needs a redesigned curriculum to start. Four small moves cover most of it:

  1. State your position in the first session. Say that you welcome AI used well and explain what you mean by well.
  2. Add a reflection section to one existing assessment. Use the prompts above, and ask about learning and productivity separately.
  3. Give process feedback on one early draft. Name what looks over-reliant on AI and invite a conversation.
  4. Model one problem live, including a failure. Show the check that caught it.

The bottleneck is usually support for staff more than willingness. Bretag and colleagues found that students perceive four assessment types as least likely to be outsourced, that these are also the types educators are least likely to set, and that educators are more likely to use them when they report positively on organisational support for teaching and learning [11]. That study concerns contract cheating before generative AI, but the mechanism it points to carries over: redesign happens where staff have time and backing. What Higher Education Needs to Do About GenAI covers the staff-capability side.

What the Evidence Does Not Yet Cover Ephemeral / ToolingKnowledge that evolves in months to a year — check for updates

The studies cited here come from adjacent settings: high-school mathematics in one school, developers learning a single library in an experiment, and experienced developers working in their own repositories. None is a university assessment in which AI use is required and reflected on. The approach on this page therefore rests on those analogies, on the assessment-design literature and on one lecturer's experience, and it hasn't been tested directly. The reflection section is also the piece most worth studying: whether students' answers track what they learned is an open question, and one that comparing reflections with unassisted performance across a whole cohort could begin to answer, with no student given less support than another.

What This Means in Practice

  • Define "well" as keeping the learner thinking. The pattern of use decides the outcome, so teach patterns and not just permissions.
  • Change the task, not only the rules. A task that needs AI and can't be completed by AI alone does work that a policy statement can't.
  • Ask about learning and productivity separately. The two often diverge, and a student's sense of which one they got is unreliable.
  • Answer suspicion with a conversation. Noticing is fallible, so choose a process that costs little when it is wrong.
  • Show the process. Modelling a check, and a mistake caught by one, teaches more than a rule about checking.
  • Provide access. If AI is required, every student needs a suitable tool.

References

  1. Higher Education Policy Institute / Kortext, "Student Generative AI Survey 2025", HEPI Policy Note 61, February 2025. https://www.hepi.ac.uk/reports/student-generative-ai-survey-2025/
  2. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakçı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. https://doi.org/10.1073/pnas.2422633122
  3. Shen, J. H., & Tamkin, A. (2026). How AI Impacts Skill Formation. arXiv:2601.20245
  4. Perkins, M., Roe, J., & Furze, L. (2025). Reimagining the Artificial Intelligence Assessment Scale (AIAS): A refined framework for educational assessment. Journal of University Teaching and Learning Practice, 22(7). https://doi.org/10.53761/rrm4y757. Level descriptions as published at aiassessmentscale.com.
  5. Corbin, T., Dawson, P., & Liu, D. (2025). Talk is cheap: why structural assessment changes are needed for a time of GenAI. Assessment & Evaluation in Higher Education, 50(7), 1087–1097. https://doi.org/10.1080/02602938.2025.2503964
  6. Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv:2507.09089
  7. Fleckenstein, J., Meyer, J., Jansen, T., Keller, S. D., Köller, O., & Möller, J. (2024). Do teachers spot AI? Evaluating the detectability of AI-generated texts among student essays. Computers and Education: Artificial Intelligence, 6, Article 100209. https://doi.org/10.1016/j.caeai.2024.100209
  8. Scarfe, P., Watcham, K., Clarke, A., & Roesch, E. (2024). A real-world test of artificial intelligence infiltration of a university examinations system: A "Turing Test" case study. PLOS ONE, 19(6), e0305354. https://doi.org/10.1371/journal.pone.0305354
  9. Collins, A., Brown, J. S., & Newman, S. E. (1989). Cognitive apprenticeship: Teaching the crafts of reading, writing, and mathematics. In L. B. Resnick (Ed.), Knowing, Learning, and Instruction: Essays in Honor of Robert Glaser (pp. 453–494). Lawrence Erlbaum Associates.
  10. Wood, D., Bruner, J. S., & Ross, G. (1976). The role of tutoring in problem solving. Journal of Child Psychology and Psychiatry, 17(2), 89–100. https://doi.org/10.1111/j.1469-7610.1976.tb00381.x
  11. Bretag, T., Harper, R., Burton, M., Ellis, C., Newton, P., van Haeringen, K., Saddiqui, S., & Rozenberg, P. (2019). Contract cheating and assessment design: exploring the relationship. Assessment & Evaluation in Higher Education, 44(5), 676–691. https://doi.org/10.1080/02602938.2018.1527892