Last updated: 2026-09-28

U
Undergraduate level

Supporting Students with GenAI: Why Trust Comes Before Policy

Most guidance on students and generative AI is a list of rules, and rules only reach students who are already listening. This page starts from a different observation, made by one lecturer over several cohorts: students who feel known, and who feel that someone expects good work from them, use AI more carefully. The observation is anecdotal, and the page marks where the anecdote ends and the evidence begins. The policy and detection questions are covered in Academic Misconduct and GenAI. This page is about the relationship those rules depend on.

"When students explain why they've worked hard on my modules without leaning on AI to do the work for them, the reason they give is usually that they don't want to disappoint me. I'm not sure I do anything beyond naturally expecting great things from my students. Presumably that comes across, unspoken, with what I suspect are parental overtones. It only works for the students I manage to build a rapport with: the ones who attend lectures, or who are interested enough to have conversations with me, online or in person." — Pat Parslow

An Observation, Not a Finding FoundationalKnowledge that endures for decades — core principles

Two kinds of evidence sit behind that quotation. One is first-person: what a lecturer sees in submitted work and in conversation. The other is reported: what students say about their own reasons. Both have limits. Students' stated reasons may not be their operative ones, and the students who talk to a lecturer are the same students who might have been careful anyway, so the observation cannot separate rapport from self-selection. Simply put, it is a hypothesis with some support behind it. The rest of the page sets out that support, and then the part of the problem it doesn't reach.

Why a Relationship Could Change How Students Use AI FoundationalKnowledge that endures for decades — core principles

Four separate strands of research point the same way, none of them about generative AI specifically.

Relationships matter at university, and are under-studied there. Hagenauer and Volet's review of teacher–student relationships in higher education covers their quality, their consequences and their antecedents, and concludes that the field is important yet under-researched, with relationships best understood as multi-dimensional and context-bound [1]. That is a reason for caution as well as interest: the evidence base for university-level rapport is thinner than for schools.

Relatedness is a motivational need. Ryan and Deci's self-determination theory names three needs, autonomy, competence and relatedness, and treats contexts that support them as those that foster self-motivation [2]. Not wanting to disappoint a particular person is a relational motive, and it works without any sanction attached. Reading "not wanting to disappoint" as an expression of relatedness is our interpretation. The theory doesn't make that claim itself.

Expectations have modest effects. Jussim and Harber's review of thirty-five years of research on teacher expectations concludes that self-fulfilling prophecies in classrooms do occur, but the effects are typically small and don't accumulate greatly, that larger effects may fall on students from stigmatised social groups, and that expectations may predict outcomes largely because they are accurate [3]. "Expecting great things" is therefore plausible as one ingredient among several, too weak to carry the whole effect. It also has to reach every student, and not just the ones who are easy to talk to.

The teaching climate is associated with cheating. In a survey of 14,086 students at eight Australian universities, Bretag and colleagues found that dissatisfaction with the teaching and learning environment was one of three variables associated with contract cheating, alongside a perception that there were many opportunities to cheat and speaking a language other than English at home. The authors suggest that universities cultivate stronger instructor–student relationships [4]. The survey predates generative AI, concerns outsourcing to third parties, and reports associations, so it supports the direction of the hypothesis and not its size.

graph LR A["Tutor expects strong work,
offers autonomy, models
their own fallibility"] --> B["Student wants to
meet that expectation"] B --> C["Student shares work
and talks about it"] C --> D["Early feedback names
over-reliance, no penalty"] D --> E["Student can be candid
about how AI was used"] E --> C E --> A F["Student never attends
or converses"] -.->|"loop cannot start"| C style F fill:#FFC857

Put together, the observation implies a reinforcing loop. Expectation, together with the autonomy and honesty about fallibility described below, gives a student a reason to care, contact makes the work visible, gentle feedback keeps contact safe, and candour about AI use feeds back into better feedback. The broken edge in the diagram is the subject of a later section.

What This Looks Like in Practice Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

Pat Parslow describes the practice behind the observation in these terms.

"I tell students that I'm in favour of using AI, but of using it well. I say that, although I'm not infallible, I do notice over-use of AI fairly reliably. I say that over-using it rather undermines the point of spending money on tuition, and that it is likely to be a problem for a graduate even if they get away with it at university, when an employer realises they don't have the skills and knowledge to do the job. I set assessments that require the use of AI, with a reflection section on its effect on learning and on productivity. If early work looks over-reliant on AI I say so in the feedback, and I don't penalise it directly unless the abuse is egregious." — Pat Parslow

Each element does relational work as well as pedagogic work. Saying at the start that AI use is welcome removes the reason to hide it, and hiding is what makes early feedback impossible. Edmondson's field study of work teams found that a shared belief that the team is safe for interpersonal risk-taking, more than ability, predicted whether members reported errors and asked for help early enough to fix them cheaply [5]. The same dynamic applies to a student deciding whether to admit that a draft leaned heavily on a model.

The argument from tuition value and employability appeals to the student's own goals, which is a different footing from a rule imposed from outside. Avoiding the AI Short-Circuit sets out the learning evidence behind that argument. Requiring AI use in assessment, covered in Teaching with GenAI, makes use discussable in the open. And feedback in place of penalty keeps the conversation going at the moment it is most useful.

The claim "I notice over-use fairly reliably" needs care, because the published research on detecting AI text by reading it looks discouraging at first sight. Fleckenstein and colleagues gave 89 pre-service and 200 experienced teachers essays that mixed student writing with ChatGPT output: neither group could identify the AI-generated texts, though experienced teachers made more differentiated and more accurate judgements, and both groups were overconfident [6]. In a blind test at a UK university, Scarfe and colleagues submitted entirely AI-written work across five undergraduate psychology modules, and 94% of it went undetected [7].

The second figure is less clear-cut than it sounds. In that study, "detected" meant flagged through the university's standard procedure for poor academic practice or misconduct. Markers were not told that AI submissions might be present, and the authors say they cannot tell whether markers who didn't flag a script had failed to notice a problem or had noticed and not reported it, whether because of time, the evidence needed, or something else [7]. The study therefore measures formal reporting, and reporting is a poor proxy for noticing.

Pat Parslow's experience, and that of colleagues, is that this gap is large in practice: the time a formal case takes, and the level of evidence needed to secure a finding, mean that relatively mild over-use doesn't get reported, whether the tool is generative AI or, before it, contract cheating. Earlier research on academic dishonesty points the same way. In a national sample of 127 US psychology instructors, the most frequent reason for overlooking behaviour that might be dishonest was insufficient evidence that cheating had occurred, followed by the time and effort of dealing with it, along with emotional and fear-related reasons [8]. That study is from 1998 and concerns a different country and technology, so it shows the pattern and not its present size.

Two limitations pull in opposite directions, and neither is settled. Formal flagging may understate how often markers notice a problem. The Scarfe submissions were unedited AI output, and the authors themselves say the 6% detection rate likely overestimates the ability to detect real-world AI cheating [7], presumably because a student who edits AI output is harder to catch. What both facts imply is that the published figures on reported cases, including those in Academic Misconduct and GenAI, count what cleared the threshold for reporting, and mild over-use sits below it.

One reading that reconciles those results with the experience above is ours, and neither study tests it: noticing that rests on knowing a student's earlier work, and on what they can say about a submission when asked, is a different act from judging an anonymous text on its style. If that is the mechanism, a relationship is what supplies the information. The design choice in the quotation also follows from the reporting gap. Formal processes are calibrated for cases that justify their cost, so mild over-use falls outside them, and a conversation is the only response proportionate to it. Feedback that says "this looks over-reliant on AI; let's talk" also costs a wrongly flagged student a conversation, where a penalty would cost them a mark and some trust. Because noticing is fallible and formal reporting is rare, a process that tolerates being wrong is worth more than one that relies on being right.

Autonomy and Fallibility Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

A second element of the approach concerns how support is offered.

"My general approach is to encourage full autonomy, and to step in with extra support only when issues become apparent or the student raises them. I expect students to build their own competence, and I try to model mine, including my failings in that regard. I suspect that has helped build the rapport too." — Pat Parslow

Autonomy is one of the three needs in self-determination theory, alongside competence and relatedness [2]. Reading the approach through that theory is our interpretation. Expecting students to build their own competence treats them as capable and self-directing, and rapport supplies relatedness, so the three needs are met together. Modelling one's own failings adds something further. Edmondson's finding that psychological safety, the shared belief that it is safe to take interpersonal risks, predicts whether people raise errors early [5] suggests a mechanism: a tutor who shows that struggle and error are normal makes it easier for a student to say "I've leaned on AI too much." That extension is also ours, since Edmondson studied work teams and not tutors and students. Getting Unstuck makes a related case for supervision that is available on demand and does not direct the work.

The approach has one dependency: support arrives when problems are apparent or raised. That works for students whose difficulties show up in submitted work or in conversation. It leaves the students who neither attend nor speak up exactly where the broken edge in the diagram above places them, and the next section is about that gap.

Where It Stops: Students Who Are Not in the Room Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The loop needs contact to start, and the students least likely to have it are the ones the research says are most at risk. Astin and Kuh find that learning and degree completion track a student's own engagement [9][10], and Tinto adds that academic and social integration are largely independent risks [11]. A student who doesn't attend and doesn't converse is out of reach of a rapport-based mechanism and is also at elevated risk on every other measure. Why Engagement Matters, and What to Do If You're Slipping covers the student-facing side of this.the loop needs a first contact

"I'm working on the engagement side. I'm hopeful that learning logs will raise engagement and make conversations with Academic Tutors easier, which would help build rapport. That's unproven so far." — Pat Parslow

The reasoning behind learning logs is plausible and its limits are known. Moon's handbook is the standard reference on learning journals in higher education [12]. A study of learning logs in secondary biology found that they encouraged reflection but didn't produce the depth of learning-strategy awareness that semi-structured interviews with the same students did [13]. On that evidence a log works best as something to talk about, and it is a poor substitute for the talking. That fits the aim of giving a tutor a concrete starting point for a conversation with a student who otherwise says little. Whether logs raise engagement in this setting has not yet been measured.

Two safeguards belong alongside any rapport-based approach. First, structure has to carry weight that relationships can't. Corbin, Dawson and Liu distinguish discursive changes, which rely on communicating rules and expectations to students, from structural changes, which alter how the assessment itself works [14]. A relationship is a discursive channel at its best. It needs assessment design underneath it that works for students who never form one, which is the subject of Teaching with GenAI.

Second, fairness can't depend on rapport. The students least likely to form a relationship may overlap with those most likely to be wrongly suspected: detectors have been criticised for flagging non-native English writers, and Bretag's survey associates a language other than English at home with contract cheating risk [4]. The same route for raising a concern, and the same route for resolving it, should apply to every student, whoever the lecturer happens to know.

A Route Through the Semester Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The existing material on this site fits together best when read along a student's path through a piece of work, since each stage raises a different question about AI.

graph LR S1["Starting
the work"] --> S2["Stuck"] S2 --> S3["Tempted
to delegate"] S3 --> S4["Uses AI"] S4 --> S5["Reflects
and discloses"] S5 --> S6["Feedback
conversation"] S6 --> S1

Starting. A student who has a scaffold, and knows it will be withdrawn, is better placed than one handed a finished structure. Scaffolding and GenAI gives a test for telling a scaffold from an answer machine. Stuck. Being stuck is where the pull towards delegation is strongest, and where a student who trusts the tutor is likeliest to say so instead. Getting Unstuck argues that raising a difficulty early is cheap and hiding it is costly, which is the same claim as Edmondson's.

Tempted. The moment of delegation is where the learning evidence bites, and Avoiding the AI Short-Circuit supplies the reasons: a homework score and retained understanding can move in opposite directions. Some students use AI as an accessibility aid, and that use must be told apart from over-reliance; AI, Accessibility, and Learning covers the distinction. Uses AI. What counts as acceptable use is set out in Academic Integrity, and it stays credible when the person applying it has been clear at the outset.

Reflects and discloses. The reflection section is where a student's account of their own use becomes visible, a subject developed in Teaching with GenAI. Feedback. The conversation that follows is where the loop closes, and where a tutor with a log and a submission in front of them has something concrete to ask about.

What Would Settle It Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The hypothesis is testable in principle, though the obvious test raises an ethical problem. Because engaged students would have been careful regardless, a fair comparison needs to break the link between rapport and self-selection, and the textbook way to do that is to give early one-to-one check-ins to a randomly chosen half of a cohort and withhold them from the other half. Withholding support from students who are paying for their education, and who may be the ones who most need it, isn't acceptable when a benefit is plausible, and any study of this kind needs ethical review before anyone is assigned to anything.

Designs exist that avoid the problem. In a stepped-wedge design every group receives the intervention and only the start date is randomised, so that some groups begin check-ins earlier in the term than others [15]. Each student is offered the same support, and the comparison comes from the weeks when one group has it and another doesn't yet. That still delays support for some students, so the stagger would need to be short, the support would need to be one that any tutor can already offer, and students would need to be told what is happening and why. Lower-risk alternatives compare outcomes across tutor groups whose existing practice already differs, or track engagement and AI-use reflections over time for a whole cohort that all receives the same support. These designs are our suggestions. Until one has been run, the strongest claim available is the modest one: relationship-based trust is a plausible, low-cost complement to policy and assessment design, and there is no evidence yet that it works without them.

What This Means in Practice

  • State your position first. Saying that AI use is welcome, provided it is used well, gives students no reason to hide it.
  • Give feedback where a penalty would end the conversation. Name over-reliance in early work and reserve sanctions for egregious cases.
  • Treat noticing as fallible. Reading a text alone is unreliable; knowing a student's earlier work and asking them about it is a different and stronger basis.
  • Offer autonomy, and show your own fallibility. Step in when problems appear or are raised, and let students see that struggle and error are ordinary, your own included.
  • Expect good work from everyone. Expectation effects are modest and fall unevenly, so the expectation has to be communicated to students who don't come to you.
  • Plan for the students you won't meet. A route to contact, such as a learning log, and an assessment design that works without rapport, are both needed.

References

  1. Hagenauer, G., & Volet, S. E. (2014). Teacher–student relationship at university: an important yet under-researched field. Oxford Review of Education, 40(3), 370–388. https://doi.org/10.1080/03054985.2014.921613
  2. Ryan, R. M., & Deci, E. L. (2000). Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. American Psychologist, 55(1), 68–78. https://doi.org/10.1037/0003-066X.55.1.68
  3. Jussim, L., & Harber, K. D. (2005). Teacher expectations and self-fulfilling prophecies: Knowns and unknowns, resolved and unresolved controversies. Personality and Social Psychology Review, 9(2), 131–155. https://doi.org/10.1207/s15327957pspr0902_3
  4. Bretag, T., Harper, R., Burton, M., Ellis, C., Newton, P., Rozenberg, P., Saddiqui, S., & van Haeringen, K. (2019). Contract cheating: a survey of Australian university students. Studies in Higher Education, 44(11). https://doi.org/10.1080/03075079.2018.1462788
  5. Edmondson, A. C. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly, 44, 350–383. https://doi.org/10.2307/2666999
  6. Fleckenstein, J., Meyer, J., Jansen, T., Keller, S. D., Köller, O., & Möller, J. (2024). Do teachers spot AI? Evaluating the detectability of AI-generated texts among student essays. Computers and Education: Artificial Intelligence, 6, Article 100209. https://doi.org/10.1016/j.caeai.2024.100209
  7. Scarfe, P., Watcham, K., Clarke, A., & Roesch, E. (2024). A real-world test of artificial intelligence infiltration of a university examinations system: A "Turing Test" case study. PLOS ONE, 19(6), e0305354. https://doi.org/10.1371/journal.pone.0305354
  8. Keith-Spiegel, P., Tabachnick, B. G., Whitley, B. E., & Washburn, J. (1998). Why professors ignore cheating: Opinions of a national sample of psychology instructors. Ethics & Behavior, 8(3), 215–227. https://doi.org/10.1207/s15327019eb0803_3
  9. Astin, A. W. (1984). Student involvement: A developmental theory for higher education. Journal of College Student Personnel, 25(4), 297–308.
  10. Kuh, G. D. (2009). The National Survey of Student Engagement: Conceptual and empirical foundations. New Directions for Institutional Research, 2009(141), 5–20. https://doi.org/10.1002/ir.283
  11. Tinto, V. (1993). Leaving College: Rethinking the Causes and Cures of Student Attrition (2nd ed.). University of Chicago Press.
  12. Moon, J. A. (2006). Learning Journals: A Handbook for Reflective Practice and Professional Development (2nd ed.). RoutledgeFalmer.
  13. Stephens, K., & Winterbottom, M. (2010). Using a learning log to support students' learning in biology lessons. Journal of Biological Education, 44(2), 72–80. https://doi.org/10.1080/00219266.2010.9656197
  14. Corbin, T., Dawson, P., & Liu, D. (2025). Talk is cheap: why structural assessment changes are needed for a time of GenAI. Assessment & Evaluation in Higher Education, 50(7), 1087–1097. https://doi.org/10.1080/02602938.2025.2503964
  15. Hussey, M. A., & Hughes, J. P. (2007). Design and analysis of stepped wedge cluster randomized trials. Contemporary Clinical Trials, 28(2), 182–191. https://doi.org/10.1016/j.cct.2006.05.007