Last updated: 2026-10-06

U
Undergraduate level

What Separates a Pass From a First: Five Levels of Literature Review

Five reviews of the same evidence, from a bare pass to a PhD-level framing, with commentary on what changes at each step

Most advice on literature reviews says what a good review looks like. It does not show the differences between levels of competence side by side. Students often recognise the difference between describing sources, synthesising them, critiquing them, and making an original contribution only after they have seen several versions of the same task. This page provides that comparison.

Every example below reviews the same body of evidence: academic studies and vendor material on AI coding assistants. The domain was chosen because it combines peer-reviewed research with industry claims about competing products, and because most computing students already have some sense of what these tools do. Using one domain throughout keeps the subject matter constant, so the comparison is about the quality of the review and not about the topic.

What changes between the five examples is the relationship between the sources. The writing does not become more sophisticated in any sense that matters, and the vocabulary is kept modest throughout. A review improves because it moves from reporting what each source says, to showing how the sources relate, to judging how they were produced, to questioning the assumptions they share.see how synthesis differs from summary

The Shared Evidence Base

The table lists the sources all five examples draw on. Each is described as its authors or publisher present it, and the notes say what kind of evidence each one provides.

SourceTypeWhat it reports
Peng et al. (2023)[8]Controlled experimentDevelopers given GitHub Copilot completed a JavaScript HTTP server task 55.8% faster than a control group
Becker et al. (2025)[3]Randomised trial16 experienced open-source developers on 246 tasks took 19% longer with AI tools allowed, though they estimated afterwards that AI had sped them up by about 20%
Vaithilingam et al. (2022)[10]User study, 24 participantsCopilot did not reliably improve completion time or success, participants preferred it as a starting point, and they found its output hard to understand, edit and debug
Pearce et al. (2022)[7]Security studyAbout 40% of 1,689 programs generated across 89 scenarios were vulnerable
Barke et al. (2023)[2]Grounded theory, 20 participantsInteraction is bimodal: an acceleration mode, where the programmer knows the next step, and an exploration mode, where they do not
Shen and Tamkin (2026)[9]Randomised learning studyDevelopers learning a new library with AI assistance showed weaker conceptual understanding, code reading and debugging, with no significant speed gain on average
Parasuraman and Manzey (2010)[6]ReviewAutomation bias occurs in both naive and expert users and is not prevented by training or instruction
GitHub (2022)[5]Vendor-sponsored researchDevelopers with Copilot completed a task 55% faster, with a 78% completion rate against 70% without it
Cursor (n.d.)[4]Vendor documentationDescribes planning, agent-based work across a codebase, and review of diffs before merging
Amazon Web Services (n.d.)[1]Vendor product pageDescribes code completion and security scanning, and claims that the scanning outperforms other publicly benchmarkable tools

The academic sources differ in method, population, and outcome. The vendor sources differ in what they choose to show. Each example below uses this same set, so a difference between two examples is a difference in analysis.how to weigh conflicting claims

Level 1: Bare Pass (approximately 40%)

Characteristics. Largely descriptive. Sources are reported one at a time. There is little comparison, little synthesis, and little critical thinking.

Several AI coding assistants are now widely available. GitHub Copilot offers code suggestions inside the editor. Cursor is an AI-first code editor with an agent that can make changes across files. Amazon Q Developer provides code completion and security scanning.

Several studies have examined productivity. Peng et al. (2023)[8] reported that developers using GitHub Copilot completed an HTTP server task 55.8% faster than a control group. Becker et al. (2025)[3] studied 16 experienced open-source developers and found that tasks took 19% longer when AI tools were allowed. GitHub (2022)[5] reported that developers with Copilot completed a task 55% faster.

Other studies have looked at different questions. Vaithilingam et al. (2022)[10] found that participants preferred Copilot, and Pearce et al. (2022)[7] found that about 40% of generated programs were vulnerable. Shen and Tamkin (2026)[9] examined learning and found that AI use impaired conceptual understanding.

These findings suggest that AI coding assistants have both benefits and risks for software development.

Analysis.

  • Strengths. The review uses real sources, and it stays relevant to the question.
  • Problems. Each source is reported separately, with no comparison of findings. The two productivity studies with opposite results are presented side by side without any acknowledgement that they disagree. Limitations are not discussed, and the methods behind the figures are not described. The conclusion could be written without reading any of the sources, and no argument emerges.

Level 2: Competent Undergraduate (2:2 / Low 2:1)

Characteristics. Sources are grouped by theme. The review begins to compare findings and shows awareness of gaps, but the critique is limited and the conclusions largely follow the authors.

Evidence on the productivity of AI coding assistants is mixed, and the size of the reported effect depends on how it is measured. Peng et al. (2023)[8] and GitHub (2022)[5] report large speed gains on a bounded task, in a controlled setting and in a vendor-sponsored study respectively. Becker et al. (2025)[3] reached the opposite conclusion for experienced developers working in mature codebases, who took 19% longer with AI tools while estimating that they had been faster. The difference may reflect task type and developer experience, although the studies do not test these explanations directly.

The commercial products also differ. GitHub Copilot is best known for completion inside the editor. Cursor is built around an agent that can plan and edit across a repository (Cursor, n.d.). Amazon Web Services (n.d.)[1] presents Amazon Q Developer as combining completion with security scanning, and claims that its scanning outperforms other tools. These differences matter because the productivity studies examine particular tasks, and none compares the three products directly.

Taken together, the evidence suggests that productivity benefits depend on context, and that comparisons between products remain limited.

Analysis.

  • What improved. The review groups sources by theme, first productivity and then products. It begins to compare findings rather than list them, and it names a gap: the products have not been compared directly. It also notes that the GitHub study is vendor-sponsored and that the vendor describes its own product.
  • What is still missing. The critique is limited. The explanation for the disagreement between Peng et al. and Becker et al. is offered as a possibility and is not examined. The conclusion follows the authors' own findings, and nothing in the review would surprise them.

Level 3: First Class Undergraduate

Characteristics. The review synthesises academic and industry evidence, identifies patterns across sources, and develops an argument of its own. It identifies a gap in the literature, and it analyses the structure of the evidence rather than only describing it.

Although most published productivity studies report gains, the evidence is less consistent than the headline figures suggest. The studies differ in population, task, and outcome: a bounded HTTP server task in one study (Peng et al., 2023[8]), experienced developers on their own repositories in another (Becker et al., 2025[3]), and a vendor-sponsored experiment in a third (GitHub, 2022). What they share is a focus on how quickly a task is completed. Vaithilingam et al. (2022)[10] found that participants did not reliably complete tasks faster or more successfully, yet they preferred Copilot, mainly because it gave them a starting point. The value developers perceive and the time they measure may therefore be different things.

The same pattern appears in the vendor material. GitHub (2022)[5] reports speed and completion rates, Amazon Web Services (n.d.)[1] emphasises the detection performance of its security scanning, and Cursor (n.d.)[4] describes planning and review workflows. Each product's evidence concentrates on the dimension it can best demonstrate. The vendor sources reviewed here report speed, completion rates, or security detection, but not long-term maintainability. That matters because Pearce et al. (2022)[7] found that about 40% of generated programs in their scenarios were vulnerable.

This suggests a gap in the evidence. Academic and commercial evaluations both privilege short-term speed and output, which leaves the effect on code quality and on developers' own understanding less examined. Shen and Tamkin (2026)[9] address the second question directly, and their finding that AI use impaired conceptual understanding and debugging points to a cost that speed measures would not capture.

Analysis.

  • What improved. The review no longer reports sources. It analyses the structure of the evidence: the three academic studies differ in population, task, and outcome, and the vendor sources concentrate on what each can show well. It draws these together into an argument that the productivity figures are less consistent than they appear, and it identifies a gap that the sources themselves leave open.
  • Why it achieves a first. The review is not summarising the literature any more. It is explaining why the literature looks as it does, and it uses the sources to make that case. The argument is the student's own, even though every piece of evidence in it belongs to someone else.

Level 4: Master's Level

Characteristics. The review critiques the methods used, questions the assumptions and definitions, and reframes the research question. It asks what the evidence can and cannot support.

The dominant account of AI coding assistants is that they raise developer productivity. That account is built from studies that measure productivity in different ways, and its apparent consistency partly reflects the choice of measure. Peng et al. (2023)[8] measure the time taken to complete a defined task. Vaithilingam et al. (2022)[10] measure usability and preference, and participants valued the tool although it did not reliably improve completion time. Becker et al. (2025)[3] measure time on real tasks in mature projects, and find the reverse of the headline result. Each study is internally coherent, but they are not directly comparable, so no single method establishes that AI speeds up development.

The vendor evidence follows the same reduction. GitHub (2022)[5] reports task time and completion rate, and its measures of flow and satisfaction rely on what developers say about their own experience. Cursor (n.d.)[4] and Amazon Web Services (n.d.)[1] describe capabilities such as agents, planning and security scanning, which the productivity studies do not measure. Products are therefore compared on what they output rather than on how the developer's work changes. Becker et al.'s participants estimated a 20% speed-up after being 19% slower, which shows that the perception feeding vendor evidence can diverge from measured outcome. Parasuraman and Manzey (2010)[6] identify a related problem: automation bias affects expert users and is not removed by training or instruction.

Viewed together, these observations suggest that much of the current evidence measures output generation rather than engineering effectiveness. The question that matters may not be whether AI speeds up coding, but which capabilities it amplifies and which it displaces during development. Answering that requires measures of quality and of the developer's own competence, which the studies reviewed here largely lack.

Analysis.

  • What improved. The review critiques the methods. It shows that the studies share an outcome, but measure it differently, so their agreement or disagreement cannot be read straight off. It questions the definition of productivity itself, and it uses the vendor material to show the same reduction in another setting. It also uses a finding about perception versus measurement to test the reliability of the evidence base.
  • Why this feels master's level. The review interrogates how knowledge about these tools is produced, not only what the studies found. It reframes the research question from whether AI speeds up coding to which capabilities it changes, and it says what kind of evidence would be needed to answer that.

Level 5: PhD-Level

Characteristics. The review challenges shared assumptions in the field, identifies epistemological issues, creates a new conceptual framing, and generates a research agenda. Its strongest contribution is a new lens through which the existing literature can be read.

The literature treats AI coding assistants as productivity technologies. This framing is shared by academic studies and vendor material alike, and it rests on the premise that the value of an assistant can be read from how quickly a defined coding task is completed. The convergence is notable because the studies disagree on the outcome (Peng et al., 2023[8]; Becker et al., 2025[3]), yet agree on the measure.

That agreement deserves scrutiny. Measuring speed on bounded tasks implicitly models software development as the production of code, with coding as its central activity. A substantial strand of software engineering research argues that requirements, design, communication and judgement are at least as important to whether a system succeeds. On that view, a comparison of GitHub Copilot, Cursor and Amazon Q evaluates a narrow slice of practice. The apparent consensus on productivity may reflect the shared measure rather than a robust finding about software engineering effectiveness.

Two findings suggest where the shared assumption is weakest. Barke et al. (2023)[2] describe programmers alternating between an acceleration mode, where they know what to do next and use the assistant to get there faster, and an exploration mode, where they are unsure how to proceed and use it to explore options. Speed measures capture the first mode and say little about the second, in which the developer's own judgement is most exercised. Shen and Tamkin (2026)[9] show that AI use can impair conceptual understanding, code reading and debugging, which are the capacities that let a developer judge output they did not write. Together, these suggest that current studies may be measuring the activity the assistant automates, while leaving unmeasured the capacities it may erode.

A more productive line of inquiry would investigate how AI assistants redistribute cognitive labour within development teams, altering not only efficiency but expertise formation, accountability and professional judgement. That question cannot be answered with the speed measures that dominate the present evidence base.

Analysis.

  • What distinguishes it. The review challenges an assumption that the studies, and the vendors, share, which is that speed is the measure of value. It identifies an epistemological problem: the convergence of measures may come from the shared model of software development rather than from the evidence. It draws on Barke et al. to show that the measure captures one mode of use and not the other, and on Shen and Tamkin to show what the measure can miss.
  • The key move. The strongest part of this review is not its summary of the literature. It is the new framing, in which the question becomes how the tools redistribute cognitive labour and what that does to expertise and judgement. That framing makes the existing studies look different, and it generates a research agenda that the field does not yet have.

Summary

LevelMain activity
Bare passDescribes sources
Competent undergraduateCompares sources
First classSynthesises sources
Master'sCritiques methods and assumptions
PhDReframes the field and generates new questions

The Central Lesson

A literature review does not improve because its writing becomes more sophisticated. It improves because the relationship between the sources becomes more sophisticated. At the bare pass, the sources sit next to each other. At the competent level, they are compared. At a first they are synthesised into a pattern. At master's level the methods that produced them are examined, and at PhD level the framing that connects them is questioned. In each case the evidence stays the same, and what changes is the thinking that puts it together.synthesis is the goal, not the method

When you review your own sources, the useful question is therefore not whether your sentences are clear, but what relationship you have shown between the sources, and whether you have tested the assumptions they share. The levels above show what that looks like in one domain, so that you can recognise the same moves in your own.

References

  1. Amazon Web Services. (n.d.). Amazon Q Developer. https://aws.amazon.com/q/developer/build
  2. Barke, S., James, M. B., & Polikarpova, N. (2023). Grounded Copilot: How programmers interact with code-generating models. Proceedings of the ACM on Programming Languages, 7(OOPSLA1), Article 78. https://doi.org/10.1145/3586030
  3. Becker, J., Rush, N., Barnes, B., & Rein, D. (2025). Measuring the impact of early-2025 AI on experienced open-source developer productivity. arXiv:2507.09089. https://arxiv.org/abs/2507.09089
  4. Cursor. (n.d.). Cursor documentation. https://cursor.com/docs
  5. GitHub. (2022, July 14). Research: Quantifying GitHub Copilot's impact on developer productivity and happiness. GitHub Blog. https://github.blog/2022-07-14-research-quantifying-github-copilots-impact-on-developer-productivity-and-happiness/
  6. Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381โ€“410.
  7. Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., & Karri, R. (2022). Asleep at the keyboard? Assessing the security of GitHub Copilot's code contributions. 2022 IEEE Symposium on Security and Privacy. arXiv:2108.09293. https://arxiv.org/abs/2108.09293
  8. Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity: Evidence from GitHub Copilot. arXiv:2302.06590. https://arxiv.org/abs/2302.06590
  9. Shen, J. H., & Tamkin, A. (2026). How AI impacts skill formation. arXiv:2601.20245. https://arxiv.org/abs/2601.20245
  10. Vaithilingam, P., Zhang, T., & Glassman, E. L. (2022). Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models. CHI Conference on Human Factors in Computing Systems Extended Abstracts. https://doi.org/10.1145/3491101.3519665