This Week in MedEd
20 open funding calls, 70 upcoming events, 72 live jobs (44 new, 15 closing this week), including:
💰 Advanced Digital Skills for AI Uptake in Health (up to €3.9M from the European Commission’s Digital Europe Programme, deadline 1 Oct 2026)
🎓 Stanford AI+HEALTH 2026 (online, 8-9 Dec 2026, abstracts due 15 Oct)
💼 Head of School of Clinical Medicine and Vice Dean (UNSW Sydney, closes 14 Oct 2026)
The full list, with deadlines, links and eligibility, is one tap away. Try everything free for 30 days.
Every study of AI marking I have read this year ends with a similar finding: some version of “the model agreed with the lecturers”. It is an understandable place to stop. Marking is slow, expensive and boring (usually), and an agreement statistic is exactly what a busy assessment lead wants to see before switching to a new process. This week I have picked three empirical studies that put AI to work as a marker or coder, and one conceptual review that explains why their headline numbers are the start of the argument rather than the end of it.
Key points
In handwritten anatomy exams, ChatGPT tracked two expert markers closely, but DeepSeek managed only 0.520, and giving Gemini and Copilot the official answer key made their agreement worse.
When faculty reviewed 536 disagreements between Claude and their own rubric marking of clinical records, they sided with the model 63.8% of the time.
Three LLMs coding the same qualitative data agreed with each other (kappa around 0.75) far more than humans agreed with each other (0.42), and their errors were more systematic.
Processing excerpts in batches rather than one at a time significantly lowered AI agreement with the human reference.
A state-of-the-science review argues that we regulate AI that helps doctors make decisions, but not AI that decides which doctors are competent, and sets out seven mechanisms that can make AI-based assessment either more or less valid.
We’ll start this week with the review article. Tolsgaard et al. wrote a state-of-the-science conceptual review in Medical Education, they work through Kane's four validity inferences (scoring, generalisation, extrapolation and implications) and ask what AI does to each1. The list of threats is specific and useful:
construct contamination, where a model rewards verbosity or formatting rather than the thing being assessed;
prompt and rubric instability, where small wording changes move scores;
domain shift and model drift across sites and versions;
automation bias, deskilling of assessors as well as learners, and unclear accountability when an AI-informed decision is wrong.
The central point is that "apparent agreement with human raters does not inherently secure the scoring inference." Their most original move is to recast seven mechanisms as continua rather than flaws: opacity can become explainability, drift can become continuous recalibration, gaming can become red-teaming, depending on design and governance. For practice, they set a minimum reporting standard for any AI-based assessment: the model and version, the prompt and rubric, the parameter settings, the date of use, and illustrative outputs. This is a conceptual review rather than a systematic one, and the authors say so, but it is the paper I would hand to any committee being asked to approve an AI marking tool.
Nahir et al. scanned the handwritten open-ended anatomy exam papers of 81 second-year dental students in Turkey and had four chatbots mark each one twice, once with no answer key and once with the full key and grading criteria2. Two anatomists, who agreed with each other almost perfectly (ICC 0.992), provided the reference.
ChatGPT came closest (ICC 0.904 without the key, 0.923 with it) and was the only model whose scores did not differ significantly from the lecturers'. The spread across models was wide: DeepSeek sat at 0.520 to 0.542, and Gemini and Copilot both lost agreement once given the answer key, while their mean scores dropped roughly nine points below the lecturers'. The conventional wisdom is that a good rubric makes marking more reliable, but in this study it made two of four AI systems harsher and less aligned. It is a single exam at a single school, and the models were asked for one first-attempt score with no follow-up, so treat it as a caution about choosing tools rather than an endorsement of any one of them.
At a Spanish medical school, Nogales et al. used Llama 3.1 and Claude 3.5 Sonnet to mark 79 student-written clinical records from simulated patient encounters against the faculty's 48-item rubric, comparing standard one-shot prompting with chain-of-thought (CoT) prompting that makes the model explain each judgement3.
On raw accuracy the two approaches for Claude were close (86.4% one-shot, 85.0% CoT), and specificity was weak at around 60%. Then three experienced professors reviewed the 536 items where Claude and the original human marks disagreed. They concluded the model had been right in 342 of them (63.8%), lifting Claude's accuracy to 94.6% and specificity to 83.3%. Most disagreements sat in the free-text history of present illness, exactly where human markers also struggle. The authors' strongest claim is not that AI marks well, but that its written reasoning exposes inconsistency in the rubric and in the markers.
This is a genuinely useful idea for anyone maintaining an OSCE or portfolio rubric. It also deserves Tolsgaard's warning about automation bias: the adjudicators could consult the model's reasoning when deciding who was right.
Rush et al. take the same logic into research, and bring the most rigorous measurement design of the four4. A large multi-institution team applied generalizability theory to deductive qualitative coding, using three current models (GPT-5.2, Claude Opus 4.5 and Gemini 3-Flash) to code 741 excerpts from an audit of AI policy documents at 146 US medical schools against a 24-subtheme framework.
Coding each excerpt independently, the LLMs agreed with the human consensus about as well as the human coders did, but batching excerpts together significantly reduced agreement (b = −0.41, p = 0.007). The striking result is how much the models agreed with each other (kappa 0.750 to 0.755) compared with how much the humans did (0.422), and how much more patterned their disagreements were (Cramér's V around 0.55 versus 0.20). Their simulations suggest that two humans plus two LLMs, taking the most common answer, would match or slightly beat two humans alone.
The authors are careful to say that higher agreement "does not by itself establish construct validity", because models trained on overlapping data may simply share the same blind spots.
These papers make an uncomfortable point about our benchmarks. In two of the three empirical studies, the human reference turned out to be the least stable part of the system. Perhaps we already know this, from our pre-AI validity studies and exam boards. When AI is introduced to the system, professors reversed a large share of their own marks when shown the model's reasoning, and human coders agreed with each other far less than the machines did.
Human-level agreement is therefore a low and moving bar. So perhaps AI (or an AI-human hybrid) system is the way to go? Nahir et al. and Nogales et al. both point to real workload and consistency gains, and Rush et al. describe a workable hybrid design.
In any case, the advice from Tolsgaard et al. is sound: name the model and version, fix and publish the prompt, keep a human in the loop who is not simply ratifying the machine, and monitor what happens to learners after the score is used. The question to ask of any AI marker is no longer whether it agrees with us, but whether we can defend the decision it helped us make.
Tolsgaard MG, Grierson L, Pusic MV, Turner L. Developing validity arguments for artificial intelligence-based assessment: Balancing affordances and threats. Medical Education. 2026. https://doi.org/10.1111/medu.70309
Nahir M, Kasap A, Bayraktar Nahir C. Comparison of artificial intelligence and lecturers' grading in the evaluation of handwritten open-ended anatomy examinations. Surgical and Radiologic Anatomy. 2026. https://doi.org/10.1007/s00276-026-03986-9
Nogales A, Denizon S, Mateos Rodriguez A, et al. Benefits of Chain-of-Thought Prompting for Clinical Record Rubric Evaluation in Undergraduate Medical Education: Experimental Evaluation Study With Medical Faculty. JMIR Medical Education. 2026. https://doi.org/10.2196/88652
Rush E, Karim MN, Yacu GS, et al. Evaluating hybrid human-LLM coding workflows for qualitative research in medical education: A generalizability study. Medical Teacher. 2026. https://doi.org/10.1080/0142159X.2026.2721362



