Skip to main navigation Skip to search Skip to main content

Large Language Models Can Generate High-Quality Pathology Multiple-Choice Questions Comparable With Questions Written by a Human Expert

Research output: Contribution to journalArticlepeer-review

Abstract

Multiple-choice questions can be effective tools to assess student and trainee performance, but the creation of these questions can be time consuming and requires expertise. To test the quality of pathology test questions created by large language models (LLMs), 100 questions on pancreas pathology were written by a human expert, and 50 questions were generated by each of 2 LLMs (ChatGPT-4.0 and Gemini 2.5 Flash). After an initial review, 16% of the multiple-choice questions generated by the 2 LLMs had to be revised through additional interactive prompting. The final set of questions was then evaluated by 190 volunteers with a variety of backgrounds and levels of expertise. We found that ChatGPT-generated—but not Gemini-generated—questions were rated as easier than human-authored questions; there were slightly more poor/unacceptable questions compared with adequate/good/excellent questions written by the LLMs than those written by the human expert (11.7% vs 10.1%; odds ratio, 1.64; 95% CI, 1.13-2.37; P = .009), but there was no difference in the proportion of questions rated good or excellent. Qualitatively, human-authored questions were thought to be most clinically realistic but felt to be more inconsistent and sometimes thought to be testing trivial points. There was no difference in the mean point biserial between human-authored and LLM-generated questions (0.31 vs 0.29; P = .56). As LLMs improve, they will form a useful tool for the efficient generation of large numbers of high-quality pathology test questions.

Original languageEnglish (US)
Article number100940
JournalModern Pathology
Volume39
Issue number1
DOIs
StatePublished - Jan 2026

Keywords

  • artificial intelligence
  • large language models
  • multiple-choice questions
  • pancreas
  • pathology

ASJC Scopus subject areas

  • Pathology and Forensic Medicine

Fingerprint

Dive into the research topics of 'Large Language Models Can Generate High-Quality Pathology Multiple-Choice Questions Comparable With Questions Written by a Human Expert'. Together they form a unique fingerprint.

Cite this