Large language models
Large language models (LLMs) are artificial-intelligence systems trained on very large collections of text to identify statistical relationships between language units and generate or transform text in response to prompts.
Large language models (LLMs) are artificial-intelligence systems trained on very large collections of text to identify statistical relationships between language units and generate or transform text in response to prompts. Most contemporary LLMs are based on transformer architectures, which use attention mechanisms to represent relationships among words or tokens across a sequence. Their outputs may include explanations, summaries, classifications, question answering, dialogue, and structured information extraction. Because an LLM generates responses from learned patterns rather than directly consulting a clinical record or verified knowledge base, its accuracy, reliability, completeness, and sensitivity to wording must be assessed for each intended use.
In medicine, LLMs are being investigated as tools for communication, documentation, anonymization, education, clinical decision support, and evidence retrieval. They may be coupled with domain-specific knowledge sources, local deployment, multimodal inputs, human review, or formal evaluation frameworks. The central medical issues are not only performance and precision, but also privacy, bias, error detection, workflow integration, clinician oversight, and the consequences of incorrect or incomplete output. The recent literature represented here is concentrated in healthcare AI evaluation and governance, including assessment of safety, adoption, anonymization, and specialty-specific performance.
- Automated Extraction of Genetic Eligibility Criteria from Clinical Trial Records Using LLMs - A Technical Case Report. PMID 42752480
- Benchmarking AI Vibe Coding for Clinical Statistical Analysis: A Structured Evaluation in Pulmonary Hypertension Research. PMID 42752488
- Large Language Models in Multidisciplinary Decision-Making for Hepatopancreatobiliary Oncology: Retrospective Comparative Feasibility Study. PMID 42753242
Where the papers sit
13 papers study large language models directly. The themes below are drawn from those 13. 1 new direction follows.
-
Clinical AI Decision Support : Clinical LLMs are moving toward benchmarked, locally deployed decision support for diagnosis, prescribing, counselling, research, and trial screening. Accuracy, completeness, safety, anonymization, and clinician oversight recur as requirements for adoption. 8 papers · 61.5%
-
Responsible AI in Healthcare : Healthcare AI is moving toward explicit frameworks for bias, calibration, safety, and accountability rather than unstructured deployment. Digital mental health applications emphasize therapist oversight and clearly bounded AI-supported functions. 3 papers · 23.1%
-
AI in Medical Education : AI is moving from novelty toward structured use in medical teaching and assessment, including Socratic inquiry and Angoff standard setting. Engagement, metacognitive gain, validity, and responsible surveillance remain central safeguards. 2 papers · 15.4%
Large language models serve as judges in medical-examination standard setting
The subject of the paper is modified Angoff standard setting for a basic medical science examination, where large language models were compared with faculty judges in determining the minimally competent candidate’s cutoff score 42715553Sep. Whereas the other papers use LLMs for clinical support, information provision, education, data extraction, anonymization, or evaluation, this study assigns them the distinct role of examination standard setters alongside human faculty. This extends LLM use from assisting with educational content or learner interaction to participating in the formal construction of assessment decisions.
Recent Findings on Large language models
Clinical AI Decision Support: Locally deployed, knowledge-augmented LLMs improved pharmacist review accuracy, sensitivity, and review time while keeping prescription data within the hospital network 42735403Sep. Quantum-enhanced hybrid architectures increased anonymization F1 score and reduced processing time for Portuguese clinical notes 42735382Sep. Clinical research and trial workflows showed strong performance on structured tasks, but variable ambiguity, missing-value handling, hallucinated genes, and incomplete statistical outputs still required validation 42752488Sep42752480Sep. Audiology evaluations found strong clinical recommendations but substantial critical errors in audiometric numerical interpretation, while reproductive counselling models showed high safety but differing completeness and relevance 42727086Sep42720754Sep. HPB oncology recommendations achieved moderate concordance with multidisciplinary team decisions, yet response stability differed substantially between models and repeated queries 42753242Sep. Health care professionals therefore favored lower-risk applications, professional oversight, transparent disclosure, auditing, and human supervision amid concerns about decision errors and algorithmic bias 42742517Sep.
Responsible AI in Healthcare: Therapist-delivered care remains centered on relational, adaptive, and accountability functions, including rupture repair, therapeutic challenge, and clinical judgment under uncertainty 42748037Sep. Digital mental health researchers are mapping evaluation constructs, instruments, procedures, and stages because standardized constructs and validated instruments remain limited; interim screening showed high reviewer calibration 42735424Sep. Bias researchers are similarly building a systematic evidence base through a manually reviewed reference dataset, an LLM-assisted screening workflow, and final decisions by human reviewers 42721473Sep. These frameworks direct LLM deployment toward explicit evaluation of safety, bias, fairness, transparency, accountability, and therapist oversight rather than unstructured clinical use.
AI in Medical Education: AI-supported Socratic inquiry may expand personalized, lower-stakes learning, but fluent questioning can produce conversational mimicry without genuine metacognitive gain 42721480Sep. The proposed Panopticon Paradox links effective tutoring data collection with possible losses of psychological safety, although its causal model remains testable rather than established 42721480Sep. In modified Angoff standard setting, LLM-generated cutoff estimates matched faculty estimates and showed higher reliability, lower rater-related variance, and similar pass rates under standardized prompting 42715553Sep. Medical education is therefore moving toward complementary LLM use with governance, curriculum integration, faculty development, and expert oversight.
Written from 13 PubMed abstracts, each one cited by PMID above. Published: 2026-09-16. Last written: 2026-09-19 by GPT. Drafted by language models from published abstracts; not medical advice.