Stigmatizing Clinical Language Biased LLM Triage Results
Large language models downgraded priority for patients based on stigmatizing medical inputs in a recent clinical trial.
Updated on Sept. 19, 2026 in Artificial Intelligence

Live Poll
Do you trust artificial intelligence to make fair triage decisions in emergency medical settings?
Researchers found that three open-weight large language models consistently reprioritized patients based on biased clinical language during triage. The study analyzed 221,556 clinical comparisons to determine how attributes like psychiatric history or frequent emergency room usage affected model performance.
Why it matters
The study demonstrates that LLMs can inadvertently mirror societal biases, posing significant safety risks in high-stakes healthcare workflows. This research provides a critical benchmark for the reliability of automated triage systems in emergency medicine.
The study evaluated Gemma, Qwen, and DeepSeek models across 18 distinct input formulations. Prioritization shifts reached up to 9.5 percent when models encountered specific stigmatizing attributes compared to matched patient controls.
The players
Gemma
An open-weight model family developed by Google based on Gemini research.
Qwen
A series of large language models developed by Alibaba Cloud for general and specialized applications.
DeepSeek
A developer of large language models focused on research and open-weight model releases.
The details
Researchers conducted an experiment at two academic emergency departments to test how artificial intelligence models handle clinical triage. By inserting specific social or stigma-related attributes into input data, the team evaluated model output for pairs of patients where one experienced deterioration within a six-hour window. The study measured whether models correctly identified the severity of the ill patient or if stigmatizing language skewed the resulting triage priority.
Timeline
September 19, 2026: The results were published in a peer-reviewed journal.
The Tech Race
This study benchmarks the clinical readiness of open-weight models against established emergency severity standards. It extends research into algorithmic accountability, forcing developers to account for human-language bias in medical AI training data.
These findings highlight the risk of relying on current LLMs for clinical decision-making or triage support without extensive localized testing. Clinicians should exercise caution as these models demonstrate measurable bias that can lead to delayed care for vulnerable patients.
The takeaway
The research confirms that language-based bias can directly degrade clinical triage performance in current open-weight models. Watch for future benchmarks from academic health centers that aim to quantify model safety for patient prioritization.
Further reading
Explore more analysis regarding the safety and accuracy of automated systems in our Artificial Intelligence section.
More information
Review the full findings in the peer-reviewed study on LLM triage bias published in Nature Digital Medicine.
Source note: This article includes information reported by Nature.
Live Poll
Do you trust artificial intelligence to make fair triage decisions in emergency medical settings?






