HomeAI in HealthBenchmarking large language models and clinicians using locally generated vignettes for primary...

Benchmarking large language models and clinicians using locally generated vignettes for primary healthcare in Kenya

Exploring the Capabilities of Large Language Models in Addressing Open-Ended Health Questions

In the evolving landscape of artificial intelligence (AI), large language models (LLMs) are increasingly being recognized for their potential to enhance various sectors, including healthcare. A recent study delves into the ability of LLMs to accurately answer open-ended health questions, particularly in resource-constrained settings. This research is crucial as it evaluates the effectiveness of AI in scenarios where access to medical expertise is limited.

Methods and Analysis

The study compared five distinct LLMs—GPT-4.1, Gemini-2.5-Flash, DeepSeek-R1, MedGemma, and o3—with Kenyan clinicians. The assessment was conducted using a randomly selected subset of 507 vignettes derived from a larger pool of 5107 clinical scenarios. These vignettes spanned 12 nursing competency categories, providing a comprehensive overview of the LLMs’ capabilities.

To ensure unbiased evaluation, responses were rated by blinded physician panels using a 5-point Likert scale. The assessment focused on 11 domains, covering aspects such as accuracy, certainty, contextual appropriateness, and communication. Bayesian ordinal logistic regression was employed to estimate the probability of achieving high-quality scores (≥4) and to facilitate pairwise comparisons between LLMs and human clinicians.

Results

The findings revealed that LLMs outperformed physicians in 9 out of 11 domains. For example, in alignment with guidelines, LLMs scored an average of 4.25-4.72 compared to 2.86 for physicians. Similar trends were observed in domains like expertise (LLMs 4.25-4.73 vs. clinicians 2.76) and logical consistency (LLMs 4.30-4.73 vs. clinicians 2.96).

In safety-related areas, LLMs also received higher ratings. Their scores ranged from 4.29 to 4.68 for minimal extent of possible harm, compared to 3.16 for clinicians. For low probability of damage, LLMs scored between 4.54 and 4.81, significantly higher than the clinicians’ 3.68.

Interestingly, both groups performed similarly regarding the inclusion of irrelevant content and avoidance of demographic bias. In terms of Bayesian models, LLMs demonstrated a >90% probability of achieving scores ≥4 in most domains, while clinicians reached similar probabilities primarily in contextual relevance and demographic/socioeconomic bias.

Among the LLMs, o3 emerged as a leader in most areas, excluding contextual relevance and demographic/socioeconomic bias. The cost of generating LLM responses ranged from $3.86 to $8.68 per model, while clinician-generated vignettes cost $3.35 each.

Conclusions

The study concludes that LLMs, when employed in controlled vignette-based tasks, offer responses that are more accurate, confident, and structured compared to those provided by physicians. These findings suggest that LLMs hold promise as complementary tools for knowledge and safety support in healthcare, particularly in resource-limited settings. The study advocates for further evaluation of LLMs in real patient interactions to assess their effectiveness, safety, and integration into clinical workflows.

For more detailed insights, the full study is available at Here.

“`

Must Read
Related News

LEAVE A REPLY

Please enter your comment!
Please enter your name here