Reliability of large language models in managing odontogenic sinusitis clinical scenarios: a preliminary multidisciplinary evaluation.

Alberto Maria Saibene, Fabiana Allevi, Christian Calvo-Henriquez, Antonino Maniaci, Miguel Mayo-Yáñez, Alberto Paderno, Luigi Angelo Vaira, Giovanni Felisati, John R Craig

European Archives of Oto-rhino-laryngology 2024 January 9

PURPOSE: This study aimed to evaluate the utility of large language model (LLM) artificial intelligence tools, Chat Generative Pre-Trained Transformer (ChatGPT) versions 3.5 and 4, in managing complex otolaryngological clinical scenarios, specifically for the multidisciplinary management of odontogenic sinusitis (ODS).

METHODS: A prospective, structured multidisciplinary specialist evaluation was conducted using five ad hoc designed ODS-related clinical scenarios. LLM responses to these scenarios were critically reviewed by a multidisciplinary panel of eight specialist evaluators (2 ODS experts, 2 rhinologists, 2 general otolaryngologists, and 2 maxillofacial surgeons). Based on the level of disagreement from panel members, a Total Disagreement Score (TDS) was calculated for each LLM response, and TDS comparisons were made between ChatGPT3.5 and ChatGPT4, as well as between different evaluators.

RESULTS: While disagreement to some degree was demonstrated in 73/80 evaluator reviews of LLMs' responses, TDSs were significantly lower for ChatGPT4 compared to ChatGPT3.5. Highest TDSs were found in the case of complicated ODS with orbital abscess, presumably due to increased case complexity with dental, rhinologic, and orbital factors affecting diagnostic and therapeutic options. There were no statistically significant differences in TDSs between evaluators' specialties, though ODS experts and maxillofacial surgeons tended to assign higher TDSs.

CONCLUSIONS: LLMs like ChatGPT, especially newer versions, showed potential for complimenting evidence-based clinical decision-making, but substantial disagreement was still demonstrated between LLMs and clinical specialists across most case examples, suggesting they are not yet optimal in aiding clinical management decisions. Future studies will be important to analyze LLMs' performance as they evolve over time.

Full text links

We have located links that may give you full text access.

Show additional links to paperHide additional links to paper

PubMed

Add to Saved Papers

Get 1-tap access

Related Resources

Haemodynamic monitoring during noncardiac surgery: past, present, and future.Karim Kouz et al.Journal of Clinical Monitoring and Computing 2024 April 31

2024 AHA/ACC/AMSSM/HRS/PACES/SCMR Guideline for the Management of Hypertrophic Cardiomyopathy: A Report of the American Heart Association/American College of Cardiology Joint Committee on Clinical Practice Guidelines.Steve R Ommen et al.Circulation 2024 May 9

Obesity pharmacotherapy in older adults: a narrative review of evidence.Alex E Henney et al.International Journal of Obesity 2024 May 7

Use of Intravenous Albumin: A Guideline from the International Collaboration for Transfusion Medicine Guidelines.Jeannie Callum et al.Chest 2024 March 5

''Myth Busting in Infectious Diseases'': A Comprehensive Review.Ali Almajid et al.Curēus 2024 March

SGLT2 Inhibitors in Kidney Diseases-A Narrative Review.Agata Gajewska et al.International Journal of Molecular Sciences 2024 May 2

For the best experience, use the Read mobile app

Get seemless 1-tap access through your institution/university

For the best experience, use the Read mobile app

All material on this website is protected by copyright, Copyright © 1994-2024 by WebMD LLC.
This website also contains material copyrighted by 3rd parties.

By using this service, you agree to our terms of use and privacy policy.

Your Privacy Choices

You can now claim free CME credits for this literature searchClaim now

Get seemless 1-tap access through your institution/university

For the best experience, use the Read mobile app

Reliability of large language models in managing odontogenic sinusitis clinical scenarios: a preliminary multidisciplinary evaluation.

Full text links

Related Resources

Trending Papers

For the best experience, use the Read mobile app