Ask the right questions: prompting strategies shape large language model performance on biliary tract cancer guideline queries
Introduction: This study evaluates how different prompting strategies affect the performance of three advanced large language models (LLMs) (GPT-4o, Claude 3.5 Sonnet, and Llama 3 70b) when answering questions about biliary tract cancer (BTC). We used European Society for Medical Oncology (ESMO) gui...
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Article (Journal) |
| Language: | English |
| Published: |
April 29 2026
|
| In: |
Digestive diseases
Year: 2026, Pages: ? |
| ISSN: | 1421-9875 |
| DOI: | 10.1159/000551433 |
| Online Access: | Verlag, lizenzpflichtig, Volltext: https://doi.org/10.1159/000551433 |
| Author Notes: | Chengpeng Li, Wei-Wei Jia, Erfan Ghanad, Hui Gao, Schaima Abdelhadi, Christoph Reißfelder, Flavius Șandra-Petrescu, Cui Yang |
| Summary: | Introduction: This study evaluates how different prompting strategies affect the performance of three advanced large language models (LLMs) (GPT-4o, Claude 3.5 Sonnet, and Llama 3 70b) when answering questions about biliary tract cancer (BTC). We used European Society for Medical Oncology (ESMO) guidelines as our reference standard. The study aimed to assess their accuracy, conciseness, evidence quality, and rates of hallucinations. Methods: We conducted a cross-sectional analysis using 40 clinical questions derived from the ESMO BTC guidelines. We tested three prompting strategies: no prompt, short prompt, and long prompt. Two independent senior physicians evaluated the responses for accuracy, conciseness, and evidence quality. Inter-rater reliability, text length of response, model performance, and hallucination rates were analyzed. Results: Prompting strategies significantly influenced LLM performance. Long prompts improved evidence quality and accuracy, especially for GPT-4o and Claude 3.5 Sonnet, while short prompts enhanced conciseness. GPT-4o exhibited superior overall performance, with higher accuracy and conciseness scores, whereas Claude 3.5 Sonnet excelled in evidence quality but generated longer responses. Llama 3 70b showed deficiencies in both accuracy and evidence quality. Hallucination rates were lowest for GPT-4o and Claude 3.5 Sonnet, but nearly 40% of their references were fabricated or misattributed. Conclusion: Prompting strategies substantially affect LLM performance in medical contexts. While GPT-4o and Claude 3.5 Sonnet demonstrate promising potential with proper prompts, the risk of hallucinations necessitates careful cross-verification. Future studies should incorporate real-world clinical scenarios to further evaluate LLM capabilities and limitations. Doctors and patients are beginning to use large language models - computer programs that can read questions and generate written answers, similar to chatbots - to search for health information. This study examined how well three such programs - ChatGPT-4o, Claude 3.5 Sonnet, and Llama 3 70b - answered questions about cancers of the bile ducts and gallbladder. The way a chatbot responds can change depending on the prompt, which is the short instruction given before asking a question. Three approaches were tested. In the first, the question was asked directly without instructions. In the second, the chatbot was told to act like an experienced doctor and to say when it did not know the answer. In the third, the chatbot was also asked to provide supporting evidence such as medical guidelines or research studies. Forty questions were taken from European clinical guidelines. Two senior cancer specialists reviewed each response and scored it for accuracy, clarity, and the quality of the evidence provided. Clear instructions improved performance for some systems. Detailed prompts increased accuracy and the quality of evidence for GPT-4o and Claude 3.5 Sonnet, while shorter prompts often produced clearer and more concise answers. Overall, GPT-4o performed best for accuracy and brevity, whereas Claude 3.5 Sonnet more often provided supporting evidence but used more words. Llama 3 70b was less reliable. All three systems sometimes produced incorrect or invented references. Because of this risk, information from chatbots should always be checked by medical professionals. |
|---|---|
| Item Description: | Gesehen am 14.07.2026 |
| Physical Description: | Online Resource |
| ISSN: | 1421-9875 |
| DOI: | 10.1159/000551433 |