🇨🇴⚖️ La Rama Judicial valida a Ariel en prueba de concepto de IA. Conoce los resultados aquí

OIT - Does a General-Purpose Large Language Model Improve Physicians Clinical Reasoning

OIT - Organización Internacional del Trabajo

Icono de documento PDF

Descargar PDF

Disponible

Detalles

Título
OIT - Does a General-Purpose Large Language Model Improve Physicians Clinical Reasoning
Autor
OIT - Organización Internacional del Trabajo
Categoría
Doctrina
Área del derecho
Laboral
Año

X Does a General-Purpose Large Language Model Improve Physicians’ Clinical Reasoning? Evidence and considerations for policy from a randomized trial in Indonesia, Kenya and the Netherlands Authors / Nicholas Rounding, Luthfi Saiful Arif, Janine Berg, Jochen Cals, Diederik De Boer, Eefje De Bont, Sander Dijksman, Ardi Findyartini, Didier Fouarge, Marie-Christine Fregin, Pawel Gmyrek, Nadia Greviana, Ralph Leijenaar, Soraiya Manji, Anastacia Mbithi, Norah Obungu, Arierta Pujitresnani, Roselyter Rianga, Diantha Soemantri, Sairabanu Mohamed Rashid Sokwalla, Sanne Steens, Lucia Velasco, Ardy Wildan, Prasandhya Astagiri Yusuf, Mark Levels

June / 2026 ILO Working Paper 175© International Labour Organization 2026 Attribution 4.0 International (CC BY 4.0) This work is licensed under the Creative Commons Attribution 4.0 International. See: https:// creativecommons.org/licenses/by/4.0/. The user is allowed to reuse, share (copy and redistribute), adapt (remix, transform and build upon the original work) as detailed in the licence. The user must clearly credit the ILO as the source of the material and indicate if changes were made to the original content. Use of the emblem, name and logo of the ILO is not permitted in connection with translations, adaptations or other derivative works. Attribution – The user must indicate if changes were made and must cite the work as follows: Rounding, N., Arif, L., Berg, J., Cals, J., De Boer, D., De Bont, E., Dijksman, S., Findyartini, A., Fouarge,

D., Fregin, M., Gmyrek, P ., Greviana, N., Leijenaar, R., Manji, S., Mbithi, A., Obungu, N., Pujitresnani, A., Rianga, R., Soemantri, D., Sokwalla, S., Steens, S., Velasco, L., Wildan, A., Yusuf, P ., Levels, M. Does a General-Purpose Large Language Model Improve Physicians’ Clinical Reasoning?: Evidence and considerations for policy from a randomized trial in Indonesia, Kenya and the Netherlands. ILO Working Paper 175. Geneva: International Labour Office, 2026.© ILO. Translations – In case of a translation of this work, the following disclaimer must be added along with the attribution: This is a translation of a copyrighted work of the International Labour Organization (ILO). This translation has not been prepared, reviewed or endorsed by the ILO and should not be considered an official ILO translation. The ILO disclaims all responsibility for its content and accuracy. Responsibility rests solely with the author(s) of the translation. Adaptations – In case of an adaptation of this work, the following disclaimer must be added along with the attribution: This is an adaptation of a copyrighted work of the International Labour Organization (ILO). This adaptation has not been prepared, reviewed or endorsed by the ILO and should not be considered an official ILO adaptation. The ILO disclaims all responsibility for its content and accuracy. Responsibility rests solely with the author(s) of the adaptation. Third-party materials – This Creative Commons licence does not apply to non-ILO copyright materials included in this publication. If the material is attributed to a third party, the user of such material is solely responsible for clearing the rights with the rights holder and for any claims of infringement. Any dispute arising under this licence that cannot be settled amicably shall be referred to arbitration in accordance with the Arbitration Rules of the United Nations Commission on International Trade Law (UNCITRAL). The parties shall be bound by any arbitration award rendered as a result of such arbitration as the final adjudication of such a dispute.

Any dispute arising under this licence that cannot be settled amicably shall be referred to arbitration in accordance with the Arbitration Rules of the United Nations Commission on International Trade Law (UNCITRAL). The parties shall be bound by any arbitration award rendered as a result of such arbitration as the final adjudication of such a dispute. For details on rights and licensing, contact: rights@ilo.org. For details on ILO publications and digital products, visit: www.ilo.org/publns.

ISBN 9789220435595 (print), ISBN 9789220435601 (web PDF), ISBN 9789220435618 (epub), ISBN 9789220435625 (html). ISSN 2708-3438 (print), ISSN 2708-3446 (digital) https://doi.org/10.54394/00033425The designations employed in ILO publications, which are in conformity with United Nations practice, and the presentation of material therein do not imply the expression of any opinion whatsoever on the part of the ILO concerning the legal status of any country, area or territory or of its authorities, or concerning the delimitation of its frontiers or boundaries. See: www.ilo. org/disclaimer. The opinions and views expressed in this publication are those of the author(s) and do not necessarily reflect the opinions, views or policies of the ILO. Reference to names of firms and commercial products and processes does not imply their endorsement by the ILO, and any failure to mention a particular firm, commercial product or process is not a sign of disapproval. Information on ILO publications and digital products can be found at: www.ilo.org/researchand-publications ILO Working Papers summarize the results of ILO research in progress, and seek to stimulate discussion of a range of issues related to the world of work. Comments on this ILO Working Paper

are welcome and can be sent to berg@ilo.org.

Authorization for publication: Caroline Fredrickson, Director, Research and Statistics department ILO Working Papers can be found at: www.ilo.org/research-and-publications/working-papers Suggested citation: Rounding, N., Arif, L., Berg, J., Cals, J., De Boer, D., De Bont, E., Dijksman, S., Findyartini, A., Fouarge, D., Fregin, M., Gmyrek, P ., Greviana, N., Leijenaar, R., Manji, S., Mbithi, A., Obungu, N., Pujitresnani, A., Rianga, R., Soemantri, D., Sokwalla, S., Steens, S., Velasco, L., Wildan, A., Yusuf, P ., Levels, M. 2026. Does a General-Purpose Large Language Model Improve Physicians’ Clinical Reasoning?: Evidence and considerations for policy from a randomized trial in Indonesia, Kenya and the Netherlands, ILO Working Paper 175 (Geneva, ILO). https://doi.org/10.54394/0003342501 ILO Working Paper 175

Abstract This paper investigates whether access to a general-purpose large language model (LLM) improves physicians’ clinical reasoning across diverse healthcare contexts, as well as the possible implications of using an LLM in healthcare settings. Using a randomized controlled trial with 249 physicians in Indonesia, Kenya, and the Netherlands, the study finds that LLM access enhances performance on standardized clinical vignettes in all three countries. The magnitude of improvement varies, with the largest gains observed in Kenya (+18%), followed by Indonesia (+10.7%) and the Netherlands (+7.2%). The results, however, reveal substantial heterogeneity. Performance distributions overlap, and some physicians with LLM access perform worse than those without, indicating that access alone does not guarantee improvement. Higher usage is associated with better outcomes, and less specialized physicians appear to benefit more, implying that LLMs may help reduce skill gaps.

The results, however, reveal substantial heterogeneity. Performance distributions overlap, and some physicians with LLM access perform worse than those without, indicating that access alone does not guarantee improvement. Higher usage is associated with better outcomes, and less specialized physicians appear to benefit more, implying that LLMs may help reduce skill gaps. Importantly, the findings emphasize that LLMs function as complements rather than substitutes for clinical expertise. However, the study identifies important risks, including automation bias, hallucinations, and context misalignment, underscoring the importance of careful integration, training, and governance. The paper concludes that while LLMs can enhance clinical reasoning, their effectiveness depends critically on how they are implemented within healthcare systems. The paper recommends that policymakers prioritize structured integration of LLMs as decision-support tools, combined with targeted training, local validation, and safeguards against automation bias rather than relying on access alone. It also emphasizes the need for investment in infrastructure, continuous monitoring, clear liability frameworks, and inclusive governance to ensure equitable, safe, and context-appropriate deployment. It argues for the importance of social dialogue in managing the process. About the authors Nicholas Rounding, Sander Dijksman, Didier Fouarge, Marie-Christine Fregin, Sanne Steens, Lucia Velasco, Mark Levels, Research Centre for Education and the Labour Market, Maastricht University. Luthfi Saiful Arif, Ardi Findyartini, Nadia Greviana, Arierta Pujitresnani, Diantha Soemantri, Prasandhya Astagiri Yusuf, Medical Education Center, Indonesian Medical Education and Research Institute, Universitas Indonesia. Janine Berg, Pawel Gmyrek, Research and Statistics Department, International Labour Organization. Jochen Cals, Eefje De Bont, Ralph Leijenaar, CAPHRI Care and Public Health Research Institute, Maastricht University. Diederik De Boer, Maastricht School of Management, Maastricht University.02 ILO Working Paper 175

Organization. Jochen Cals, Eefje De Bont, Ralph Leijenaar, CAPHRI Care and Public Health Research Institute, Maastricht University. Diederik De Boer, Maastricht School of Management, Maastricht University.02 ILO Working Paper 175 Soraiya Manji, Anastacia Mbithi, Norah Obungu, Roselyter Rianga, Sairabanu Mohamed Rashid Sokwalla, Aga Khan University Hospital, Nairobi. Ardy Wildan, Division of Endocrinology, Metabolism, and Diabetes, Department of Internal Medicine, Universitas Indonesia.03 ILO Working Paper 175 Abstract 01 About the authors 01 X Introduction 06 X 1 What Doctors Do: Physicians’ tasks and the scope for automation 07 X 2 Testing the potential of Generative AI to improve clinical reasoning: A randomized control trial (RCT) in Indonesia, Kenya and the Netherlands 10 2.1. The design of the RCT 11 X 3 Quantitative Findings 14 3.1. Effect of LLM Access on Cross-Country Differences 15 3.2. LLM Vignette Performance 16 3.3. Relationship between LLM Usage on Vignette Performance 16 3.4. Relationship between Specialisation, LLM use and Vignette Performance 17 X 4 Doctors’ perception of the AI assistance 19 X 5 AI use by doctors: Policy considerations 24 X Conclusion 30 Annex 31 References 33 Table of contents04 ILO Working Paper 175 List of Figures Figure 1. Distribution of Potential Task Automation Scores in 6-Digit Medical Occupations in Poland 08 Figure 2. Performance results of physicians with LLM access compared with those without LLM access, Indonesia, Kenya and the Netherlands 14 Figure 3. Physician performance with LLM depending on degree of usage 17 Figure 4. Performance of physicians in internal medicine and other specializations 18

Poland 08 Figure 2. Performance results of physicians with LLM access compared with those without LLM access, Indonesia, Kenya and the Netherlands 14 Figure 3. Physician performance with LLM depending on degree of usage 17 Figure 4. Performance of physicians in internal medicine and other specializations 18 Figure 5. Responses to “my occupation is at risk of being replaced by Generative AI” 20 Figure 6. Responses to “Generative AI will make workers in my occupation more productive” 20 Figure 7. Responses to “using Generative AI at work will reduce my stress” 21 Figure 8. Responses to “using the chatbot improved my performance in this experiment”, Indonesia and Kenya only 22 Figure 9. Responses to “using the chatbot was useful in providing diagnosis”, Indonesia and Kenya only 2305 ILO Working Paper 175 List of Tables Table 1. Baseline Characteristics 11 Table 2. Scores of those doctors “without access” and those “with access” 15 Table 3. Cross-country differences in scores 15 Table A.1. Potential task automation score and justification, Generalist Medical Practitioner (ISCO 2211) 3106 ILO Working Paper 175 X Introduction By 2030, the World Health Organization estimates a shortfall of roughly 11 million healthcare workers, with the burden concentrated in low and lower-middle-income countries (WHO, 2024). Beyond high patient-to-provider ratios, many health systems face an uneven distribution of medical expertise and limited access to diagnostic technologies. These constraints contribute to persistent deficiencies in the quality of care. Physicians are often required to make complex diagnostic and treatment decisions under time pressure, with incomplete information and limited specialist support, increasing the risk of misdiagnosis, delayed treatment, and inappropriate management. Against this backdrop, attention has increasingly turned to whether recent advances in generative Artificial Intelligence (GenAI) can support clinical work. LLMs have been demonstrated to have expert-level medical knowledge (Singhal et al. 2022; 2023). Indeed, empirical studies have

ate management. Against this backdrop, attention has increasingly turned to whether recent advances in generative Artificial Intelligence (GenAI) can support clinical work. LLMs have been demonstrated to have expert-level medical knowledge (Singhal et al. 2022; 2023). Indeed, empirical studies have shown that LLMs can aid doctors in core clinical tasks, such as clinical reasoning and the generation of differential diagnoses, and LLMs have performed strongly in simulated clinical settings (Brodeur et al., 2025; Cabral et al., 2024, McDuff et al., 2025; Nori et al., 2023; Singhal et al., 2025). Recent research indicates that LLMs can enhance quality of care by supporting physicians’ clinical decision-making in both diagnostic and management tasks (Everett et al., 2025; Goh et al., 2024; 2025; Korom et al., 2025). While the capacity of LLMs to improve physicians’ quality of care seems well established, it remains uncertain whether these effects are influenced by cultural differences in clinical reasoning (Findyartini et al., 2016; Karunaratne et al., 2025). This matters because innovations in Generative AI, and specifically Large Language Models (LLMs), have been proposed for application worldwide, as tools that could strengthen healthcare delivery by assisting physicians in their work (Ali et al., 2023, Lam, 2023; Tripathi et al, 2025). This paper examines whether access to a large language model improves physicians’ clinical reasoning performance in a controlled experimental setting. Using standardized clinical vignettes, we conduct a randomized controlled trial with physicians in Indonesia, Kenya, and the Netherlands. By comparing performance with and without model access across these three contexts, the study provides evidence on the potential of GenAI to augment clinical reasoning and on the extent to which such tools may help reduce or reinforce existing differences in quality of care in different cultural and institutional contexts. We also consider how physicians feel about using the tool and some policy implications from integrating LLMs into diagnostics.07 ILO Working Paper 175

which such tools may help reduce or reinforce existing differences in quality of care in different cultural and institutional contexts. We also consider how physicians feel about using the tool and some policy implications from integrating LLMs into diagnostics.07 ILO Working Paper 175 X 1 What Doctors Do: Physicians’ tasks and the scope for automation

One of the principal approaches in assessing the possible effects of technology on jobs is to analyse the ability of a specific technology to automate tasks within an occupation. Such an approach is based on the insight that jobs are best understood as “bundles of tasks” (Autor, 2015). As such, the question of whether task automation leads to job automation depends on the extent to which tasks within an occupation can be automated and the importance of such automatable tasks to the occupation (Autor and Thompson 2025). Jobs evolve over time, thus automation of tasks within an occupation does not necessarily imply the disappearance of the job itself, particularly if such tasks are peripheral or if there are other tasks within the occupation that continue to require human input (Autor et al. 2024). Following this approach, Gmyrek et al. (2025) use machine learning and human judgement to derive scores of exposure to GenAI for the tasks of internationally standardized ISCO-08 occupations.1 Each of the 436 4-digit occupations in the ISCO-08 structure is classified on an exposure gradient that ranges from no or minimal exposure to gradient 4 (highest exposure to potential automation). The assessment is a theoretical estimate of what tasks can currently be performed using the technology rather than its application in practice, which may be constrained by an interplay of factors, including inadequate infrastructure or skills, high costs, competing organizational priorities, or even disappointing performance of the technology. The occupation of “generalist medical practitioner” (ISCO 2211), is defined in ISCO-08 as family and primary care doctors who “diagnose, treat and prevent illness, disease, injury and other physical and mental impairments and maintain general health in humans through application

The occupation of “generalist medical practitioner” (ISCO 2211), is defined in ISCO-08 as family and primary care doctors who “diagnose, treat and prevent illness, disease, injury and other physical and mental impairments and maintain general health in humans through application of the principles and procedures of modern medicine” (ILO, 2012).2 Specialist physicians (2212) are classified separately at the 4-digit level, while national classifications will often elaborate significantly more detailed levels. For example, the Polish 6-digit system distinguishes more than 80 physician occupations, the majority corresponding to specialised fields. The task descriptions associated with physicians’ work in ISCO-08 therefore reflect a set of generic, cross-cutting medical activities rather than the full procedural specificity of individual specialisations. They provide an aggregated representation of core clinical, diagnostic and administrative functions that characterise medical practice at the international classification level. Despite substantial variation in clinical focus and procedural intensity, diagnostic reasoning and treatment decision-making, remain core components of work across these roles. The average potential task automation score for generalist medical practitioner (2211) at 0.29 is relatively low, compared with other ISCO-08 occupations in the ILO index. Dispersion across tasks is also limited (standard deviation of 0.1), indicating that exposure to GenAI is not driven 1 Each task within an occupation is assigned a score ranging from 0 to 1, where 0 is no possibility for automating the specific task with generative AI technology and 1 represents full automation potential. The task scores are then averaged and based on the average score and its standard deviation (i.e., the range of scores of each of the tasks within an occupation). See Gmyrek et al. (2025) for methodological details. 2 The ISCO-08 classification is hierarchical, organised into major groups (1-digit), sub-major groups (2-digit), minor groups (3-digit), and unit groups (4-digit). Health professionals are located within Major Group 2 (Professionals), specifically under Sub-major Group

methodological details. 2 The ISCO-08 classification is hierarchical, organised into major groups (1-digit), sub-major groups (2-digit), minor groups (3-digit), and unit groups (4-digit). Health professionals are located within Major Group 2 (Professionals), specifically under Sub-major Group 22 (Health Professionals). Within this structure, separate 4-digit unit groups distinguish Generalist Medical Practitioners (2211) from Specialist Medical Practitioners (2212) and other health occupations.08 ILO Working Paper 175 by a small subset of highly automatable tasks but rather reflects the overall structure of medical work. (See Appendix Table 1).3 A more granular view, using the detailed 6-digit classification from the Polish 6-digit system, helps clarify this structure (Figure 1). Across physician occupations, tasks cluster into three broad functional categories: diagnostic and clinical reasoning tasks; administrative and documentation-related tasks; and manual or procedural tasks. These categories exhibit distinct exposure patterns. Administrative and information-processing activities tend to receive relatively higher scores, while diagnostic and treatment decision-making tasks, as well as manual and embodied procedures, consistently fall at the lower end of the distribution. By contrast, tasks central to medical practice (physical examinations, surgical and clinical procedures, diagnosis, treatment planning, and patient interaction) consistently receive low scores in the ILO framework, reflecting the need for judgment, ethical responsibility, and human engagement. Even where AI can contribute to diagnostic support, referrals, counselling, or preventive planning, its role remains complementary rather than substitutive. The lowest scores are generally assigned to manual and procedural tasks that depend on hands-on clinical practice and in-person patient interaction. X Figure 1. Distribution of Potential Task Automation Scores in 6-Digit Medical Occupations in Poland As a result of this scoring distribution, most medical professions are not classified as exposed to the risk of automation by the ILO. Potential augmentation benefits, where identified, are concentrated primarily in administrative and documentation-related tasks. Nevertheless, even if GenAI cannot substitute for the core clinical functions that are central to the profession, there remains

the risk of automation by the ILO. Potential augmentation benefits, where identified, are concentrated primarily in administrative and documentation-related tasks. Nevertheless, even if GenAI cannot substitute for the core clinical functions that are central to the profession, there remains scope for the technology to support decision-making. Many recent investments in healthcare AI have been focused on developing such tools, including AI-enabled medical devices that assist diagnosis, particularly in imaging-heavy specialties. Regulators have been keen to support these developments; for example, the U.S. FDA approved 223 AI-enabled medical devices in 2023, primarily in the areas of detection and clinical interpretation support (Stanford HAI, 2025).4 3 Appendix Table 1 presents the scores for 10 tasks associated with this occupation, based on the ILO methodology. 4 There have also been commercial investments in scheduling, revenue-cycle and prior-authorization automation, and patient engagement or education assistants. Most of these investments have been framed as improving efficiency, standardization, and access to information rather than replacing clinical judgment (Stults et al, 2025; Kunze et al. 2025).09 ILO Working Paper 175 At the same time, important questions arise when considering the deployment of such tools in developing countries. Some of the strongest claims regarding AI in health concern its potential to expand access in contexts characterized by weak healthcare coverage and low physician density. In such settings, partial automation or decision-support systems are often presented as a means of extending diagnostic capacity, reducing misclassification, lowering unnecessary testing costs, and improving quality of care. However, most currently deployed AI-supported healthcare systems have been developed and validated in high-income countries, raising questions about their ability to perform well in lower-middle-income countries (Riang’a et al., 2025). Such concerns relate not only to differences in disease profiles and clinical practice patterns, but also to language, infrastructure, and data availability. Tools optimized for one regulatory, epidemiological, or linguistic context may not translate seamlessly to another. Whether these expectations are realistic remains an open empirical question. The analysis presented here does not assess health system infrastructure or regulatory integration. Instead, it

availability. Tools optimized for one regulatory, epidemiological, or linguistic context may not translate seamlessly to another. Whether these expectations are realistic remains an open empirical question. The analysis presented here does not assess health system infrastructure or regulatory integration. Instead, it focuses on a narrower but central issue: whether access to a large language model improves individual physicians’ clinical reasoning performance in controlled vignette-based assessments across three countries. While this approach does not capture the full complexity of real-world deployment, it provides evidence on a core dimension of medical work.10 ILO Working Paper 175 X 2 Testing the potential of Generative AI to improve clinical reasoning: A randomized control trial (RCT) in Indonesia, Kenya and the Netherlands

Clinical reasoning lies at the heart of physicians’ work. It involves seeing and hearing the patient, probing for vital information, coming to a diagnosis, and managing the outcomes. Effective clinical reasoning is essential for high quality of care. Yet, clinical reasoning is a highly complex, multifaceted and idiosyncratic skill (Pelaccia et al., 2011). How physicians end up with their final differential diagnosis is of some debate. Drawing on dual process theory, most models of clinical reasoning distinguish between two separate modes: the intuitive mode and the analytical mode. With the intuitive mode, physicians draw on their experience to make links between the problem they are presented with and patterns stored in long-term memory. The analytical mode relies on hypothetico-deduction, a process where diagnostic hypotheses are generated and then tested by the physician. These two models are not mutually exclusive and often used in conjunction (Eva, 2005). Utilising both modes during clinical reasoning processes increases the performance of the physicians compared to employing either one alone (Ark et al., 2006; Norman et al., 1999;

Norman & Eva, 2010). The relationship between LLMs and clinical reasoning is not yet entirely clear. LLMs can potentially employ both analytical and intuitive reasoning approaches. By identifying statistically important correlations between patient information and potential diagnoses, LLMs rely on largescale pattern recognition to identify diagnoses. An overreliance on pattern recognition and data retrieval could limit LLMs abilities in the face of logical errors and changing contexts, requiring human oversight (Christof & Armoundas, 2025). Evidence suggests that LLMs demonstrate biases toward inflexible pattern matching, rather than engaging in flexible reasoning (Kim et al., 2025). However, advanced prompting strategies such as chain-of-thought or retrieval augmented generation may be able to mimic certain aspects of the reasoning process (Lewis et al., 2021; Wei et al., 2023).5 Additionally, LLM models are available that are explicitly trained to respond to medical questions, with some studies suggesting that the most up-to-date medical trained models can outperform the generalist models such as GPT-4o or GPT-o3 (e.g., see McDuff et al, 2025; Wang et al, 2025), but other studies, including a benchmark review, finding that generalist foundation models (such as GPT-4o) outperformed specialist trained models (Nori et al, 2023). In either case, LLMs matched or exceeded medical student performance when tested on international, diverse clinical vignette datasets designed to assess how new information adjusts diagnostic or therapeutic judgments under uncertainty (McCoy et al., 2025). Whether through pattern recognition or intuitive reasoning, LLMs offer an opportunity to augment physician clinical reasoning. Despite this augmentation potential, there are significant risks associated with LLM use. For example, LLMs may provide inappropriate and poor-quality answers, as illustrated in attempts of diagnosing complex urology cases (Cocci et al., 2025). Other studies show that it failed to 5 Chain-of-thought (CoT) prompting is a prompt engineering technique that instructs AI models to break down complex problems into

example, LLMs may provide inappropriate and poor-quality answers, as illustrated in attempts of diagnosing complex urology cases (Cocci et al., 2025). Other studies show that it failed to 5 Chain-of-thought (CoT) prompting is a prompt engineering technique that instructs AI models to break down complex problems into intermediate, step-by-step logical steps before delivering a final answer; Retrieval-Augmented Generation (RAG) is the process of optimizing the output of a large language model, so it references an authoritative knowledge base outside of its training data sources before generating a response.11 ILO Working Paper 175 accurately diagnose patients across various pathologies, performing significantly worse than physicians (Hager et al., 2025). One issue is that LLMs’ outputs are biased by the type of medical data they are trained on, which reflect the characteristics of the individuals in the training data (Zhang et al., 2024). When information on patients’ disease history is biased, LLMs are equally likely as human doctors to reach fallacious conclusions (Schmidt et al., 2024). Another reason for caution is that hallucination is an intrinsic feature of LLM models (Yao et al., 2023) that seriously hampers LLM’s ability to reliably process and communicate medical information (Gilbert, Kather and Hogan, 2024). Furthermore, LLMs can introduce new biases into the clinical reasoning process. Automation b

Estás viendo una vista previa

Lee el documento completo con Ariel

Este es un fragmento de uno de los más de 1.2 millones de documentos de la biblioteca de Ariel. Crea tu cuenta para leerlo completo, descargarlo y consultarlo con Ariel, que siempre te lleva a la fuente exacta: Ariel NO alucina.

Consultar sobre este documento ...