Skip to main content
Back to AI Intelligence Briefing
LOG Standards Daily BriefingJuly 29, 2026

Urgent Patient Safety Concerns Mount for Clinical AI: New Reports Highlight Severe Harms, Automation Bias, and Regulatory Gaps

Today's briefing from LOG Standards highlights escalating concerns regarding the safety and reliability of AI-enabled clinical decision support (AI-CDS) systems. Recent studies reveal a troubling incidence of severe harm, with large language models (LLMs) demonstrating significant error rates and a propensity for omissions in critical recommendations. A 2026 State of Clinical AI report found potential severe harm in up to 22% of evaluated cases, predominantly due to errors of omission, while the Stanford–Harvard NOHARM benchmark corroborated these findings, showing leading models producing severely harmful clinical recommendations in 11.8–14.6% of cases. Compounding these risks is the pervasive issue of automation bias among clinicians. Randomized trials indicate that physicians exhibit substantial over-reliance on erroneous AI recommendations, leading to a significant degradation in diagnostic accuracy. This underscores the critical need for 'human-in-the-loop' governance, where clinicians actively oversee AI outputs to mitigate bias and ensure patient safety. While the FDA continues its total product lifecycle oversight for AI/ML medical devices, recent guidance updates have broadened exclusions for certain CDS software, potentially reducing regulatory scrutiny for some AI-enabled tools. In response to these challenges, expert consensus reports are advocating for standardized validation, national-level safety monitoring, and the establishment of a National Healthcare AI-CDS Safety Reporting Clearinghouse to track adverse events and bias-related harms. The Joint Commission has also launched its voluntary Responsible Use of AI in Healthcare (RUAIH) certification, aiming to establish rigorous processes for AI governance in healthcare settings.

This is an original LOG Standards editorial briefing based on the day's reported developments. It is intended for general information and does not constitute clinical, legal, or regulatory advice.
The landscape of clinical artificial intelligence (AI) is marked by both rapid innovation and increasingly urgent patient safety concerns, as evidenced by a series of recent reports. A 2026 State of Clinical AI report, analyzing 31 large language models (LLMs) used for clinical decision support, found potential severe harm in up to 22% of evaluated cases. A significant majority of these harms (77%) were attributed to errors of omission, such as failing to recommend critical tests or treatments. This alarming finding was reinforced by the Stanford–Harvard NOHARM benchmark, which evaluated the same 31 medical LLMs and reported that even leading systems produced severely harmful clinical recommendations in 11.8–14.6 cases per 100, with the worst models exceeding 40 severe errors per 100 cases. Over three-quarters of these harmful errors also stemmed from omissions, raising serious questions about patient safety as these models are integrated into clinical workflows. Further underscoring these risks, a Nature study assessing an LLM-based clinical decision support system in African primary care settings identified hallucinated or clinically inaccurate content in 3.4% of responses. Potentially harmful outputs were dominated by inappropriate medication recommendations (46.2%) and omissions of critical differential diagnoses (31.6%). These findings demonstrate that while such systems can aid clinicians, they pose nontrivial patient safety risks that necessitate rigorous safeguards, ongoing monitoring, and careful local validation before routine use. A critical factor exacerbating these inherent AI limitations is automation bias. The 2026 State of Clinical AI report highlights randomized trial evidence indicating that physicians exhibit substantial over-reliance on erroneous AI recommendations, significantly degrading diagnostic accuracy. This phenomenon underscores the vital importance of 'human-in-the-loop' governance, a principle supported by UC Davis research demonstrating that maintaining human review over AI outputs is critical to reducing bias and protecting patient safety in clinical decision support tools. Clinicians must actively oversee AI-generated recommendations rather than deferring to them, to prevent harmful automation bias and ensure equitable care. In response to these escalating concerns, expert consensus is coalescing around the need for more robust governance and oversight. Recommendations for AI-Enabled Clinical Decision Support, published in the Journal of the American Medical Informatics Association, advocate for standardized validation and certification, national-level safety monitoring, and comprehensive user training. The authors emphasize that AI-CDS can reflect and amplify dataset and development biases related to race, gender, and socioeconomic status, calling for a National Healthcare AI-CDS Safety Reporting Clearinghouse to make adverse events and bias-related harms part of the public record. Complementing these efforts, The Joint Commission has launched the voluntary Responsible Use of AI in Healthcare (RUAIH) certification, designed to help hospitals and health systems establish rigorous processes for validating, monitoring, and overseeing AI in patient care. Regulatory bodies are also adapting, albeit with some nuanced approaches. The FDA's regulation of AI/ML medical devices remains centered on a total product lifecycle approach, with ongoing performance expectations. However, a January 2026 update to Clinical Decision Support guidance has narrowed what counts as a medical device, expanding categories of AI-enabled CDS tools that fall outside FDA device regulation. This shift, alongside updated digital health guidance, signals a potentially more hands-off approach to certain digital health product regulations, which could impact how developers bring products to market and the level of oversight they face. Beyond clinical decision support, the use of AI in mental health presents a complex picture. While a meta-analysis found small but significant benefits from generative AI mental health chatbots for symptoms like depression and anxiety, the authors stress high variability and notable safety and ethical concerns requiring stronger clinical oversight. Expert consensus from the National Academy of Medicine concludes that while general-purpose chatbots can reliably refer users to resources, they are not suitable replacements for licensed therapy, performing poorly at crisis intervention and diagnosis. These findings, coupled with observations of users treating chatbots as ongoing companions, raise critical questions about dependency, clinical safety, and the need for policy recommendations such as banning clinical misrepresentation and restricting the use of minors’ data for personalization. LOG Standards emphasizes that the rapid deployment of AI in healthcare necessitates a proactive and comprehensive approach to governance. The prevalence of severe errors, automation bias, and the complex regulatory landscape demand that all stakeholders prioritize patient safety through rigorous validation, continuous monitoring, and robust ethical frameworks. The emerging certifications and policy discussions are crucial steps, but their effectiveness will hinge on widespread adoption and a commitment to transparency regarding AI performance and potential harms.
LOG Standards

About LOG Standards

LOG Standards provides an independent accreditation signal for healthcare AI. Our AI Intelligence Briefing is published daily, tracking developments in AI safety, AI in medicine, mental health AI, clinical AI governance, and regulatory policy.