Generative Artificial Intelligence (AI) is rapidly advancing into clinical workflows, fundamentally reshaping healthcare delivery. Recent reviews in Nature Medicine highlight its growing application in diagnostic support, documentation, and personalized treatment planning, with emerging evidence of benefits alongside significant concerns regarding bias, reliability, and governance. The AMIE system, a conversational clinical AI for disease management, has demonstrated non-inferiority to primary care physicians in management reasoning, excelling in treatment and investigation precision and adherence to clinical guidelines. Similarly, generative AI chatbots are showing promise in mental health, with users reporting positive experiences and a meta-analysis indicating a small-to-moderate, statistically significant reduction in negative mental health outcomes such as depression and anxiety. These developments suggest a future where AI acts as a powerful adjunct to clinical practice, enhancing efficiency and potentially improving patient outcomes. Despite these promising advancements, critical safety and reliability concerns persist. The NOHARM benchmark, developed by Stanford and Harvard researchers, exposed severe clinical errors in medical AI models, with even high-performing models generating severely harmful recommendations in 11.8–14.6% of cases. Notably, 76.6% of these errors were due to omissions, raising alarms about automation bias and patient safety. Further, a Nature Medicine study evaluating an LLM-based clinical decision support (CDS) system in African primary care found 3.4% of AI responses contained hallucinated, fabricated, or clinically inaccurate content, with inappropriate medication recommendations and critical differential diagnosis omissions being prevalent harmful outputs. These findings underscore the profound patient safety risks associated with deploying unvalidated or inadequately monitored AI systems, particularly in resource-constrained environments. Bias remains a pervasive challenge across the entire medical AI pipeline, from data collection to deployment. Imbalanced datasets, provider biases in labels, and real-world implementation issues can lead AI clinical decision support systems to systematically misdiagnose or undertreat marginalized groups, directly impacting patient safety and equity. While some studies, such as one involving GPT-4-based assistance, suggest carefully designed AI can improve decision quality without exacerbating demographic bias, rigorous safeguards are still paramount. The LOG Standards emphasizes that addressing bias is not merely an ethical imperative but a foundational component of AI safety and trustworthiness. In response to these challenges, a robust governance framework is rapidly evolving. The FDA and EMA have jointly issued ten principles for good AI practice across the medicines lifecycle, aiming to harmonize international standards for evidence generation, safety monitoring, and regulatory decision-making. A report to the FDA advocates for mandatory postmarket performance monitoring, training-data transparency, and demographic subgroup evaluation for approved AI tools, calling for community-sourced regulation involving patients and underrepresented groups. ARPA-H has also launched a significant initiative to develop the first FDA-authorized agentic AI system for clinical care, focusing on novel models for AI oversight, lifecycle management, and safety assurance for autonomous clinical decision-making tools. States are also moving forward with their own healthcare AI governance, with over 40 bills across 25 US states targeting AI use in clinical contexts, including requirements for human oversight and restrictions on AI in mental health treatment. The FDA's updated guidance for clinical decision support software and consumer wearables further clarifies regulatory boundaries, signaling a risk-based oversight approach. The National Academy of Medicine cautions that while AI mental health chatbots can offer basic support, they should not be used for diagnosis or crisis intervention, documenting harmful behaviors and proposing safeguards such as banning crisis management by chatbots and restricting minors' data use. From the LOG Standards perspective, these developments highlight the critical need for comprehensive accreditation and continuous monitoring. The emergence of general-purpose chatbots outperforming specialized clinical AI tools on real-world physician questions raises concerns about unvalidated use and underscores the necessity for clear regulatory pathways and validation standards for all AI tools used in clinical settings. As EHR-integrated clinical AI agents become more sophisticated, translating clinician intent into actionable operations, robust safety protocols, oversight mechanisms, and seamless workflow integration become non-negotiable. LOG Standards will continue to advocate for rigorous clinical trials, transparent documentation, and a national safety monitoring and reporting clearinghouse for AI-related adverse events, ensuring that innovation in clinical AI is matched by an unwavering commitment to patient safety and equitable care.