The landscape of clinical AI is marked by rapid innovation and increasing capabilities, particularly in diagnostic and clinical reasoning domains. Recent studies from Harvard and Stanford highlight AI's potential to significantly impact patient care. For instance, a Harvard-led study published in *Science* demonstrated an OpenAI reasoning model outperforming physicians in emergency department triage and several clinical reasoning tasks, identifying correct or closely related diagnoses in 67% of cases compared to 50-55% for human doctors. Similarly, another Harvard Medical School report suggested that large language models are sufficiently advanced to warrant rigorous clinical testing, having outperformed physicians across multiple complex clinical reasoning tasks. While these findings are promising, they also underscore the critical gap between impressive benchmark performance and the challenges of safe, equitable deployment in real-world clinical settings. Despite these advancements, significant concerns persist regarding the safety, fairness, and reliability of clinical AI. A joint Stanford–Harvard review of over 500 medical AI studies revealed that a substantial portion relied on exam-style questions rather than real patient data, with very few addressing bias, fairness, or uncertainty. This lack of rigorous, outcome-focused evaluation is problematic, especially given instances where clinicians following incorrect AI recommendations worsened patient outcomes. Furthermore, a JAMA study alarmingly found that biased AI decision-support tools could reduce clinicians’ diagnostic accuracy by more than 11 percentage points, even when detailed explanations were provided, indicating that explanations alone are insufficient to mitigate the risks of biased AI. The proliferation of AI in mental health care presents a unique set of challenges. While generative AI shows promise in patient-centered care and can offer support in settings with access gaps, as noted by Harvard's coverage on chatbot therapy, there are serious warnings. Stanford researchers and other reports caution that AI therapy chatbots may produce biased, stigmatizing, or even dangerous responses, potentially encouraging self-harm or worsening distress. The regulatory scramble around therapy chatbots, with states moving to restrict their presentation as therapists, highlights the urgent need for evidence-based validation and explicit safeguards for fairness, algorithmic drift monitoring, and workflow integration to prevent undermining communication or exacerbating inequities. In response to these burgeoning capabilities and risks, there is a strong and consistent call for robust governance and regulatory oversight. The FDA is actively tightening its lifecycle oversight, evidenced by new draft guidance on AI-enabled device software functions, which emphasizes transparency, bias evaluation, and robust plans for managing model updates. Reports to the FDA are critically examining the current regulatory framework, warning of gaps in safety evaluations and post-market surveillance, and advocating for a specialized regulatory body, independent AI auditing, and mandatory disclosure of training data. Experts are proposing comprehensive frameworks for responsible AI deployment. The Journal of the American Medical Informatics Association (JAMIA) suggests leveraging the Clinical Laboratory Improvement Amendments (CLIA) framework as a model for regulating healthcare AI, advocating for coordinated 'collaborative governance' that links centralized regulators, local health system oversight, and standardized performance evaluation. Other recommendations, including a JAMIA advance article, span training data governance, explainability, validation, certification, and continuous monitoring, with explicit attention to data privacy, fairness, and regulatory oversight to prevent biased or unsafe tools from harming patients. The World Health Organization (WHO) has also issued guidance on the ethics and governance of large multimodal models, reinforcing the global consensus on the need for responsible AI development. From the LOG Standards perspective, the current environment necessitates immediate and coordinated action. The consistent findings of potential bias, the limitations of current evaluation methodologies, and the risks of over-reliance on AI underscore the critical importance of independent accreditation and robust lifecycle governance. We advocate for mandatory subgroup performance reporting, standardized fairness assessments, and continuous post-market surveillance to detect and mitigate algorithmic drift. Healthcare stakeholders, including developers, providers, and regulators, must collaborate to establish clear, enforceable standards that prioritize patient safety, ensure equitable outcomes, and build trust in clinical AI systems. The illusion of safety, as one report warns, must be replaced by verifiable, outcome-focused evidence of safe and effective AI integration.