Spertus on HEART SCORE: AI Safety, PROs, and Billion-Dollar Impact

Listen to this article · 8 min listen

The rapid integration of artificial intelligence into healthcare promises transformative changes, yet it simultaneously elevates critical questions around safety, efficacy, and clinical reliability. For clinical informaticists and clinicians alike, the foundational challenge lies in establishing robust validation frameworks that go beyond technical performance metrics to truly capture patient impact. This is where the emphasis on patient-reported outcomes (PROs) becomes not just beneficial, but essential for assessing AI safety, particularly as articulated by leading figures like John Spertus.

The Indispensable Role of Patient-Reported Outcomes in AI Safety Assessment

The development and deployment of AI in clinical settings necessitates a paradigm shift in how we define and measure success. Traditionally, clinical trials have focused on objective endpoints: mortality, rehospitalization rates, laboratory values. While these remain crucial, they often fail to capture the nuances of a patient’s lived experience, their functional status, quality of life, and overall well-being. This gap is precisely what PROs are designed to fill, offering a patient-centered measure of AI safety that complements traditional clinical endpoints. John Spertus, a distinguished figure in cardiovascular outcomes research and a professor at UMKC, has long championed the integration of PROs into clinical practice and research. His work underscores the profound insight that can be gained by directly asking patients about their symptoms, functional limitations, and perceptions of care. In the context of AI, this means assessing whether an AI-powered diagnostic tool, a treatment recommendation system, or a monitoring device actually improves how a patient feels and functions, not just how it alters a biomarker. For instance, an AI might accurately predict a cardiac event, but if the subsequent intervention, partly guided by that AI, leaves the patient with significantly diminished quality of life due to unforeseen side effects or psychological distress, then the overall “safety” and “benefit” of that AI must be re-evaluated through the patient’s lens. The American Heart Association, a key organization in driving cardiovascular health standards, has also increasingly emphasized patient-centered outcomes, aligning with the philosophy that care should be holistic and reflective of patient priorities. When AI systems are designed to assist in complex clinical decision-making, their impact on patient comfort, adherence to treatment, and overall satisfaction becomes a direct measure of their clinical utility and safety. Ignoring PROs risks creating AI solutions that are technically brilliant but clinically irrelevant or, worse, detrimental to patient well-being. The relationship here is clear: patient-reported outcomes provide a patient-centered measure of AI safety that complements clinical endpoints.

Bridging Technical Validation with Human Experience: Harlan Krumholz’s Perspective

The call for integrating PROs into AI validation is echoed by other prominent voices in healthcare innovation, including Harlan Krumholz. Krumholz, known for his work on healthcare quality and outcomes, is affiliated with Yale University and serves as the director of the Yale New Haven Hospital Center for Outcomes Research and Evaluation (CORE). He advocates for a comprehensive approach to evaluating new technologies, emphasizing not just clinical effectiveness but also patient experience and value. For AI, this translates into a demand for evidence that demonstrates not only diagnostic accuracy or predictive power but also how these translate into tangible, positive impacts on patients’ lives. Consider an AI algorithm designed to identify patients at high risk for a particular condition. While its accuracy in identifying these patients is a primary metric, equally important are the patient-reported outcomes related to the subsequent management. Does the AI-driven intervention reduce anxiety? Does it improve their ability to perform daily activities? Does it enhance their understanding of their condition? If the AI leads to overdiagnosis, unnecessary procedures, or increased patient burden without commensurate improvements in quality of life, its safety profile is compromised, regardless of its statistical precision. This holistic view is particularly pertinent given the potential for algorithmic bias to disproportionately affect certain patient populations. If an AI system, due to biases in its training data, leads to suboptimal care or adverse experiences for specific demographic groups, PROs can serve as an early warning system. Patients from these groups might report higher levels of dissatisfaction, increased symptom burden, or poorer functional status, signaling a safety concern that might be missed by purely objective clinical metrics. This makes PROs a crucial component of equitable and safe AI deployment.

Regulatory Frameworks and the Imperative for PROs

The regulatory landscape for AI in healthcare is rapidly evolving, with bodies like the FDA actively developing frameworks to ensure the safety and effectiveness of these technologies. The FDA’s Software as a Medical Device (SaMD) Framework, for instance, provides a structured approach to assessing AI-driven medical software. The FDA’s evolving guidance for AI/ML-based SaMD, including its emphasis on a Total Product Lifecycle (TPLC) approach and Predetermined Change Control Plans (PCCPs), explicitly incorporates real-world evidence (RWE) and post-market surveillance, thereby creating a clear opening for PROs. FDA guidance on SaMD clinical validation As AI models are deployed and continuously learn, monitoring their long-term impact on patients becomes paramount. PROs collected as part of RWE can provide invaluable feedback loops, allowing developers and regulators to identify unforeseen adverse effects or areas where the AI’s performance deviates from its intended patient benefit. This continuous assessment, informed by the patient’s voice, is essential for maintaining trust and ensuring that AI tools remain safe and beneficial over their lifecycle. The integration of PROs into AI validation aligns with the broader move towards value-based care, where patient outcomes and experiences are central to healthcare delivery. For clinical informaticists and clinicians, this means advocating for AI solutions that are not only technically sound but also demonstrably improve the patient journey. Data point DP09 supports the notion that patient-reported outcomes are a critical metric for evaluating the real-world impact and safety of AI in clinical settings.

Conclusion: The Future of Clinically Reliable AI is Patient-Centric

The journey towards clinically reliable AI in healthcare is not solely a technological one; it is fundamentally about ensuring patient safety and improving human health. The insights from experts like John Spertus and Harlan Krumholz underscore that for AI to be truly safe and effective, its impact must be measured through the eyes of the patient. Patient-reported outcomes offer a unique and indispensable lens, providing a patient-centered measure of AI safety that complements clinical endpoints. As we navigate the complexities of FDA pathways and peer-review standards for AI, the consistent integration of PROs will be a hallmark of truly validated and trustworthy AI health tools. This patient-centric approach is not merely an add-on; it is a foundational requirement for building safe AI in healthcare standards that genuinely serve those they are designed to help. American Heart Association position on patient-reported outcomes

Frequently Asked Questions

Why are Patient Reported Outcomes (PROs) essential for assessing AI safety in healthcare?

PROs are essential because traditional objective endpoints often miss the patient’s lived experience, functional status, and quality of life. They provide a patient-centered measure of AI safety that complements clinical endpoints, ensuring AI solutions genuinely improve how a patient feels and functions. This helps re-evaluate AI safety and benefit if interventions guided by AI lead to diminished quality of life.

What is the primary challenge in integrating AI into healthcare for clinical informaticists and clinicians?

The primary challenge is establishing robust validation frameworks that go beyond technical performance metrics to truly capture patient impact. This involves assessing whether AI-powered tools improve patient well-being, not just alter biomarkers or achieve statistical precision. Ignoring PROs risks creating AI solutions that are technically brilliant but clinically irrelevant or detrimental to patient well-being.

How do PROs contribute to identifying potential algorithmic bias in AI systems?

If an AI system, due to biases in its training data, leads to suboptimal care or adverse experiences for specific demographic groups, PROs can serve as an early warning system. Patients from these groups might report higher levels of dissatisfaction, increased symptom burden, or poorer functional status. This signals a safety concern that might be missed by purely objective clinical metrics, making PROs a crucial component of equitable and safe AI deployment.

How do regulatory bodies like the FDA incorporate PROs into their frameworks for AI in healthcare?

The FDA’s evolving guidance for AI/ML-based Software as a Medical Device (SaMD), including its emphasis on a Total Product Lifecycle (TPLC) approach and Predetermined Change Control Plans (PCCPs), explicitly incorporates real-world evidence (RWE) and post-market surveillance. This creates a clear opening for PROs to be collected and used to monitor the long-term impact of AI models on patients.

Editorial Team

The editorial team behind Clinical AI Standards Hub.