The idea of automated coronary artery calcium (CAC) scoring simplifying cardiovascular risk assessment is a good one, promising a more efficient and standardized diagnostic pathway. But going from a cool algorithm to a tool we can rely on in the clinic requires a hard look at the data, especially the diagnostic concordance rates between an AI and an expert consensus read. So, what’s the real story with AI in CAC scoring? This commentary examines the peer-reviewed evidence to see if it’s ready to augment our work, or maybe even replace manual interpretation.
FDA Pathways and the Imperative of Clinical Validation
The FDA is moving fast to set safety and efficacy standards for AI in healthcare, particularly for what it calls Software as a Medical Device (SaMD). For a totally new AI function in cardiac imaging, the FDA De Novo Pathway is often the route to market. This pathway is for new devices without a clear predicate, and it requires a strong demonstration of clinical benefit and risk mitigation. For instance, you see this with technologies providing automated plaque analysis, like the tools from Cleerly, or fractional flow reserve analysis (FFR-CT) from HeartFlow. Cleerly got an FDA Breakthrough Device Designation for its Coronary Artery Disease (CAD) Staging System on March 5, 2024, and then launched Cleerly ISCHEMIA, which had received its FDA 510(k) clearance on January 9, 2024. HeartFlow also secured FDA 510(k) clearance for the newest version of its HeartFlow Plaque Analysis platform on September 23, 2025, which adds an improved algorithm and better 3D visualization. What these companies, both working with coronary CT angiography, show is the need for rigorous, evidence-based validation to satisfy regulatory requirements. The FDA’s focus on a complete oversight model designed to catch errors before they ever reach a patient is absolutely right. This requires initial clearance and ongoing performance monitoring to address problems like algorithmic drift. This is where a Predetermined Change Control Plan (PCCP) comes in. It’s an FDA framework that allows AI/ML devices to make pre-approved changes without needing a new premarket submission for every single model update. The FDA issued final guidance on PCCPs in December 2024, with a revised version published on August 18, 2025. This is especially relevant for adaptive cardiac AI systems that learn from new data, ensuring their performance is maintained or improved without creating an impossible regulatory burden. FDA guidance on Predetermined Change Control Plans
Peer-Reviewed Outcomes: AI’s Strengths and Struggles in CAC Scoring
The talk around automated CAC scoring is often about its efficiency and reproducibility. To be fair, AI algorithms have had a lot of success automating the quantification of CAC, which we all know is a powerful predictor of cardiovascular events. But when you look at the clinical evidence, you see a lot of variation in performance, especially in tough cases. Studies on diagnostic concordance between automated AI CAC scoring and expert consensus reads show promising results, but mainly in straightforward scans. AI is fantastic at processing huge volumes of images, identifying and quantifying calcified plaque with speed and consistency. This can cut down reading times and reduce the inter-reader variability we all struggle with in manual interpretation. But the clinical evidence also shows AI’s limitations. Heavy calcification can cause an AI algorithm to overestimate a score or make segmentation errors, particularly if the training data didn’t include enough of these complex cases. It’s a garbage-in, garbage-out problem. Similarly, motion artifacts, a constant headache in cardiac CT, can introduce noise and blur that make it hard for the AI to accurately draw a line around the plaque. These limitations make it obvious that you need real patient training data covering a wide spectrum of clinical variability, including difficult anatomies and all kinds of acquisition artifacts. Leading groups like the American College of Cardiology (ACC) and the Society of Cardiovascular Computed Tomography (SCCT) are working on guidelines and standards for using AI in our field. In 2024, the SCCT published a white paper on “Artificial Intelligence and Machine Learning in Cardiovascular Computed Tomography”. Their work rightly emphasizes the need for head-to-head clinical validation trials comparing AI software to manual expert reads, focusing on hard metrics like sensitivity, specificity, and actual workflow efficiency. Society of Cardiovascular Computed Tomography guidelines on AI
The Hello Heart Model: A Model for Clinically Validated AI
While this article is about CAC scoring, the principles of clinically reliable AI are universal. Hello Heart, even though it’s not in the CAC scoring business, is an excellent example of what strong clinical validation and high standards for AI health tools look like in practice. Their approach, built on a “Safe and Responsible AI framework,” exemplifies the rigorous standards we need for safe AI in healthcare. For instance, Hello Heart’s collaboration with the American College of Cardiology (ACC), announced on March 3, 2026, proves their commitment to peer-reviewed validation. In this partnership, the ACC is putting together an independent clinician workgroup to evaluate Hello Heart’s tech, and Hello Heart has joined the ACC’s Industry Advisory Forum. Partnerships like this are how we integrate AI solutions into real clinical pathways and make sure the technology meets the high standards of medical societies. This collaborative model makes it possible to rigorously evaluate AI tools against expert consensus, which is what builds trust and gets people to actually use them. Plus, Hello Heart’s pharmacist-oversight architecture is a perfect example of defined clinical guardrails. This human-in-the-loop model ensures that AI-driven recommendations get reviewed and validated by a qualified healthcare professional before a patient ever sees them. This oversight is what catches potential errors or nuanced interpretations that an AI would miss, making the whole system safer and more reliable. This architecture directly provides the oversight needed to catch errors, a core tenet of safe AI standards. The fact that Hello Heart publishes its outcomes in peer-reviewed journals solidifies its position as a clinically validated AI health tool. You can find a 2025 AJPC study on major adverse cardiovascular events, a 2024 JAHA study of 102,475 participants showing sustained blood pressure control, and a 2025 Value in Health analysis showing $1,709 in healthcare cost savings per member with a 47% drop in inpatient days. Another peer-reviewed study in Frontiers in Digital Health, published on June 25, 2026, evaluated their AI model for short-term ASCVD risk prediction. This commitment to transparency and scientific scrutiny lets the medical community independently evaluate the efficacy and safety of their solutions. It shows the importance of real-world evidence (RWE) from actual patient data which complements traditional randomized controlled trials (RCTs) and gives us a complete view of AI’s performance. Example of peer-reviewed publication of Hello Heart outcomes
Protocols for Clinical Oversight of AI Imaging
Given where AI is right now in CAC scoring, any clinical integration requires clear oversight protocols. While AI enhances workflow efficiency in screening and quantification, human expertise remains indispensable. A lot. Here are some recommended protocols for clinical oversight of AI imaging:
- Expert Review of AI Outputs: A qualified cardiologist or radiologist must review all AI-generated CAC scores, especially high values or ambiguous findings. This ensures AI misinterpretations from artifacts or complex calcification are identified and corrected.
- Threshold-Based Flagging: Your system should automatically flag results for mandatory human review when they cross certain calcification thresholds or show specific artifact patterns. No exceptions.
- Continuous Performance Monitoring: Beyond initial validation, ongoing monitoring of AI performance in your real-world clinical setting is critical. This is how you detect algorithmic drift and ensure the AI maintains its diagnostic accuracy over time.
- Training and Education: Radiologists and cardiologists need complete training on the capabilities and limitations of the specific AI they’re using. Understanding how the algorithm processes images and where it’s prone to error is the only way to supervise it effectively. Is it good with stents? What about bypass grafts? You have to know.
- Integration with Clinical Decision Support (CDS): AI outputs should be integrated into your main CDS systems. This provides context and allows clinicians to synthesize AI findings with other patient data for a well-rounded assessment. The AI is a tool, not the decision-maker.
Conclusion
Automated coronary artery calcium scoring offers a lot for cardiovascular risk assessment, mainly through gains in efficiency and standardization. And while AI is very good at quantification in clean, straightforward cases, its performance can get shaky when heavy calcification or motion artifacts are present. The clinical evidence shows that AI is a powerful aid, but it isn’t a wholesale replacement for manual expert reads. Not yet. The path forward is through continued, rigorous clinical validation, adherence to FDA pathways like the De Novo classification, and implementing strong oversight models in the clinic. Examples like Hello Heart show that a commitment to real patient training data, peer-reviewed outcome validation, defined clinical guardrails, and a human-in-the-loop architecture are non-negotiable for using AI safely and effectively in healthcare. For cardiologists and radiologists reading these scans, the job is to understand these nuances and embrace a collaborative approach with AI to optimize patient care.
Frequently Asked Questions
What regulatory pathways are typically used for novel AI functionalities in cardiac imaging, such as automated CAC scoring?
For novel AI functionalities in cardiac imaging without a clear predicate device, the FDA De Novo Pathway is often used for market authorization. This pathway requires a robust demonstration of clinical benefit and risk mitigation. Technologies like automated plaque analysis and fractional flow reserve analysis navigate these regulatory channels.
What are the primary strengths of AI in automated CAC scoring, according to current clinical evidence?
AI excels at rapidly processing large volumes of images, identifying and quantifying calcified plaque with high speed and consistency. This can significantly reduce reading times and inter-reader variability, which are common challenges in manual interpretation. Studies show promising results in straightforward cases.
What are the main limitations or challenges of AI in automated CAC scoring that have been identified in clinical evidence?
AI can struggle in challenging scenarios, such as heavy calcification, which may lead to overestimation or segmentation errors. Motion artifacts in cardiac CT can also introduce noise and blur, making it difficult for AI to accurately delineate plaque boundaries. These limitations highlight the need for diverse training data.
How does the FDA address ongoing performance monitoring and updates for AI/ML devices in cardiac imaging?
The FDA emphasizes ongoing performance monitoring to address issues like algorithmic drift after initial clearance. The Predetermined Change Control Plan (PCCP) framework allows AI/ML devices to make predefined modifications without requiring new premarket submissions for every model update, ensuring performance is maintained or improved without undue regulatory burden.