Blog

Clinical AI a long way from calling the shots – study

Just two per cent of AI tools currently in clinical trials are capable of participating in semi-autonomous or closed-loop treatment decision-making, according to a new analysis.

The data suggests that the reality of AI in healthcare remains a long way behind the hype.

The study, published on arXiv, identified 8,532 AI-related clinical trials registered at ClinicalTrials.gov, making it the largest multidimensional characterisation of the clinical AI pipeline assembled to date.

Each trial was classified across seven dimensions: clinical function, data modality, specialty, AI integration and autonomy, workflow position, translational maturity and epistemic role.

The results show that clinical AI has moved beyond retrospective algorithm development and is now generating prospective evidence at scale. However, the data also reveal that AI in healthcare has not yet reached the point of being able to answer one of the most important questions for clinicians – how should I treat this patient?

There are also major imbalances in specialty coverage, geographic representation, translational maturity and evidential quality.

Imaging-based AI remains the dominant modality, accounting for 29 per cent of registered AI trials, but there has been a rapid expansion in trials involving text and natural language processing (NLP)-based systems. NLP-based trials increased seven-fold from 36 in 2018 to 252 in 2025.

An increasing number of trials are focused on prognosis through risk stratification or personalised prediction, particularly in oncology, cardiology and haematology.

Trials of treatment recommendation systems remain relatively uncommon. Just 768 trials are evaluating AI in this setting, accounting for 9 per cent of the total. Only 184 trials involved Level 4 semi-autonomous or closed-loop AI, around 130 of which focused on glucose management.

“This gap is particularly important because prescriptive AI (i.e. systems that actively recommend interventions, doses, or management strategies) likely represents the area of greatest potential clinical impact and weakest prospective evidence base,” the study states.

It attributes this deficit to under-developed regulatory, ethical and liability frameworks needed to support such systems at scale.

Most trials were designed for retrospective validation or silent prospective evaluation, suggesting that much of the field is still generating algorithmic rather than clinical evidence.

The paper notes that “a model that performs well retrospectively or during silent deployment may demonstrate technical competence, but it has not yet shown that it changes clinician behaviour, improves patient outcomes, or justifies deployment costs”.

Additionally, 35 per cent of trials had fewer than 100 participants, which the study says is generally insufficient to demonstrate clinical benefit, subgroup robustness or reliable safety estimates.

The specialty distribution is also markedly unbalanced and does not reflect global disease burden. ENT and otolaryngology along with oncology account for nearly 39 per cent of all trials.

The paper notes that progress in medical AI appears to be following “the path of least data resistance”.

It says these specialties “generate large, standardised imaging datasets that are ideally suited for deep learning. In contrast, rheumatology, haematology, infectious diseases, nephrology, and medical genetics collectively account for fewer than 7 per cent of trials.”

The paper concludes that “most current AI trials remain poorly positioned to answer the questions that matter most to clinicians, regulators, and payers: whether these systems improve outcomes, reduce costs, perform equitably across populations, and remain reliable over time as clinical environments evolve”.

About the author

Asonblog

Add Comment

Click here to post a comment