Blog

New method puts AI decisions to the human test

A new methodology has been developed to measure whether artificial intelligence systems are making decisions to the same standard as qualified human professionals.

The Above-the-Loop Governance methodology uses a sample of AI decisions and compares them with decisions made independently by human experts on the same cases.

The approach is set out in a research paper authored by medical doctor and healthcare executive Dr Andrew Rochford, who is also founder of technology company Trakked.

Under the methodology, human experts assess selected cases without seeing the AI’s decision. Their responses form what the paper describes as a “living benchmark,” which is then used to measure how closely the AI is tracking against professional judgement.

Where AI performance falls below the benchmark for a particular class of decision, decision-making authority would return to humans until the system could again demonstrate the required level of performance.

Dr Rochford said organisations were increasingly using AI for consequential decisions without having a reliable way to demonstrate that those decisions were sound.

“This methodology has been developed because organisations are handing consequential decisions to AI faster than they can prove those decisions are sound,” he said.

“We don’t trust a surgeon or an accountant to make decisions just because they’re clever. We trust them because of everything standing behind them: the training, the registration, the accountability when it goes wrong, and because that trust can be taken away.”

The methodology is designed to avoid requiring a human to review every AI decision. Instead, it uses statistical sampling to assess a proportion of decisions against the human benchmark.

The comparison produces a score using agreement statistics drawn from clinical research and quality science, with sampled cases, expert assessments and resulting actions retained as an auditable record.

The approach also measures whether the human experts themselves agree. Where qualified experts cannot reliably reach the same decision, this would trigger further review and could indicate that the decision is not suitable for delegation to AI.

Dr Rochford said the approach drew on methods already used in areas including clinical trials, financial audits and quality control.

“AI regulations can set the standard, but they can’t tell an organisation whether it’s meeting that standard today, in its own systems,” he said.

“I trained in medicine, where we never test a treatment on every patient, we test a properly drawn sample and monitor it for as long as it’s prescribed.”

The paper proposes three levels of decision-making authority. Decisions where experts themselves cannot consistently agree would remain with humans, while other AI decisions would initially operate alongside human decision-making and continue to be measured against it.

AI could eventually be given conditional authority to make some decisions where sufficient evidence had accumulated to demonstrate its performance, but that authority would be withdrawn if performance deteriorated.

The methodology could be applied to clinical decision-making as well as areas including lending, insurance, hiring and legal services.

In healthcare, the paper says a sample of AI-generated clinical decisions could be compared with decisions made by qualified clinicians, allowing organisations to identify when an AI system begins to move away from current clinical judgement.

The paper has been reviewed by academics from the University of Sydney’s Centre for AI, Trust and Governance, while its statistical methodology has been independently reviewed by University of Sydney Business School lecturer Dr Bradley Rava.

Centre for AI, Trust and Governance co-director Professor Terry Flew said trust was becoming a central issue as organisations increased their use of AI.

“Public confidence is earned when organisations can demonstrate that AI-enabled decisions are accountable, transparent and consistently aligned with professional standards,” Prof Flew said.