Understanding and managing the AI Paradox in Healthcare
In the fast-evolving landscape of healthcare technology, artificial intelligence presents as both our most promising ally and our most worrisome bedfellow. We can all see the amazing capabilities of AI in everyday life and, rightly, expect that these same capabilities can serve us in our professional lives and public services too. We are constantly exposed to TV, film and social media depictions of the art of the possible and it doesn’t take a PhD in computing science to see how these capabilities could be applied in ways that radically improve our working environments…and yet at the same time we just can’t bring ourselves to fully trust this technology.
This is especially true in healthcare. Patients and professionals alike express enthusiasm about AI’s potential to revolutionize healthcare, yet harbour profound scepticism about its implementation. They want the benefits without the perceived risks. They crave the efficiency but fear the loss of human touch.
This paradox isn’t surprising. Healthcare, at its core, is built on trust. The patient-clinician relationship remains one of the most sacred bonds in our society, and introducing AI into this equation naturally triggers caution. So, is this caution warranted, or just the natural consequence of our sacrosanct treatment of personal health data?
Intelligent stupidity
AI is not new; it is an evolution of mathematical and technological innovation that we have all been exposed to for a long time. However, the exponential pace of this evolution has left us applying our lived experience of AI in its infancy to AI in its current ‘teenage’ existence. This makes us quite reasonably over-cautious. Who can remember sending a text in error on our early mobile phones using the predictive text feature that in many cases simply resulted in an incoherent sentence but on the odd occasion mortifyingly mixed up the name of the recipient or the content of the message? But, hey, nobody died, right?
Advanced AI now has the ability to perform complex tasks with superhuman efficiency but with that comes the ability to make even more monumental errors. There are countless examples of this ‘intelligent stupidity’, with my favourite being the AI powered camera system designed to track football movements during an Inverness Caledonian Thistle FC match in 2020. The AI system mistakenly fixated on the bald head of the linesman instead of the ball, which meant the viewers were treated to 90 minutes of touchline cranium watching. Both the amusing and the very serious mishaps of AI will be amplified relative to the billions of occasions when AI does not make a mistake and in some respects, AI is its own worst enemy, designing algorithms that promote its failure over its successes on social media. The key is striking a balance so that we are neither indifferent to adverse events (I’m reminded of Steve Coogan’s Pool Attendant character in The Day Today ), nor overly cautious to the extent that we never fully realise the true potential of AI’s evolution.
AI is just Maths…
One of the most common misconceptions about AI is treating it purely as a technology or digital entity rather than what it truly is – applied mathematics. This distinction matters enormously for adoption. If we are to mitigate the risk of adverse impact of AI adoption, then we really need the people who understand what’s going on under the hood.
The need for AI should always be identified by clinicians, front line practitioners or operational managers who understand the healthcare requirements or business problems that need addressing. However, the decision to adopt particular AI solutions belongs in the realm of those who understand the models being used, their biases and the consequences of their use. The professional discipline where this mathematical and statistical expertise sits is not very well defined in health and care systems and it can be among some of the clinical, academic, digital or analytical workforce but the Chief Data and Analytics Officer or equivalent should definitely be where you start. Getting this subject matter expertise to rubberstamp your decision to adopt will ensure that although no adoption is without risk you are at least going in eyes wide open so that risk treatment and mitigation can be undertaken. Meanwhile, mobilization, implementation, and user interface design can lean on the wisdom of clinical and digital teams drawing on their user experience experts.
Is it really that important to have the maths and stats geeks green-light your models prior to deployment? Afterall, LLMs (Large Language Models) have become pretty ubiquitous, and they all do a similar job, right? Are the differences between models like Gemini 2.0, Claude 3, GPT-4, or DeepSeek-R1 such that selecting one over another increases a likelihood of harm? The answer is…probably not. They are all trained on massive amounts of similar data. But they are not trained on exactly the same data, and not all of that data is appropriate to the healthcare tasks we are putting it to. If we are to retain public trust in the use of patient data, we must be able to at least have a go at answering the question “why are you using this one (Llama 3.1) over that one (Qwen 2.5)?” Most people I speak to involved in LLM use can’t tell you why the decision was taken to use one over another or even if they understand the options or differences.
On its own LLM adoption of one of the major players such as those listed above, won’t present too much of a risk on a day to day, case-by-case basis but for me the concern is a naive, layered application of AI that compounds inequality and risk.
There is the old adage that ‘history is the best predictor of the future’ and AI is the living embodiment of that philosophy. It can only predict based on what it has learnt. It is widely acknowledged that the history of disease acquisition and healthcare provision is one that is full of systemic bias and inequality . The corresponding data we collect will mirror this and then for good measure, year on year, we add a slightly nuanced political and ideological skew that reflects democracy in action. In a phrase often attributed to management guru, Peter Drucker, “what gets measured, gets managed” and the consequence is that we have perfected a self-reinforcing system where inequality is hard-baked in. If we are not careful, AI becomes our unwitting accomplice in the further disenfranchisement of sections of society.
The language and cultural bias that may be exhibited in the use of the wrong LLM is a small risk to carry but when combined with other AI capabilities that each carry a small flaw, the compounded effect can be devastating. One prominent healthcare leader recently lauded the success rate of AI imaging tools for spotting cancer malignancy in 98% of cases, improving on even the best human clinician rates. This is and should be celebrated but… what can you tell me about the 2% of cases where AI gets it wrong. I’ve not investigated the stats but I wouldn’t be at all surprised if it is those with darker skin, those who may be culturally inhibited from presenting early and for whom there isn’t much training data… in other words those same people who already experience exclusion due to AI such as LLMs on account of English not being their first language; or those who have socio-cultural predisposition to multimorbidity where the AI is only being trained on single condition status. Compounded inequality matters. In many ways the most important and informative data in any analysis is the data which you are not looking at. Clinicians, with their slightly lower success rates and internal biases in the image and pattern recognition space are at least in a position to know what they don’t know. It’s questionable how much the GenAI we are using is able to accommodate this appreciation.
“But we only do stuff that’s evidence based!”
“We’ve tested all this stuff, peer reviewed it!”
Well yes…and no. In many cases the pace of adoption is exceeding the pace of our ability to robustly test our creations. How many initiatives are we rolling out without any form of robust evaluation. Even where evaluation has been conducted it’s not always with the rigour we’d wish for. Oddy et al , researched risk stratification models in use across the world to understand their level of evaluation rigour. The results were quite eye opening with many models actually demonstrating elements of adverse impact and very few being validated in populations outside their population of origin (predominantly the US) despite widespread adoption of these tools internationally. The US, a base source of much AI training data, has a particular health system construction which can mask a homogeneity within the population data that isn’t present in a different system such as the UK. When it comes to the rollout of evidence.. what’s good for the goose isn’t always good for the gander. Although there may be some changes now that see a greater level of equivalence or comparison between the health policies of different countries, in data terms we should always be alert to the long-term analytical consequences of historical differences. For example, the healthcare policies and practices 10 or 20 years ago, such as patient dumping, might still have a material impact on the health status and outcomes of patients or a population today.
This temporal data aspect of AI use also highlights the importance of being able to recognise and accommodate the pace of change in this field.
Looking Beyond Today’s Problems
Perhaps our biggest AI challenge isn’t a technological one but conceptual one. We’re caught in the trap of using cutting-edge tools to solve yesterday’s problems. We build AI systems to tackle documentation backlogs, appointment scheduling, and basic triage – all worthy goals, but hardly transformative. What about tomorrow’s healthcare landscape? The real revolution awaits in image analysis combined with computational simulation (e.g. InSilico Trials), genomic medicine, wearable health monitoring, and whole-person care models. Yet the computational and storage requirements for these applications – and their associated costs – seem to be afterthoughts in our current planning.
Investing in the basics right now – establishing proper data infrastructure, building public trust, developing in-house expertise, and creating ethical frameworks – are as important as the adoption of specific AI solutions themselves and will allow these more advanced applications to develop much more rapidly and organically when the time comes.
A prescription for successful AI adoption.
This article has been full of cautionary tale, but it is important that a responsible but low appetite for risk does not lead to total inertia. There is a way to balance risk with innovation success in the adoption of AI. Perhaps Ironically, it is actually with people and not the computers that this achieved.
Here is my simple 5-point framework to successful AI adoption:
• Be clear about what you mean by AI – Do you mean robots and devices; do you mean image recognition and pattern analysis; do you mean complicated maths; or do you mean the ability to process huge amounts of data really quickly
• Be clear about what you want AI to do – replace humans, augment humans or do things humans can’t do. This is essential if you are going to work out what to stop, otherwise you just double run everything and waste precious resource
• Get people who know what’s going on under the hood to green-light your decision to adopt so that you can do it with confidence and in the knowledge of where the pitfalls are
• Get the data basics right – AI is fuelled by data so put the right amount of effort and resource into data capture, integrity and quality so that the AI that sits on top will evolve organically and with validity
• Don’t over-invest in using AI to solve the problems of today because tomorrow they will be yesterday’s problems. Today’s health problems are described in data and analysis that fits the system of a decade ago – it’s transactional and financial events based. We are already entering an age which is genetically described, person-centred and using real-time surveillance.
I mentioned that the prescription here lies in the people rather than the tech but do we have enough of the right people in the NHS. Those who understand the mathematical underpinnings of AI? Those who can strategically and effectively describe the business problem to be solved? Those with the right translational communication skills? My honest reflection would be…no – at least not yet. We face a significant expertise gap and we need a strategic approach to building this expertise, potentially through the recognition of Data/AI Leaders (as opposed to leaving this to digital and IT leadership); partnerships with academic institutions; creating dedicated AI fellowships and dedicated knowledge mobilisation roles, and establishing centres of excellence that can disseminate knowledge.
The AI paradox in healthcare ultimately resolves to this: technology moves at the speed of innovation, but implementation moves at the speed of trust. By recognizing AI as applied mathematics rather than just another technology solution, we can build the expertise, infrastructure, and public confidence necessary to truly transform healthcare for the better.
Matt Hennessey,
Chief Intelligence and Analytics Officer at NHS Greater Manchester
Honorary Senior Research Fellow at University of Manchester
Ref
[1] https://youtu.be/O8YfgxF3APY?si=6DUbXSB2aeqr4Wce
[2] Marmot M. The health gap: the challenge of an unequal world. Lancet. 2015 Dec 12;386(10011):2442-4. doi: 10.1016/S0140-6736(15)00150-6. Epub 2015 Sep 9. PMID: 26364261.
[3] Oddy C, Zhang J, Morley J, Ashrafian H. Promising algorithms to perilous applications: a systematic review of risk stratification tools for predicting healthcare utilisation. BMJ Health Care Inform. 2024 Jun 19;31(1):e101065. doi: 10.1136/bmjhci-2024-101065. PMID: 38901863; PMCID: PMC11191805.
