latentbrief
← Back to editorials

Editorial · AI Safety

AI Diagnostic Tool Fails to Impress Clinicians

3h ago3 min brief

The hype surrounding AI diagnostic tools is starting to wear off as clinicians are realizing these systems are not the game changers they were promised to be. A recent study found that non-experts tended to trust AI-generated diagnostic advice even when it was incorrect, while clinicians were more likely to recognize the AI's mistakes. This raises serious concerns about the reliability of these tools and their potential to do more harm than good.

The study tested non-experts and primary care providers in skin disease diagnosis, with and without the help of different explainable AI systems. The results showed that non-experts' diagnostic accuracy improved, but it was largely due to deference to the AI system. They trusted the AI's explanations whether they were right or wrong, and found explanations more convincing when they were vague or generic. On the other hand, clinicians were not tripped up by incorrect AI assistance and performed best when given only a model's prediction, with no accompanying explanation. This highlights the need for AI systems to be designed with users in mind, taking into account their level of expertise and potential biases.

The limitations of AI diagnostic tools are further highlighted by a lawsuit against an AI company, alleging that its chatbot's medical advice nearly caused a patient's death. The chatbot had dismissed the patient's symptoms as minor and advised him to stay immobile, which led to a prolonged period of inactivity that exacerbated his condition. This case underscores the distinction between AI's research potential and its unsupervised use for medical advice. While AI may be able to match or exceed human doctors in diagnostic accuracy within controlled settings, it is not yet ready to be used as a sole source of medical guidance.

The research on AI diagnosis shows that these systems can be excellent diagnosticians in controlled settings, but their performance in real-world medical cases is often poor. A study pitted an AI chatbot against human doctors in clinical cases and found that the chatbot posted a high score, but was also flatly incorrect more often than the human residents. This raises serious concerns about the potential for AI to cause harm if used as a sole source of medical guidance. Clinicians are right to be skeptical of these tools, and it is time for AI companies to take responsibility for the potential consequences of their products.

As we move forward, it is clear that AI diagnostic tools need to be redesigned with users in mind, taking into account their level of expertise and potential biases. We need to develop explainability methods that encourage critical thinking rather than overreliance on the model. The potential consequences of getting this wrong are too great to ignore. We need to stop pretending that AI diagnostic tools are ready for prime time and take a step back to reevaluate their limitations and potential risks. Only then can we start to build AI systems that truly improve healthcare outcomes, rather than putting patients at risk.

Editorial perspective - synthesised analysis, not factual reporting.

If you liked this

More editorials.