AI healthcare can’t beat a specialist

By News EditorPublished On: September 9, 2024Last Updated: September 9, 2024

The capabilities, limitations and risks of generative AI are currently a topic of major interest, with widely varying predictions about the roles that language models may one day be able to fill. An area that’s frequently brought up in this regard is healthcare, and this study is certainly food for thought.

Researches from several institutions, including the School of Clinical Medicine at the University of Cambridge, tasked several language models with answering a large number of ophthalmology-based multiple-choice questions taken from a medical textbook.

The same questions were shown to eye doctors, junior eye doctors, and junior doctors who haven’t yet picked a specialty, with the latter group intended to roughly correspond to the level of ophthalmology knowledge that might be expected from a GP. The full results are published here.

To summarise, the language model performance varied widely but the best results were from a model called GPT-4, which answered 69 per cent of the questions correctly. That was significantly better than the unspecialised doctors (43 per cent on average) and slightly better than the ophthalmology trainees (59 per cent), but worse than the ophthalmologists (76 per cent on average, with the highest mark being 90 per cent).

As Dr Arun Thirunavukarasu, who led the study, suggests in the pull quote below, one could speculate from this that language models won’t replace specialists, but may one day be suitable for a role in triage – determining as a first port of call whether a case is serious enough to be referred to a specialist for an expert opinion, in the same way as a GP does currently.

On the other hand, it strikes me that multiple-choice textbook questions are likely to play to GPT-4’s strengths. The researchers note that the questions weren’t used as part of the language model’s training data, but nonetheless it wouldn’t be surprising if other textbooks had similar questions that the AI might have encountered during training.

And perhaps more obviously, receiving a written summary of a patient’s condition is very different from being confronted with a real-life patient and examining their eyes yourself, which a language model wouldn’t seem to stand much chance of doing.

Textbook questions are usually written in a way that’s intended to lead the reader towards the correct answer, whereas in reality everyone’s eyes are different and human judgment would seem an essential ingredient in a correct diagnosis. As the authors note, “Examination performance is an unvalidated indicator of clinical aptitude”.

Still, this may be a hint at what could one day be possible as AI continues to develop.

Half of FDA-approved AI medical devices not trained on real patient data

AI will surpass human brains once we crack the ‘neural code’

Cookie	Duration	Description
__cfduid	1 month	The cookie is used by cdn services like CloudFare to identify individual clients behind a shared IP address and apply security settings on a per-client basis. It does not correspond to any user ID in the web application and does not store any personally identifiable information.
__hssrc	session	This cookie is set by Hubspot. According to their documentation, whenever HubSpot changes the session cookie, this cookie is also set to determine if the visitor has restarted their browser. If this cookie does not exist when HubSpot manages cookies, it is considered a new session.
cookielawinfo-checkbox-advertisement	1 year	The cookie is set by GDPR cookie consent to record the user consent for the cookies in the category "Advertisement".
cookielawinfo-checkbox-analytics	1 year	This cookies is set by GDPR Cookie Consent WordPress Plugin. The cookie is used to remember the user consent for the cookies under the category "Analytics".
cookielawinfo-checkbox-necessary	1 year	This cookie is set by GDPR Cookie Consent plugin. The cookies is used to store the user consent for the cookies in the category "Necessary".
cookielawinfo-checkbox-performance	1 year	This cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Performance".

Cookie	Duration	Description
__hssc	30 minutes	This cookie is set by HubSpot. The purpose of the cookie is to keep track of sessions. This is used to determine if HubSpot should increment the session number and timestamps in the __hstc cookie. It contains the domain, viewCount (increments each pageView in a session), and session start timestamp.
tve_leads_unique	1 month	This cookie is set by the provider Thrive Themes. This cookie is used to know which optin form the visitor has filled out when subscribing a newsletter.

Cookie	Duration	Description
__hstc	1 year 24 days	This cookie is set by Hubspot and is used for tracking visitors. It contains the domain, utk, initial timestamp (first visit), last timestamp (last visit), current timestamp (this visit), and session number (increments for each subsequent session).
_ga	2 years	This cookie is installed by Google Analytics. The cookie is used to calculate visitor, session, campaign data and keep track of site usage for the site's analytics report. The cookies store information anonymously and assign a randomly generated number to identify unique visitors.
_gid	1 day	This cookie is installed by Google Analytics. The cookie is used to store information of how visitors use a website and helps in creating an analytics report of how the wbsite is doing. The data collected including the number visitors, the source where they have come from, and the pages viisted in an anonymous form.
hubspotutk	1 year 24 days	This cookie is used by HubSpot to keep track of the visitors to the website. This cookie is passed to Hubspot on form submission and used when deduplicating contacts.

Cookie	Duration	Description
cookielawinfo-checkbox-functional	1 year	The cookie is set by GDPR cookie consent to record the user consent for the cookies in the category "Functional".
cookielawinfo-checkbox-others	1 year	No description
lfuuid	9 years 11 months	Third party (Lead Forensics) cookie which enables us to track visitor behaviour on our site. Tracking is performed anonymously until a user identifies themselves by submitting a form.
tl_554_555_1	1 month	No description
tl_554_605_2	1 month	No description
tlf_1	5 days	No description