The Times Australia
The Times World News

.
The Times Real Estate

.

Tests that diagnose diseases are less reliable than you’d expect. Here’s why

  • Written by Adrian Barnett, Professor of Statistics, Queensland University of Technology
Tests that diagnose diseases are less reliable than you’d expect. Here’s why

You feel unwell, and visit your doctor. They ask some questions and take some blood for testing; a few days later they call to say you have been diagnosed with a disease.

What are the chances you actually have the disease? For some common diagnostic tests, the answer is surprisingly low.

Few medical tests are 100% accurate. Part of the reason is that people are inherently variable, but many tests are also built on limited or biased samples of patients – and our own work has shown researchers may deliberately exaggerate[1] the effectiveness of new tests.

None of this means we should stop trusting diagnostic tests, but a better understanding of their strengths and weaknesses is essential if we want to use them wisely.

People are variable

An example of a widely used imperfect test is prostate-specific antigen (PSA) screening, which measures the level of a particular protein in the blood as an indicator of prostate cancer.

The test catches an estimated 93% of cancers – but it has a very high false positive rate, as around 80% of men with a positive result do not actually have cancer. For those in the 80%, the result creates unnecessary stress[2] and likely further testing including painful biopsies.

Read more: Prostate cancer testing: has the bubble burst?[3]

Rapid antigen tests for COVID-19 are another widely used imperfect test. A review of these tests[4] found that, of people without symptoms but with a positive test result, only 52% actually had COVID.

Among people with COVID symptoms and a positive result, the accuracy of the tests rose to 89%. This shows how a test’s performance cannot be summarised by a single number and depends on individual context.

Why aren’t diagnostic tests perfect? One key reason is that people are variable. A high temperature for you, for example, might be perfectly normal for someone else. For blood tests, many extraneous factors can influence the results, such as the time of day or how recently you have eaten.

Even the ubiquitous blood pressure test can be inaccurate[5]. Results can vary depending on whether the cuff is a good fit for your arm, if you have your legs crossed, and if you’re talking when the test is done.

Small samples and statistical skullduggery

There’s an enormous amount of research on new diagnostic models. New models frequently make the headlines as “medical breakthroughs”, such as how your handwriting could detect Parkinson’s disease[6], how your pharmacy loyalty card could detect ovarian cancer earlier[7], or how eye movements could detect schizophrenia[8].

But living up to the headlines is often a different story.

Many diagnostic models are developed based on small sample sizes. A review[9] found half of diagnostic studies used just over 100 patients. It is hard to get a true picture of the accuracy of a diagnostic test from such small samples.

For accurate results, the patients who use the test should be similar to those who were used to develop the test. For example, the widely used Framingham Risk Score for identifying people at high risk of heart disease was developed in the United States and is known to perform poorly[10] in Aboriginal and Torres Strait Islander people.

Similar disparities in accuracy have been found for “polygenic risk scores”. These combine information on thousands of genes to predict disease risk, but were developed in European populations and perform poorly in non-European populations[11].

Recently, we identified another important problem: researchers have exaggerated the accuracy of some models[12] to gain journal publications.

There are many ways to exaggerate the performance of a test, such as dropping hard-to-predict patients from the sample. Some tests are also not truly predictive, as they include information from the future, such as a predictive model of infection[13] that includes whether the patient had been prescribed antibiotics.

Read more: Elizabeth Holmes: Theranos scandal has more to it than just toxic Silicon Valley culture[14]

Perhaps the most extreme example of exaggerating the power of a diagnostic test was the Theranos scandal[15], in which a finger-prick blood test supposed to diagnose multiple health conditions attracted hundreds of millions of dollars from investors. This was too good to be true – and the mastermind has now been convicted of fraud.

Big data can’t make tests perfect

In the era of precision medicine and big data, it seems appealing to combine tens or hundreds of pieces of information about a patient – perhaps using machine learning or artificial intelligence – to provide highly accurate predictions. However, the promise is so far outstripping the reality.

One study[16] estimated 80,000 new prediction models were published between 1995 and 2020. That’s around 250 new models every month.

Are these models transforming healthcare? We see no sign of it – and if they really were having a big impact, surely we wouldn’t need such a steady stream of new models.

For many diseases there are data problems that no amount of sophisticated modelling can fix, such as measurement errors or missing data that make accurate predictions impossible.

Some diseases or illnesses are likely inherently random, and involve complex chains of events which a patient cannot describe and no model could predict. Examples might include injuries or previous illnesses that happened to a patient decades ago, which they cannot recall and are not in their medical notes.

Diagnostic tests will never be perfect. Acknowledging their imperfections will enable doctors and their patients to have an informed discussion about what a result means – and most importantly, what to do next.

References

  1. ^ deliberately exaggerate (bmcmedicine.biomedcentral.com)
  2. ^ creates unnecessary stress (theconversation.com)
  3. ^ Prostate cancer testing: has the bubble burst? (theconversation.com)
  4. ^ review of these tests (www.cochrane.org)
  5. ^ can be inaccurate (www.ama-assn.org)
  6. ^ handwriting could detect Parkinson’s disease (www.jpost.com)
  7. ^ detect ovarian cancer earlier (www.theguardian.com)
  8. ^ eye movements could detect schizophrenia (www.abdn.ac.uk)
  9. ^ A review (www.bmj.com)
  10. ^ perform poorly (pubmed.ncbi.nlm.nih.gov)
  11. ^ perform poorly in non-European populations (www.nature.com)
  12. ^ the accuracy of some models (bmcmedicine.biomedcentral.com)
  13. ^ predictive model of infection (www.statnews.com)
  14. ^ Elizabeth Holmes: Theranos scandal has more to it than just toxic Silicon Valley culture (theconversation.com)
  15. ^ Theranos scandal (theconversation.com)
  16. ^ study (osf.io)

Read more https://theconversation.com/tests-that-diagnose-diseases-are-less-reliable-than-youd-expect-heres-why-213359

The Times Features

An Introduction to Complete Hip Replacement Surgery

Hip replacement or total hip arthroplasty is a relatively common medical procedure to regain mobility and bring an end to incessant pain in victims of extreme pain in the hip joi...

2 in 3 Melbourne Families Are Downsizing—But Not for the Reason You Think, Says Big Stuff Movers

MELBOURNE, AUSTRALIA — [16-05-25] — In a city known for its vibrant culture and sprawling suburbs, a quiet revolution is underway. According to recent internal data from Big Stuf...

Runway With a Hug: Gary Bigeni’s Colourful Comeback

By Cesar Ocampo Photographer | AFW 2025 Some designers you photograph once, admire from afar, and move on. But others — like Gary Bigeni — pull you in and never let go. Not becaus...

Tassie’s best pie enters NSW with the launch National Pies’ new fresh range

Fresh from Tasmanian Bakeries in Hobart, National Pies has just delivered Tassie’s best-selling pie to the ready meals aisles of Woolworths stores across NSW.  The delicious roll o...

IORDANES SPYRIDON GOGOS RUNWAY | AFW 2025

Fifth Collection by ISG | Words + Photography by Cesar Ocampo Some runway shows are about the clothes. Others are about the culture they carry. With Iordanes Spyridon Gogos, it’s ...

AJE Resort ‘26 — “IMPRESSION”

Photographed by Cesar Ocampo | AFW 2025 Day 3, Barangaroo Pier Pavilion There are runways, and then there are moments. Aje’s Resort ‘26 collection, IMPRESSION, wasn’t just a fashi...

Times Magazine

Senior of the Year Nominations Open

The Allan Labor Government is encouraging all Victorians to recognise the valuable contributions of older members of our community by nominating them for the 2025 Victorian Senior of the Year Awards.  Minister for Ageing Ingrid Stitt today annou...

CNC Machining Meets Stage Design - Black Swan State Theatre Company & Tommotek

When artistry meets precision engineering, incredible things happen. That’s exactly what unfolded when Tommotek worked alongside the Black Swan State Theatre Company on several of their innovative stage productions. With tight deadlines and intrica...

Uniden Baby Video Monitor Review

Uniden has released another award-winning product as part of their ‘Baby Watch’ series. The BW4501 Baby Monitor is an easy to use camera for keeping eyes and ears on your little one. The camera is easy to set up and can be mounted to the wall or a...

Top Benefits of Hiring Commercial Electricians for Your Business

When it comes to business success, there are no two ways about it: qualified professionals are critical. While many specialists are needed, commercial electricians are among the most important to have on hand. They are directly involved in upholdin...

The Essential Guide to Transforming Office Spaces for Maximum Efficiency

Why Office Fitouts MatterA well-designed office can make all the difference in productivity, employee satisfaction, and client impressions. Businesses of all sizes are investing in updated office spaces to create environments that foster collaborat...

The A/B Testing Revolution: How AI Optimized Landing Pages Without Human Input

A/B testing was always integral to the web-based marketing world. Was there a button that converted better? Marketing could pit one against the other and see which option worked better. This was always through human observation, and over time, as d...

LayBy Shopping