Introduction
Artificial intelligence (AI), and in particular the family of machine learning methods known as deep learning, has moved from experimental proof of concept toward clinical evaluation across a range of diagnostic tasks over the past decade. Convolutional neural networks trained on large labelled datasets can now detect diabetic retinopathy, classify skin lesions, and identify abnormalities on chest radiographs at a level that, under controlled conditions, approaches that of specialist clinicians (Gulshan et al., 2016; Esteva et al., 2017; Topol, 2019). For a health system such as Australia’s, which faces an ageing population, rising demand for diagnostic imaging, and persistent workforce shortages in rural and remote regions, the prospect of accurate and scalable diagnostic support is appealing (Australian Institute of Health and Welfare [AIHW], 2024). Enthusiasm in the research literature has, however, frequently outpaced rigorous evidence of clinical benefit, and questions of accuracy, equity, workflow integration, and regulation remain unresolved.
This review synthesises peer-reviewed literature on the application of AI to medical diagnosis and considers its implications for the Australian setting. The discussion is organised around six themes: diagnostic imaging; clinical decision support; diagnostic accuracy relative to clinicians; implementation and workflow; bias, equity, and explainability; and regulation. The aim is to characterise both the demonstrated capabilities of diagnostic AI and the limitations that temper its translation into routine care. Figure 1 illustrates the diagnostic pipeline that frames the analysis, and Table 1 summarises eight representative studies that recur throughout the review.
Search strategy
Relevant literature was identified through PubMed, Scopus, IEEE Xplore, and Google Scholar, using combinations of the terms “artificial intelligence”, “deep learning”, “machine learning”, “diagnosis”, “medical imaging”, and “clinical decision support”. The search was restricted to English-language, peer-reviewed publications between 2015 and 2025, a window chosen because it captures the rapid expansion of deep learning in medicine that followed breakthroughs in general image recognition. Systematic reviews, meta-analyses, and external validation studies were prioritised over single-centre reports, and the reference lists of key papers were screened for additional sources. Australian regulatory and policy material was drawn from the Therapeutic Goods Administration (TGA), the Australian Digital Health Agency, the AIHW, the Royal Australian and New Zealand College of Radiologists (RANZCR), and the National Health and Medical Research Council (NHMRC) to situate the international evidence within a local context.
Table 1: Summary of eight representative studies of artificial intelligence in medical diagnosis, 2016 to 2021.
| Author and year | Context | Method | Key finding |
|---|---|---|---|
| Gulshan et al. (2016) | Diabetic retinopathy screening (ophthalmology) | Convolutional neural network trained on approximately 128,000 retinal fundus images | Sensitivity and specificity for referable retinopathy comparable to ophthalmologists |
| Esteva et al. (2017) | Skin cancer classification (dermatology) | Network trained on roughly 130,000 clinical images, tested against dermatologists | Dermatologist-level accuracy in distinguishing malignant from benign lesions |
| Rajpurkar et al. (2018) | Chest radiograph interpretation (radiology) | CheXNeXt model compared with practising radiologists across fourteen pathologies | Matched radiologist performance on most pathologies with faster reporting |
| Liu et al. (2019) | Multiple imaging domains (systematic review) | Meta-analysis of eighty-two diagnostic accuracy studies of deep learning versus clinicians | Broadly equivalent accuracy, but few externally validated or prospective studies |
| Obermeyer et al. (2019) | Population-health risk stratification | Analysis of a widely used commercial risk-prediction algorithm | Systematic racial bias understated the illness burden of Black patients |
| McKinney et al. (2020) | Breast cancer screening (mammography) | AI system evaluated on United Kingdom and United States screening datasets | Reduced both false positives and false negatives relative to human readers |
| Nagendran et al. (2020) | Study design and reporting (systematic review) | Review of design, reporting standards, and claims of deep learning studies | Few randomised trials, incomplete reporting, and frequent overstatement of results |
| Ghassemi et al. (2021) | Explainability in clinical AI | Critical review of explainable AI methods for health care | Current explainability methods are unreliable for individual clinical decisions |
Artificial intelligence in diagnostic imaging
Medical imaging has been the most productive proving ground for diagnostic AI, partly because images provide structured, high-dimensional input that suits deep neural networks and partly because large annotated datasets are comparatively accessible. The literature in this theme is dominated by radiology, pathology, and dermatology, and it is these subspecialties that supply most of the studies summarised in Table 1.
Radiology and ophthalmology
Two influential validation studies established the credibility of the field. Gulshan et al. (2016) trained a convolutional neural network on approximately 128,000 retinal fundus photographs and reported sensitivity and specificity for referable diabetic retinopathy that were comparable to those of experienced ophthalmologists. In chest imaging, Rajpurkar et al. (2018) compared the CheXNeXt model with practising radiologists across fourteen pathologies and found that the algorithm matched clinician performance on most conditions while producing reports substantially faster. Screening mammography has attracted particular attention because of its high volume and its inter-reader variability; McKinney et al. (2020) evaluated an AI system on United Kingdom and United States screening datasets and reported reductions in both false positives and false negatives relative to human readers. These results suggest genuine potential to support high-throughput screening, although each study relied on retrospective data.
Pathology and dermatology
Deep learning has also been applied to the classification of tissue and skin. Esteva et al. (2017) trained a network on roughly 130,000 clinical images and demonstrated dermatologist-level accuracy in distinguishing malignant from benign lesions, a finding that stimulated interest in teledermatology and consumer-facing triage tools. In pathology, digitised whole-slide images have enabled models that detect nodal metastases and grade tumours, although the pooled evidence indicates that performance is highly dependent on staining protocols and scanner characteristics (Liu et al., 2019). Across imaging subspecialties, a recurring caution is that models trained at one institution often degrade when applied to images acquired elsewhere, a problem of generalisability that is central to the sections that follow.
Clinical decision support
Beyond image interpretation, AI has been positioned as a form of clinical decision support that integrates diverse data to assist diagnostic reasoning. Topol (2019) describes a convergence of human and machine intelligence in which algorithms surface patterns across pathology results, physiology, and the electronic record that would be difficult for a clinician to detect unaided. Applications include the early prediction of sepsis and acute deterioration from vital-sign trends, risk stratification for chronic disease, and triage tools that prioritise urgent cases. Importantly, decision-support systems do not necessarily render a diagnosis; more often they flag risk or direct attention, leaving the diagnostic judgment with the clinician. This distinction matters for accountability and for regulatory classification. The value of such systems depends heavily on the quality and representativeness of the underlying data, and, as the analysis of bias below indicates, a poorly specified target variable can encode existing inequities rather than correct them (Obermeyer et al., 2019).
Diagnostic accuracy relative to clinicians
A central question in the literature is whether diagnostic AI genuinely matches or exceeds clinicians. The most rigorous synthesis remains the systematic review and meta-analysis by Liu et al. (2019), which examined eighty-two studies comparing deep learning with health professionals in medical imaging. The authors concluded that diagnostic performance was broadly equivalent, but they issued a strong methodological warning: very few studies used externally validated data or prospective designs, and most compared algorithms against clinicians under artificial conditions rather than in practice. Nagendran et al. (2020) reinforced this concern in a review of design and reporting standards, finding that randomised trials were rare, reporting was frequently incomplete, and claims of superiority were often overstated relative to the evidence. As summarised in Table 1, headline accuracy figures should therefore be interpreted cautiously; strong retrospective performance does not guarantee benefit once a model is embedded in a clinical workflow with real patients, incomplete data, and time pressure.
Implementation and workflow
Evidence of diagnostic accuracy is necessary but not sufficient for clinical value, because a model must also fit the workflows and information systems in which clinicians operate. The Australian Digital Health Agency (2023) has identified interoperable data and the maturation of national infrastructure, including My Health Record, as prerequisites for the safe deployment of data-driven tools. In practice, implementation raises risks that laboratory studies rarely capture. Automation bias, in which clinicians defer to an algorithmic output even when it is wrong, and alert fatigue, in which frequent low-value notifications are ignored, can both undermine the intended benefit. The literature increasingly recommends prospective silent trials, in which a model runs in the background without influencing care, so that its real-world behaviour can be observed before it is permitted to affect decisions. For Australia, where the AIHW (2024) reports enduring maldistribution of the specialist workforce, thoughtful integration could extend diagnostic capacity to regional services, but only if tools are embedded without adding to clinician workload.
Bias, equity, and explainability
Concerns about fairness and transparency have become prominent as diagnostic AI approaches deployment. Obermeyer et al. (2019) demonstrated that a widely used population-health algorithm systematically understated the illness burden of Black patients because it used historical health expenditure as a proxy for need, thereby reproducing the effects of unequal access. The lesson generalises: models learn the biases present in their training data, and datasets that under-represent particular groups will tend to perform worse for those groups. In the Australian context this has direct implications for Aboriginal and Torres Strait Islander peoples and for rural communities, whose data may be scarce in the predominantly metropolitan and international datasets on which many models are trained. Transparency offers only partial reassurance. Ghassemi et al. (2021) argue that current explainability methods provide a false hope, because post hoc explanations are often unstable and cannot be relied upon to justify an individual clinical decision. Such findings support a governance emphasis on rigorous validation and ongoing monitoring rather than on explanation alone.
Regulation and governance
Regulatory frameworks have evolved to treat many diagnostic algorithms as software based medical devices. In Australia, the TGA (2021) reformed its regulation of software based medical devices, clarifying classification, evidence requirements, and manufacturer obligations, and situating these tools within the Australian Register of Therapeutic Goods. Professional bodies have complemented statutory regulation with practice standards; RANZCR (2021) has published standards of practice and ethical principles for AI in radiology and radiation oncology that emphasise clinician oversight, transparency, and accountability. At a system level, the NHMRC (2023) has articulated ethical expectations for the use of AI in health care, including consent, equity, and continued human responsibility for clinical decisions. A distinctive regulatory challenge concerns adaptive algorithms that continue to learn after deployment, since a device that changes over time complicates the assumption of a fixed, approved product and heightens the importance of post-market surveillance.
Synthesis and gaps
Taken together, the literature describes a field with strong retrospective performance but comparatively weak evidence of clinical benefit. Diagnostic imaging models can match specialists under controlled conditions, yet few studies demonstrate improved patient outcomes in prospective, externally validated settings (Liu et al., 2019; Nagendran et al., 2020). Several gaps are apparent. First, there is a shortage of randomised and prospective trials that measure outcomes rather than accuracy alone. Second, generalisability across institutions, equipment, and populations is inconsistently reported, which is critical for a country with diverse and geographically dispersed populations. Third, equity has been under-examined, particularly for First Nations and rural Australians. Fourth, explainability remains immature, so trust must rest on validation and monitoring rather than on interpretation of a model’s reasoning. Finally, the governance of adaptive systems is unsettled. Notably, most pivotal studies were conducted outside Australia, so their transferability to local patient mixes, imaging equipment, and care pathways cannot be assumed and requires local evaluation.
Implications for Australian healthcare
The Australian evidence base for diagnostic AI is still emerging, and the international literature must be interpreted with local circumstances in mind. Workforce maldistribution and the growing burden of diagnostic imaging suggest that decision support could be valuable in regional and remote services, provided that models are validated on Australian data and integrated through trusted infrastructure such as My Health Record (Australian Digital Health Agency, 2023; AIHW, 2024). Equity considerations are pressing, because without representative datasets, deployment risks widening rather than narrowing the health gap for Aboriginal and Torres Strait Islander peoples. A layered governance approach, combining TGA regulation of software based medical devices, RANZCR practice standards, and NHMRC ethical guidance, provides a foundation, but it requires mechanisms for continuous post-market monitoring and for auditing performance across subpopulations. Realistic implementation should therefore proceed through carefully evaluated pilots with clinician oversight rather than through wholesale adoption.
Conclusion
The literature published between 2015 and 2025 establishes that AI can perform selected diagnostic tasks with an accuracy approaching that of specialist clinicians, most convincingly in medical imaging. The same body of work, however, exposes a persistent gap between retrospective performance and demonstrated clinical benefit, alongside unresolved problems of generalisability, equity, explainability, and the regulation of adaptive systems. For Australia, the opportunity lies in using diagnostic AI to support an overstretched and unevenly distributed workforce, but realising that opportunity depends on local validation, representative data, careful workflow integration, and a governance framework anchored in the TGA, RANZCR, and NHMRC. Diagnostic AI is best understood at present as a tool that augments clinical judgment rather than one that replaces it, and future research should prioritise prospective, outcome-focused, and Australian evaluations.
References
Australian Digital Health Agency. (2023). National Digital Health Strategy 2023-2028. Australian Digital Health Agency.
Australian Institute of Health and Welfare. (2024). Australia’s health 2024: In brief. AIHW.
Esteva, A., Kuprel, B., Novoa, R. A., Ko, J., Swetter, S. M., Blau, H. M., & Thrun, S. (2017). Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542(7639), 115-118.
Ghassemi, M., Oakden-Rayner, L., & Beam, A. L. (2021). The false hope of current approaches to explainable artificial intelligence in health care. The Lancet Digital Health, 3(11), e745-e750.
Gulshan, V., Peng, L., Coram, M., Stumpe, M. C., Wu, D., Narayanaswamy, A., & Webster, D. R. (2016). Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA, 316(22), 2402-2410.
Liu, X., Faes, L., Kale, A. U., Wagner, S. K., Fu, D. J., Bruynseels, A., & Denniston, A. K. (2019). A comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: A systematic review and meta-analysis. The Lancet Digital Health, 1(6), e271-e297.
McKinney, S. M., Sieniek, M., Godbole, V., Godwin, J., Antropova, N., Ashrafian, H., & Shetty, S. (2020). International evaluation of an AI system for breast cancer screening. Nature, 577(7788), 89-94.
Nagendran, M., Chen, Y., Lovejoy, C. A., Gordon, A. C., Komorowski, M., Harvey, H., & Maruthappu, M. (2020). Artificial intelligence versus clinicians: Systematic review of design, reporting standards, and claims of deep learning studies. BMJ, 368, m689.
National Health and Medical Research Council. (2023). The ethical use of artificial intelligence in health care. NHMRC.
Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447-453.
Rajpurkar, P., Irvin, J., Ball, R. L., Zhu, K., Yang, B., Mehta, H., & Lungren, M. P. (2018). Deep learning for chest radiograph diagnosis: A retrospective comparison of the CheXNeXt algorithm to practising radiologists. PLOS Medicine, 15(11), e1002686.
Royal Australian and New Zealand College of Radiologists. (2021). Standards of practice for artificial intelligence in radiology and radiation oncology. RANZCR.
Therapeutic Goods Administration. (2021). Regulation of software based medical devices. Australian Government Department of Health.
Topol, E. J. (2019). High-performance medicine: The convergence of human and artificial intelligence. Nature Medicine, 25(1), 44-56.