2.1 Diagnostic Error as a Systems Problem
Diagnosis has long rested on clinical observation, pattern recognition, and accumulated heuristics, and in many respects it still does. What has changed, perhaps faster than many anticipated, is the degree to which these traditional skills are now supplemented by data-intensive approaches oriented toward precision and patient safety (Hirosawa & Shimizu, 2025). The NASEM report is usually credited with much of the recent momentum. In the decade since its publication, international research networks have expanded considerably, and their attention has shifted toward the systemic roots of diagnostic failure: cognitive bias, breakdowns in communication, and the fragmentation of care across settings and providers (Shimizu et al., 2025).
Earlier epidemiological work had already pointed in this direction. The outpatient error estimates reported by Singh et al. (2014) were simply too large to be explained away as isolated individual lapses, and later patient-safety syntheses came to treat diagnosis as a property of the system as much as a skill of the clinician (Hall et al., 2020). For laboratories, this reframing carries a particular edge. Errors can originate before, during, or after the analytical phase, and many of them remain invisible to the clinician who eventually acts on the result. It is partly for this reason, we suspect, that the laboratory has come to be seen as a natural home for predictive quality control and, by extension, for AI.
2.2 From Rule-Based Systems to Large Language Models
AI now sits close to the centre of this conversation. Its intellectual lineage is usually traced to Alan Turing’s 1950 inquiry into whether machines could think and to the 1956 Dartmouth Conference, where the field acquired its name (Niu et al., 2025). Early medical AI systems were largely rule-based, built on explicit if–then logic that was transparent but brittle. Contemporary systems look quite different. They draw on ML, deep convolutional neural networks (CNNs), vision transformers, and LLMs such as GPT-4 and PaLM (Al-Antari, 2023; Niu et al., 2025). Beneath the newer architectures, however, much of what is actually deployed in pathology and laboratory medicine still depends on supervised learning, and its reliability hinges on data-preprocessing decisions that are seldom discussed but are often decisive (Albahra et al., 2023).
Generative models have widened the range of possible applications and, it seems, the range of possible failures too. Egli (2023) asked whether LLMs might represent the next revolution for clinical microbiology; the answer so far looks like a qualified “perhaps.” GPT-4-based agents, for example, have been tested for detecting AMR mechanisms, yet their outputs still appear to require careful expert verification before they could be trusted in routine use (Giske et al., 2024). Broader reviews describe a similar ambivalence: real potential for diagnostics and workflow support, shadowed by unresolved ethical questions and uneven public trust (Bekbolatova et al., 2024; Sallam et al., 2025).
2.3 Laboratory Medicine, In Vitro Diagnostics, and Consolidation
The appeal of AI for diagnostic safety is fairly intuitive. Algorithms can process large volumes of structured laboratory and imaging data, automate complex pattern recognition, and, in principle, reduce (though hardly eliminate) some of the cognitive biases that shape human judgement (Hirosawa & Shimizu, 2025; Niu et al., 2025). Because IVD informs so large a share of clinical decisions (Wurcel et al., 2019), even modest gains in error detection might carry disproportionate downstream effects. Clinical decision support systems built on laboratory data are one obvious route, although Daly et al. (2026) caution that their promise depends heavily on how data quality, missingness, and interoperability are handled in practice. Work in cardiology, orthopaedics, and oncology suggests, too, that such tools can support diagnosis and prevention when they are embedded thoughtfully in care pathways rather than appended to them (Guzman-Garcia et al., 2025).
Structural change in the laboratory sector adds another layer. Vandenberg et al. (2020) describe how the consolidation of clinical microbiology services into large, centralised facilities has coincided with the introduction of transformative technologies. Consolidation may well bring efficiencies of scale. It also concentrates risk, since a miscalibrated algorithm in a high-volume hub reaches far more patients than an isolated error at a single bench. The performance figures reported for these applications are summarised in Table 1 and examined in Section 4.
2.4 Computational Pathology in Cutaneous Melanoma
Cutaneous melanoma is among the most aggressive skin malignancies, and the accuracy and timing of its histopathological evaluation bear directly on survival (Venturi et al., 2025). Conventional light microscopy, however, is complicated by marked biological heterogeneity and by diagnostic criteria that leave room for interpretation. Inter-observer discordance among experts is reported at roughly 10% to 25%, most pronounced in borderline melanocytic lesions, spitzoid neoplasms, and severely dysplastic naevi (Venturi et al., 2025). The spread of whole-slide imaging (WSI) converted glass slides into high-resolution digital assets, a shift that may seem unremarkable but is precisely what made AI-assisted evaluation practical (Serag et al., 2019; Venturi et al., 2025). The resulting computational pipeline, running from digitisation and image quality control through classification, feature extraction, immune mapping, molecular inference, and explainable morphometry to tumour board decision-making, is summarised in Figure 1.
Deep learning architectures, particularly CNNs and multiple-instance learning (MIL) models, have shown diagnostic performance comparable to, and in some studies better than, that of expert dermatopathologists in distinguishing melanoma from benign naevi. Pooled estimates place sensitivity at approximately 89% to 92% and specificity at 90% to 94%, while hybrid models report areas under the receiver operating characteristic curve (AUCs) of 0.96 to 0.98 (Venturi et al., 2025). These are striking figures. Still, they derive largely from curated, retrospective datasets, a caveat that recurs throughout this literature. Classification, moreover, is only part of the picture. Segmentation frameworks such as U-Net and Mask R-CNN are increasingly used to extract the parameters on which staging depends; automated Breslow thickness measurements agree closely with manual ones, and automated detection of mitoses and ulceration offers more objective inputs for American Joint Committee on Cancer (AJCC) staging (Venturi et al., 2025).
AI pipelines appear especially well suited to characterising the tumour microenvironment. Graph-based deep learning can map tumour-infiltrating lymphocytes (TILs) by density and proximity to tumour nests, and the resulting metrics seem to track transcriptomic immune signatures and might help anticipate response to immune checkpoint blockade (Venturi et al., 2025). A related, rather more provocative line of work, often called “molecular histopathology,” attempts to infer genetic alterations directly from haematoxylin and eosin (H&E) slides. The groundwork was laid outside dermatology, when Coudray et al. (2018) showed that deep learning could predict several common mutations from non-small cell lung cancer histopathology images. In melanoma, algorithms predicting BRAF V600 status reach accuracies of roughly 75% to 85%, apparently by detecting subtle morphological correlates such as pleomorphism and architectural disorder (Serag et al., 2019; Venturi et al., 2025). That level is not enough to replace next-generation sequencing (NGS). It might, however, serve as a sensible triage step before formal molecular testing.
Clinical hesitation toward opaque algorithms is understandable, and recent work has responded by favouring interpretable morphometric pipelines. Veronesi et al. (2025) offer a useful example: their nuclei-level ML framework extracted more than six million nuclei from WSIs, evaluated 44 geometric and spatial variables, and reached a diagnostic accuracy of 90.4%. What is arguably more important than the headline accuracy is that the model’s predictions map onto familiar histopathological criteria, such as nuclear pleomorphism, irregular spacing, and anisotropy. Explainable systems of this kind are beginning to find a place in Molecular Tumor Boards (MTBs), where multimodal models can support teams facing grey-zone lesions such as melanocytic tumours of uncertain malignant potential (MELTUMPs) and atypical spitzoid lesions (Venturi et al., 2025). Whether such tools will resolve this ambiguity, or merely describe it with greater precision, is still an open question.

Figure 1. AI-enabled computational pathology workflow for cutaneous melanoma, from slide digitisation to Molecular Tumor Board decision support. The diagram traces eight sequential stages: (1) digitisation of H&E slides into whole-slide images; (2) pre-analytical image quality control and stain normalisation (e.g., HistoQC, PathProfiler, DeepFocus); (3) CNN- and MIL-based diagnostic classification; (4) segmentation of staging features such as Breslow thickness, mitoses, and ulceration; (5) graph-based mapping of tumour-infiltrating lymphocytes; (6) inference of BRAF V600 status from morphology as a triage step before sequencing; (7) explainable nuclei-level morphometry; and (8) multimodal review at the Molecular Tumor Board, where the final decision remains human. Reported performance figures are shown in italics within each stage. The lower panel highlights three threats that condition every stage: batch effects, reliance on retrospective evidence, and automation bias. Synthesised from Browning et al. (2024), Coudray et al. (2018), Serag et al. (2019), Venturi et al. (2025), and Veronesi et al. (2025).
2.5 Neurocardiology and the Heart–Brain Axis
Neurocardiology concerns the bidirectional relationship between the central nervous and cardiovascular systems, and it is commonly organised around two axes. The heart–brain axis covers cardiac sources of cerebrovascular events; the brain–heart axis covers secondary cardiac injury following acute neurological insult (Basem et al., 2025). AI has become an increasingly important tool along both, mainly by accelerating diagnosis and sharpening prevention. The principal applications, with the data modalities and performance figures reported for each, are mapped in Figure 2 and tabulated in Table 4.
Atrial fibrillation (AF) is the leading cause of cardioembolic stroke, and its paroxysmal, often silent course means it is easily missed on a routine electrocardiogram (ECG). Deep neural networks offer a partial workaround. In a widely cited retrospective analysis, Attia et al. (2019) showed that an AI-enabled ECG algorithm identified patients with AF from a single sinus-rhythm ECG with an AUC of approximately 0.87, rising to about 0.90 when all ECGs in the preceding 31-day window were considered. Later work suggests that such models may outperform conventional scores such as CHA₂DS₂-VASc and forecast incident AF over one year with AUCs near 0.85 (Basem et al., 2025). The capability is also moving out of the hospital: FDA-cleared wearable and handheld devices have reported community-screening sensitivities as high as 98.5% and specificities of 91.4% (Basem et al., 2025). This matters for embolic stroke of undetermined source (ESUS), where Choi et al. (2024) found that a deep learning model applied to sinus-rhythm ECGs could identify patients with previously undiagnosed AF; across cryptogenic stroke cohorts, reported AUCs range from roughly 0.81 to 0.88 (Basem et al., 2025).
Along the brain–heart axis, acute events such as ischaemic stroke and aneurysmal subarachnoid haemorrhage can trigger catecholamine surges that lead to myocardial injury, arrhythmia, and Takotsubo syndrome (Basem et al., 2025). Telling Takotsubo syndrome apart from acute myocardial infarction (AMI) is difficult at the bedside. Laumer et al. (2022) trained a temporal CNN on echocardiograms that separated the two with an AUC of 0.79 and an accuracy of 74.8%, outperforming a panel of senior cardiologists (AUC 0.71; accuracy 64.4%). Cardiac magnetic resonance appears to offer a richer substrate still; Cau et al. (2023) reported that ML models combining atrial and ventricular strain with parametric mapping could diagnose Takotsubo cardiomyopathy with considerable accuracy. Not every application is image-based, either. Natural language processing (NLP) of unstructured electronic medical record (EMR) text has classified stroke subtypes with about 80% agreement with expert neurologists (Basem et al., 2025), which is respectable, though it also implies that roughly one case in five would still need human adjudication.

Figure 2. Principal AI applications across the heart–brain and brain–heart axes of neurocardiology. The upper band defines the two axes: cardiac sources of cerebrovascular events (heart → brain, blue) and secondary cardiac injury after acute neurological insult (brain → heart, red). Each row beneath traces an application from its input data (grey), through the AI model and clinical task (blue or red), to the performance reported in the literature (green); the amber box marks the underlying clinical problem of distinguishing Takotsubo syndrome from acute myocardial infarction. Applications include occult atrial fibrillation detection from sinus-rhythm ECGs, wearable community screening, post-ESUS risk stratification, valvular and heart-failure detection, NLP-based stroke subtyping, Takotsubo differentiation, MI rule-out, and mortality forecasting. The footer summarises the performance gradient from acute binary to prognostic tasks. Synthesised from Attia et al. (2019), Basem et al. (2025), Cau et al. (2023), Choi et al. (2024), and Laumer et al. (2022).
2.6 Pathogen Genomics and AI in Clinical Microbiology
In clinical microbiology, the most consequential change of the past decade has arguably been genomic rather than algorithmic, although the two are increasingly entangled. Gador-Whyte et al. (2026) describe how WGS has moved from reference laboratories into routine use, supporting susceptibility prediction for Mycobacterium tuberculosis, outbreak investigation for Clostridioides difficile, and national surveillance of carbapenemase-producing Klebsiella clones. AI and ML add a further layer of interpretation. Ali and Muhammad (2023) argue that ML-based AMR prediction is promising but limited by uneven training data, inconsistent phenotypic reference standards, and the difficulty of interpreting model outputs, challenges that seem to stand between proof of concept and practical implementation.
A systematic review by Baddal et al. (2024) gives a sense of the breadth now covered: deep learning to detect false positives in SARS-CoV-2 RT-PCR fluorescence curves, gradient-boosted models predicting minimum inhibitory concentrations for Salmonella, CNN-based identification of Pseudomonas aeruginosa strains, U-Net detection of Bacillus anthracis in tissue, and even the computational discovery of a structurally novel antibiotic candidate, halicin. These applications are summarised in Table 2. They are not equally mature. Some, such as WGS-based tuberculosis susceptibility testing, have already displaced conventional methods in certain public-health laboratories (Gador-Whyte et al., 2026), whereas others remain experimental. The infrastructure required, as noted earlier, is also unevenly distributed, and consolidation of laboratory services may either ease or worsen that disparity depending on how it is managed (Vandenberg et al., 2020).
2.7 Quality Control, Distribution Shift, and Algorithmic Bias
For all the enthusiasm, moving clinical AI into routine practice runs into substantial technical, legal, and operational barriers (Hirosawa & Shimizu, 2025; Zhang et al., 2026). Perhaps the most persistent is the tacit assumption that a single, uncalibrated model will perform uniformly across institutions. Multicentre evaluations repeatedly suggest otherwise. In a retrospective cohort of 969,292 admissions across nine facilities, Schnetler et al. (2025) found that baseline sepsis-prediction models produced excessive false alerts at sites not represented in training, a pattern they attributed to unmitigated site-specific distribution shifts and to location-dependent variation in care.
Digital pathology faces a comparable problem in the form of batch effects. Differences between scanners, variation in H&E staining intensity, and out-of-focus artefacts can lower CNN classification accuracy by 10 to 23 percentage points (Venturi et al., 2025). In response, automated quality-control tools such as PathProfiler, HistoQC, and DeepFocus are increasingly built into laboratory workflows to flag unusable regions and normalise staining before diagnostic inference; Browning et al. (2024), for instance, evaluated PathProfiler within a working diagnostic pathology service. The problem is not confined to images. Vocal biomarkers appear sensitive to ambient acoustic noise in clinical environments, which argues for standardised signal-quality preprocessing before such tools are relied upon (Augusto, 2025). The broader point, made forcefully by Zhang et al. (2026), is that without standardised evaluation methods these failure modes are difficult even to detect, let alone compare across studies.
2.8 Human Factors, Automation Bias, and Ethics
Introducing AI into clinical workflows also brings human-factor dynamics that are not always easy to predict. Survey data suggest that 32% of frontline laboratory staff worry about job displacement through automation, compared with 25% of laboratory managers (Niu et al., 2025). The gap is modest, but it hints that anxiety is distributed unevenly across professional hierarchies. A more immediate clinical risk is automation bias, the tendency to accept algorithmic output uncritically even when it conflicts with clinical judgement or the underlying data (Niu et al., 2025; Venturi et al., 2025). The danger seems greatest in exactly those grey-zone cases where AI is most tempting to consult.
Explainable components, such as feature-attribution heatmaps and calibrated confidence scores, are often proposed as a remedy, on the grounds that they support meaningful human-in-the-loop oversight without undermining clinical autonomy (Niu et al., 2025). Serag et al. (2019) frame the goal differently, as “intelligence augmentation” rather than automation, with repetitive and visually taxing work offloaded to the model while judgement stays with the clinician. Ethical concerns run alongside these practical ones. Privacy and consent obligations constrain how laboratory data can be shared and reused (Daly et al., 2026), and public attitudes toward AI in healthcare remain mixed (Bekbolatova et al., 2024). Whether explanations actually change clinician behaviour, rather than simply reassuring users, has so far received less empirical scrutiny than one might hope.
2.9 Regulatory Governance and Standardised Evaluation
Regulators have begun to respond with formal frameworks. The European Union’s AI Act, published in the Official Journal in July 2024 and in force from August 2024, adopts a risk-based approach under which most AI diagnostic tools fall into the high-risk tier, bringing obligations relating to data governance, bias mitigation, human oversight, transparency, and post-market surveillance (Gasser, 2023; Niu et al., 2025). In the United States, the Food and Drug Administration (FDA) oversees such systems through its Software as a Medical Device (SaMD) action plan, which emphasises Total Product Lifecycle (TPLC) oversight for algorithms that continue to learn after deployment (Augusto, 2025; Niu et al., 2025).
Alongside these regulatory instruments, structured tools have emerged to audit the accuracy and completeness of AI-generated health information before clinical use, most notably the METRICS checklist and the CLEAR tool (Niu et al., 2025). How these layers fit together, linking the barriers described above to laboratory-level mitigations, regulatory instruments, and a longer-term improvement agenda, is sketched in Figure 3. The structural barriers themselves are catalogued in Table 3.

Figure 3. Integrated framework linking structural barriers, laboratory-level mitigations, and governance instruments to a five-domain roadmap for validated AI practice. The left column lists recurring structural barriers identified in this review (see Table 3); the centre column pairs each with a mitigation that can be applied at laboratory level, such as federated learning, site-specific calibration, automated quality control, explainable outputs, interoperability standards, and real-world validation. The right column summarises the governance and evaluation instruments that frame these mitigations, including the EU AI Act, the FDA Software as a Medical Device Total Product Lifecycle approach, and the METRICS and CLEAR evaluation tools. The lower panel maps the five domains of the diagnostic-excellence agenda (research, education, practice improvement, patient engagement, and policy) to a 2035 horizon, and the dashed loop indicates that post-market surveillance feeds back into recalibration. Synthesised from Augusto (2025), Browning et al. (2024), Daly et al. (2026), Gasser (2023), Niu et al. (2025), Schnetler et al. (2025), Shimizu et al. (2025), and Zhang et al. (2026).
2.10 Gaps in the Existing Literature
Several gaps stand out from this body of work. Much of the performance evidence is retrospective and drawn from curated datasets, so it says relatively little about behaviour under real-world distribution shift (Schnetler et al., 2025; Venturi et al., 2025). Studies of explainability tend to describe what an explanation looks like rather than whether it improves decisions. Pathogen-genomics research is concentrated in well-resourced settings, leaving the question of decentralised implementation largely open (Ali & Muhammad, 2023; Gador-Whyte et al., 2026). And there is, as yet, little that brings accuracy data, barriers, and governance together within a single frame that a laboratory director could use. The present review was designed with these gaps in mind.