A Trustworthy Multimodal Deep Learning Framework for Robust Disease Diagnosis Under Real-World Domain Shifts
Published in Avrmitra Journal of Engineering, Technology and Applied Sciences, this peer-reviewed open access research paper addresses methodological advancements and experimental results in its subject area.
Abstract
Deep learning systems for disease diagnosis are increasingly built on several complementary data streams, such as medical images and structured clinical records, yet they are almost always validated on data that resemble the training set. In deployment, acquisition hardware, measurement protocols, patient case-mix and test availability all change, and a model that was accurate and well calibrated at the development site can become silently wrong and over-confident elsewhere. This paper presents TMDF, a Trustworthy Multimodal Deep learning Framework that targets robustness under such domain shifts while remaining honest about what it does not know. TMDF combines (i) modality-specific encoder ensembles trained under feature-level domain randomisation, (ii) a shrinkage-regularised, label-free test-time adaptation of input statistics, (iii) per-modality temperature calibration, (iv) an uncertainty-weighted product-of-experts fusion that naturally down-weights unreliable or absent modalities, and (v) a trust layer that supports selective prediction and unseen-disease flagging. Because large multi-site clinical corpora with controllable shifts are not available for systematic stress-testing, we evaluate on a transparent, fully reproducible synthetic multimodal benchmark in which the severity of imaging-acquisition shift, clinical-measurement shift, class-prevalence shift, missing modalities, corruptions and an unseen disease class can be dialled independently; five random seeds, five shift levels, five baselines and six ablations are reported. In-distribution accuracy is on par with the strongest baseline (94.0% versus 94.3%), but under the most severe shift TMDF retains 85.4% accuracy against 78.1% for the best-matched augmented early-fusion baseline (+7.2 percentage points), loses 8.7 points from the clean to the most shifted setting versus 16.8 for plain early fusion, tolerates a missing modality with accuracy of 84.6% (imaging absent) and 83.6% (clinical data absent), and raises unseen-disease detection AUROC from 0.652 to 0.717. The ablations also expose limitations: naive re-standardisation of all modalities under prevalence shift is harmful, and source-fitted temperature scaling does not reliably improve calibration under shift. We discuss these findings candidly and outline the evidence required before clinical use.
Author Affiliations & Contributions
Open Access & Reproducibility Statement
This article is published under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. Anyone is free to read, download, copy, distribute, print, search, or link to the full texts of these articles for any lawful purpose without financial or technical barriers. All experimental code, datasets, and benchmark results are preserved in public academic archives.