TY - GEN
T1 - Patient-Centric Statistical Multi-Modal Fusion for Medical Diagnosis
T2 - 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025
AU - Choi, Seo Yeon
AU - Lee, Kyungsu
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Deep learning (DL) has led to substantial progress in medical image analysis, particularly for disease classification. However, the integration of patient-specific attributes, such as age, body mass index (BMI), and lifestyle factors with radiomics and raw imaging data remains a key challenge in the development of personalized diagnostic models. To alleviate this, in this research, we propose a novel multi-modal framework, denoted as Statistically Coherent Network (SCN), which jointly models imaging, radiomic, and patient metadata through a structured multi-space latent representation. SCN facilitates distributional coherence across subpopulations by leveraging a newly devised statistics-based loss in conjunction with a triplet loss, thereby aligning feature distributions among clinically similar cohorts. This statistical alignment using T-test facilitates more interpretable and robust representation learning across heterogeneous patient groups. We evaluate SCN on four clinically diverse tasks, including breast cancer (mammography), obstructive sleep apnea (CT), rotator cuff tear (MRI), and Cormack-Lehane grading (X-ray), and demonstrate the consistent improvements over conventional single-space and multi-modal baselines. The experimental results highlight the importance of explicitly incorporating patient metadata, in terms of multimodal learning, to enhance model generalizability and clinical relevance.
AB - Deep learning (DL) has led to substantial progress in medical image analysis, particularly for disease classification. However, the integration of patient-specific attributes, such as age, body mass index (BMI), and lifestyle factors with radiomics and raw imaging data remains a key challenge in the development of personalized diagnostic models. To alleviate this, in this research, we propose a novel multi-modal framework, denoted as Statistically Coherent Network (SCN), which jointly models imaging, radiomic, and patient metadata through a structured multi-space latent representation. SCN facilitates distributional coherence across subpopulations by leveraging a newly devised statistics-based loss in conjunction with a triplet loss, thereby aligning feature distributions among clinically similar cohorts. This statistical alignment using T-test facilitates more interpretable and robust representation learning across heterogeneous patient groups. We evaluate SCN on four clinically diverse tasks, including breast cancer (mammography), obstructive sleep apnea (CT), rotator cuff tear (MRI), and Cormack-Lehane grading (X-ray), and demonstrate the consistent improvements over conventional single-space and multi-modal baselines. The experimental results highlight the importance of explicitly incorporating patient metadata, in terms of multimodal learning, to enhance model generalizability and clinical relevance.
KW - deep learning
KW - dicom
KW - multimodal
KW - radiomics
KW - representation learning
UR - https://www.scopus.com/pages/publications/105035229921
U2 - 10.1109/ICCVW69036.2025.00238
DO - 10.1109/ICCVW69036.2025.00238
M3 - Conference paper
AN - SCOPUS:105035229921
T3 - Proceedings - 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025
SP - 2273
EP - 2284
BT - Proceedings - 2025 IEEE/CVF International Conference on Computer Vision Workshops, ICCV-W 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 19 October 2025 through 20 October 2025
ER -