TY - GEN
T1 - Spatial Attention-Guided Prompt Learning for Anomaly Segementation in Medical Images
AU - Lee, Haeyun
AU - Lee, Kyungsu
AU - Hwang, Jae Youn
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Accurate segmentation of anomalies in medical images is critical for early diagnosis and effective treatment planning. Although recent vision-language frameworks such as MediCLIP have shown promise in few-shot settings, they often suffer from limited domain specificity and reduced localization accuracy when applied to heterogeneous medical data. In this work, we propose a fully supervised anomaly segmentation framework that integrates learnable text prompts with a novel Organic Spatial Attention Adapter. Unlike the baseline adapter in MediCLIP, our design employs an Inter-Layer Gated Fusion mechanism to adaptively integrate semantic features from different transformer depths, combined with a lightweight spatial attention module to enhance the localization of fine-grained abnormalities. The adapter also includes a target channel projection for downstream compatibility and residual normalization to ensure stable training. Leveraging curated prompt sets of normal and abnormal anatomical concepts, the model aligns multimodal embeddings for robust anomaly detection. Experiments on the BMAD benchmark - covering BraTS2021, BTCV+LiTs, and RESC datasets - demonstrate that our method achieves state-of-the-art performance in both image-level AUROC (I-AUROC) and pixel-level AUROC (P-AUROC), with notable improvements in pixel-level localization accuracy. These results validate the effectiveness of our architecture in enhancing domain adaptation and anomaly segmentation across diverse medical imaging modalities.
AB - Accurate segmentation of anomalies in medical images is critical for early diagnosis and effective treatment planning. Although recent vision-language frameworks such as MediCLIP have shown promise in few-shot settings, they often suffer from limited domain specificity and reduced localization accuracy when applied to heterogeneous medical data. In this work, we propose a fully supervised anomaly segmentation framework that integrates learnable text prompts with a novel Organic Spatial Attention Adapter. Unlike the baseline adapter in MediCLIP, our design employs an Inter-Layer Gated Fusion mechanism to adaptively integrate semantic features from different transformer depths, combined with a lightweight spatial attention module to enhance the localization of fine-grained abnormalities. The adapter also includes a target channel projection for downstream compatibility and residual normalization to ensure stable training. Leveraging curated prompt sets of normal and abnormal anatomical concepts, the model aligns multimodal embeddings for robust anomaly detection. Experiments on the BMAD benchmark - covering BraTS2021, BTCV+LiTs, and RESC datasets - demonstrate that our method achieves state-of-the-art performance in both image-level AUROC (I-AUROC) and pixel-level AUROC (P-AUROC), with notable improvements in pixel-level localization accuracy. These results validate the effectiveness of our architecture in enhancing domain adaptation and anomaly segmentation across diverse medical imaging modalities.
KW - Fully Supervised Learning
KW - Medical Anomaly Segmentation
KW - Prompt Learning
KW - Spatial Attention Adapter
KW - Vision-Language Models
UR - https://www.scopus.com/pages/publications/105033219238
U2 - 10.1109/AIxMHC65380.2025.00025
DO - 10.1109/AIxMHC65380.2025.00025
M3 - Conference paper
AN - SCOPUS:105033219238
T3 - Proceedings - 2025 2nd International Conference on Artificial Intelligence for Medicine, Health and Care, AIxMHC 2025
SP - 94
EP - 99
BT - Proceedings - 2025 2nd International Conference on Artificial Intelligence for Medicine, Health and Care, AIxMHC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2nd International Conference on Artificial Intelligence for Medicine, Health and Care, AIxMHC 2025
Y2 - 13 October 2025 through 15 October 2025
ER -