TY - GEN
T1 - Dual Contrastive Pre-training with Heatmap-Based Segmentation-Guided Attention for Balanced Multi-Class Skin Lesion Classification
AU - Pak, Jiwon
AU - Ko, Jaeeun
AU - Lee, Junghoon
AU - Kim, Sungmin
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Accurate classification of skin lesions is essential for early diagnosis and treatment planning. However, severe class imbalance in dermatological datasets hinders the effective training of multi-class classification models. To address this challenge, we propose an end-to-end framework combining dual contrastive learning with segmentation-guided attention. Our model uses a ResNet18-based U-Net encoder, pretrained with Self-Supervised Contrastive Learning (SSCL) and Supervised Contrastive Learning (SCL). The U-Net decoder generates a spatial attention map that leverages segmentation information to identify lesion boundaries. This segmentation-guided attention map is element-wise multiplied with the original image to create lesion-focused input for classification. This enhanced input is then processed by a classification head for final diagnosis. Evaluated on SLICE-3D and HAM10000 datasets, the proposed method achieved 72.19% accuracy, 72.96% weighted F1-score, and 88.39% macro AUC. Ablation studies confirm the effectiveness of both segmentation and attention, as well as the synergy of the dual contrastive strategy. The framework demonstrates robust and balanced performance, making it clinically applicable for the skin lesion classification.
AB - Accurate classification of skin lesions is essential for early diagnosis and treatment planning. However, severe class imbalance in dermatological datasets hinders the effective training of multi-class classification models. To address this challenge, we propose an end-to-end framework combining dual contrastive learning with segmentation-guided attention. Our model uses a ResNet18-based U-Net encoder, pretrained with Self-Supervised Contrastive Learning (SSCL) and Supervised Contrastive Learning (SCL). The U-Net decoder generates a spatial attention map that leverages segmentation information to identify lesion boundaries. This segmentation-guided attention map is element-wise multiplied with the original image to create lesion-focused input for classification. This enhanced input is then processed by a classification head for final diagnosis. Evaluated on SLICE-3D and HAM10000 datasets, the proposed method achieved 72.19% accuracy, 72.96% weighted F1-score, and 88.39% macro AUC. Ablation studies confirm the effectiveness of both segmentation and attention, as well as the synergy of the dual contrastive strategy. The framework demonstrates robust and balanced performance, making it clinically applicable for the skin lesion classification.
KW - contrastive learning
KW - heatmap
KW - medical image analysis
KW - segmentation-guided attention
KW - skin lesion classification
UR - https://www.scopus.com/pages/publications/105033040763
UR - https://www.scopus.com/pages/publications/105033040763#tab=citedBy
U2 - 10.1109/ICRCV67407.2025.11349230
DO - 10.1109/ICRCV67407.2025.11349230
M3 - Conference contribution
AN - SCOPUS:105033040763
T3 - 2025 7th International Conference on Robotics and Computer Vision, ICRCV 2025
SP - 132
EP - 136
BT - 2025 7th International Conference on Robotics and Computer Vision, ICRCV 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 7th International Conference on Robotics and Computer Vision, ICRCV 2025
Y2 - 24 October 2025 through 26 October 2025
ER -