A Comparative Evaluation of Class Activation Mapping-Based Explainable Artificial Intelligence Methods for Target Detection in Optical Remote Sensing Using YOLOv5
Journal of Applied Remote Sensing, cilt.1, sa.1, ss.1-48, 2026 (Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 1 Sayı: 1
- Basım Tarihi: 2026
- Dergi Adı: Journal of Applied Remote Sensing
- Derginin Tarandığı İndeksler: Scopus
- Sayfa Sayıları: ss.1-48
- Kocaeli Üniversitesi Adresli: Evet
Özet
Explainable artificial intelligence (XAI) has become a critical component in the deployment of deep learning–based object detection models, particularly in remote sensing applications where complex scenes, heterogeneous targets, and safety critical decisions demand transparency and interpretability. Among XAI techniques, Class Activation Mapping (CAM) based methods are widely adopted; however, their relative effectiveness in object detection settings, particularly under varying scene complexity and target density, remains insufficiently explored. In this study, a systematic qualitative and quantitative evaluation of CAM based XAI methods for target detection in optical remote sensing imagery is conducted, using the YOLOv5 architecture as the underlying detector. The evaluated methods are Grad-CAM, Grad-CAM++, XGrad-CAM, Eigen-CAM, Score-CAM, KPCA-CAM, Shapley-CAM, and Finer-CAM. Experiments are performed on the DIOR dataset, which provides a diverse set of object classes and scene configurations representative of real-world high-resolution remote sensing imagery. To quantitatively evaluate CAM-based explanations, a masking based evaluation framework is employed, consisting of two primary strategies (M1 and M2), and complemented by an additional entropy-based analysis (M3). The M1 and M2 strategies evaluate the relevance and sufficiency of CAM highlighted regions by selectively removing or preserving salient areas and analyzing the resulting changes in detection performance. In addition, an adapted M3 strategy based on Bernoulli entropy is introduced to analyze changes in the Bernoulli entropy derived from detector confidence scores under CAM-guided perturbations, providing a complementary assessment of the spatial concentration of CAM explanations around annotated targets. The experimental results reveal substantial variation in the behavior of the evaluated CAM methods, demonstrating that no single explanation approach is universally optimal across all evaluation criteria. Instead, CAM selection should consider application requirements and be supported by quantitative evaluation rather than relying solely on qualitative visualization. The proposed evaluation framework provides a systematic basis for comparing and selecting CAM-based explanation methods for remote sensing target detection.