|
Title |
Quantitative Evaluation of Grad-CAM Interpretability for CNN-Based Concrete Structural Damage Classification Using Bounding Box Reference Regions
|
|
Authors |
김일순(Il Sun Kim) ; 양은익(Eun Ik Yang) |
|
DOI |
https://doi.org/10.11112/jksmi.2026.30.4.126 |
|
Keywords |
경계 상자(Bounding box); 합성곱신경망(CNN); 콘크리트 손상; Grad-CAM; 공간 정합성 Bounding box; CNN; Concrete damage; Grad-CAM; Spatial correspondence |
|
Abstract |
Damage inspection of concrete structures is important for ensuring structural safety and durability, and recent studies have actively investigated deep learning-based automated damage recognition. In this study, the classification performance, repeated-training-based stability, and Grad-CAM-based interpretability of ResNet-50, GoogLeNet, and MobileNetV2 were comprehensively analyzed for concrete structural damage images. An AI-Hub damage image dataset was used for three damage types: crack, efflorescence, and rebar exposure. The spatial correspondence between JSON annotation-based bounding box reference regions and Grad-CAM activation regions was quantitatively evaluated using IoU and CAM Precision. The results showed that ResNet-50 achieved the highest average Macro F1-score and Grad-CAM-based interpretability metrics among the three models, while MobileNetV2 demonstrated classification performance comparable to that of ResNet-50 despite its lightweight architecture. In addition, some cases showed that Grad-CAM activation regions did not spatially correspond to the bounding box-based reference regions despite high softmax prediction confidence. Therefore, to evaluate the applicability of CNN-based structural damage classification models, it is necessary to consider not only classification performance but also repeated-training-based stability, Grad-CAM-based interpretability, and model complexity.
|