| Title |
Risk Classification of Construction Accidents and SHAP-Based Derivation of High-Risk Scenarios |
| Authors |
한재영(Han, Jae-Young) ; 고태규(Ko, Tae-Gyu) ; 김기남(Kim, Ki-Nam) ; 이민재(Lee, Min-Jae) |
| DOI |
https://doi.org/10.5659/JAIK.2026.42.7.365 |
| Keywords |
Construction Safety; SHAP; Machine Learning; SHAP-IQ; Multi-class Classification |
| Abstract |
This study developed a multi-class classification model to categorize construction accident risks into three levels using the Hazard Profile
dataset from the Korea Authority of Land & Infrastructure Safety. The dataset included 98,224 accident records containing structured
categorical data such as construction type, accident location, and damage type. Among the evaluated models, CatBoost demonstrated the
strongest performance across key metrics, including a Balanced Accuracy of 0.781, a Macro-F1 score of 0.758, and a Micro-F1 score of
0.805. In particular, the model achieved the highest recall rate for high-risk cases at 0.833, effectively reducing false negatives for severe
incidents. A SHAP-based feature importance analysis identified work process, accident location in both subcategories and midcategories, and
human injury as the most influential factors in the model’s predictions. In addition, feature importance was examined by risk level, and a
SHAP-IQ interaction analysis was conducted specifically for high-risk accidents. The analysis visualized and evaluated 688 combinations that
positively contributed to high-risk predictions, identifying recurring patterns and presenting the five most significant scenarios based on
average interaction values. By examining the combined effects of occurrence frequency and feature interactions, the study proposes a
data-driven analytical approach for targeted safety management, including the selection of key risk factors and the establishment of priorities
at construction sites. |