| Title |
Performance Analysis of Text Embedding-Based Automatic Matching of Construction Work Item Codes in Bills of Quantities |
| Authors |
윤영채(Yun, Yeong-Chae) ; 윤석헌(Yun, Seok-Heon) |
| DOI |
https://doi.org/10.5659/JAIK.2026.42.7.391 |
| Keywords |
Construction Work Code; Bill of Quantities (BoQ); Text Embedding; Automatic Matching |
| Abstract |
Construction bills of quantities (BoQs) continuously incorporate new knowledge such as emerging construction methods and materials;
however, systematic updating and management of these items remain limited. As a result, BoQ items are often recorded as unstructured text,
creating difficulties in collecting and standardizing construction data. To address this issue, this study explores text embedding techniques that
capture the semantic meaning of unstructured BoQ item descriptions. An automatic construction work code matching framework is developed
to identify correspondences between BoQ item texts and standard construction work codes. Using the same matching algorithm environment,
different text embedding methods are compared to evaluate performance. Four embedding methods are assessed under identical experimental
conditions: TF-IDF, Word2Vec, FastText, and Sentence-BERT. The experimental dataset consists of item descriptions collected from real
construction BoQs, and performance is evaluated in terms of accuracy and runtime. The results show that TF-IDF achieves the highest
matching accuracy across all code levels, while Word2Vec demonstrates stable performance with faster execution time. In contrast, FastText
and Sentence-BERT show lower accuracy overall. The proposed framework and analysis provide insights for improving the standardization of
BoQ data and enhancing the utilization of construction work codes. |