EventsThe 1st International Online Conference on Buildings
Published
This submission belongs to the session S5. Construction Management, and Computers & Digitization of the event The 1st International Online Conference on Buildings
Published date
24 Oct, 2023
Academic Editor
author-avatarJun WANG
Citation
Luís Jacques de Sousa, João Poças Martins, Luís Sanhudo, Tackling the Data Sourcing Problem in Construction Procurement with File Scraping Algorithms, in Proceedings of The 1st International Online Conference on Buildings, 24 October–26 October 2023, MDPI: Basel, Switzerland, doi: 10.3390/IOCBD2023-15190
Share
Email
Facebook
Twitter
LinkedIn

Tackling the Data Sourcing Problem in Construction Procurement with File Scraping Algorithms

image
1. CONSTRUCT/GEQUALTEC, FEUP DEC, Portugal
2. BUILT CoLAB – Collaborative Laboratory for the Future Built Environment
3. BUILT CoLAB – Collaborative Laboratory for the Future Built Environment, Portugal
Abstract

The Architecture, Engineering, and Construction (AEC) sector is observed to have a lower adoption rate of machine learning (ML) tools compared to other industries that share similar characteristics. A significant contributing factor to this lower adoption rate is the limited availability of data, as ML techniques rely on large datasets to train algorithms effectively. However, the construction process generates substantial data that provide a detailed characterisation of the project. This inclination towards generating abundant data in the Construction sector contradicts ML developers' prevailing challenge in sourcing sufficient data within the AEC industry.

In the specific case of Portuguese Construction Procurement, public construction projects are mandatorily submitted to online, open-source repositories. However, the consultation and extraction of procurement files is decentralised and not automated, making data agglomeration difficult and time-consuming.

In this sense, this paper presents a data-scraping algorithm to scrape construction procurement repositories to develop an ML-ready dataset of training data for ML and Natural Language Processing (NLP) algorithms focused on the Construction sector's procurement phase. This tool automatically scrapes procurement repositories, developing a procurement file dataset comprising bills of quantities (BoQ) and project specifications.

In future studies, the dataset will be processed into a standardised format suitable for NLP BOQ task-matching algorithms. These matching algorithms will aim to automate construction budgeting for tender proposal purposes.

Keywords
Construction
Public Procurement
Contract Awarding
Scraping Algorithm
Database
Artificial Intelligence
Machine Learning
Natural Language Processing
Manuscript
Oral Presentation
Building Information Modeling (BIM) Implementation in Public–Private Partnership (PPP) Projects
Research on Asymmetrical Reinforced Concrete Low-Rise Frames Under Multiple Seismic Events