Leveraging deep learning segmentation techniques and connected component analysis to automate high-level cost estimates of facade retrofits using 2D images

María Escalada

Spain

Higher Technical School of Architecture of San Sebastián, University of the Basque Country image/svg+xml

|

Accepted: 2024-12-13

|

Published: 2024-12-13

DOI: https://doi.org/10.4995/vitruvio-ijats.2024.22421
Funding Data

Downloads

Keywords:

deep learning, semantic segmentation, connected component analysis, high-level cost estimates, facade rehabilitations

Supporting agencies:

University of the Basque Country (UPV/EHU)

Higher Technical School of Architecture of San Sebastián

Abstract:

Deep learning semantic segmentation techniques applied to 2D facade images hold a great promise in several domains that go far beyond model generation, mainly if the data used are front-parallel or orthonormal photographs. However, effective applications in the field of built heritage have not been adequately explored, largely due to the absence of multidisciplinary teams that include architecture professionals as early as the dataset creation stage. The aim of this research is to introduce a holistic view in order to demonstrate the practical usefulness of state-of-the-art segmentation models to automate high-level cost estimates of urbanscale residential building facade rehabilitations when combined with a connected component analysis. To achieve this, a scalable bottom-up approach is formulated in five simple phases, encompassing both data science and architecture expertise. This strategy seeks to improve the accuracy of analyses at early stages when limited information on constructions is available and there is a significant cost uncertainty, and therefore to optimise the strategies used by construction stakeholders involved in economic feasibility studies
and decision-making processes.

Show more Show less

References:

Berg, A., Grabler, F., & Malik, J. (2007). Parsing Images of Architectural Scenes. 2007 IEEE 11th International Conference on Computer Vision. 1-8. https://doi.org/10.1109/ICCV.2007.4409091

Dai, M., Ward, W., Meyers, G., Densley, D., & Mayfield, M. (2021). Residential building facade segmentation in the urban environment. Building and Environment, 199, 107921. https://doi.org/10.1016/j.buildenv.2021.107921

Čech, J., & Radim, S. (2009). Languages for Constrained Binary Segmentation Based on Maximum A Posteriori Probability Labeling. International Journal of Imaging Systems & Technology, 19, 69-79. https://doi.org/10.1002/ima.20181

Chen, L-C., Zhu, Y., Papandreou, G., Schroff, F., & Adam, H. (2018). Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. arXiv (Cornell University). https://doi.org/10.1007/978-3-030-01234-2_49

Detlefsen, N., Borovec, J., Schock, J., Harsh, A., Koker, T., Di Liello, L., Stancl, D., Quan, C., Grechkin, M. & Falcon, W. (2022). TorchMetrics - Measuring Reproducibility in PyTorch. The Journal of Open Source Software, 7, 70. https://doi.org/10.21105/joss.04101

Dugast, A., Parizet, I., & Fleury, M. (1990). Dictionnaire par noms d’architectes des constructions élevées à Paris aux XIXe et XXe siècles, période 1876-1899: notices 1 à 1340. Vol. I. Bulletin monumental. Institut d’Histoire de Paris.

Femiani, J., Para, W., Mitra, N., & Wonka, P. (2018). Facade Segmentation in the Wild. arXiv (Cornell University). https://doi.org/10.48550/arXiv.1805.08634

Fröhlich, B., Rodner, E., & Denzler, J. (2010). A Fast Approach for Pixelwise Labeling of Facade Images. 20th International Conference on Pattern Recognition, 3029-3032. https://doi.org/10.1109/ICPR.2010.742

Gadde, R., Marlet, R., & Paragios, N. (2016). Learning grammars for architecture-specific facade parsing. International Journal of Computer Vision, 117(3), 290–316. https://doi.org/10.1007/s11263-016-0887-4

Iakubovskii, P. (2019). Segmentation Models Pytorch. GitHub. Available at https://github.com/qubvel/segmentation_models.pytorch (accessed 13 April 2024).

Iglovikov, V., & Shvets, A. (2018). TernausNet: U-Net with VGG11 Encoder Pre-Trained on ImageNet for Image Segmentation. arXiv (Cornell University). https://doi.org/10.48550/arXiv.1801.05746

Jampani, V., Gadde, R., & Gehler, P. (2015). Facade segmentation. Max Planck Institute for Intelligent Systems. http://ps-old.is.tue.mpg.de/project/Facade_Segmentation (accessed 14 May 2024).

Kelly, T., Femiani, J., Wonka, P., & Mitra, N. (2017). BigSUR: Large-scale Structured Urban Reconstruction. Transactions on Graphics, 36(6), 204. https://doi.org/10.1145/3130800.3130823

Koziński, M., Gadde, R., Zagoruyko, S., Obozinski, G., & Marlet, R. (2015). A MRF shape prior for facade parsing with occlusions. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2820-2828. https://doi.org/10.1109/CVPR.2015.7298899

Laisney, F., & Koltirine, R. (1988). Règle et règlement. La question du règlement dans l’évolution de l’urbanisme parisien, 1600-1902. Research report [WWW document]. Ecole Nationale Supérieure d’Architecture de Paris-Belleville. Available at https://hal.science/hal-01903202 (accessed 20 May 2024).

Liu, H., Zhang, J., Zhu, J., & Hoi, S. (2017). DeepFacade: A Deep Learning Approach to Facade Parsing. Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence. https://doi.org/10.24963/ijcai.2017/320

Mairie de Paris. (2023). Protections patrimoniales: 5ème arrondissement. Plan Local D’Urbanisme De Paris [WWW document]. URL http://pluenligne.paris.fr/plu/sites-plu/site_statique_55/pages/page_1182.html (accessed 14 May 2024).

Martinovic, A. (n.d.). Angelo Martinovic, Ph.D. URL http://martinovi.ch/ (accessed 14 May 2024).

Martinovic, A., Mathias, M., Weissenberg, J., & Van Gool, L. (2012). A Three-Layered approach to facade parsing. Lecture Notes in Computer Science, 7578, 416–429. https://doi.org/10.1007/978-3-642-33786-4_31

Martinovic, A., & Van Gool, L. (2013). Bayesian grammar learning for inverse procedural modeling. IEEE Conference on Computer Vision and Pattern Recognition, 201-208. https://doi.org/10.1109/CVPR.2013.33

Mathias, M., Martinovic, A., & Van Gool, L. (2016). ATLAS: A three-layered approach to facade parsing. International Journal of Computer Vision 118(1), 22-48. https://doi.org//10.1007/s11263-015-0868-z

Müller, P., Zeng, G., Wonka, P., & Van Gool, L. (2007). Image-based procedural modeling of facades. ACM Transactions on Graphics, 26(3), 85. https://doi.org/10.1145/1275808.1276484

Musialski, P., Wonka, P., Aliaga, D., Wimmer, M., Van Gool, L., & Purgathofer, W. (2013) A survey of urban reconstruction. Computer Graphics Forum, 32(6), 146-177. https://doi.org/10.1111/cgf.12077

OpenCV (n.d.) OpenCV: Structural Analysis and shape Descriptors. Available at https://docs.opencv.org/3.4/d3/dc0/group__imgproc__shape.html (accessed 14 May 2024).

Pantoja, B., Swamy, V., & Sakota, M. (2020) Extracting Masonry Building Facades through Polygon Image Segmentation. EPFL. Available at https://www.epfl.ch/labs/mlo/wp-content/uploads/2021/05/crpmlcourse-paper782.pdf (accessed 14 May 2024).

Paris (n.d.) Les voies de Paris : dénominations et numéros d’immeubles. Available at https://www.paris.fr/pages/les-voies-de-parisdenominations-et-numeros-d-immeubles-7550 (accessed 20 May 2024)

Rahmani, K., Huang, H., & Mayer, H. (2017). Facade segmentation with a structured random forest. ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, Volume IV-1/W1, 175-181. https://doi.org/10.5194/isprs-annals-IV-1-W1-175-2017

Riemenschneider, H., Krispel, U., Thaller, W., Donoser, M., Havemann, S., Fellner, D., & Bischof, H. (2012). Irregular lattices for complex shape grammar facade parsing. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), 1640-1647. https://doi.org/10.1109/CVPR.2012.6247857

Ripperda, N., & Brenner, C. (2006). Reconstruction of Facade Structures Using a Formal Grammar and RjMCMC. Pattern Recognition. Vol. 4174, 750-759. https://doi.org/10.1007/11861898_75

Simon, L., Teboul, O., Koutsourakis, P., & Paragios, N. (2011). Random Exploration of the Procedural Space for Single-View 3D Modeling of Buildings. Int J Comput Vis., 93, 253–271. https://doi.org/10.1007/s11263-010-0370-6

Shapiro, S.C. (ed.) (1992). Encyclopedia of Artificial Intelligence, 2nd edn., Vol. II. John Wiley & Sons, New York.

Schmitz, M. & Mayer, H. (2016). A convolutional network for semantic facade segmentation and interpretation. The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences. XLI-B3, 709–715. https://doi.org/10.5194/isprsarchives-xli-b3-709-2016

TorchVision maintainers and contributors (2016). TorchVision: PyTorch’s Computer Vision library. GitHub repository. Available at https://github.com/pytorch/vision (accessed 14 May 2024).

Teboul, O. (2011). Shape grammar parsing : application to image-based modeling. PhD thesis [WWW document]. Ecole Centrale Paris. Available at https://theses.hal.science/tel-00628906 (accessed 14 May 2024).

Tylecek, R. (2013). The CMP Facade Database (Version 1.1). Research report [WWW document]. Available at https://cmp.felk.cvut.cz/~tylecr1/facade/CMP_facade_DB_2013.pdf (accessed 14 May 2024)

VarCity ETHZ (2017). VarCity - The Video - semantic and dynamic city modelling from images. Available at https://www.youtube.com/watch?v=6pjEs84DR6Q (accessed 14 May 2024).

Zhang, G., Pan, Y., & Zhang, L. (2022). Deep learning for detecting building façade elements from images considering prior knowledge. Automation in Construction, 133. https://doi.org/10.1016/j.autcon.2021.104016

Zhuo, X., Tian, J., & Fraundorfer, F. (2023). Cross field-based segmentation and learning-based vectorization for rectangular windows. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 16, 431-448. https://doi.org/10.1109/JSTARS.2022.3218767

Show more Show less