Modelado del conductor en vehículos autónomos: La importancia de la fusión multimodal
Enviado: 30-01-2026
|Aceptado: 22-05-2026
|Publicado: 25-05-2026
Derechos de autor 2026 Raúl Fernández Matellán, David Martin Gomez, Arturo de la Escalera Hueso

Esta obra está bajo una licencia internacional Creative Commons Atribución-NoComercial-CompartirIgual 4.0.
Descargas
Palabras clave:
Control compartido, Redes neuronales, Interacción humano-vehículo, Cooperación y nivel de automatización, Automatización y diseño centrado en el ser humano, Modelado fisiológico, Fusión de información y sensores, Vehículos autónomos
Agencias de apoyo:
Resumen:
En los niveles SAE 2–3, la seguridad de la conducción automatizada depende de que el conductor pueda retomar el control ante una solicitud de traspaso de control, lo que exige estimar su disponibilidad. Aunque existen enfoques unimodales basados en visión o fisiología, suelen ser vulnerables a degradaciones y su comparación es difícil por la asimetría de dimensionalidad y por la falta de análisis de robustez e interpretabilidad. Evaluamos sobre el dataset TD2D, una arquitectura multimodal basada en autocodificadores que proyecta cada modalidad a una representación común y compara fusión latente y fusión tardía frente a baselines unimodales. La fusión a nivel latente con reducción moderada (PCA-100) ofrece el mejor equilibrio entre clases y, mediante inyección de fallos en el latente, se observa que la degradación del vídeo impacta más que la fisiología, sugiriendo dominancia visual. Estos resultados respaldan la fusión multimodal, pero señalan la necesidad de integraciones equilibradas y modelos explicables que permitan auditar la contribución de cada modalidad.
Citas:
Akiba, T., Sano, S., Yanase, T., Ohta, T., Koyama, M., 2019. Optuna: A nextgeneration hyperparameter optimization framework. In: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. DOI: https://doi.org/10.1145/3292500.333070
Bank, D., Koenigstein, N., Giryes, R., 2023. Autoencoders. Machine learning for data science handbook: data mining and knowledge discovery handbook, 353–374. DOI: https://doi.org/10.1007/978-3-031-24628-9 16
Barka, R. E., Politis, I., 2024. Driving into the future: A scoping review of smartwatch use for real-time driver monitoring. Transportation Research Interdisciplinary Perspectives 25, 101098. DOI: https://doi.org/10.1016/j.trip.2024.101098
Berahmand, K., Daneshfar, F., Salehi, E. S., Li, Y., Xu, Y., 2024. Autoencoders and their applications in machine learning: a survey. Artificial intelligence review 57 (2), 28. DOI: https://doi.org/10.1007/s10462-023-10662-6
Brown, A., Tulkens, J., Mattelin, M., Sanglet, T., Dhuyvetters, B., 2026. Remote photoplethysmography for health assessment: a review informed by intelliprove technology. Frontiers in Digital Health Volume 7 - 2025. DOI: https://doi.org/10.3389/fdgth.2025.1667423
Chu, S., Xia, M., Yuan, M., Liu, X., Sepp¨anen, T., Zhao, G., Shi, J., 2025. Codephys: Robust video-based remote physiological measurement through latent codebook querying. IEEE Journal of Biomedical and Health Informatics 29 (7), 4932–4945. DOI: https://doi.org/10.1109/JBHI.2025.3540134
Dargahi Nobari, K., Bertram, T., 2024. A multimodal driver monitoring benchmark dataset for driver modeling in assisted driving automation. Scientific data 11 (1), 327. DOI: https://doi.org/10.1038/s41597-024-03137-y
Dontoh, A., Ivey, S., Sirbaugh, L., Danyo, A., Aboah, A., 2025. Visual dominance and emerging multimodal approaches in distracted driving detection: A review of machine learning techniques. arXiv preprint arXiv:2505.01973. DOI: https://doi.org/10.48550/arXiv.2505.01973
Du, G., Li, T., Li, C., Liu, P. X., Li, D., 2021. Vision-based fatigue driving recognition method integrating heart rate and facial features. IEEE Transactions on Intelligent Transportation Systems 22 (5), 3089–3100. DOI: https://doi.org/10.1109/TITS.2020.2979527
Du, G., Zhang, L., Su, K., Wang, X., Teng, S., Liu, P. X., 2022. A multimodal fusion fatigue driving detection method based on heart rate and perclos. IEEE Transactions on Intelligent Transportation Systems 23 (11), 21810–21820. DOI: https://doi.org/10.1109/TITS.2022.3176973
Eckmann, J.-P., Kamphorst, S. O., Ruelle, D., 1995. Recurrence plots of dynamical systems. In: Turbulence, strange attractors and chaos. World Scientific, pp. 441–445. DOI: https://doi.org/10.1142/9789812833709 0030
European Parliament and Council of the European Union, 2016. Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 (General Data Protection Regulation). Official Journal of the European Union, L 119, 4 May 2016, pp. 1–88. CELEX: 32016R0679.
Fernández Matellán, R., Martín Gómez, D., De la Escalera Hueso, A., sept 2025. Fusión de imágenes y señales fisiológicas para modelar al conductor. Jornadas de Automática 46. DOI: https://doi.org/10.17979/ja-cea.2025.46.12121
Fujiwara, K., Iwamoto, H., Hori, K., Kano, M., 2024. Driver drowsiness detection using r-r interval of electrocardiogram and self-attention autoencoder. IEEE Transactions on Intelligent Vehicles 9 (1), 2956–2965. DOI: https://doi.org/10.1109/TIV.2023.3308575
Gjoreski, M., Gams, M. Z., Lustrek, M., Genc, P., Garbas, J.-U., Hassan, T., 2020. Machine learning and end-to-end deep learning for monitoring driver distractions from physiological and visual signals. IEEE Access 8, 70590–70603. DOI: https://doi.org/10.1109/ACCESS.2020.2986810
Han, D. W., Chung, H., Cao, Y., Zhou, F., Molnar, L., Robert, L. P., Tilbury, D. M., Yang, X. J., 2025. A systematic review of metrics measuring takeover performance in conditionally automated driving. International Journal of Human–Computer Interaction 0 (0), 1–30. DOI: https://doi.org/10.1080/10447318.2025.2552863
Hazmoune, S., Bougamouza, F., 2024. Using transformers for multimodal emotion recognition: Taxonomies and state of the art review. Engineering Applications of Artificial Intelligence 133, 108339. DOI: https://doi.org/10.1016/j.engappai.2024.108339
He, K., Chen, X., Xie, S., Li, Y., Doll´ar, P., Girshick, R., 2022. Masked autoencoders are scalable vision learners. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 16000–16009.
Hwang, J., Choi, W., Lee, J., Kim, W., Rhim, J., Kim, A., 2025. a dataset on takeover during distracted l2 automated driving. Scientific Data 12 (1), 539. DOI: https://doi.org/10.1038/s41597-025-04781-8
Kingma, D. P., Salimans, T., Welling, M., 2015. Variational dropout and the local reparameterization trick. Advances in neural information processing systems 28.
Lee, H., Lee, J., Shin, M., 2019. Using wearable ecg/ppg sensors for driver drowsiness detection based on distinguishable pattern of recurrence plots. Electronics 8 (2). DOI: 10.3390/electronics8020192
Lim, C., Liang, N., Villarreal, R. T., Prakah-Asante, K. O., Pitts, B. J., Yu, D., 2025. Multimodal prediction of situation awareness during automated driving: a gaze and eeg-based approach. Ergonomics 0 (0), 1–30, pMID: 41091109. DOI: https://doi.org/10.1080/00140139.2025.2571792
Maleki Varnosfaderani, S., Shaikh, M. R., Forouzanfar, M., 2025. A comprehensive review of unobtrusive biosensing in intelligent vehicles: Sensors, algorithms, and integration challenges. Bioengineering 12 (6). DOI: https://doi.org/10.3390/bioengineering12060669
Marwan, N., Carmen Romano, M., Thiel, M., Kurths, J., 2007. Recurrence plots for the analysis of complex systems. Physics Reports 438 (5), 237–329. DOI: https://doi.org/10.1016/j.physrep.2006.11.001
Meteier, Q., Capallera, M., de Salis, E., Angelini, L., Carrino, S., Widmer, M., Abou Khaled, O., Mugellini, E., Sonderegger, A., 2023. A dataset on the physiological state and behavior of drivers in conditionally automated driving. Data in Brief 47, 109027. DOI: https://doi.org/10.1016/j.dib.2023.109027
Muthuswamy, A., Dewan, M. A. A., Murshed, M., Parmar, D., 2023. Driver distraction classification using deep convolutional autoencoder and ensemble learning. IEEE Access 11, 71435–71448. DOI: https://doi.org/10.1109/ACCESS.2023.3293110
Ni, J., Zhao, Z., Shen, C., Tong, H., Song, D., Cheng, W., Luo, D., Chen, H., 2025. Harnessing vision models for time series analysis: A survey. arXiv preprint arXiv:2502.08869. DOI: https://doi.org/10.48550/arXiv.2502.08869
Ortega, J. D., Kose, N., Ca˜nas, P., Chao, M.-A., Unnervik, A., Nieto, M., Otaegui, O., Salgado, L., 2020. Dmd: A large-scale multi-modal driver monitoring dataset for attention and alertness analysis. In: European Conference on Computer Vision. Springer, pp. 387–405. DOI: https://doi.org/10.1007/978-3-030-66823-5 23
Peng, Y., Deng, H., Xiang, G., Wu, X., Yu, X., Li, Y., Yu, T., 2024. A multisource fusion approach for driver fatigue detection using hysiological signals and facial image. IEEE Transactions on Intelligent Transportation Systems 25 (11), 16614–16624. DOI: https://doi.org/10.1109/TITS.2024.3420409
Pihlgren, G. G., Sandin, F., Liwicki, M., 2020. Improving image autoencoder embeddings with perceptual loss. In: 2020 International Joint Conference on Neural Networks (IJCNN). pp. 1–7. DOI: https://doi.org/10.1109/IJCNN48605.2020.9207431
Roitberg, A., Peng, K., Marinov, Z., Seibold, C., Schneider, D., Stiefelhagen, R., 2022. A comparative analysis of decision-level fusion for multimodal driver behaviour understanding. In: 2022 IEEE Intelligent Vehicles Symposium (IV). pp. 1438–1444. DOI: https://doi.org/10.1109/IV51971.2022.9827426
SAE International, 2021. Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles. Tech. Rep. J3016 202104, SAE International. DOI: https://doi.org/10.4271/J3016 202104
Schiappa, M. C., Rawat, Y. S., Shah, M., 2023. Self-supervised learning for videos: A survey. ACM Computing Surveys 55 (13s), 1–37. DOI: https://doi.org/10.1145/357792
Sekadakis, M., Yannis, G., 2025. Systematic review and meta-analysis of takeover time from automated driving at sae levels 2 and 3 to manual control. Transportation Research Part F: Traffic Psychology and Behaviour 113, 263–306. DOI: https://doi.org/10.1016/j.trf.2025.04.003
Sharma, P. K., Chakraborty, P., 2024. A review of driver gaze estimation and application in gaze behavior understanding. Engineering Applications of Artificial Intelligence 133, 108117. DOI: https://doi.org/10.1016/j.engappai.2024.108117
Smyth, J., Chen, H., Donzella, V., Woodman, R., 2021. Public acceptance of driver state monitoring for automated vehicles: Applying the utaut framework. Transportation Research Part F: Traffic Psychology and Behaviour 83, 179–191. DOI: https://doi.org/10.1016/j.trf.2021.10.003
Stahlschmidt, S. R., Ulfenborg, B., Synnergren, J., 01 2022. Multimodal deep learning for biomedical data fusion: a review. Briefings in Bioinformatics 23 (2), bbab569. DOI: https://doi.org/10.1093/bib/bbab569
Tong, Z., Song, Y.,Wang, J.,Wang, L., 2022. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. In: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A. (Eds.), Advances in Neural Information Processing Systems. Vol. 35. Curran Associates, Inc., pp. 10078–10093.
Weaver, B. W., DeLucia, P. R., 2022. A systematic review and meta-analysis of takeover performance during conditionally automated driving. Human Factors 64 (7), 1227–1260, pMID: 33307821. DOI: https://doi.org/10.1177/0018720820976476
Xiang, G., Yao, S., Deng, H.,Wu, X.,Wang, X., Xu, Q., Yu, T.,Wang, K., Peng, Y., 2024. A multi-modal driver emotion dataset and study: Including facial expressions and synchronized physiological signals. Engineering Applications of Artificial Intelligence 130, 107772. DOI: https://doi.org/10.1016/j.engappai.2023.107772
Zhao, M., Beurier, G.,Wang, H.,Wang, X., 2021. Exploration of driver posture monitoring using pressure sensors with lower resolution. Sensors 21 (10). DOI: https://doi.org/10.3390/s21103346
Zhu, J., Ma, Y., Yang, H., Lv, C., Zhang, Y., Hao, S., 2025. Driver’s hand trajectory-guided takeover intention inference in conditionally automated driving. IEEE Transactions on Vehicular Technology 74 (10), 15331–15342. DOI: https://doi.org/10.1109/TVT.2025.3568449




