Identificación facial automática en secuencias de vídeo empleando redes neuronales y aprendizaje profundo

|

Aceptado: 22-05-2026

|

Publicado: 25-05-2026

DOI: https://doi.org/10.4995/riai.2026.25454
Datos de financiación

Descargas

Palabras clave:

Procesamiento de imágenes, Redes neuronales, Aprendizaje máquina, Técnicas de inteligencia artificial, Visión por computador

Agencias de apoyo:

Esta investigación no contó con financiación

Resumen:

En una cadena de televisión se manejan grandes volúmenes de material audiovisual que deben ser organizados y catalogados para integrarse en el archivo histórico y poder reutilizarse en el futuro, aportando así valor añadido. Tradicionalmente, esta tarea recaía en el departamento de documentación, que la realizaba de forma manual. Disponer de un sistema automático capaz de detectar, reconocer y catalogar a las personas que aparecen en las secuencias de vídeo agilizaría el trabajo documental. Este trabajo presenta un sistema de inteligencia artificial basado en redes neuronales profundas capaz de detectar y reconocer personas específicas en las imágenes de informativos televisivos. Para su desarrollo, se elaboró un conjunto de datos compuesto de 18476 imágenes, centrado principalmente en figuras políticas de ámbito nacional. El sistema identifica automáticamente a los individuos en la escena mediante la red YOLOv8-seg y, posteriormente, lleva a cabo su reconocimiento utilizando el clasificador que ofrece mayor fiabilidad. Con este fin, se evaluaron siete arquitecturas de redes neuronales adaptadas a esta tarea, siendo DenseNet-169 la que mostró el mejor rendimiento promedio en las pruebas realizadas. Los resultados obtenidos confirman la viabilidad del sistema y abren el camino para futuras investigaciones.

Ver más Ver menos

Citas:

Asensi-González, R., 2024. Reconocimiento del rostro humano en imágenes de informativos televisivos mediante redes convolucionales profundas, Trabajo de Fin de Máster en Investigación en Ingeniería de Software y Sistemas Informáticos, Universidad Nacional de Educación a Distancia, Madrid.

Asensi-González, R., Herrera, P.J. 2025. Reconocimiento facial en informativos televisivos mediante redes convolucionales profundas. XLVI Jornadas de Automática 46. DOI: 10.17979/ja-cea.2025.46.12046

Boutrus, F., Damer, N., Fang, M., Kirchbuchner, F. Kuijper, A., 2021. MixFaceNets: extremely efficient face recognition networks. IEEE

International Joint Conference on Biometrics (IJCB), pp. 1-8. DOI: 10.1109/IJCB52358.2021.9484374

Chen, S., Liu, Y., Gao, X, Han, Z., 2018. MobileFaceNets: efficient CNNs for accurate real-time face verification on mobile devices. In: Zhou, J., et al. Biometric Recognition. CCBR 2018. Lecture Notes in Computer Science. Vol. 10996. Springer, Cham, pp. 428-438. DOI: 10.1007/978-3-319-97909-0_46

Chollet, F., 2017. Xception: deep learning with depthwise separable convolutions. arXiv. DOI: 10.48550/arXiv.1610.02357

Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., Fei-Fei, L., 2009. ImageNet: a large-scale hierarchical image database. IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, pp. 248-255. DOI: 10.1109/CVPR.2009.5206848

Deng, J., Guo, J., Yang, J., Xue, N., Kotsia, I., Zafeirius, S., 2022. ArFace: Additive angular margin loss for deep face recognition. arXiv. DOI: 10.48550/arXiv.1801.07698

Everingham, M., Van Gool, L., Williams, C.K.I., Winn, J., Zisserman, A, 2010. The pascal visual object classes (VOC) challenge. International Journal of Computer Vision 88, 303–338. DOI: 10.1007/s11263-009-0275-4

Everingham, M., Eslami, S.M.A., Van Gool, L., Williams, C.K.I., Winn, J., Zisserman, A, 2015. The pascal visual object classes challenge: a retrospective. International Journal of Computer Vision 111, 98-136. DOI: 10.1007/s11263-014-0733-5

Fahlman, S., Lebiere, C., 1990. The Cascade-Correlation Learning Architecture. In D. S. Touretzy (Ed), Advances in neural information processing systems (Vol. 2, pp. 524-532). Morgan Kaufman.

Fukushima, K., 1980. Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biological Cybernetics 36, 193–202. DOI: 10.1007/BF00344251

Girshick R., Donahue, J., Darrell, T., Malik, J., 2013. R-CNN rich feature hierarchies for accurate object detection and semantic segmentation. arXiv. DOI: 10.48550/arXiv.1311.2524

Girshick, R., 2014. Fast R-CNN. IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, pp. 1440-1448, DOI: 10.1109/ICCV.2015.169

Guo, Y., Zhang, L., Hu, Y., He, X., Gao, J., 2016. MS-Celeb-1M: A dataset and benchmark for large-scale face recognition. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (Eds) Computer Vision – ECCV 2016. Lecture Notes in Computer Science, vol. 9907, Springer, Cham, pp 87–102. DOI: 10.1007/978-3-319-46487-9_6

Hariharan, B., Arbeláez, P., Girshick, R., Malik, J., 2014. Simultaneous detection and segmentation. In D. Fleet, T. Pajdla, B. Schiele, T. Tuytelaars (Eds.), Computer Vision-ECCV 2014, pp. 297-312. Springer. DOI:10.1007/978-3-319-10584-0_20

He, K, Zhang, X, Ren, A., Sun, J., 2015. Deep residual learning for image recognition. arXiv. DOI: 10.48550/arXiv.1512.03385

He, K., Gkioxari, G., Dollár, P., Girshick, R., 2017. Mask R-CNN. IEEE International Conference on Computer Vision (ICCV), Venice, Italy, pp. 2980-2988. DOI: 10.1109/ICCV.2017.322

Howard, A., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T. Andreetto, A., 2017. MobileNets: efficient convolutional neural networks for mobile vision applications. arXiv. DOI: 10.48550/arXiv.1704.04861

Huang, G., Liu, Z., Van Der Maaten, L., Weinberger, K. Q., 2017. Densely connected convolutional networks. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, pp. 2261-2269. DOI: 10.1109/CVPR.2017.243

Huang, G. B., Ramesh, M., Berg, T., Learned-Miller. E., 2007. Labeled faces in the wild: a database for studying face recognition in unconstrained environments. University of Massachusetts, Amherst, Technical Report 07-49.

Huang, R., Pedoeem, J., Chen, C., 2018. YOLO-LItE: A real-time object detection algorithm optimized for non-GPU computers. arXiv. DOI: 10.48550/arXiv.1811.05588

Hubel, D. H., Wiesel, T. N. (1959). Receptive fields of single neurones in the cat's striate cortex. The Journal of Physiology, 148, 574-591. DOI: 10.1113/jphysiol.1959.sp006308

Hubel, D.H., Wiesel, T. N. (1968). Receptive Field and Functional Arquitecture of Monkey Striate Cortex. J. Physiol. (1968), 195, pp. 215-243.

Jocher, G., Qiu, J., Chaurasia, A., 2023. Ultralytics YOLO (Version 8.0.0). https://github.com/ultralytics/ultralytics (Accedido 30 abril 2025).

Kim, M., Jain, A. K., Liu, X., 2023. AdaFace: Quality adaptatibe margin for face recognition. arXiv. DOI: 10.48550/arXiv.2204.00964

Krizhevsky, A., Sutskever, I. Hinton, G.E., 2012. ImageNet classification with deep convolutional neural networks. Neural Information Processing Systems, 25. DOI: 10.1145/3065386

LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R.E., Hubbard, W., Jackel, L. D. 1989. Backpropagation Applied to Handwritten Zip Code Recognition. Neural Computation, 1(4), pp. 541-551. DOI: 10.1162/neco.1989.1.4.541

LeCun, Y., Boser, B., Denker, J. S., Howard, R. E., Habbard, W., Jackel, L. D., Henderson, D., 1990. Handwritten digit recognition with a back-propagation network. Advances in neural information processing systems 2.Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, pp. 396-404. DOI: 10.5555/109230.109279

LeCun, Y., Botton, L., Bengio, Y., Haffner, P. 1998. Gradient-based Learning Applied to document Recognition. Proceedings of the IEEE, 86(11), pp. 2278-2324. DOI: 10.1109/5.726791

Li, J., Wang, Y., Wan, C., Tai, Y., Qian, J., Yang, J., Wang, C., 2019. DSFD: dual shot face detector. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, pp. 5055-5064. DOI: 10.1109/CVPR.2019.00520

Lin, T. -Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Zitnick, C. L., Dollár, P., 2015. Microsoft COCO: common objects in context. arXiv. DOI: 10.48550/arXiv.1405.0312

Lu, C. Tang, X. 2014. Surpassing Human-level Face Verification Performance on LFW with GaussianFace. arXiv. DOI: 10.48550/arXiv.1404.3840

Nech, A., Kemelmacher-Shlizerman, I., 2017. Level playing field for million scale face recognition. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, pp. 3406-3415. DOI: 10.1109/CVPR.2017.363

Pajares, G., Herrera, P. J., Besada, E., 2021. Aprendizaje profundo. RC Libros Editorial, Madrid.

Redmon, J., Divvala, S., Girshick, R., Fahardi, A., 2016. You only look once: Unified, real-time object detection. arXiv. DOI: 10.48550/arXiv.1506.02640

Redmon, J., Farhadi, A., 2017. YOLO9000: Better, faster, stronger. arXiv. DOI: 10.48550/arXiv.1612.08242

Redmon, J., Farhadi, A., 2018. YOLOv3: An incremental Improvement. arXiv. DOI: 10.48550/arXiv.1804.02767

Ren S., He K., Girshick, R., Sun J., 2015. Faster R-CNN: towards real-time object detection with region proposal networks. In: Proceedings of the 29th International Conference on Neural Information Processing Systems (NIPS'15), Vol. 1. MIT Press, Cambridge, MA, USA, pp. 91–99. DOI: 10.5555/2969239.2969250

Schroff, F., Kalenichenko, D., Philbin, J., 2015. FaceNet: a unified embedding for face recognition and clustering. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, pp. 815-823. DOI: 10.1109/CVPR.2015.7298682

Sermanet, P., LeCun, Y., 2012. Convolutional neural networks applied to house numbers digit classification. In Proceedings of the 2012 International Conference on Pattern Recognition (ICPR), pp. 3288-3291. IEEE.

Simonyan, K. Zisserman, A., 2014. Very deep convolutional networks for large-scale image recognition. In: 3rd International Conference on Learning Representations (ICLR 2015), San Diego, pp. 1-14. DOI: 10.48550/arXiv.1409.1556

Szegedy, C., Liu, W., Jia, Y., Sermanet, P., 2014. Going deeper with convolutions. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, pp. 1-9. DOI: 10.1109/CVPR.2015.7298594

Tang, X., Du, D. K., He, Z, Liu, J., 2018. PyramidBox: a context-assisted single shot face detector. In: 15th European Conference on Computer Vision (ECCV 2018), Munich, Germany, Proceedings, Part IX. Springer-Verlag, Berlin, Heidelberg, pp. 812-828. DOI: 10.1007/978-3-030-01240-3_49

Tian, Y., Ye, Q., Doermann, D., 2025. YOLOv12: Attention-Centric Real-Time Object Detectors. arXiv. DOI: 10.48550/arXiv.2502.12524

Ultralytics. 2023. Ultralytics YOLOv8-Documentación. Disponible en: https://docs.ultralytics.com/models/yolov8/

Ultralytics. 2026. Ultralytics YOLO26-Documentación. Disponible en: https://docs.ultralytics.com/models/yolo26/

Wolf, L., Hassner, T., Maoz, I., 2011. Face recognition in unconstrained videos with matched background similarity. In: IEEE Conf. on Computer Vision and Pattern Recognition, Colorado Springs, CO, USA, pp. 529-534. DOI: 10.1109/CVPR.2011.5995566

Zeiler, M., Fergus, R., 2014. Visualizing and understanding convolutional networks. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (Eds), 13th European Conference on Computer Vision (ECCV 2014), Lecture Notes in Computer Science, vol 8689, Springer, Cham. DOI: 10.1007/978-3-319-10590-1_53

Ver más Ver menos