Identificación facial automática en secuencias de vídeo empleando redes neuronales y aprendizaje profundo
Enviado: 30-01-2026
|Aceptado: 22-05-2026
|Publicado: 25-05-2026
Derechos de autor 2026 Ricardo Asensi González, Pedro Javier Herrera Caro, Manuel Jesús Gómez Zotano

Esta obra está bajo una licencia internacional Creative Commons Atribución-NoComercial-CompartirIgual 4.0.
Descargas
Palabras clave:
Procesamiento de imágenes, Redes neuronales, Aprendizaje máquina, Técnicas de inteligencia artificial, Visión por computador
Agencias de apoyo:
Resumen:
En una cadena de televisión se manejan grandes volúmenes de material audiovisual que deben ser organizados y catalogados para integrarse en el archivo histórico y poder reutilizarse en el futuro, aportando así valor añadido. Tradicionalmente, esta tarea recaía en el departamento de documentación, que la realizaba de forma manual. Disponer de un sistema automático capaz de detectar, reconocer y catalogar a las personas que aparecen en las secuencias de vídeo agilizaría el trabajo documental. Este trabajo presenta un sistema de inteligencia artificial basado en redes neuronales profundas capaz de detectar y reconocer personas específicas en las imágenes de informativos televisivos. Para su desarrollo, se elaboró un conjunto de datos compuesto de 18476 imágenes, centrado principalmente en figuras políticas de ámbito nacional. El sistema identifica automáticamente a los individuos en la escena mediante la red YOLOv8-seg y, posteriormente, lleva a cabo su reconocimiento utilizando el clasificador que ofrece mayor fiabilidad. Con este fin, se evaluaron siete arquitecturas de redes neuronales adaptadas a esta tarea, siendo DenseNet-169 la que mostró el mejor rendimiento promedio en las pruebas realizadas. Los resultados obtenidos confirman la viabilidad del sistema y abren el camino para futuras investigaciones.
Citas:
Asensi-González, R., 2024. Reconocimiento del rostro humano en imágenes de informativos televisivos mediante redes convolucionales profundas, Trabajo de Fin de Máster en Investigación en Ingeniería de Software y Sistemas Informáticos, Universidad Nacional de Educación a Distancia, Madrid.
Asensi-González, R., Herrera, P.J. 2025. Reconocimiento facial en informativos televisivos mediante redes convolucionales profundas. XLVI Jornadas de Automática 46. DOI: 10.17979/ja-cea.2025.46.12046
Boutrus, F., Damer, N., Fang, M., Kirchbuchner, F. Kuijper, A., 2021. MixFaceNets: extremely efficient face recognition networks. IEEE
International Joint Conference on Biometrics (IJCB), pp. 1-8. DOI: 10.1109/IJCB52358.2021.9484374
Chen, S., Liu, Y., Gao, X, Han, Z., 2018. MobileFaceNets: efficient CNNs for accurate real-time face verification on mobile devices. In: Zhou, J., et al. Biometric Recognition. CCBR 2018. Lecture Notes in Computer Science. Vol. 10996. Springer, Cham, pp. 428-438. DOI: 10.1007/978-3-319-97909-0_46
Chollet, F., 2017. Xception: deep learning with depthwise separable convolutions. arXiv. DOI: 10.48550/arXiv.1610.02357
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., Fei-Fei, L., 2009. ImageNet: a large-scale hierarchical image database. IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, pp. 248-255. DOI: 10.1109/CVPR.2009.5206848
Deng, J., Guo, J., Yang, J., Xue, N., Kotsia, I., Zafeirius, S., 2022. ArFace: Additive angular margin loss for deep face recognition. arXiv. DOI: 10.48550/arXiv.1801.07698
Everingham, M., Van Gool, L., Williams, C.K.I., Winn, J., Zisserman, A, 2010. The pascal visual object classes (VOC) challenge. International Journal of Computer Vision 88, 303–338. DOI: 10.1007/s11263-009-0275-4
Everingham, M., Eslami, S.M.A., Van Gool, L., Williams, C.K.I., Winn, J., Zisserman, A, 2015. The pascal visual object classes challenge: a retrospective. International Journal of Computer Vision 111, 98-136. DOI: 10.1007/s11263-014-0733-5
Fahlman, S., Lebiere, C., 1990. The Cascade-Correlation Learning Architecture. In D. S. Touretzy (Ed), Advances in neural information processing systems (Vol. 2, pp. 524-532). Morgan Kaufman.
Fukushima, K., 1980. Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biological Cybernetics 36, 193–202. DOI: 10.1007/BF00344251
Girshick R., Donahue, J., Darrell, T., Malik, J., 2013. R-CNN rich feature hierarchies for accurate object detection and semantic segmentation. arXiv. DOI: 10.48550/arXiv.1311.2524
Girshick, R., 2014. Fast R-CNN. IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, pp. 1440-1448, DOI: 10.1109/ICCV.2015.169
Guo, Y., Zhang, L., Hu, Y., He, X., Gao, J., 2016. MS-Celeb-1M: A dataset and benchmark for large-scale face recognition. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (Eds) Computer Vision – ECCV 2016. Lecture Notes in Computer Science, vol. 9907, Springer, Cham, pp 87–102. DOI: 10.1007/978-3-319-46487-9_6
Hariharan, B., Arbeláez, P., Girshick, R., Malik, J., 2014. Simultaneous detection and segmentation. In D. Fleet, T. Pajdla, B. Schiele, T. Tuytelaars (Eds.), Computer Vision-ECCV 2014, pp. 297-312. Springer. DOI:10.1007/978-3-319-10584-0_20
He, K, Zhang, X, Ren, A., Sun, J., 2015. Deep residual learning for image recognition. arXiv. DOI: 10.48550/arXiv.1512.03385
He, K., Gkioxari, G., Dollár, P., Girshick, R., 2017. Mask R-CNN. IEEE International Conference on Computer Vision (ICCV), Venice, Italy, pp. 2980-2988. DOI: 10.1109/ICCV.2017.322
Howard, A., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T. Andreetto, A., 2017. MobileNets: efficient convolutional neural networks for mobile vision applications. arXiv. DOI: 10.48550/arXiv.1704.04861
Huang, G., Liu, Z., Van Der Maaten, L., Weinberger, K. Q., 2017. Densely connected convolutional networks. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, pp. 2261-2269. DOI: 10.1109/CVPR.2017.243
Huang, G. B., Ramesh, M., Berg, T., Learned-Miller. E., 2007. Labeled faces in the wild: a database for studying face recognition in unconstrained environments. University of Massachusetts, Amherst, Technical Report 07-49.
Huang, R., Pedoeem, J., Chen, C., 2018. YOLO-LItE: A real-time object detection algorithm optimized for non-GPU computers. arXiv. DOI: 10.48550/arXiv.1811.05588
Hubel, D. H., Wiesel, T. N. (1959). Receptive fields of single neurones in the cat's striate cortex. The Journal of Physiology, 148, 574-591. DOI: 10.1113/jphysiol.1959.sp006308
Hubel, D.H., Wiesel, T. N. (1968). Receptive Field and Functional Arquitecture of Monkey Striate Cortex. J. Physiol. (1968), 195, pp. 215-243.
Jocher, G., Qiu, J., Chaurasia, A., 2023. Ultralytics YOLO (Version 8.0.0). https://github.com/ultralytics/ultralytics (Accedido 30 abril 2025).
Kim, M., Jain, A. K., Liu, X., 2023. AdaFace: Quality adaptatibe margin for face recognition. arXiv. DOI: 10.48550/arXiv.2204.00964
Krizhevsky, A., Sutskever, I. Hinton, G.E., 2012. ImageNet classification with deep convolutional neural networks. Neural Information Processing Systems, 25. DOI: 10.1145/3065386
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R.E., Hubbard, W., Jackel, L. D. 1989. Backpropagation Applied to Handwritten Zip Code Recognition. Neural Computation, 1(4), pp. 541-551. DOI: 10.1162/neco.1989.1.4.541
LeCun, Y., Boser, B., Denker, J. S., Howard, R. E., Habbard, W., Jackel, L. D., Henderson, D., 1990. Handwritten digit recognition with a back-propagation network. Advances in neural information processing systems 2.Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, pp. 396-404. DOI: 10.5555/109230.109279
LeCun, Y., Botton, L., Bengio, Y., Haffner, P. 1998. Gradient-based Learning Applied to document Recognition. Proceedings of the IEEE, 86(11), pp. 2278-2324. DOI: 10.1109/5.726791
Li, J., Wang, Y., Wan, C., Tai, Y., Qian, J., Yang, J., Wang, C., 2019. DSFD: dual shot face detector. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, pp. 5055-5064. DOI: 10.1109/CVPR.2019.00520
Lin, T. -Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Zitnick, C. L., Dollár, P., 2015. Microsoft COCO: common objects in context. arXiv. DOI: 10.48550/arXiv.1405.0312
Lu, C. Tang, X. 2014. Surpassing Human-level Face Verification Performance on LFW with GaussianFace. arXiv. DOI: 10.48550/arXiv.1404.3840
Nech, A., Kemelmacher-Shlizerman, I., 2017. Level playing field for million scale face recognition. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, pp. 3406-3415. DOI: 10.1109/CVPR.2017.363
Pajares, G., Herrera, P. J., Besada, E., 2021. Aprendizaje profundo. RC Libros Editorial, Madrid.
Redmon, J., Divvala, S., Girshick, R., Fahardi, A., 2016. You only look once: Unified, real-time object detection. arXiv. DOI: 10.48550/arXiv.1506.02640
Redmon, J., Farhadi, A., 2017. YOLO9000: Better, faster, stronger. arXiv. DOI: 10.48550/arXiv.1612.08242
Redmon, J., Farhadi, A., 2018. YOLOv3: An incremental Improvement. arXiv. DOI: 10.48550/arXiv.1804.02767
Ren S., He K., Girshick, R., Sun J., 2015. Faster R-CNN: towards real-time object detection with region proposal networks. In: Proceedings of the 29th International Conference on Neural Information Processing Systems (NIPS'15), Vol. 1. MIT Press, Cambridge, MA, USA, pp. 91–99. DOI: 10.5555/2969239.2969250
Schroff, F., Kalenichenko, D., Philbin, J., 2015. FaceNet: a unified embedding for face recognition and clustering. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, pp. 815-823. DOI: 10.1109/CVPR.2015.7298682
Sermanet, P., LeCun, Y., 2012. Convolutional neural networks applied to house numbers digit classification. In Proceedings of the 2012 International Conference on Pattern Recognition (ICPR), pp. 3288-3291. IEEE.
Simonyan, K. Zisserman, A., 2014. Very deep convolutional networks for large-scale image recognition. In: 3rd International Conference on Learning Representations (ICLR 2015), San Diego, pp. 1-14. DOI: 10.48550/arXiv.1409.1556
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., 2014. Going deeper with convolutions. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, pp. 1-9. DOI: 10.1109/CVPR.2015.7298594
Tang, X., Du, D. K., He, Z, Liu, J., 2018. PyramidBox: a context-assisted single shot face detector. In: 15th European Conference on Computer Vision (ECCV 2018), Munich, Germany, Proceedings, Part IX. Springer-Verlag, Berlin, Heidelberg, pp. 812-828. DOI: 10.1007/978-3-030-01240-3_49
Tian, Y., Ye, Q., Doermann, D., 2025. YOLOv12: Attention-Centric Real-Time Object Detectors. arXiv. DOI: 10.48550/arXiv.2502.12524
Ultralytics. 2023. Ultralytics YOLOv8-Documentación. Disponible en: https://docs.ultralytics.com/models/yolov8/
Ultralytics. 2026. Ultralytics YOLO26-Documentación. Disponible en: https://docs.ultralytics.com/models/yolo26/
Wolf, L., Hassner, T., Maoz, I., 2011. Face recognition in unconstrained videos with matched background similarity. In: IEEE Conf. on Computer Vision and Pattern Recognition, Colorado Springs, CO, USA, pp. 529-534. DOI: 10.1109/CVPR.2011.5995566
Zeiler, M., Fergus, R., 2014. Visualizing and understanding convolutional networks. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (Eds), 13th European Conference on Computer Vision (ECCV 2014), Lecture Notes in Computer Science, vol 8689, Springer, Cham. DOI: 10.1007/978-3-319-10590-1_53




