Resumen de vídeo centrado en la identidad mediante fusión jerárquica de características biométricas, de apariencia y corporales 3D

|

Aceptado: 30-06-2026

|

Publicado: 07-07-2026

DOI: https://doi.org/10.4995/riai.2026.25683
Datos de financiación

Descargas

Palabras clave:

Resumen de vídeo, Seguimiento de múltiples objetos, Fusión de características, Percepción centrada en el ser humano, Visión por computador, Aprendizaje profundo

Agencias de apoyo:

LUCIA (Lucha contra el Cibercrimen mediante la aplicación de Inteligencia Artificial)

Resumen:

Este trabajo presenta un algoritmo de resumen de vídeo basado en el seguimiento de múltiples objetos (MOT) y
reidentificación de personas (ReID). La propuesta integra embeddings faciales (AdaFace), pose (4D-Humans/SMPL) y apariencia visual (TransReID). Estas representaciones guían una asignación jerárquica de identidad y el seguimiento mediante anclaje bidireccional, recuperando trayectorias ante oclusiones severas o baja calidad visual y generando finalmente un conjunto compacto de resúmenes por identidad. Para seleccionar fotogramas clave se utiliza una ponderación multifactorial que permite optimizar la claridad biométrica, interacción social y dinámica de movimiento; mientras que la Supresión Adaptativa de No Máximos (A-NMS) asegura la diversidad temporal. La evaluación en un conjunto de datos propio demuestra la estabilidad del seguimiento, logrando un IDF1 del 97.89 % y un MOTA del 95.79 %. Además, nuestro resumen mejora el obtenido mediante la selección Top-K, aumentando la diversidad visual en un 146 %, la cobertura temporal en un 89 % y la recuperabilidad de información en un 3.5 %.

Ver más Ver menos

Citas:

Alaa, T., Mongy, A., Bakr, A., Diab, M., Gomaa, W., 2024. Video Summarization Techniques: A Comprehensive Review. https://doi.org/10.48550/ARXIV.2410.04449

Alomar, K., Aysel, H.I., Cai, X., 2025. CNNs, RNNs and Transformers in human action recognition: a survey and a hybrid model. Artif Intell Rev 58, 387. https://doi.org/10.1007/s10462-025-11388-3

Apostolidis, E., Adamantidou, E., Metsai, A.I., Mezaris, V., Patras, I., 2021a. AC-SUM-GAN: Connecting Actor-Critic and Generative Adversarial Networks for Unsupervised Video Summarization. IEEE Trans. Circuits Syst. Video Technol. 31, 3278–3292. https://doi.org/10.1109/TCSVT.2020.3037883

Apostolidis, E., Adamantidou, E., Metsai, A.I., Mezaris, V., Patras, I., 2021b. Video Summarization Using Deep Neural Networks: A Survey. Proc. IEEE 109, 1838–1863. https://doi.org/10.1109/JPROC.2021.3117472

Apostolidis, E., Balaouras, G., Mezaris, V., Patras, I., 2021c. Combining Global and Local Attention with Positional Encoding for Video Summarization, in: 2021 IEEE International Symposium on Multimedia (ISM). Presented at the 2021 IEEE International Symposium on Multimedia (ISM), IEEE, Naple, Italy, pp. 226–234. https://doi.org/10.1109/ISM52913.2021.00045

Argaw, D.M., Yoon, S., Heilbron, F.C., Deilamsalehy, H., Bui, T., Wang, Z., Dernoncourt, F., Chung, J.S., 2024. Scaling Up Video Summarization Pretraining with Large Language Models, in: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Presented at the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Seattle, WA, USA, pp. 8332–8341. https://doi.org/10.1109/CVPR52733.2024.00796

Baskurt, K.B., Samet, R., 2019. Video synopsis: A survey. Computer Vision and Image Understanding 181, 26–38. https://doi.org/10.1016/j.cviu.2019.02.004

Bernardin, K., Stiefelhagen, R., 2008. Evaluating Multiple Object Tracking Performance: The CLEAR MOT Metrics. EURASIP Journal on Image and Video Processing 2008, 1–10. https://doi.org/10.1155/2008/246309

Bewley, A., Ge, Z., Ott, L., Ramos, F., Upcroft, B., 2016. Simple online and realtime tracking, in: 2016 IEEE International Conference on Image Processing (ICIP). Presented at the 2016 IEEE International Conference on Image Processing (ICIP), IEEE, Phoenix, AZ, USA, pp. 3464–3468. https://doi.org/10.1109/ICIP.2016.7533003

Bhute, M.M., Tare, S.S., S, S.R., 2025. Query-Driven Video Summarization for Long Video Footage Analysis Using Faster-RCNN and Determinantal Point Processes. Procedia Computer Science 258, 3989–3999. https://doi.org/10.1016/j.procs.2025.04.650

Biswas, R., Chaves, D., Fernández-Robles, L., Fidalgo, E., Alegre, E., 2021. A Video Summarization Approach to Speed-up the Analysis of Child Sexual Exploitation Material, in: XLII JORNADAS DE AUTOMÁTICA : LIBRO DE ACTAS. Servizo de Publicacións da UDC, pp. 648–654. https://doi.org/10.17979/spudc.9788497498043.648

Danesh Pazho, A., Alinezhad Noghre, G., Rahimi Ardabili, B., Neff, C., Tabkhi, H., 2023. CHAD: Charlotte Anomaly Dataset, in: Gade, R., Felsberg, M., Kämäräinen, J.-K. (Eds.), Image Analysis, Lecture Notes in Computer Science. Springer Nature Switzerland, Cham, pp. 50–66. https://doi.org/10.1007/978-3-031-31435-3_4

Dendorfer, P., Os̆ep, A., Milan, A., Schindler, K., Cremers, D., Reid, I., Roth, S., Leal-Taixé, L., 2021. MOTChallenge: A Benchmark for Single-Camera Multiple Target Tracking. Int J Comput Vis 129, 845–881. https://doi.org/10.1007/s11263-020-01393-0

Deng, J., Guo, J., Xue, N., Zafeiriou, S., 2019. ArcFace: Additive Angular Margin Loss for Deep Face Recognition, in: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Presented at the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Long Beach, CA, USA, pp. 4685–4694. https://doi.org/10.1109/CVPR.2019.00482

Ester, M., Kriegel, H.-P., Sander, J., Xu, X., 1996. A density-based algorithm for discovering clusters in large spatial databases with noise, in: Proceedings of the Second International Conference on Knowledge Discovery and Data Mining, KDD’96. AAAI Press, Portland, Oregon, pp. 226–231.

Goel, S., Pavlakos, G., Rajasegaran, J., Kanazawa, A., Malik, J., 2023. Humans in 4D: Reconstructing and Tracking Humans with Transformers, in: 2023 IEEE/CVF International Conference on Computer Vision (ICCV). Presented at the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE, Paris, France, pp. 14737–14748. https://doi.org/10.1109/ICCV51070.2023.01358

Guan, Z., Wang, Z., Zhang, G., Li, L., Zhang, M., Shi, Z., Jiang, N., 2025a. Multi-object tracking review: retrospective and emerging trend. Artif Intell Rev 58, 235. https://doi.org/10.1007/s10462-025-11212-y

Guan, Z., Wang, Z., Zhang, G., Li, L., Zhang, M., Shi, Z., Jiang, N., 2025b. Multi-object tracking review: retrospective and emerging trend. Artif Intell Rev 58, 235. https://doi.org/10.1007/s10462-025-11212-y

Gujar, P., 2024. How Marketers Can Improve Short-Form Video Ads In The Age Of GenAI. URL https://www.forbes.com/councils/forbestechcouncil/2024/09/12/how-marketers-can-improve-short-form-video-ads-in-the-age-of-genai/ (accessed 5.23.25).

Guo, J., Deng, J., Lattas, A., Zafeiriou, S., 2021. Sample and Computation Redistribution for Efficient Face Detection. https://doi.org/10.48550/ARXIV.2105.04714

Gygli, M., Grabner, H., Riemenschneider, H., Van Gool, L., 2014. Creating Summaries from User Videos, in: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (Eds.), Computer Vision – ECCV 2014, Lecture Notes in Computer Science. Springer International Publishing, Cham, pp. 505–520. https://doi.org/10.1007/978-3-319-10584-0_33

Hassan, S., Mujtaba, G., Rajput, A., Fatima, N., 2023. Multi-object tracking: a systematic literature review. Multimed Tools Appl 83, 43439–43492. https://doi.org/10.1007/s11042-023-17297-3

He, K., Zhang, X., Ren, S., Sun, J., 2016. Deep Residual Learning for Image Recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).

He, S., Luo, H., Wang, P., Wang, F., Li, H., Jiang, W., 2021. TransReID: Transformer-based Object Re-Identification, in: 2021 IEEE/CVF International Conference on Computer Vision (ICCV). Presented at the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE, Montreal, QC, Canada, pp. 14993–15002. https://doi.org/10.1109/ICCV48922.2021.01474

Hsu, T.-C., Liao, Y.-S., Huang, C.-R., 2023. Video Summarization With Spatiotemporal Vision Transformer. IEEE Trans. on Image Process. 32, 3013–3026. https://doi.org/10.1109/TIP.2023.3275069

Jiang, L., Lan, L., 2025. MVS-SupCon: multimodal video summarization with supervised contrastive learning, in: Déniz, L.G., Meng, H., Carli, R. (Eds.), Fifth International Conference on Image Processing and Intelligent Control (IPIC 2025). Presented at the Fifth International Conference on Image Processing and Intelligent Control, SPIE, Qingdao, China, p. 75. https://doi.org/10.1117/12.3074773

Kim, M., Jain, A.K., Liu, X., 2022. AdaFace: Quality Adaptive Margin for Face Recognition, in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Presented at the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, New Orleans, LA, USA, pp. 18729–18738. https://doi.org/10.1109/CVPR52688.2022.01819

Lee, M.J., Gong, D., Cho, M., 2025. Video Summarization with Large Language Models, in: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 18981–18991. https://doi.org/10.1109/CVPR52734.2025.01768

Li, H., Klabjan, D., Utke, J., 2024. Unsupervised Video Summarization via Iterative Training and Simplified GAN, in: Proceedings of the Asian Conference on Computer Vision (ACCV). pp. 1585–1601.

Li, H., Zhu, Y., Shang, Z., Wang, Z., Wu, X., 2025. A Comprehensive Survey on Video Summarization: Challenges and Advances. IEEE Trans. Circuits Syst. Video Technol. 1–1. https://doi.org/10.1109/TCSVT.2025.3596006

Li, Q., Chen, J., Xie, Q., Han, X., 2023. Video summarization for event-centric videos. Neural Networks 161, 359–370. https://doi.org/10.1016/j.neunet.2023.01.047

Li, S., Ren, H., Xie, X., Cao, Y., 2025. A Review of Multi‐Object Tracking in Recent Times. IET Computer Vision 19, e70010. https://doi.org/10.1049/cvi2.70010

Li, X., Wang, Z., Lu, X., 2018. Video Synopsis in Complex Situations. IEEE Trans Image Process 27, 3798–3812. https://doi.org/10.1109/TIP.2018.2823420

Li, Z., Yang, L., 2021. Weakly Supervised Deep Reinforcement Learning for Video Summarization With Semantically Meaningful Reward, in: 2021 IEEE Winter Conference on Applications of Computer Vision (WACV). Presented at the 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), IEEE, Waikoloa, HI, USA, pp. 3238–3246. https://doi.org/10.1109/WACV48630.2021.00328

Liu, T., Meng, Q., Huang, J.-J., Vlontzos, A., Rueckert, D., Kainz, B., 2022. Video Summarization Through Reinforcement Learning With a 3D Spatio-Temporal U-Net. IEEE Trans. on Image Process. 31, 1573–1586. https://doi.org/10.1109/TIP.2022.3143699

Luiten, J., Os̆ep, A., Dendorfer, P., Torr, P., Geiger, A., Leal-Taixé, L., Leibe, B., 2021. HOTA: A Higher Order Metric for Evaluating Multi-object Tracking. Int J Comput Vis 129, 548–578. https://doi.org/10.1007/s11263-020-01375-2

Maggiolino, G., Ahmad, A., Cao, J., Kitani, K., 2023. Deep OC-Sort: Multi-Pedestrian Tracking by Adaptive Re-Identification, in: 2023 IEEE International Conference on Image Processing (ICIP). Presented at the 2023 IEEE International Conference on Image Processing (ICIP), IEEE, Kuala Lumpur, Malaysia, pp. 3025–3029. https://doi.org/10.1109/ICIP49359.2023.10222576

Meena, P., Kumar, H., Kumar Yadav, S., 2023. A review on video summarization techniques. Engineering Applications of Artificial Intelligence 118, 105667. https://doi.org/10.1016/j.engappai.2022.105667

Mirjalili, M., Alegre Gutiérrez, E., Fidalgo Fernández, E., González Castro, V., Tanveer, W., 2025. Human-Centric Video Summarization via Identity-Aware Tracking. JA-CEA. https://doi.org/10.17979/ja-cea.2025.46.12249

Narasimhan, M., Rohrbach, A., Darrell, T., 2021. Clip-it! language-guided video summarization. Advances in neural information processing systems 34, 13988–14000.

Narwal, P., Duhan, N., Bhatia, K.K., 2025. Dynamic and personalized video summarization towards sports entertainment. Entertainment Computing 55, 100999. https://doi.org/10.1016/j.entcom.2025.100999

Narwal, P., Duhan, N., Kumar Bhatia, K., 2022. A comprehensive survey and mathematical insights towards video summarization. Journal of Visual Communication and Image Representation 89, 103670. https://doi.org/10.1016/j.jvcir.2022.103670

Otani, M., Nakashima, Y., Rahtu, E., Heikkilä, J., 2019. Rethinking the Evaluation of Video Summaries, in: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Presented at the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Long Beach, CA, USA, pp. 7588–7596. https://doi.org/10.1109/CVPR.2019.00778

Qaroush, A.M., Jubran, M., Olayyan, Q., 2025. A novel supervised framework for multi-video summarization: Addressing dataset biases and enhancing feature representation. Neurocomputing 653, 131180. https://doi.org/10.1016/j.neucom.2025.131180

Ramos, W., Silva, M., Araujo, E., Moura, V., Oliveira, K., Marcolino, L.S., Nascimento, E.R., 2023. Text-Driven Video Acceleration: A Weakly-Supervised Reinforcement Learning Method. IEEE Trans. Pattern Anal. Mach. Intell. 45, 2492–2504. https://doi.org/10.1109/TPAMI.2022.3157198

Ristani, E., Solera, F., Zou, R., Cucchiara, R., Tomasi, C., 2016. Performance Measures and a Data Set for Multi-target, Multi-camera Tracking, in: Hua, G., Jégou, H. (Eds.), Computer Vision – ECCV 2016 Workshops, Lecture Notes in Computer Science. Springer International Publishing, Cham, pp. 17–35. https://doi.org/10.1007/978-3-319-48881-3_2

Saini, P., Kumar, K., Kashid, S., Saini, A., Negi, A., 2023. Video summarization using deep learning techniques: a detailed analysis and investigation. Artif Intell Rev 56, 12347–12385. https://doi.org/10.1007/s10462-023-10444-0

Sharghi, A., Laurel, J.S., Gong, B., 2017. Query-Focused Video Summarization: Dataset, Evaluation, and a Memory Network Based Approach, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).

Shoitan, R., Moussa, M.M., Gharghory, S.M., Elnemr, H.A., Cho, Y.-I., Abdallah, M.S., 2023. User Preference-Based Video Synopsis Using Person Appearance and Motion Descriptions. Sensors 23, 1521. https://doi.org/10.3390/s23031521

Sultani, W., Chen, C., Shah, M., 2018. Real-World Anomaly Detection in Surveillance Videos, in: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 6479–6488. https://doi.org/10.1109/CVPR.2018.00678

Tang, Y., Bi, J., Xu, S., Song, L., Liang, S., Wang, T., Zhang, D., An, J., Lin, J., Zhu, R., Vosoughi, A., Huang, C., Zhang, Z., Liu, P., Feng, M., Zheng, F., Zhang, J., Luo, P., Luo, J., Xu, C., 2025. Video Understanding with Large Language Models: A Survey. IEEE Trans. Circuits Syst. Video Technol. 1–1. https://doi.org/10.1109/TCSVT.2025.3566695

Tank, C., 2023. The Data Bonanza: Exploring the Untapped Riches of Video Surveillance. Daten & Wissen. URL https://datenwissen.com/blog/the-data-bonanza-exploring-the-untapped-riches-of-video-surveillance/ (accessed 5.23.25).

Varghese, R., M., S., 2024. YOLOv8: A Novel Object Detection Algorithm with Enhanced Performance and Robustness, in: 2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS). Presented at the 2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS), IEEE, Chennai, India, pp. 1–6. https://doi.org/10.1109/ADICS58448.2024.10533619

Wistia, 2025. State of Video Report: Video Marketing Statistics for 2025. URL https://wistia.com/learn/marketing/video-marketing-statistics (accessed 10.15.25).

Wojke, N., Bewley, A., Paulus, D., 2017. Simple online and realtime tracking with a deep association metric, in: 2017 IEEE International Conference on Image Processing (ICIP). Presented at the 2017 IEEE International Conference on Image Processing (ICIP), IEEE, Beijing, pp. 3645–3649. https://doi.org/10.1109/ICIP.2017.8296962

Wu, G., Lin, J., Silva, C.T., 2022. Intentvizor: Towards generic query guided interactive video summarization, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10503–10512.

Xiao, S., Zhao, Z., Zhang, Z., Yan, X., Yang, M., 2020. Convolutional Hierarchical Attention Network for Query-Focused Video Summarization. AAAI 34, 12426–12433. https://doi.org/10.1609/aaai.v34i07.6929

Yale Song, Vallmitjana, J., Stent, A., Jaimes, A., 2015. TVSum: Summarizing web videos using titles, in: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Presented at the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Boston, MA, USA, pp. 5179–5187. https://doi.org/10.1109/CVPR.2015.7299154

Yuan, Y., Zhang, J., 2023. Unsupervised Video Summarization via Deep Reinforcement Learning With Shot-Level Semantics. IEEE Trans. Circuits Syst. Video Technol. 33, 445–456. https://doi.org/10.1109/TCSVT.2022.3197819

Zhang, Y., Sun, P., Jiang, Y., Yu, D., Weng, F., Yuan, Z., Luo, P., Liu, W., Wang, X., 2022. ByteTrack: Multi-object Tracking by Associating Every Detection Box, in: Avidan, S., Brostow, G., Cissé, M., Farinella, G.M., Hassner, T. (Eds.), Computer Vision – ECCV 2022, Lecture Notes in Computer Science. Springer Nature Switzerland, Cham, pp. 1–21. https://doi.org/10.1007/978-3-031-20047-2_1

Zhang, Y., Zhu, P., Zheng, T., Yu, P., Wang, J., 2024. Surveillance video synopsis framework base on tube set. Journal of Visual Communication and Image Representation 98, 104057. https://doi.org/10.1016/j.jvcir.2024.104057

Zhao, B., Gong, M., Li, X., 2023. AudioVisual Video Summarization. IEEE Trans. Neural Netw. Learning Syst. 34, 5181–5188. https://doi.org/10.1109/TNNLS.2021.3119969

Ver más Ver menos