Video description method with fusion of instance-aware temporal features

Junbin Huang; He Yan; Lingkun Liu; Yuhan Liu

doi:10.1117/12.3000765

ScienceGate Book Chapters

JOURNAL ARTICLE

Video description method with fusion of instance-aware temporal features

Junbin Huang He Yan Lingkun Liu Yuhan Liu

Year: 2023

DOI: 10.1117/12.3000765

Get Full-Text PDF Get Analytical Report

Abstract

There are still challenges in the field of video understanding today, especially how to use natural language to describe the visual content in videos. Existing video encoder-decoder models struggle to extract deep semantic information and effectively understand the complex contextual semantics in a video sequence. Furthermore, different visual elements in the video contribute differently to the generation of video text descriptions. In this paper, we propose a video description method that fuses instance-aware temporal features. We extract local features of instances on the temporal sequence to enhance perception of temporal instances. We also employ spatial attention to perform weighted fusion of temporal features. Finally, we use bidirectional long short-term memory networks to encode the contextual semantic information of the video sequence, thereby helping to generate higher quality descriptive text. Experimental results on two public datasets demonstrate that our method achieves good performance on various evaluation metrics.

Keywords:

Computer science Encoder Semantics (computer science) ENCODE Artificial intelligence Field (mathematics) Sequence (biology) Perception Information retrieval

Metrics

Cited By

0.18

FWCI (Field Weighted Citation Impact)

Refs

0.41

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Multimodal Machine Learning Applications

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Video Analysis and Summarization

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Advanced Image and Video Retrieval Techniques

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Video description method with fusion of instance-aware temporal features

Abstract

Metrics

Citation History

Topics

Related Documents

Hybrid Instance-Aware Temporal Fusion for Online Video Instance Segmentation

Video instance segmentation based on temporal feature fusion

STFormer: Spatial-Temporal-Aware Transformer for Video Instance Segmentation

Temporal Based Instance-Level Fusion for Video Object Detection

IAST: Instance Association Relying on Spatio-Temporal Features for Video Instance Segmentation