Hierarchical Time-Aware Summarization with an Adaptive Transformer for Video Captioning

Leonardo Vilela Cardoso; Silvio Jamil F. Guimarães; Zenilton K. G. Patrocínio

doi:10.1142/s1793351x23640031

ScienceGate Book Chapters

JOURNAL ARTICLE

Hierarchical Time-Aware Summarization with an Adaptive Transformer for Video Captioning

Leonardo Vilela Cardoso Silvio Jamil F. Guimarães Zenilton K. G. Patrocínio

Year: 2023 Journal: International Journal of Semantic Computing Vol: 17 (04)Pages: 569-592 Publisher: World Scientific

DOI: 10.1142/s1793351x23640031

Get Full-Text PDF Get Analytical Report

Abstract

A coherent description is an ultimate goal regarding video captioning via a couple of sentences because it might also affect the consistency and intelligibility of the generated results. In this context, a paragraph describing a video is affected by the activities used to both produce its specific narrative and provide some clues that can also assist in decreasing textual repetition. This work proposes a model, named Hierarchical time-aware Summarization with an Adaptive Transformer (HSAT), that uses a strategy to enhance the frame selection reducing the amount of information that needed to be processed along with attention mechanisms to enhance a memory-augmented transformer. This new approach increases the coherence among the generated sentences, assessing data importance (about the video segments) contained in the self-attention results and uses that to improve readability using only a small fraction of time spent by the other methods. The test results show the potential of this new approach as it provides higher coherence among the various video segments, decreasing the repetition in the generated sentences and improving the description diversity in the ActivityNet Captions dataset.

Keywords:

Computer science Closed captioning Automatic summarization Transformer Paragraph Readability Sentence Speech recognition Natural language processing Artificial intelligence Voltage World Wide Web

Metrics

Cited By

0.73

FWCI (Field Weighted Citation Impact)

Refs

0.65

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Video Analysis and Summarization

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Multimodal Machine Learning Applications

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Natural Language Processing Techniques

Physical Sciences → Computer Science → Artificial Intelligence

Hierarchical Time-Aware Summarization with an Adaptive Transformer for Video Captioning

Abstract

Metrics

Citation History

Topics

Related Documents

Hierarchical Time-Aware Approach for Video Summarization

SkimCap: A Transformer-Based Video Captioning Method with Adaptive Attention and Hierarchical Skimming Features

Topic-aware video summarization using multimodal transformer

Memory-enhanced hierarchical transformer for video paragraph captioning

Hierarchical Boundary-Aware Neural Encoder for Video Captioning