JOURNAL ARTICLE

Hierarchical Time-Aware Summarization with an Adaptive Transformer for Video Captioning

Leonardo Vilela CardosoSilvio Jamil F. GuimarãesZenilton K. G. Patrocínio

Year: 2023 Journal:   International Journal of Semantic Computing Vol: 17 (04)Pages: 569-592   Publisher: World Scientific

Abstract

A coherent description is an ultimate goal regarding video captioning via a couple of sentences because it might also affect the consistency and intelligibility of the generated results. In this context, a paragraph describing a video is affected by the activities used to both produce its specific narrative and provide some clues that can also assist in decreasing textual repetition. This work proposes a model, named Hierarchical time-aware Summarization with an Adaptive Transformer (HSAT), that uses a strategy to enhance the frame selection reducing the amount of information that needed to be processed along with attention mechanisms to enhance a memory-augmented transformer. This new approach increases the coherence among the generated sentences, assessing data importance (about the video segments) contained in the self-attention results and uses that to improve readability using only a small fraction of time spent by the other methods. The test results show the potential of this new approach as it provides higher coherence among the various video segments, decreasing the repetition in the generated sentences and improving the description diversity in the ActivityNet Captions dataset.

Keywords:
Computer science Closed captioning Automatic summarization Transformer Paragraph Readability Sentence Speech recognition Natural language processing Artificial intelligence Voltage World Wide Web

Metrics

4
Cited By
0.73
FWCI (Field Weighted Citation Impact)
8
Refs
0.65
Citation Normalized Percentile
Is in top 1%
Is in top 10%

Citation History

Topics

Video Analysis and Summarization
Physical Sciences →  Computer Science →  Computer Vision and Pattern Recognition
Multimodal Machine Learning Applications
Physical Sciences →  Computer Science →  Computer Vision and Pattern Recognition
Natural Language Processing Techniques
Physical Sciences →  Computer Science →  Artificial Intelligence
© 2026 ScienceGate Book Chapters — All rights reserved.