ReFormer: The Relational Transformer for Image Captioning

Xuewen Yang; Yingru Liu; Xin Wang

doi:10.1145/3503161.3548409

ScienceGate Book Chapters

JOURNAL ARTICLE

ReFormer: The Relational Transformer for Image Captioning

Xuewen Yang Yingru Liu Xin Wang

Year: 2022 Journal: Proceedings of the 30th ACM International Conference on Multimedia Pages: 5398-5406

DOI: 10.1145/3503161.3548409

Get Full-Text PDF Get Analytical Report

Abstract

Image captioning is shown to be able to achieve a better performance by using scene graphs to represent the relations of objects in the image. The current captioning encoders generally use a Graph Convolutional Net (GCN) to represent the relation information and merge it with the object region features via concatenation or convolution to get the final input for sentence decoding. However, the GCN-based encoders in the existing methods are less effective for captioning due to two reasons. First, using the image captioning as the objective (i.e., Maximum Likelihood Estimation) rather than a relation-centric loss cannot fully explore the potential of the encoder. Second, using a pre-trained model instead of the encoder itself to extract the relationships is not flexible and cannot contribute to the explainability of the model. To improve the quality of image captioning, we propose a novel architecture ReFormer- a RElational transFORMER to generate features with relation information embedded and to explicitly express the pair-wise relationships between objects in the image. ReFormer incorporates the objective of scene graph generation with that of image captioning using one modified Transformer model. This design allows ReFormer to generate not only better image captions with the benefit of extracting strong relational image features, but also scene graphs to explicitly describe the pair-wise relationships. Experiments on publicly available datasets show that our model significantly outperforms state-of-the-art methods on image captioning and scene graph generation.

Keywords:

Closed captioning Computer science Transformer Artificial intelligence Scene graph Encoder Graph Decoding methods Sentence Image (mathematics) Computer vision Natural language processing Theoretical computer science Rendering (computer graphics) Algorithm Voltage

Metrics

Cited By

4.07

FWCI (Field Weighted Citation Impact)

Refs

0.95

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Multimodal Machine Learning Applications

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Advanced Image and Video Retrieval Techniques

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Domain Adaptation and Few-Shot Learning

Physical Sciences → Computer Science → Artificial Intelligence

ReFormer: The Relational Transformer for Image Captioning

Abstract

Metrics

Citation History

Topics

Related Documents

Relational-Convergent Transformer for image captioning

Relational Graph Reasoning Transformer for Image Captioning

Relational Attention with Textual Enhanced Transformer for Image Captioning

Image Captioning with Relational Knowledge

Rotary transformer for image captioning