Retrieval-augmented Image Captioning

Rita Ramos; Desmond Elliott; Bruno Martins

doi:10.18653/v1/2023.eacl-main.266

ScienceGate Book Chapters

JOURNAL ARTICLE

Retrieval-augmented Image Captioning

Rita Ramos Desmond Elliott Bruno Martins

Year: 2023 Pages: 3666-3681

DOI: 10.18653/v1/2023.eacl-main.266

Get Full-Text PDF Get Analytical Report

Abstract

Inspired by retrieval-augmented language generation and pretrained Vision and Language (V&L) encoders, we present a new approach to image captioning that generates sentences given the input image and a set of captions retrieved from a datastore, as opposed to the image alone. The encoder in our model jointly processes the image and retrieved captions using a pretrained V&L BERT, while the decoder attends to the multimodal encoder representations, benefiting from the extra textual evidence from the retrieved captions. Experimental results on the COCO dataset show that image captioning can be effectively formulated from this new perspective. Our model, named EXTRA, benefits from using captions retrieved from the training dataset, and it can also benefit from using an external dataset without the need for retraining. Ablation studies show that retrieving a sufficient number of captions (e.g., k=5) can improve captioning quality. Our work contributes towards using pretrained V&L encoders for generative tasks, instead of standard classification tasks.

Keywords:

Closed captioning Computer science Encoder Artificial intelligence Image (mathematics) Set (abstract data type) Language model Natural language processing Retraining Image retrieval Generative grammar Speech recognition Information retrieval

Metrics

Cited By

5.46

FWCI (Field Weighted Citation Impact)

Refs

0.95

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Multimodal Machine Learning Applications

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Domain Adaptation and Few-Shot Learning

Physical Sciences → Computer Science → Artificial Intelligence

Advanced Image and Video Retrieval Techniques

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Retrieval-augmented Image Captioning

Abstract

Metrics

Citation History

Topics

Related Documents

Retrieval-Augmented Transformer for Image Captioning

Understanding Retrieval Robustness for Retrieval-augmented Image Captioning

Towards Retrieval-Augmented Architectures for Image Captioning

Retrieval-augmented prompts for text-only image captioning

HRACap: Lightweight Image Captioning via Hierarchical Retrieval-Augmented Prompt