Image Captioning using CNN and Attention Based Transformer

Deepa Mulimani; Prakashgoud Patil; Nagaraj Chaklabbi

doi:10.56155/978-81-955020-2-8-14

ScienceGate Book Chapters

BOOK-CHAPTER

Image Captioning using CNN and Attention Based Transformer

Deepa Mulimani Prakashgoud Patil Nagaraj Chaklabbi

Year: 2023 Soft Computing Research Society eBooks Pages: 157-166

DOI: 10.56155/978-81-955020-2-8-14

Get Full-Text PDF Get Analytical Report

Abstract

Image captioning is a technique for generating sentences that describe a scenario captured in photos. It can identify objects in a picture and carries out a few processes with the goal of locating the image’s most crucial parts. Algorithms now have the ability to generate text in the context of natural phrases that accurately describe an image. To extract image visual features, this work employs a pre-trained Convolution Neural Network (CNN) viz. EfficientNetB0, and then uses Transformer Encoder and Decoder to construct an appropriate caption. The model is trained using the Flickr8k dataset. The findings back up the model’s capacity to understand and produce text from pictures. The evaluation metric is the BLEU (bilingual evaluation understudy) score. The model obtains the image description, converts into text, and then into a voice. For visually impaired people who are unable to grasp visuals, image description is the ideal approach.

Keywords:

Closed captioning Computer science Transformer Artificial intelligence GRASP Convolutional neural network Image (mathematics) Encoder Computer vision Speech recognition Natural language processing Pattern recognition (psychology) Engineering

Metrics

Cited By

2.14

FWCI (Field Weighted Citation Impact)

Refs

0.88

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Multimodal Machine Learning Applications

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Subtitles and Audiovisual Media

Social Sciences → Arts and Humanities → Language and Linguistics

Image Captioning using CNN and Attention Based Transformer

Abstract

Metrics

Citation History

Topics

Related Documents

Image captioning using transformer-based double attention network

Automated Image Captioning Using Transformer-Based Visual Attention Networks

Improving scene text image captioning using transformer-based multilevel attention

Spatial Cross-Attention for Transformer-Based Image Captioning

Attention-based transformer model for Arabic image captioning