BOOK-CHAPTER

Image Captioning using CNN and Attention Based Transformer

Abstract

Image captioning is a technique for generating sentences that describe a scenario captured in photos. It can identify objects in a picture and carries out a few processes with the goal of locating the image’s most crucial parts. Algorithms now have the ability to generate text in the context of natural phrases that accurately describe an image. To extract image visual features, this work employs a pre-trained Convolution Neural Network (CNN) viz. EfficientNetB0, and then uses Transformer Encoder and Decoder to construct an appropriate caption. The model is trained using the Flickr8k dataset. The findings back up the model’s capacity to understand and produce text from pictures. The evaluation metric is the BLEU (bilingual evaluation understudy) score. The model obtains the image description, converts into text, and then into a voice. For visually impaired people who are unable to grasp visuals, image description is the ideal approach.

Keywords:
Closed captioning Computer science Transformer Artificial intelligence GRASP Convolutional neural network Image (mathematics) Encoder Computer vision Speech recognition Natural language processing Pattern recognition (psychology) Engineering

Metrics

4
Cited By
2.14
FWCI (Field Weighted Citation Impact)
32
Refs
0.88
Citation Normalized Percentile
Is in top 1%
Is in top 10%

Citation History

Topics

Multimodal Machine Learning Applications
Physical Sciences →  Computer Science →  Computer Vision and Pattern Recognition
Subtitles and Audiovisual Media
Social Sciences →  Arts and Humanities →  Language and Linguistics

Related Documents

JOURNAL ARTICLE

Image captioning using transformer-based double attention network

Hashem ParvinAhmad Reza Naghsh‐NilchiHossein Mahvash Mohammadi

Journal:   Engineering Applications of Artificial Intelligence Year: 2023 Vol: 125 Pages: 106545-106545
JOURNAL ARTICLE

Improving scene text image captioning using transformer-based multilevel attention

Swati SrivastavaHimanshu Sharma

Journal:   Journal of Electronic Imaging Year: 2023 Vol: 32 (03)
JOURNAL ARTICLE

Attention-based transformer model for Arabic image captioning

Israa Al BadarnehRana Husni Al MahmoudBassam HammoOmar S. Al-Kadi

Journal:   Neural Computing and Applications Year: 2025 Vol: 37 (20)Pages: 15501-15533
© 2026 ScienceGate Book Chapters — All rights reserved.