ViTOL: Vision Transformer for Weakly Supervised Object Localization

Saurav Gupta; Sourav Lakhotia; Abhay Rawat; Rahul Tallamraju

doi:10.1109/cvprw56347.2022.00455

ScienceGate Book Chapters

JOURNAL ARTICLE

ViTOL: Vision Transformer for Weakly Supervised Object Localization

Saurav Gupta Sourav Lakhotia Abhay Rawat Rahul Tallamraju

Year: 2022 Journal: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) Pages: 4100-4109

DOI: 10.1109/cvprw56347.2022.00455

Get Full-Text PDF Get Analytical Report

Abstract

Weakly supervised object localization (WSOL) aims at predicting object locations in an image using only image-level category labels. Common challenges that image classification models encounter when localizing objects are, (a) they tend to look at the most discriminative features in an image that confines the localization map to a very small region, (b) the localization maps are class agnostic, and the models highlight objects of multiple classes in the same image and, (c) the localization performance is affected by background noise. To alleviate the above challenges we introduce the following simple changes through our proposed method ViTOL. We leverage the vision-based transformer for self-attention and introduce a patch-based attention dropout layer (p-ADL) to increase the coverage of the localization map and a gradient attention rollout mechanism to generate class-dependent attention maps. We conduct extensive quantitative, qualitative and ablation experiments on the ImageNet-1K and CUB datasets. We achieve state-of-the-art MaxBoxAcc-V2 localization scores of 70.47% and 73.17% on the two datasets respectively. Code is available on https://github.com/Saurav-31/ViTOL.

Keywords:

Artificial intelligence Computer science Discriminative model Leverage (statistics) Pattern recognition (psychology) Computer vision Contextual image classification Image (mathematics)

Metrics

Cited By

1.86

FWCI (Field Weighted Citation Impact)

Refs

0.89

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Advanced Neural Network Applications

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Domain Adaptation and Few-Shot Learning

Physical Sciences → Computer Science → Artificial Intelligence

Multimodal Machine Learning Applications

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

ViTOL: Vision Transformer for Weakly Supervised Object Localization

Abstract

Metrics

Citation History

Topics

Related Documents

Token Masking Transformer for Weakly Supervised Object Localization

CLIP-Driven Transformer for Weakly Supervised Object Localization

Task-Aware Weakly Supervised Object Localization With Transformer

Multiscale Vision Transformer With Deep Clustering-Guided Refinement for Weakly Supervised Object Localization

Reperceive Global Vision of Transformer for Remote Sensing Images Weakly Supervised Object Localization