Context-aware Scene Graph Generation with Seq2Seq Transformers

Yichao Lu; Himanshu Rai; Jason S. Chang; B. A. Knyazev; Guangwei Yu; Shashank Shekhar; Graham W. Taylor; Maksims Volkovs

doi:10.1109/iccv48922.2021.01563

ScienceGate Book Chapters

JOURNAL ARTICLE

Context-aware Scene Graph Generation with Seq2Seq Transformers

Yichao Lu Himanshu Rai Jason S. Chang B. A. Knyazev Guangwei Yu Shashank Shekhar Graham W. Taylor Maksims Volkovs

Year: 2021 Journal: 2021 IEEE/CVF International Conference on Computer Vision (ICCV) Pages: 15911-15921

DOI: 10.1109/iccv48922.2021.01563

Get Full-Text PDF Get Analytical Report

Abstract

Scene graph generation is an important task in computer vision aimed at improving the semantic understanding of the visual world. In this task, the model needs to detect objects and predict visual relationships between them. Most of the existing models predict relationships in parallel assuming their independence. While there are different ways to capture these dependencies, we explore a conditional approach motivated by the sequence-to-sequence (Seq2Seq) formalism. Different from the previous research, our proposed model predicts visual relationships one at a time in an autoregressive manner by explicitly conditioning on the already predicted relationships. Drawing from translation models in NLP, we propose an encoder-decoder model built using Transformers where the encoder captures global context and long range interactions. The decoder then makes sequential predictions by conditioning on the scene graph constructed so far. In addition, we introduce a novel reinforcement learning-based training strategy tailored to Seq2Seq scene graph generation. By using a self-critical policy gradient training approach with Monte Carlo search we directly optimize for the (mean) recall metrics and bridge the gap between training and evaluation. Experimental results on two public benchmark datasets demonstrate that our Seq2Seq learning approach achieves strong empirical performance, outperforming previous state-of-the-art, while remaining efficient in terms of training and inference time. Full code for this work is available here: https://github.com/layer6ai-labs/SGG-Seq2Seq.

Keywords:

Computer science Reinforcement learning Encoder Artificial intelligence Transformer Inference Machine learning Graph Theoretical computer science

Metrics

Cited By

3.86

FWCI (Field Weighted Citation Impact)

Refs

0.96

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Multimodal Machine Learning Applications

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Advanced Image and Video Retrieval Techniques

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Human Pose and Action Recognition

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Context-aware Scene Graph Generation with Seq2Seq Transformers

Abstract

Metrics

Citation History

Topics

Related Documents

Multi-modal Context-Aware Network for Scene Graph Generation

Scene Graph Generation With Hierarchical Context

Scene Graph Generation with Geometric Context

Uncertainty-Aware Scene Graph Generation

Distribution-aware network with context and entity attention for scene graph generation