Cross-Modal Adaptive Dual Association for Text-to-Image Person Retrieval

D. M. Lin; Yi-Xing Peng; Jingke Meng; Wei‐Shi Zheng

doi:10.1109/tmm.2024.3355644

ScienceGate Book Chapters

JOURNAL ARTICLE

Cross-Modal Adaptive Dual Association for Text-to-Image Person Retrieval

D. M. Lin Yi-Xing Peng Jingke Meng Wei‐Shi Zheng

Year: 2024 Journal: IEEE Transactions on Multimedia Vol: 26 Pages: 6609-6620 Publisher: Institute of Electrical and Electronics Engineers

DOI: 10.1109/tmm.2024.3355644

Get Full-Text PDF Get Analytical Report

Abstract

Text-to-image person re-identification (ReID) aims to retrieve images of a person based on a given textual description. The key challenge is to learn the relations between detailed information from visual and textual modalities. Existing work focuses on learning a latent space to narrow the modality gap and further build local correspondences between two modalities. However, these methods assume that image-to-text and text-to-image associations are modality-agnostic, resulting in suboptimal associations. In this work, we demonstrate the discrepancy between image-to-text association and text-to-image association and proposecross-modal adaptive dual association (CADA) to build fine bidirectional image-text detailed associations. Our approach features a decoder-based adaptive dual association module that enables full interaction between visual and textual modalities, enabling bidirectional and adaptive cross-modal correspondence associations. Specifically, this paper proposes a bidirectional association mechanism: Association of text Tokens to image Patches (ATP) and Association of image Regions to text Attributes (ARA). We adaptively model the ATP based on the fact that aggregating cross-modal features based on mistaken associations will lead to feature distortion. For modeling the ARA, since attributes are typically the first distinguishing cues of a person, we explore attribute-level associations by predicting the masked text phrase using the related image region. Finally, we learn the dual associations between texts and images, and the experimental results demonstrate the superiority of our dual formulation. The code used in this article will be made publicly available at https://github.com/LinDixuan/CADA .

Keywords:

Computer science Association (psychology) Dual (grammatical number) Modality (human–computer interaction) Modalities Artificial intelligence Identification (biology) Feature (linguistics) Image (mathematics) Distortion (music) Key (lock) Modal Information retrieval Pattern recognition (psychology) Natural language processing Linguistics

Metrics

Cited By

16.43

FWCI (Field Weighted Citation Impact)

Refs

0.99

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Video Surveillance and Tracking Methods

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Face recognition and analysis

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Advanced Image and Video Retrieval Techniques

Physical Sciences → Computer Science → Computer Vision and Pattern Recognition

Cross-Modal Adaptive Dual Association for Text-to-Image Person Retrieval

Abstract

Metrics

Citation History

Topics

Related Documents

Learning Hierarchical Cross-modal Association with Intra-modal Context for Text-Image Person Retrieval

Cross-modal Shared Concept Learning for Text-to-Image Person Retrieval

Cross-modal Collaborative Representation Learning for Text-to-Image Person Retrieval

Cross-modal Collaborative Representation Learning for Text-to-Image Person Retrieval

Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching