JOURNAL ARTICLE

Learning Pixel Affinity Pyramid for Arbitrary-Shaped Text Detection

Zilong FuHongtao XieShancheng FangYuxin WangMengting XingYongdong Zhang

Year: 2022 Journal:   ACM Transactions on Multimedia Computing Communications and Applications Vol: 19 (1s)Pages: 1-24   Publisher: Association for Computing Machinery

Abstract

Arbitrary-shaped text detection in natural images is a challenging task due to the complexity of the background and the diversity of text properties. The difficulty lies in two aspects: accurate separation of adjacent texts and sufficient text feature representation. To handle these problems, we consider text detection as instance segmentation and propose a novel text detection framework, which jointly learns semantic segmentation and a pixel affinity pyramid in a unified fully convolutional network. Specifically, the pixel affinity pyramid is proposed to encode multi-scale instance affiliation relationships of pixels, which is not only robust to varying shapes of text but also provides an accurate boundary description for separating closely located texts. In the inference phase, a simple but effective post-processing is presented to reconstruct text instances from the semantic segmentation results under the guidance of the learned pixel affinity pyramid, achieving good accuracy and efficiency. Furthermore, to enhance the representation of text features in the neural network, two modules — the Region Enhancement Module (REM) and Attentional Fusion Module (AFM) — are proposed. The REM models the semantic correlations of regional features to enhance the features from the text area, which effectively suppresses false-positive detection. The AFM adaptively fuses multi-scale textual information through an attention mechanism to obtain abundant text semantic features, which benefits multi-sized text detection. Extensive ablation experiments are conducted demonstrating the effectiveness of the REM and AFM. Evaluation results on standard benchmarks, including Total-Text, ICDAR2015, SCUT-CTW1500, and MSRA-TD500, show that our method surpasses most existing text detectors and achieves state-of-the-art performance, denoting its superior capability in detecting arbitrary-shaped texts.

Keywords:
Computer science Pyramid (geometry) Segmentation Artificial intelligence Pattern recognition (psychology) Pixel Feature (linguistics) Representation (politics) Inference Convolutional neural network Text detection Natural language processing Image (mathematics) Mathematics

Metrics

17
Cited By
2.10
FWCI (Field Weighted Citation Impact)
67
Refs
0.86
Citation Normalized Percentile
Is in top 1%
Is in top 10%

Citation History

Topics

Handwritten Text Recognition Techniques
Physical Sciences →  Computer Science →  Computer Vision and Pattern Recognition
Vehicle License Plate Recognition
Physical Sciences →  Engineering →  Media Technology
Image Retrieval and Classification Techniques
Physical Sciences →  Computer Science →  Computer Vision and Pattern Recognition

Related Documents

JOURNAL ARTICLE

Arbitrary-shaped text detection with adaptive convolution and path enhancement pyramid network

Qi ChengGuodong WangQian DongBin Wei

Journal:   Multimedia Tools and Applications Year: 2020 Vol: 79 (39-40)Pages: 29225-29242
BOOK-CHAPTER

Bidirectional Regression for Arbitrary-Shaped Text Detection

Tao ShengZhouhui Lian

Lecture notes in computer science Year: 2021 Pages: 187-201
JOURNAL ARTICLE

Fast Arbitrary Shaped Scene Text Detection via Text Discriminator

Chengbin ZengChunli Song

Journal:   Journal of Physics Conference Series Year: 2021 Vol: 2025 (1)Pages: 012014-012014
JOURNAL ARTICLE

Arbitrary-Shaped Text Detection With Adaptive Text Region Representation

Xiufeng JiangShugong XuShunqing ZhangShan Cao

Journal:   IEEE Access Year: 2020 Vol: 8 Pages: 102106-102118
© 2026 ScienceGate Book Chapters — All rights reserved.