Real-Time Audio-Visual End-To-End Speech Enhancement

Zirun Zhu; Hemin Yang; Min Tang; Ziyi Yang; Şefik Emre Eskimez; Huaming Wang

doi:10.1109/icassp49357.2023.10094724

ScienceGate Book Chapters

JOURNAL ARTICLE

Real-Time Audio-Visual End-To-End Speech Enhancement

Zirun Zhu Hemin Yang Min Tang Ziyi Yang Şefik Emre Eskimez Huaming Wang

Year: 2023 Pages: 1-5

DOI: 10.1109/icassp49357.2023.10094724

Get Full-Text PDF Get Analytical Report

Abstract

Audio-visual speech enhancement (AV-SE) methods utilize auxiliary visual cues to enhance speakers' voices. Therefore, technically they should be able to outperform the audio-only speech enhancement (SE) methods. However, there are few works in the literature on an AV-SE system that can work in real time on a CPU. In this paper, we propose a low-latency real-time audio-visual end-to-end enhancement (AV-E3Net) model based on the recently proposed end-to-end enhancement network (E3Net). Our main contribution includes two aspects: 1) We employ a dense connection module to solve the performance degradation caused by the deep model structure. This module significantly improves the model's performance on the AV-SE task. 2) We propose a multi-stage gating-and-summation (GS) fusion module to merge audio and visual cues. Our results show that the proposed model provides better perceptual quality and intelligibility than the baseline E3net model with a negligible computational cost increase.

Keywords:

End-to-end principle Computer science Speech recognition Speech enhancement Audio visual End user Computer vision Multimedia Artificial intelligence Noise reduction World Wide Web

Metrics

Cited By

1.61

FWCI (Field Weighted Citation Impact)

Refs

0.80

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Speech and Audio Processing

Physical Sciences → Computer Science → Signal Processing

Advanced Adaptive Filtering Techniques

Physical Sciences → Engineering → Computational Mechanics

Speech Recognition and Synthesis

Physical Sciences → Computer Science → Artificial Intelligence

Real-Time Audio-Visual End-To-End Speech Enhancement

Abstract

Metrics

Citation History

Topics

Related Documents

Visual Context-driven Audio Feature Enhancement for Robust End-to-End Audio-Visual Speech Recognition

End-to-End Audio-Visual Speech Recognition for Overlapping Speech

End-To-End Audio-Visual Speech Recognition with Conformers

An Improved End-to-End Audio-Visual Speech Recognition Model

End-to-End Multi-Person Audio/Visual Automatic Speech Recognition