Multi-Level Feature Dynamic Fusion Neural Radiance Fields for Audio-Driven Talking Head Generation

Wenchao Song; Qiong Liu; Yanchao Liu; Pengzhou Zhang; Juan Cao

doi:10.3390/app15010479

ScienceGate Book Chapters

JOURNAL ARTICLE

Multi-Level Feature Dynamic Fusion Neural Radiance Fields for Audio-Driven Talking Head Generation

Wenchao Song Qiong Liu Yanchao Liu Pengzhou Zhang Juan Cao

Year: 2025 Journal: Applied Sciences Vol: 15 (1)Pages: 479-479 Publisher: Multidisciplinary Digital Publishing Institute

DOI: 10.3390/app15010479

Get Full-Text PDF Get Analytical Report

Abstract

Audio-driven cross-modal talking head generation has experienced significant advancement in the last several years, and it aims to generate a talking head video that corresponds to a given audio sequence. Out of these approaches, the NeRF-based method can generate videos featuring a specific person with more natural motion compared to the one-shot methods. However, previous approaches failed to distinguish the importance of different regions, resulting in the loss of information-rich region features. To alleviate the problem and improve video quality, we propose MLDF-NeRF, an end-to-end method for talking head generation, which can achieve better vector representation through multi-level feature dynamic fusion. Specifically, we designed two modules in MLDF-NeRF to enhance the cross-modal mapping ability between audio and different facial regions. We initially developed a multi-level tri-plane hash representation that uses three sets of tri-plane hash networks with varying resolutions of limitation to capture the dynamic information of the face more accurately. Then, we introduce the idea of multi-head attention and design an efficient audio-visual fusion module that explicitly fuses audio features with image features from different planes, thereby improving the mapping between audio features and spatial information. Meanwhile, the design helps to minimize interference from facial areas unrelated to audio, thereby improving the overall quality of the representation. The quantitative and qualitative results indicate that our proposed method can effectively generate talk heads with natural actions and realistic details. Compared with previous methods, it performs better in terms of image quality, lip sync, and other aspects.

Keywords:

Computer science Head (geology) Artificial intelligence Fusion Feature (linguistics) Computer vision Speech recognition Geology

Metrics

Cited By

10.65

FWCI (Field Weighted Citation Impact)

Refs

0.91

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Speech and Audio Processing

Physical Sciences → Computer Science → Signal Processing

Speech Recognition and Synthesis

Physical Sciences → Computer Science → Artificial Intelligence

Music and Audio Processing

Physical Sciences → Computer Science → Signal Processing

Multi-Level Feature Dynamic Fusion Neural Radiance Fields for Audio-Driven Talking Head Generation

Abstract

Metrics

Citation History

Topics

Related Documents

Dynamic Region Fusion Neural Radiance Fields for Audio-Driven Talking Head Generation

Emotional Semantic Neural Radiance Fields for Audio-Driven Talking Head

ERF: Generating Hierarchical Editing Audio-Driven Talking Head Using Dynamic Neural Radiance Fields

AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis

Decoupled Two-Stage Talking Head Generation via Gaussian-Landmark-Based Neural Radiance Fields