JOURNAL ARTICLE

CTNeRF: Cross-time Transformer for dynamic neural radiance field from monocular video

Abstract

The goal of our work is to generate high-quality novel views from monocular videos of complex and dynamic scenes. Prior methods, such as DynamicNeRF, have shown impressive performance by leveraging time-varying dynamic radiation fields. However, these methods have limitations when it comes to accurately modeling the motion of complex objects, which can lead to inaccurate and blurry renderings of details. To address this limitation, we propose a novel approach that builds upon a recent generalization NeRF, which aggregates nearby views onto new viewpoints. However, such methods are typically only effective for static scenes. To overcome this challenge, we introduce a module that operates in both the time and frequency domains to aggregate the features of object motion. This allows us to learn the relationship between frames and generate higher-quality images. Our experiments demonstrate significant improvements over state-of-the-art methods on dynamic scene datasets. Specifically, our approach outperforms existing methods in terms of both the accuracy and visual quality of the synthesized views. Our code is available on https://github.com/xingy038/CTNeRF.

Keywords:
Radiance Computer science Artificial intelligence Monocular Transformer Computer vision Remote sensing Geology Engineering Electrical engineering Voltage

Metrics

6
Cited By
4.91
FWCI (Field Weighted Citation Impact)
53
Refs
0.91
Citation Normalized Percentile
Is in top 1%
Is in top 10%

Citation History

Topics

Optical Imaging and Spectroscopy Techniques
Health Sciences →  Medicine →  Radiology, Nuclear Medicine and Imaging
Visual perception and processing mechanisms
Life Sciences →  Neuroscience →  Cognitive Neuroscience
Infrared Thermography in Medicine
Health Sciences →  Medicine →  Radiology, Nuclear Medicine and Imaging
© 2026 ScienceGate Book Chapters — All rights reserved.