DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models

Zhengfu He; Tianxiang Sun; Qiong Tang; Kuanning Wang; Xuanjing Huang; Xipeng Qiu

doi:10.18653/v1/2023.acl-long.248

ScienceGate Book Chapters

JOURNAL ARTICLE

DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models

Zhengfu He Tianxiang Sun Qiong Tang Kuanning Wang Xuanjing Huang Xipeng Qiu

Year: 2023 Pages: 4521-4534

DOI: 10.18653/v1/2023.acl-long.248

Get Full-Text PDF Get Analytical Report

Abstract

We present DiffusionBERT, a new generative masked language model based on discrete dif- fusion models. Diffusion models and many pre- trained language models have a shared training objective, i.e., denoising, making it possible to combine the two powerful models and enjoy the best of both worlds. On the one hand, dif- fusion models offer a promising training strat- egy that helps improve the generation quality. On the other hand, pre-trained denoising lan- guage models (e.g., BERT) can be used as a good initialization that accelerates convergence. We explore training BERT to learn the reverse process of a discrete diffusion process with an absorbing state and elucidate several designs to improve it. First, we propose a new noise schedule for the forward diffusion process that controls the degree of noise added at each step based on the information of each token. Sec- ond, we investigate several designs of incorpo- rating the time step into BERT. Experiments on unconditional text generation demonstrate that DiffusionBERT achieves significant improve- ment over existing diffusion models for text (e.g., D3PM and Diffusion-LM) and previous generative masked language models in terms of perplexity and BLEU score. Promising re- sults in conditional generation tasks show that DiffusionBERT can generate texts of compa- rable quality and more diverse than a series of established baselines.

Keywords:

Computer science Perplexity Language model Initialization Artificial intelligence Generative grammar Security token Noise (video) Process (computing) Generative model Machine learning

Metrics

Cited By

14.56

FWCI (Field Weighted Citation Impact)

Refs

0.99

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Topic Modeling

Physical Sciences → Computer Science → Artificial Intelligence

Natural Language Processing Techniques

Physical Sciences → Computer Science → Artificial Intelligence

Speech Recognition and Synthesis

Physical Sciences → Computer Science → Artificial Intelligence

DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models

Abstract

Metrics

Citation History

Topics

Related Documents

Masked Diffusion Language Models with Frequency-Informed Training

Improving Text Style Transfer Using Masked Diffusion Language Models with Inference-Time Scaling

BartSmiles: Generative Masked Language Models for Molecular Representations

Simple and Effective Masked Diffusion Language Models

Graph Generative Models Evaluation with Masked Autoencoder