Improving Neural Machine Translation Through Code‐Mixed Data Augmentation

Ramakrishna Appicharla; Kamal Gupta; Asif Ekbal; Pushpak Bhattacharyya

doi:10.1111/coin.70033

ScienceGate Book Chapters

JOURNAL ARTICLE

Improving Neural Machine Translation Through Code‐Mixed Data Augmentation

Ramakrishna Appicharla Kamal Gupta Asif Ekbal Pushpak Bhattacharyya

Year: 2025 Journal: Computational Intelligence Vol: 41 (2) Publisher: Wiley

DOI: 10.1111/coin.70033

Get Full-Text PDF Get Analytical Report

Abstract

ABSTRACT This paper studies neural machine translation (NMT) of code‐mixed (CM) text. Specifically, we generate synthetic CM data and how it can be used to improve the translation performance of NMT through the data augmentation strategy. We conduct experiments on three data augmentation approaches viz. CM‐Augmentation, CM‐Concatenation, and Multi‐Encoder approaches, and the latter two approaches are inspired by document‐level NMT, where we use synthetic CM data as context to improve the performance of the NMT models. We conduct experiments on three language pairs, viz. Hindi–English, Telugu–English and Czech–English. Experimental results demonstrate that the proposed approaches significantly improve performance over the baseline model trained without data augmentation and over the existing data augmentation strategies. The CM‐Concatenation model attains the best performance.

Keywords:

Computer science Machine translation Natural language processing Artificial intelligence Translation (biology) Speech recognition Code (set theory) Programming language Biology

Metrics

Cited By

4.82

FWCI (Field Weighted Citation Impact)

Refs

0.92

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Natural Language Processing Techniques

Physical Sciences → Computer Science → Artificial Intelligence

Topic Modeling

Physical Sciences → Computer Science → Artificial Intelligence

Speech Recognition and Synthesis

Physical Sciences → Computer Science → Artificial Intelligence

Improving Neural Machine Translation Through Code‐Mixed Data Augmentation

Abstract

Metrics

Citation History

Topics

Related Documents

Robust Data Augmentation for Neural Machine Translation through EVALNET

Training Data Augmentation for Code-Mixed Translation

Corpus Augmentation for Improving Neural Machine Translation

Improving the Punjabi-Hindi Braille Neural Machine Translation through Syntax Augmentation

AdMix: A Mixed Sample Data Augmentation Method for Neural Machine Translation