JOURNAL ARTICLE

Using noisy bilingual data for statistical machine translation

Abstract

SMT systems rely on sufficient amount of parallel corpora to train the translation model. This paper investigates possibilities to use word-to-word and phrase-to-phrase translations extracted not only from clean parallel corpora but also from noisy comparable corpora. Translation results for a Chinese to English translation task are given.

Keywords:
Machine translation Computer science Phrase Natural language processing Artificial intelligence Translation (biology) Example-based machine translation Word (group theory) Task (project management) Parallel corpora Synchronous context-free grammar Transfer-based machine translation Machine translation software usability Rule-based machine translation Speech recognition Linguistics Engineering

Metrics

15
Cited By
1.15
FWCI (Field Weighted Citation Impact)
6
Refs
0.82
Citation Normalized Percentile
Is in top 1%
Is in top 10%

Citation History

Topics

Natural Language Processing Techniques
Physical Sciences →  Computer Science →  Artificial Intelligence
Topic Modeling
Physical Sciences →  Computer Science →  Artificial Intelligence
Algorithms and Data Compression
Physical Sciences →  Computer Science →  Artificial Intelligence
© 2026 ScienceGate Book Chapters — All rights reserved.