Abstract

Treebank is an important resource for both research and application of natural language processing. For Vietnamese, we still lack of such kind of corpora. This paper presents up-to-date results of a project for Vietnamese treebank con-struction. Since Vietnamese is an isolated language and has no word delimiter, there are many ambiguities in sentence analysis. We systematically applied a lot of linguistic techniques to handle such ambiguities. Annotators are supported by automatic-labeling tools and a tree-editor tool. Raw texts are extracted from Youth, an online Vietnamese daily newspaper. The current annotation agreement is around 90 percent. 1

Keywords:
Treebank Vietnamese Computer science Natural language processing Artificial intelligence Annotation Newspaper Sentence Linguistics Political science

Metrics

99
Cited By
6.86
FWCI (Field Weighted Citation Impact)
10
Refs
0.98
Citation Normalized Percentile
Is in top 1%
Is in top 10%

Citation History

Topics

Natural Language Processing Techniques
Physical Sciences →  Computer Science →  Artificial Intelligence
Topic Modeling
Physical Sciences →  Computer Science →  Artificial Intelligence
Text and Document Classification Technologies
Physical Sciences →  Computer Science →  Artificial Intelligence
© 2026 ScienceGate Book Chapters — All rights reserved.