\nIn this paper, aiming at the characteristics of Chinese text classification, using the ICTCLAS(Chinese lexical analysis system of Chinese academy of sciences) for document segmentation, and for data cleaning and filtering the Stop words, using the information gain and document frequency feature selection algorithm to document feature selection. Based on this, based on the Naive Bayesian algorithm implemented text classifier , and use Chinese corpus of Fudan University has carried on the experiment and analysis on the system.\n
Peng LiuHui-han ZhaoJia-yu TengYanyan YangYafeng LiuZongwei Zhu
Xinyi SunLiming DuYanshuang Wang