当前位置: 首页 > 文章 > 基于随机森林的文本分类模型研究 农业图书情报学报 2016 (11) 50-54
Position: Home > Articles > Research on text Classification Model Based on Random Forests Journal of Library and Information Science in Agriculture 2016 (11) 50-54

基于随机森林的文本分类模型研究

作  者:
罗新
单  位:
华南理工大学工商管理学院
关键词:
文本分类;随机森林;CART树
摘  要:
文本分类作为处理大量文本数据的关键技术,可以在较大程度上解决"信息爆炸"所带来的问题。Breiman提出的随机森林算法具有良好的泛化性和鲁棒性、对噪声不敏感、能处理连续属性的特点,很适合用来建立文本分类模型。笔者将随机森林算法尝试性引入文本分类领域,构建基于随机森林的文本分类模型,并在标准文本测试集Reuters-21578进行测试和比较,结果表明:(1)该模型可以较好地应用于文本分类;(2)与基于CART、REPTree和J48的文本分类模型的结果相比较,基于随机森林的文本分类模型的效果最好,F1-Measure达到了0.777;(3)基于随机森林的文本分类模型操作方便、直观有效、评价结果可靠,为文本分类研究提供了新思路。
译  名:
Research on text Classification Model Based on Random Forests
作  者:
LUO Xin;School of Business Administration,South China University of Technology;
关键词:
Random forests;;Text classification;;CART(Classification and Regression Tree)
摘  要:
Text classification is the key technology for processing large amount of text data. It can solve the information explosion problem in a certain extent. Random forests algorithm proposed by Breiman has the characteristics of good generalization and robustness, insensitivity for noise and ability in dealing with continuous attributes, which is very suitable for the establishment of text classification model. This paper attempted to construct the text classification model based on random forests algorithm, and compared with the text categorization model Reuters-21578 to verify the model's validity and accuracy for classification. Results showed: this model could be applied in text classification well; compared with the results of CART, REPTree and J48 it models, it had the best effect, whose F1-Measure was 0.777; it had easy, intuitive and effective operation, and reliable results, which provided new idea for text classification research.

相似文章

计量
文章访问数: 8
HTML全文浏览量: 0
PDF下载量: 0

所属期刊

推荐期刊