creators_name: Guan, Hu
creators_name: Zhou, Jingyu
creators_name: Guo, Minyi
type: conference_item
datestamp: 2009-04-06 19:08:56
lastmod: 2009-04-14 04:37:06
metadata_visibility: show
title: A Class-Feature-Centroid Classifier for Text Categorization
ispublished: pub
full_text_status: public
pres_type: paper
abstract: Automated text categorization is an important technique for many web applications, such as document indexing, document filtering, and cataloging web resources. Many different approaches have been proposed for the automated text categorization problem. Among them, centroid-based approaches have the advantages of short training time and testing time due to its computational efficiency. As a result, centroid-based classifiers have been widely used in many web applications. However, the accuracy of centroid-based classifiers is inferior to SVM, mainly because centroids found during construction are far from perfect locations. We design a fast Class-Feature-Centroid (CFC) classifier for multi-class, single-label text categorization. In CFC, a centroid is built from two important class distributions: inter-class term index and inner-class term index. CFC proposes a novel combination of these indices and employs a denormalized cosine measure to calculate the similarity score between a text vector and a centroid. Experiments on the Reuters-21578 corpus and 20-newsgroup email collection show that CFC consistently outperforms the state-of-the-art SVM classifiers on both micro-F1 and macro-F1 scores. Particularly, CFC is more effective and robust than SVM when data is sparse.
date: 2009-04
pagerange: 201-201
event_title: 18th International World Wide Web Conference
event_location: Madrid, Spain
event_dates: April 20th-24th, 2009
event_type: conference
refereed: TRUE
citation: Guan, Hu <http://www2009.eprints.org/view/author/Guan=3AHu=3A=3A.html> and Zhou, Jingyu <http://www2009.eprints.org/view/author/Zhou=3AJingyu=3A=3A.html> and Guo, Minyi <http://www2009.eprints.org/view/author/Guo=3AMinyi=3A=3A.html> (2009) A Class-Feature-Centroid Classifier for Text Categorization. In: 18th International World Wide Web Conference, April 20th-24th, 2009, Madrid, Spain.
document_url: http://www2009.eprints.org/21/1/p201.pdf