Skip to main navigation Skip to search Skip to main content

Incorporating virtual relevant documents for learning in text categorization

  • NII (National Institute of Informatics)

Research output: Contribution to conferenceChapterpeer-review

Abstract

This paper proposes a virtual relevant document technique in the learning phase for text categorization. The method uses a simple transformation of relevant documents, i.e. making virtual documents by combining document pairs in the training set. The virtual document produced by this method has the enriched term vector space, with greater weights for the terms that co-occur in two relevant documents. The experimental results showed a significant improvement over the baseline, which proves the usefulness of the proposed method: 71% improvement on TREC-11 filtering test collection and 11% improvement on Reuters-21578 test set for the topics with less than 100 relevant documents in the micro average F1. The result analysis indicates that the addition of virtual relevant documents contributes to the steady improvement of the performance.

Original languageEnglish
Title of host publicationLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
EditorsTengku Mohd Tengku Sembok, Halimah Badioze Zaman, Hsinchun Chen, Shalini R. Urs, Sung Hyon Myaeng
PublisherSpringer Verlag
Pages62-72
Number of pages11
ISBN (Electronic)9783540206088
DOIs
StatePublished - 2003

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume2911
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Fingerprint

Dive into the research topics of 'Incorporating virtual relevant documents for learning in text categorization'. Together they form a unique fingerprint.

Cite this