Skip to main navigation Skip to search Skip to main content

Near-optimal feature selection for large databases

  • J. Yang*
  • , S. Ólafsson
  • *Corresponding author for this work
  • Iowa State University

Research output: Contribution to journalJournal articlepeer-review

Abstract

We analyse a new optimization-based approach for feature selection that uses the nested partitions method for combinatorial optimization as a heuristic search procedure to identify good feature subsets. In particular, we show how to improve the performance of the nested partitions method using random sampling of instances. The new approach uses a two-stage sampling scheme that determines the required sample size to guarantee convergence to a near-optimal solution. This approach therefore also has attractive theoretical characteristics. In particular, when the algorithm terminates in finite time, rigorous statements can be made concerning the quality of the final feature subset. Numerical results are reported to illustrate the key results, and show that the new approach is considerably faster than the original nested partitions method and other feature selection methods.Journal of the Operational Research Society (2009) 60, 1045-1055. doi:10.1057/palgrave.jors.2602651; published online 20 August 2008.

Original languageEnglish
Pages (from-to)1045-1055
Number of pages11
JournalJournal of the Operational Research Society
Volume60
Issue number8
DOIs
StatePublished - 2009.08

Keywords

  • Combinatorial optimization
  • Data mining
  • Feature selection
  • Scalability

Quacquarelli Symonds(QS) Subject Topics

  • Business & Management Studies
  • Mathematics
  • Statistics & Operational Research
  • Data Science

Fingerprint

Dive into the research topics of 'Near-optimal feature selection for large databases'. Together they form a unique fingerprint.

Cite this