Beibei, Zou and Xuesong, Ma and Bettina, Kemme and Glen, Newton and Doina, Precup (2006) Data Mining Using Relational Database Management Systems. [Preprint]
Full text available as:
|
PDF
77Kb |
Abstract
Software packages providing a whole set of data mining and machine learning algorithms are attractive because they allow experimentation with many kinds of algorithms in an easy setup. However, these packages are often based on main-memory data structures, limiting the amount of data they can handle. In this paper we use a relational database as secondary storage in order to eliminate this limitation. Unlike existing approaches, which often focus on optimizing a single algorithm to work with a database backend, we propose a general approach, which provides a database interface for several algorithms at once. We have taken a popular machine learning software package, Weka, and added a relational storage manager as back-tier to the system. The extension is transparent to the algorithms implemented in Weka, since it is hidden behind Weka’s standard main-memory data structure interface. Furthermore, some general mining tasks are transfered into the database system to speed up execution. We tested the extended system, refered to as WekaDB, and our results show that it achieves a much higher scalability than Weka, while providing the same output and maintaining good computation time.
Item Type: | Preprint |
---|---|
Additional Information: | Supported by the National Science and Engineering Council (NSERC), the Cnada Foundation for Innovation (CFI) and the National Research Council (NRC) Canada Institute for Scientific and Technical Information (CISTI). |
Keywords: | data mining, machine learning, data structures, WEKA |
Subjects: | Computer Science > Machine Learning |
ID Code: | 4851 |
Deposited By: | Newton, Glen |
Deposited On: | 04 May 2006 |
Last Modified: | 11 Mar 2011 08:56 |
References in Article
Select the SEEK icon to attempt to find the referenced article. If it does not appear to be in cogprints you will be forwarded to the paracite service. Poorly formated references will probably not work.
Metadata
- ASCII Citation
- Atom
- BibTeX
- Dublin Core
- EP3 XML
- EPrints Application Profile (experimental)
- EndNote
- HTML Citation
- ID Plus Text Citation
- JSON
- METS
- MODS
- MPEG-21 DIDL
- OpenURL ContextObject
- OpenURL ContextObject in Span
- RDF+N-Triples
- RDF+N3
- RDF+XML
- Refer
- Reference Manager
- Search Data Dump
- Simple Metadata
- YAML
Repository Staff Only: item control page