Abstract
This paper presents a detailed study of technologies based on Hadoop and MapReduce available over the cloud for large-scale data mining and predictive analytics. Although some studies may have shown that cloud technologies relying on the MapReduce framework do not perform as well as parallel database management systems, e.g., with ad hoc queries and interactive applications, MapReduce has still been widely used by many organizations for big data storage and analytics. A number of MapReduce based tools are broadly available over the cloud. In this work we explore the Apache Hive data warehousing solution and particularly its Mahout data mining libraries for predictive analytics. We present results in the context of text classification, recommender systems and decision support. We develop prototype tools in these areas and discuss our outcomes from the study useful to researchers and other professionals in cloud computing and application domains. To the best of our knowledge, ours is among the first few in-depth studies on Mahout with application prototypes available for use.
Original language | English |
---|---|
Pages | 607-612 |
Number of pages | 6 |
DOIs | |
State | Published - 2013 |
Event | 2013 13th IEEE International Conference on Data Mining Workshops, ICDMW 2013 - Dallas, TX, United States Duration: 7 Dec 2013 → 10 Dec 2013 |
Other
Other | 2013 13th IEEE International Conference on Data Mining Workshops, ICDMW 2013 |
---|---|
Country/Territory | United States |
City | Dallas, TX |
Period | 7/12/13 → 10/12/13 |
Keywords
- Cloud computing
- Data mining
- Mahout
- Predictive analytics