A Survey on Clustering Techniques for Mining Big Data

Authors

  • Prachi Surwade Author
  • Prof. Satish S. Banait Author

Keywords:

Data mining, Big data, Clustering techniques, Big data analytics.

Abstract

Clustering technique is mining process in which whole dataset is divided in to meaningful subclasses. It is unsupervised classification of data items (feature vectors (FV) observations, or patterns unsupervised) into teams (cluster). Clustering process is very useful in numerous pattern classification, grouping, exploratory pattern analysis, machine learning, document retrieval, image segmentation and decision making. “Big data” is nothing but the datasets whose size in on many facets the power of typical software package data tools to efficiently capture, store, manage and analyze. That is we tend to outline massive information in terms of being bigger than meticulous verity of thousands of gigabyte (terabyte). Big data is useful for populace and corporations but in some cases it is difficult to store and it is also time consuming. So, one of the ways to overcome this problem is to improve the clustering methods, however it suffers from high computational complexity. Data mining is the technique in which helpful information and hidden relationship among data is extracted, but the traditional data mining approaches cannot be directly used for big data due to their inherent complexity. Key objective is to introduce a simple general overview of data clustering categorizations for big data. And also explain and summarizing some of the related work for it. This paper presents a theoretical overview of some of current clustering techniques used for analyzing big data.

References

[1] Btissam Zerhari, Ayoub Ait Lahcen, Salma Mouline1, “Big Data Clustering: Algorithms and Challenges”, CONFERENCE PAPER · MAY 2015F. Chung, Spectral Graph Theory. Providence, RI, USA: American Mathematical Society, 1997.

[2] S. Suthaharan, M. Alzahrani, ``Labelled data collection for anomaly detection in wireless sensor networks,'' in Proc. 6th Int. Conf. Intell. Sensors, Sensor Netw. Inform. Process. (ISSNIP), Dec. 2010, pp. 269_274.

[3] A.K. Jain, M.N. Murty, and P.J. Flynn, “Data Clustering: A Review,” ACM Computing Surveys, vol. 31, no. 3, pp. 264-323, Sept. 1999.

[4] A. Katal, M. Wazid and R.H. Goudar, “Big data: Issues, challenges, tools and goodpractices,” Contemporary Computing (IC3), 2013 Sixth International Conference on,IEEE, 2013.

[5] R. Xu and D. Wunsch, “Survey of clustering algorithms.,” IEEE transactions on neural networks / a publication of the IEEE Neural Networks Council, vol. 16, no. 3, pp. 645-78, May. 2005.

[6] 12 F. Bu, Z. Chen, Q. Zhang, and X. Wang, “Incomplete Big Data Clustering Algorithm Using Feature Selection and Partial Distance,” InDigital Home (ICDH), 5th International Conference on. IEEE, p. 263-266, 2014.

[7] B. J. Kim, “A Classifier for Big Data,” In Convergence and Hybrid Information Technology. Springer Berlin Heidelberg, p. 505-512, 2012.

[8] Manasi N. Joshi,” Parallel K - Means Algorithm on Distributed Memory Multiprocessors” Spring 2003 Computer Science Department University of Minnesota, Twin Cities

[9] K. Stoffel and A. Belkoniene, “Parallel k/h-means clustering for large data sets,” In Euro-Par’99 Parallel Processing. Springer Berlin Heidelberg, p. 1451-1454, 1999

[10] Donald Miner and Adam Shook” MapReduce Design Patterns”, Printed in the United States of America.

[11] Y. Zhao, Y. Chen, Z. Liang, S. Yuan, and Y. Li, “Big Data Processing with Probabilistic Latent Semantic Analysis on MapReduce, International Conference on Cyber-Enabled Distributed Computing and Knowledge Discovery, 162 – 166, 2014.

[12] K. Younghoon, S. Kyuseok, K. Min-Soeng, L. June Sup, “DBCUREMR: An efficient density-based clustering algorithm for large data using MapReduce,” Information Systems, vol. 42, p. 15-35, 2014.

[13] T. Zhang, R. Ramakrishnan, and M. Livny, ``BIRCH: An efficient data clustering method for very large databases,'' in Proc. ACM SIGMOD Rec., Jun. 1996, vol. 25, no. 2, pp. 103_114.

[14] S. Guha, R. Rastogi, and K. Shim, ``Cure: An ef_cient clustering algorithm for large databases,'' in Proc. ACMSIGMOD Rec., Jun. 1998, vol. 27, no. 2.

[15] G. Karypis, E.-H. Han, and V. Kumar, ``Chameleon: Hierarchical clustering using dynamic modelling,'' IEEE Comput., vol. 32, no. 8, pp. 68_75, Aug. 1999.

[16] S. Guha, R. Rastogi, and K. Shim, ``Rock: A robust clustering algorithm for categorical attributes,'' Inform. Syst., vol. 25, no. 5, pp. 345_366, 2000.

[17] M. Dutta, A. Kakoti Mahanta and A.K. Pujari, QROCK: A quick version of the ROCK algorithm for clustering of categorical data, Pattern Recognition Letters, 26 (2005), 2364-2373.

[18] R. T. Ng and J. Han, ``CLARANS: A method for clustering objects for spatial data mining,'' IEEE Trans. Knowl. Data Eng. (TKDE), vol. 14, no. 5, pp. 1003_1016, Sep./Oct. 2002.

[19] ALSABTI K., RANKA S., SINGH V., “ An Efficient k-means Clustering Algorithm, Proc. First Workshop High Performance Data Mining, 1998.

[20] ]Ester M., Kriegel H.-P., Sander J., Xu X.: “A Density- Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise”, Proc. 2cnd Int. Conf. On Knowledge Discovery and Data Mining, Portland, Oregon, 1996, AAAI Press, 1996.

[21] A. Hinneburg and D. A. Keim, ‘‘Optimal grid-clustering: Towards breaking the curse of dimensionality in high-dimensional clustering,’’ in Proc. 25th Int. Conf. Very Large Data Bases (VLDB), 1999, pp. 506–517.

[22] Dhillon , Xu , Stoffle]and A. Belkonience. “Parallel k-Means clustering for large datasets”. Proceedings of EuroPar -1999.

Downloads

Published

2016-10-30

How to Cite

A Survey on Clustering Techniques for Mining Big Data. (2016). International Journal of Advanced Research in Science, Management and Technology, 2(5), 1-11. https://ijarsmt.in/ijarsmt/article/view/41

Most read articles by the same author(s)

Similar Articles

41-50 of 58

You may also start an advanced similarity search for this article.