Dimensionality Reduction using Clustering Techniques
Keywords:
Clustering algorithms, Big data, Nearest neighbours, OutliersAbstract
Clustering is often an activity involving discovering homogeneous sets of the particular studied physical objects. Recently, numerous analysts use a considerable desire for establishing clustering algorithms. Clustering plays a significant position in lots of info mining apps including, computational biology, spatial repository apps, data collection, word mining, CRM, healthcare diagnostics, controlled info query, promoting, and also net research. “Big data” is actually speaking about terabytes and also pet bytes involving info. Major info is actually complicated because of its a few essential traits including quantity, speed, range, variability and also complication. Thus massive info is actually difficult to address utilizing typical resources and also strategies. You will find so many issues inside clustering strategies, therefore a few of the issues is actually how you can procedure the results and also massive info is actually clustered inside scaled-down data format, Clustering formula are afflicted by security issue, set involving sole and also numerous degree clustering. A crucial concern inside clustering is actually that individuals do not have sooner know-how relating to info. Likewise number of feedback boundaries including variety of most adjacent neighbours, variety of clusters inside these types of algorithms creates clustering some sort of complicated activity. The leading aim is to review and also review the present clustering algorithms, effect involving dimensionality lessening and also working with outliers.
References
[1] J. MacQueen, ‘‘Some methods for classification and analysis of multivariate observations,’’ in Proc. 5th Berkeley Symp. Math. Statist. Probab., Berkeley, CA, USA, 1967, pp. 281–297.
[2] A. P. Dempster; N. M. Laird; D. B. Rubin, Maximum Likelihood from Incomplete Data via the EM Algorithm, Journal of the Royal Statistical Society. Series B (Methodological), Vol. 39, No. 1. (1977), pp. 1-38
[3] J. C.Bezdek, R.Ehrlich, and W.Full,‘‘FCM: Thefuzzy c-means clustering algorithm,’’ Comput. Geosci., vol. 10, nos. 2–3, pp. 191–203, 1984.
[4] D. H. Fisher, ‘‘Knowledge acquisition via incremental conceptual clustering,’’ Mach. Learn., vol. 2, no. 2, pp. 139–172, Sep. 1987.
[5] A.K.Jainand R.C.Dubes, Algorithms for Clustering Data.Upper Saddle River, NJ, USA: Prentice-Hall, 1988.
[6] R.T.Ng andJ.Han, ‘‘Efficient and effective clustering methods for spatial data mining,’’ in Proc. Int. Conf. Very Large Data Bases (VLDB), 1994, pp. 144–155
[7] T. Zhang, R. Ramakrishnan, and M. Livny, ‘‘BIRCH: An efficient data clustering method for very large databases,’’in Proc. ACMSIGMOD Rec., Jun. 1996, vol. 25, no. 2, pp. 103–114
[8] Ester M., Kriegel H.-P., Sander J., Xu X.: “A Density- Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise”, Proc. 2cnd Int. Conf. On Knowledge Discovery and Data Mining, Portland, Oregon, 1996, AAAI Press, 1996.
[9] Z.Huang,‘‘A fast clustering algorithm to cluster very large categorical datasets in data mining,’’in Proc. SIGMOD Workshop Res.Issues Data Mining Knowl. Discovery, 1997, pp. 1–8.
[10] X. Xu, M. Ester, H.-P. Kriegel, and J. Sander, ‘‘A distribution-based clustering algorithm fo rmining in large spatial ldatabases,’’inProc.14thIEEE Int. Conf. Data Eng. (ICDE), Feb. 1998, pp. 324–331.
[11] S. Guha, R. Rastogi, and K. Shim, ‘‘CURE: An efficient clustering algorithm For large databases,’’inProc. CMSIGMODRec., Jun.1998,vol.27,no.2, pp. 73–84
[12] G. Sheikholeslami, S. Chatterjee, and A. Zhang, ‘‘Wavecluster: A multi resolution clustering approach for very large spatial databases,’’ in Proc. Int. Conf. Very Large Data Bases (VLDB), 1998, pp. 428–439.
[13] A. Hinneburg and D. A. Keim, ‘‘Optimal grid-clustering: Towards breaking the curse of dimensionality in high-dimensional clustering,’’ in Proc. 25th Int. Conf. Very Large Data Bases (VLDB), 1999, pp. 506–517.
[14] G.Karypis,E.-H.Han,andV.Kumar,‘‘Chameleon:Hierarchicalclustering using dynamic modelling,’’ IEEE Comput., vol. 32, no. 8, pp. 68–75, Aug. 1999.
[15] S. Guha, R. Rastogi, and K. Shim, ‘‘Rock: A robust clustering algorithm for categorical attributes,’’ Inform.Syst., vol.25,no.5,pp.345–366,2000.
[16] R. T. Ng and J. Han, ‘‘CLARANS: A method for clustering objects for spatial data mining,’’IEEE Trans. Knowl. Data Eng.(TKDE),vol.14,no.5, pp. 1003–1016, Sep./Oct. 2002.
[17] A. N. Mahmood, C. Leckie, and P. Udaya, ‘‘ECHIDNA: Efficient clustering of hierarchical data for network traffic analysis,’’ in Proc. 5th Int. IFIP-TC6 Conf. Netw. Technol., Services, Protocols Perform. Comput. Commun. Netw. Mobile Wireless Commun. Syst. (NETWORKING), 2006, pp. 1092–1098.
[18] A. Fahad, N. alshatri, Z. Tari, A. Alamri, I. Khalil, A. Y. Zomaya, S. Foufou, and A. Bouras, “A survey of clustering algorithms for big data: taxonomy and empirical analysis”, IEEE Transactions on emerging topics in computing, vol 2, no. 3, Sept 2014.
Downloads
Published
Issue
Section
Categories
License

This work is licensed under a Creative Commons Attribution 4.0 International License.
This work is licensed under a Creative Commons Attribution 4.0 International License.
Under this license, authors retain ownership of the copyright for their articles. By submitting to the International Journal of Advanced Research in Science, Management, and Technology (IJARSMT), authors grant the journal the right of first publication. Users are free to share, copy, and redistribute the material in any medium or format, and to adapt, remix, transform, and build upon the material for any purpose, including commercially, provided that appropriate credit is given to the original author(s) and the journal, a link to the license is provided, and any changes made are indicated.
