Dimensionality Reduction using Clustering Techniques

Authors

  • Snehal D. Borase Author
  • Prof. Satish S. Banait Author

Keywords:

Clustering algorithms, Big data, Nearest neighbours, Outliers

Abstract

Clustering is often an activity involving discovering homogeneous sets of the particular studied physical objects. Recently, numerous analysts use a considerable desire for establishing clustering algorithms. Clustering plays a significant position in lots of info mining apps including, computational biology, spatial repository apps, data collection, word mining, CRM, healthcare diagnostics, controlled info query, promoting, and also net research. “Big data” is actually speaking about terabytes and also pet bytes involving info. Major info is actually complicated because of its a few essential traits including quantity, speed, range, variability and also complication. Thus massive info is actually difficult to address utilizing typical resources and also strategies. You will find so many issues inside clustering strategies, therefore a few of the issues is actually how you can procedure the results and also massive info is actually clustered inside scaled-down data format, Clustering formula are afflicted by security issue, set involving sole and also numerous degree clustering. A crucial concern inside clustering is actually that individuals do not have sooner know-how relating to info. Likewise number of feedback boundaries including variety of most adjacent neighbours, variety of clusters inside these types of algorithms creates clustering some sort of complicated activity. The leading aim is to review and also review the present clustering algorithms, effect involving dimensionality lessening and also working with outliers.

References

[1] J. MacQueen, ‘‘Some methods for classification and analysis of multivariate observations,’’ in Proc. 5th Berkeley Symp. Math. Statist. Probab., Berkeley, CA, USA, 1967, pp. 281–297.

[2] A. P. Dempster; N. M. Laird; D. B. Rubin, Maximum Likelihood from Incomplete Data via the EM Algorithm, Journal of the Royal Statistical Society. Series B (Methodological), Vol. 39, No. 1. (1977), pp. 1-38

[3] J. C.Bezdek, R.Ehrlich, and W.Full,‘‘FCM: Thefuzzy c-means clustering algorithm,’’ Comput. Geosci., vol. 10, nos. 2–3, pp. 191–203, 1984.

[4] D. H. Fisher, ‘‘Knowledge acquisition via incremental conceptual clustering,’’ Mach. Learn., vol. 2, no. 2, pp. 139–172, Sep. 1987.

[5] A.K.Jainand R.C.Dubes, Algorithms for Clustering Data.Upper Saddle River, NJ, USA: Prentice-Hall, 1988.

[6] R.T.Ng andJ.Han, ‘‘Efficient and effective clustering methods for spatial data mining,’’ in Proc. Int. Conf. Very Large Data Bases (VLDB), 1994, pp. 144–155

[7] T. Zhang, R. Ramakrishnan, and M. Livny, ‘‘BIRCH: An efficient data clustering method for very large databases,’’in Proc. ACMSIGMOD Rec., Jun. 1996, vol. 25, no. 2, pp. 103–114

[8] Ester M., Kriegel H.-P., Sander J., Xu X.: “A Density- Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise”, Proc. 2cnd Int. Conf. On Knowledge Discovery and Data Mining, Portland, Oregon, 1996, AAAI Press, 1996.

[9] Z.Huang,‘‘A fast clustering algorithm to cluster very large categorical datasets in data mining,’’in Proc. SIGMOD Workshop Res.Issues Data Mining Knowl. Discovery, 1997, pp. 1–8.

[10] X. Xu, M. Ester, H.-P. Kriegel, and J. Sander, ‘‘A distribution-based clustering algorithm fo rmining in large spatial ldatabases,’’inProc.14thIEEE Int. Conf. Data Eng. (ICDE), Feb. 1998, pp. 324–331.

[11] S. Guha, R. Rastogi, and K. Shim, ‘‘CURE: An efficient clustering algorithm For large databases,’’inProc. CMSIGMODRec., Jun.1998,vol.27,no.2, pp. 73–84

[12] G. Sheikholeslami, S. Chatterjee, and A. Zhang, ‘‘Wavecluster: A multi resolution clustering approach for very large spatial databases,’’ in Proc. Int. Conf. Very Large Data Bases (VLDB), 1998, pp. 428–439.

[13] A. Hinneburg and D. A. Keim, ‘‘Optimal grid-clustering: Towards breaking the curse of dimensionality in high-dimensional clustering,’’ in Proc. 25th Int. Conf. Very Large Data Bases (VLDB), 1999, pp. 506–517.

[14] G.Karypis,E.-H.Han,andV.Kumar,‘‘Chameleon:Hierarchicalclustering using dynamic modelling,’’ IEEE Comput., vol. 32, no. 8, pp. 68–75, Aug. 1999.

[15] S. Guha, R. Rastogi, and K. Shim, ‘‘Rock: A robust clustering algorithm for categorical attributes,’’ Inform.Syst., vol.25,no.5,pp.345–366,2000.

[16] R. T. Ng and J. Han, ‘‘CLARANS: A method for clustering objects for spatial data mining,’’IEEE Trans. Knowl. Data Eng.(TKDE),vol.14,no.5, pp. 1003–1016, Sep./Oct. 2002.

[17] A. N. Mahmood, C. Leckie, and P. Udaya, ‘‘ECHIDNA: Efficient clustering of hierarchical data for network traffic analysis,’’ in Proc. 5th Int. IFIP-TC6 Conf. Netw. Technol., Services, Protocols Perform. Comput. Commun. Netw. Mobile Wireless Commun. Syst. (NETWORKING), 2006, pp. 1092–1098.

[18] A. Fahad, N. alshatri, Z. Tari, A. Alamri, I. Khalil, A. Y. Zomaya, S. Foufou, and A. Bouras, “A survey of clustering algorithms for big data: taxonomy and empirical analysis”, IEEE Transactions on emerging topics in computing, vol 2, no. 3, Sept 2014.

Downloads

Published

2016-04-30

How to Cite

Dimensionality Reduction using Clustering Techniques. (2016). International Journal of Advanced Research in Science, Management and Technology, 2(2), 1-7. https://ijarsmt.in/ijarsmt/article/view/30

Most read articles by the same author(s)

Similar Articles

1-10 of 48

You may also start an advanced similarity search for this article.