Relevance Feature Discovery for Text Mining using Dynamic Approach

Authors

  • Salve Bhavna B. Author
  • Prof. N. V. Alone Author

Keywords:

Text Mining, Text Feature Extraction, Text Classification, Hierarchical Agglomerative clustering, Term polarity, Term frequency.

Abstract

Text mining techniques helps users to find useful information from a large amount of digital text documents on a Web or databases engines. It is therefore crucial that a good text mining model should retrieve the information that meets users’ needs within a relatively efficient time frame. Traditional Information Retrieval (IR) has the same goal of automatically retrieving relevant documents as many as possible while filtering out non-relevant ones at the same time. It was intended to generate dynamic models to classify multiple topics in a collection of documents. A fundamental supposition for these approaches is that the documents in the collection are all about one topic. To guarantee the quality of discovered relevance feature in text documents for describing the user preferences is very crucial and challenging task because of large scale terms and data patterns. Most existing text mining and classification methods adopt term based approaches which suffer from the problem of polysemy and synonymy. We all believe in hypothesis that the pattern based methods performs better than term based ones in describing user preferences. The challenging issue is large scale patterns remains as hard problem in text mining. To deal with the above mentioned limitations and problems, proposed model presents the dynamic approach for Relevance Feedback Discovery by classifying terms into different categories dynamically and updating term weights and their distribution in patterns efficiently by improving the performance of text mining. The proposed model significantly outperform both Term based Methods and Pattern based methods. To evaluate the effectiveness of the proposed model TREC data collection, Reuters Corpus Volume 1 and Reuters-21578 are used.

References

[1] C. Buckley, G. Salton, and J. Allan, “The effect of adding relevance information in a relevance feedback environment,” in Proc. Annu. Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval, 1994, pp. 292–300.

[2] Y. Li, A. Algarni, and N. Zhong, “Mining positive and negative patterns for relevance feature discovery,” in Proc. ACM SIGKDD Knowl. Discovery Data Mining, 2010, pp. 753–762.

[3] M. Seno and G. Karypis, “Slpminer: An algorithm for finding frequent sequential patterns using length-decreasing support constraint,” in Proc. 2nd IEEE Conf. Data Mining, 2002, pp. 418–425.

[4] Z. Xu, and R. Akella, “Active relevance feedback for difficult queries,” in Proc. ACM Conf. Inf. Knowl. Manage., 2008, pp. 459–468.

[5] S. Zhu, X. Ji, W. Xu, and Y. Gong, “Multi-labelled classification using maximum entropy method,” in Proc. Annu. Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval, 2005, pp. 1041–1048.

[6] Q. Song, J. Ni, and G. Wang, “A fast clustering-based feature subset selection algorithm for high-dimensional data,” in IEEE Trans. Knowl. Data Eng., vol. 25, no. 1, pp. 1–14, Jan. 2013.

[7] . Shehata, F. Karray, and M. Kamel, “A concept-based model for enhancing text categorization,” in Proc. ACM SIGKDD Knowl. Discovery Data Mining, 2007, pp. 629–637.

[8] M. Aghdam, N. Ghasem-Aghaee, and M. Basiri, “Text feature selection using ant colony optimization,” in Expert Syst. Appl., vol. 36, pp. 6843–6853, 2009.

[9] Algarni and Y. Li, “Mining specific features for acquiring user information needs,” in Proc. Pacific Asia Knowl. Discovery Data Mining, 2013, pp. 532–543.

[10] Algarni, Y. Li, and Y. Xu, “Selected new training documents to update user profile,” in Proc. Int. Conf. Inf. Knowl. Manage., 2010, pp. 799–808.

[11] . Azam and J. Yao, “Comparison of term frequency and document frequency based feature selection metrics in text categorization,” Expert Syst. Appl., vol. 39, no. 5, pp. 4760–4768, 2012.

[12] R. Bekkerman and M. Gavish, “High-precision phrase-based document classification on a modern scale,” in Proc. 11th ACM SIGKDD Knowl. Discovery Data Mining, 2011, pp. 231–239.

[13] Blum and P. Langley, “Selection of relevant features and examples in machine learning,” Artif. Intell., vol. 97, nos. 1/2, pp. 245– 271, 1997.

[14] Buckley, G. Salton, and J. Allan, “The effect of adding relevance information in a relevance feedback environment,” in Proc. Annu. Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval, 1994, pp. 292–300.

[15] G. Cao, J.-Y. Nie, J. Gao, and S. Robertson, “Selecting good expansion terms for pseudo-relevance feedback,” in Proc. Annu. Int. ACM SIGIR Conf. Res. Develop. Inf. Retrieval, 2008, pp. 243–250.

[16] G. Chandrashekar and F. Sahin, “Asurvey on feature selection methods,” in Comput. Electr. Eng., vol. 40, pp. 16–28, 2014.

[17] Croft, D. Metzler, and T. Strohman, Search Engines: Information Retrieval in Practice. Reading, MA, USA: Addison-Wesley, 2009.

[18] F. Debole and F. Sebastiani, “An analysis of the relative hardness of Reuters-21578 subsets,” J. Amer. Soc. Inf. Sci. Technol., vol. 56, no. 6, pp. 584–596, 2005.

Downloads

Published

2015-11-30

How to Cite

Relevance Feature Discovery for Text Mining using Dynamic Approach. (2015). International Journal of Advanced Research in Science, Management and Technology, 1(5), 1-6. https://ijarsmt.in/ijarsmt/article/view/14

Similar Articles

31-40 of 99

You may also start an advanced similarity search for this article.