Combining TF-IDF Features for Classifying Amazon Product Review in Positive and Negative Classes |
Author(s): |
| Anita More , Sri Aurobindo Institute of Technology Indore; Priyanshu Jadon, Sri Aurobindo Institute of Technology Indore; Dr. Durgesh Kumar Mishra, Sri Aurobindo Institute of Technology Indore |
Keywords: |
| Amazon Product Review, NLP, Supervised Learning, Feature Selection, Text Classification |
Abstract |
|
The text mining and their techniques are useful in various real world applications such as medical information retrieval, analysis of text for identifying frauds and managing customers in business intelligence applications. In these applications the data mining and natural language process help in processing the data and obtaining the required patterns. In this presented work text mining techniques are used for classifying the Amazon product reviews. This analysis may helpful for buyers to make their purchasing decision; additionally it is helpful for providing the feedback of product to the product manufacturer. In this context the Amazon product review dataset from Kaggle is used. First the data set is preprocessed to filter the stop words and special characters form the text data. In next step two different feature selection techniques has been implemented first the Natural Language Processing (NLP) parser based Part of Speech Tagging (POS) Tagging technique is used for feature extraction and then the Term Frequency-Inverted Document Frequency (TF-IDF) based feature selection technique is used to select potential keywords from the reviews. Further both the features are combined and then the data features (combined) are used for creating training and testing datasets. The training dataset is further used with the Support vector machine (SVM) for training and after training the created test dataset is used for validation of the performance. The experiments are carried out with the different size of samples in increasing size. Additionally the performance in terms of precision, recall, F1-score, time and memory has been measured. The performance of the model demonstrates higher accuracy and less resource consumption. Finally based on the experimental study some future extension of the work is also proposed. |
Other Details |
|
Paper ID: IJSRDV9I30230 Published in: Volume : 9, Issue : 3 Publication Date: 01/06/2021 Page(s): 250-255 |
Article Preview |
|
|
|
|
