Measuring Similarity Between Text Documents for Information Retrieval |
Author(s): |
| Geetanjali Gupta , Jawaharlal Institute of Technology Borawan Khargone (M.P.) India 451228; Mr. Kapil Shah , Jawaharlal Institute of Technology Borawan Khargone (M.P.) India 451228 |
Keywords: |
| Document, Text, Similarity, Preprocessing, terms |
Abstract |
|
There are several parameters by which similarity can be evaluated. The first categories of similarity evaluation is based on the document size and structure the length of the document, the number of paragraphs, number of sentences, average number of characters per word, average number of words per sentence etc. The second category is based on “styleâ€, whether the contents have been written in the first person conversational style or in the third person and so on. Thirdly, similarity can be based on the set of words used in the document. The fourth category of similarity is “content Similarity†which reflects to what extent the contents of the two documents are alike. This category is adopted throughout this thesis wherever similarity is talked of hereafter. The similarity between two documents is computed by any one of the several similarity measures based on the two corresponding feature vectors, e.g. cosine, dice, and jacquard measure. In this paper we measure similarity between texts documents using terms and token. A Document represented in a 3-Dimensional term vector space. There are several similarity coefficient are used to compare similarity. We used Euclidean distance to check the similarity. |
Other Details |
|
Paper ID: IJSRDV8I90131 Published in: Volume : 8, Issue : 9 Publication Date: 01/12/2020 Page(s): 337-341 |
Article Preview |
|
|
|
|
