Scanning World Wide Web documents with the vector space model

2006 ◽  
Vol 42 (2) ◽  
pp. 690-699 ◽  
Author(s):  
Cheryl Aasheim ◽  
Gary J. Koehler
Author(s):  
CHOOCHART HARUECHAIYASAK ◽  
MEI-LING SHYU ◽  
SHU-CHING CHEN

Due to the explosive growth of available information on the World Wide Web (WWW), users have suffered from the information overload. To alleviate this problem, there is a need for an intelligent tool to help the users screening and filtering for interesting and useful information. In this paper, a method of automatically identifying topics for Web documents via a classification technique is proposed. Topic identification can be applied as a filtering tool for recommender systems to prune down the number of documents to within some particular topics. We adopt the fuzzy association concept as a machine learning technique to classify the documents into some predefined categories or topics. Our approach is compared to the vector space model with the cosine coefficient using the data sets collected from three different Web portals: Yahoo!, Open Directory Project and Excite. The results show that our approach yields higher classification accuracy compared to the vector space model.


Author(s):  
Anthony Anggrawan ◽  
Azhari

Information searching based on users’ query, which is hopefully able to find the documents based on users’ need, is known as Information Retrieval. This research uses Vector Space Model method in determining the similarity percentage of each student’s assignment. This research uses PHP programming and MySQL database. The finding is represented by ranking the similarity of document with query, with mean average precision value of 0,874. It shows how accurate the application with the examination done by the experts, which is gained from the evaluation with 5 queries that is compared to 25 samples of documents. If the number of counted assignments has higher similarity, thus the process of similarity counting needs more time, it depends on the assignment’s number which is submitted.


2018 ◽  
Vol 9 (2) ◽  
pp. 97-105
Author(s):  
Richard Firdaus Oeyliawan ◽  
Dennis Gunawan

Library is one of the facilities which provides information, knowledge resource, and acts as an academic helper for readers to get the information. The huge number of books which library has, usually make readers find the books with difficulty. Universitas Multimedia Nusantara uses the Senayan Library Management System (SLiMS) as the library catalogue. SLiMS has many features which help readers, but there is still no recommendation feature to help the readers finding the books which are relevant to the specific book that readers choose. The application has been developed using Vector Space Model to represent the document in vector model. The recommendation in this application is based on the similarity of the books description. Based on the testing phase using one-language sample of the relevant books, the F-Measure value gained is 55% using 0.1 as cosine similarity threshold. The books description and variety of languages affect the F-Measure value gained. Index Terms—Book Recommendation, Porter Stemmer, SLiMS Universitas Multimedia Nusantara, TF-IDF, Vector Space Model


1985 ◽  
Vol 8 (2) ◽  
pp. 253-267
Author(s):  
S.K.M. Wong ◽  
Wojciech Ziarko

In information retrieval, it is common to model index terms and documents as vectors in a suitably defined vector space. The main difficulty with this approach is that the explicit representation of term vectors is not known a priori. For this reason, the vector space model adopted by Salton for the SMART system treats the terms as a set of orthogonal vectors. In such a model it is often necessary to adopt a separate, corrective procedure to take into account the correlations between terms. In this paper, we propose a systematic method (the generalized vector space model) to compute term correlations directly from automatic indexing scheme. We also demonstrate how such correlations can be included with minimal modification in the existing vector based information retrieval systems.


Sign in / Sign up

Export Citation Format

Share Document