Adaptive document clustering based on query-based similarity [An article from: Information Processing and Management]

Name: Adaptive document clustering based on query-based similarity [An article from: Information Processing and Management]
Author: S.H. Na, I.S. Kang, J.H. Lee
ISBN: 978B000PDYKC2

Author S.H. Na, I.S. Kang, J.H. Lee

Publisher Elsevier

Shop on Amazon — pick your country

🇺🇸 USA 🇨🇦 Canada 🇬🇧 UK 🇩🇪 Germany 🇫🇷 France 🇮🇳 India

7.95 USD

Buy New on Amazon 🇺🇸

Available for download now

Book Details

Author(s) S.H. Na, I.S. Kang, J.H. Lee

Publisher Elsevier

ISBN / ASIN B000PDYKCG

ISBN-13 978B000PDYKC2

Availability Available for download now

Sales Rank #99,999,999

Marketplace United States 🇺🇸

Description

This digital document is a journal article from Information Processing and Management, published by Elsevier in 2007. The article is delivered in HTML format and is available in your Amazon.com Media Library immediately after purchase. You can view it with any web browser.

Description:
In information retrieval, cluster-based retrieval is a well-known attempt in resolving the problem of term mismatch. Clustering requires similarity information between the documents, which is difficult to calculate at a feasible time. The adaptive document clustering scheme has been investigated by researchers to resolve this problem. However, its theoretical viewpoint has not been fully discovered. In this regard, we provide a conceptual viewpoint of the adaptive document clustering based on query-based similarities, by regarding the user's query as a concept. As a result, adaptive document clustering scheme can be viewed as an approximation of this similarity. Based on this idea, we derive three new query-based similarity measures in language modeling framework, and evaluate them in the context of cluster-based retrieval, comparing with K-means clustering and full document expansion. Evaluation result shows that retrievals based on query-based similarities significantly improve the baseline, while being comparable to other methods. This implies that the newly developed query-based similarities become feasible criterions for adaptive document clustering.