TY - GEN
T1 - Near duplicate detection in an academic digital library
AU - Williams, Kyle
AU - Giles, C. Lee
N1 - Copyright:
Copyright 2014 Elsevier B.V., All rights reserved.
PY - 2013
Y1 - 2013
N2 - The detection and potential removal of duplicates is desirable for a number of reasons, such as to reduce the need for unnecessary storage and computation, and to provide users with uncluttered search results. This paper describes an investigation into the application of scalable simhash and shingle state of the art duplicate detection algorithms for detecting near duplicate documents in the CiteSeerX digital library. We empirically explored the duplicate detection methods and evaluated their performance and application to academic documents and identified good parameters for the algorithms. We also analyzed the types of near duplicates identified by each algorithm. The highest F-scores achieved were 0.91 and 0.99 for the simhash and shingle-based methods respectively. The shingle-based method also identified a larger variety of duplicate types than the simhash-based method.
AB - The detection and potential removal of duplicates is desirable for a number of reasons, such as to reduce the need for unnecessary storage and computation, and to provide users with uncluttered search results. This paper describes an investigation into the application of scalable simhash and shingle state of the art duplicate detection algorithms for detecting near duplicate documents in the CiteSeerX digital library. We empirically explored the duplicate detection methods and evaluated their performance and application to academic documents and identified good parameters for the algorithms. We also analyzed the types of near duplicates identified by each algorithm. The highest F-scores achieved were 0.91 and 0.99 for the simhash and shingle-based methods respectively. The shingle-based method also identified a larger variety of duplicate types than the simhash-based method.
UR - http://www.scopus.com/inward/record.url?scp=84887373179&partnerID=8YFLogxK
UR - http://www.scopus.com/inward/citedby.url?scp=84887373179&partnerID=8YFLogxK
U2 - 10.1145/2494266.2494312
DO - 10.1145/2494266.2494312
M3 - Conference contribution
AN - SCOPUS:84887373179
SN - 9781450317894
T3 - DocEng 2013 - Proceedings of the 2013 ACM Symposium on Document Engineering
SP - 91
EP - 94
BT - DocEng 2013 - Proceedings of the 2013 ACM Symposium on Document Engineering
PB - Association for Computing Machinery
T2 - 2013 ACM Symposium on Document Engineering, DocEng 2013
Y2 - 10 September 2013 through 13 September 2013
ER -