토큰 크기 및 출현 빈도에 기반한 웹 페이지 유사도
Web Page Similarity based on Size and Frequency of Tokens
- 한국IT서비스학회
- 한국IT서비스학회지
- 한국IT서비스학회지 제11권 제4호
-
2012.12263 - 275 (13 pages)
- 21
It is becoming hard to maintain web applications because of high complexity and duplication of web pages. However, most of research about code clone is focusing on code hunks. and their target is limited to a specific language. Thus, we propose GSIM, a language-independent statistical approach to detect similar pages based on scarcity and frequency of customized tokens. The tokens, which can be obtained from pages splitted by a set of given separators, are defined as atomic elements lor calculating similarity between two pages. In this paper, the domain definition for web applications and algorithms for collecting tokens, making matrics, calculating similarity are given. We also conducted experiments on open source codes for evaluation, with our GSIM tool. The results show the applicability of the proposed method and the effects of parameters such as threshold, toughness, length of tokens, on their quality and performance.
Abstract
1. 서론
2. 관련 연구
3. GSIM 정의
4. 실험 결과
5. 결론
참고문헌
저자소개
(0)
(0)