상세검색
최근 검색어 전체 삭제
다국어입력
즐겨찾기0
학술저널

토큰 크기 및 출현 빈도에 기반한 웹 페이지 유사도

Web Page Similarity based on Size and Frequency of Tokens

  • 21
110650.jpg

It is becoming hard to maintain web applications because of high complexity and duplication of web pages. However, most of research about code clone is focusing on code hunks. and their target is limited to a specific language. Thus, we propose GSIM, a language-independent statistical approach to detect similar pages based on scarcity and frequency of customized tokens. The tokens, which can be obtained from pages splitted by a set of given separators, are defined as atomic elements lor calculating similarity between two pages. In this paper, the domain definition for web applications and algorithms for collecting tokens, making matrics, calculating similarity are given. We also conducted experiments on open source codes for evaluation, with our GSIM tool. The results show the applicability of the proposed method and the effects of parameters such as threshold, toughness, length of tokens, on their quality and performance.

Abstract

1. 서론

2. 관련 연구

3. GSIM 정의

4. 실험 결과

5. 결론

참고문헌

저자소개

(0)

(0)

로딩중