Patent Document Similarity Based on Image Analysis Using the SIFT-Algorithm and OCR-Text

박정범; Thomas Mandl; 김도완

doi:10.5392/IJoC.2017.13.4.070

ACOMS+ 및 학술지 리포지터리 설명회

한국과학기술정보연구원(KISTI) 서울분원 대회의실(별관 3층)
2024년 07월 03일(수) 13:30

사전등록 바로가기

오늘 하루 그만보기

권한신청
P-ISSN1738-6764
E-ISSN2093-7504
KCI

홈으로

OA 정책

ISSN : 1738-6764

논문 상세

이전 다음

논문 투고

Vol.13 No.4

Citation Share

Patent Document Similarity Based on Image Analysis Using the SIFT-Algorithm and OCR-Text

INTERNATIONAL JOURNAL OF CONTENTS / INTERNATIONAL JOURNAL OF CONTENTS, (P)1738-6764; (E)2093-7504

2017, v.13 no.4, pp.70-79

https://doi.org/10.5392/IJoC.2017.13.4.070

박정범 (배재대학교)
Thomas Mandl (University of Hildesheim)
김도완 (배재대학교)

박정범, Thomas, M. , & 김도완. (2017). . INTERNATIONAL JOURNAL OF CONTENTS, 13(4), 70-79, https://doi.org/10.5392/IJoC.2017.13.4.070

복사

Abstract

Images are an important element in patents and many experts use images to analyze a patent or to check differences between patents. However, there is little research on image analysis for patents partly because image processing is an advanced technology and typically patent images consist of visual parts as well as of text and numbers. This study suggests two methods for using image processing; the Scale Invariant Feature Transform(SIFT) algorithm and Optical Character Recognition(OCR). The first method which works with SIFT uses image feature points. Through feature matching, it can be applied to calculate the similarity between documents containing these images. And in the second method, OCR is used to extract text from the images. By using numbers which are extracted from an image, it is possible to extract the corresponding related text within the text passages. Subsequently, document similarity can be calculated based on the extracted text. Through comparing the suggested methods and an existing method based only on text for calculating the similarity, the feasibility is achieved. Additionally, the correlation between both the similarity measures is low which shows that they capture different aspects of the patent content.

keywords: Patent Similarity, Image Processing, Information Retrieval, Correlation Coefficient, SIFT, OpenIMAJ, OCR, Tess4j.

바로가기메뉴

논문 상세

Vol.13 No.4

Patent Document Similarity Based on Image Analysis Using the SIFT-Algorithm and OCR-Text

Abstract

INTERNATIONAL JOURNAL OF CONTENTS