An Analytical Study on Automatic Classification of Domestic Journal articles Using Random Forest

대표적인 앙상블 기법으로서 랜덤포레스트(RF)를 문헌정보학 분야의 학술지 논문에 대한 자동분류에 적용하였다. 특히, 국내 학술지 논문에 주제 범주를 자동 할당하는 분류 성능 측면에서 트리 수, 자질선정, 학습집합 크기 등 주요 요소들에 대한 다각적인 실험을 수행하였다. 이를 통해, 실제 환경의 불균형 데이터세트(imbalanced dataset)에 대하여 랜덤포레스트(RF)의 성능을 최적화할 수 있는 방안을 모색하였다. 결과적으로 국내 학술지 논문의 자동분류에서 랜덤포레스트(RF)는 트리 수 구간 100〜1000(C)과 카이제곱통계량(CHI)으로 선정한 소규모의 자질집합(10%), 대부분의 학습집합(9〜10년)을 사용하는 경우에 가장 좋은 분류 성능을 기대할 수 있는 것으로 나타났다.

keywords: 자동분류, 자동주석, 디지털 큐레이션, 학술지 논문, 랜덤포레스트(RF), 복수-범주 분류, 불균형 데이터, 자질선정, automatic classification, automatic annotation, digital curation, journal articles, random forest (RF), multi-label classification, imbalanced data, feature selection

Abstract

Random Forest (RF), a representative ensemble technique, was applied to automatic classification of journal articles in the field of library and information science. Especially, I performed various experiments on the main factors such as tree number, feature selection, and learning set size in terms of classification performance that automatically assigns class labels to domestic journals. Through this, I explored ways to optimize the performance of random forests (RF) for imbalanced datasets in real environments. Consequently, for the automatic classification of domestic journal articles, Random Forest (RF) can be expected to have the best classification performance when using tree number interval 100〜1000(C), small feature set (10%) based on chi-square statistic (CHI), and most learning sets (9-10 years).

keywords: 자동분류, 자동주석, 디지털 큐레이션, 학술지 논문, 랜덤포레스트(RF), 복수-범주 분류, 불균형 데이터, 자질선정, automatic classification, automatic annotation, digital curation, journal articles, random forest (RF), multi-label classification, imbalanced data, feature selection

바로가기메뉴

논문 상세

Vol.36 No.2

랜덤포레스트를 이용한 국내 학술지 논문의 자동분류에 관한 연구

An Analytical Study on Automatic Classification of Domestic Journal articles Using Random Forest

초록

Abstract

정보관리학회지