정보관리학회지, 한국정보관리학회

11

김희영(연세대학교 일반대학원 문헌정보학과) ; 박지홍(연세대학교 문헌정보학과) 2022, Vol.39, No.1, pp.1-15 https://doi.org/10.3743/KOSIM.2022.39.1.001

초록보기

초록

본 연구는 약물 연구 분야에 속하는 특허 사이에 나타나는 지식의 흐름을 살펴보고 이들 간의 영향력을 파악해보기 위해 특허데이터에서 나타나는 인용 관계를 분석하였다. 특허데이터의 수집은 Google Patents에서 진행하였다. 약물 연구와 관련된 특허 문서를 검색하여 상위 25개의 출원인을 선정하였고, 이를 바탕으로 출원인 사이에서의 인용 관계를 알아보고 각 출원인의 각 문서에 대한 피인용빈도와 순위를 활용하여 h-지수와 h-지수의 파생지표들의 값을 계산하여 비교하였다. 분석 결과를 종합하면, ‘Pfizer, MIT, Abbott’ 등의 출원인이 약물 연구 분야에서 영향력이 높은 출원인으로 드러났다. 5개의 계량서지학적 지표 중에서 g-지수와 hS-지수가 서로 유사한 결과를 보여주었고, 총인용빈도, 최대인용빈도, CPP의 순위를 가장 잘 반영하는 지표로 나타났다. 또한, 총인용빈도, CPP, 최대인용빈도 순으로 5개의 계량서지학적 지표와의 상관관계가 높았다. 한편, 기존의 특허 출원인의 기술적 영향력을 나타내는 것으로 알려진 지표인 CPP만으로는 정확한 비교가 어려운 경우도 나타났다.

Abstract

This study analyzes the relationship of citations appearing in the patent data to understand knowledge transfers and impacts between patent documents in the field of pharmaceutical research. Patent data were collected from a website, Google Patents. The top 25 assignees were selected by searching for patent documents related to pharmaceutical research. We identify the citation relationships between assignees, then calculate and compare the values of h-index and derived indicators by using the number of citations and rank for each document of each assignee. As a result, in the case of pharmaceutical research, the assignee, such as ‘Pfizer, MIT, and Abbott’ shows a high impact. Among the five bibliometric indicators, the g-index and hS-index show similar results, and the indicators are the most related to the rankings of Total Citation Frequency, Cites per Patents, and Maximum Citation Frequency. In addition, it is highly related to the five indicators in the order of Total Citation Frequency, Cites per Patents, and Maximum Citation Frequency. In some cases, it is difficult to make an accurate comparison with Cites per Patents alone, which is previously known to indicate the technological influence of patent assignees.

12

리뷰 정보를 활용한 이용자의 선호요인 식별에 관한 연구

송성전(독립연구자) ; 심지영(연세대학교 대학도서관발전연구소) 2022, Vol.39, No.3, pp.311-336 https://doi.org/10.3743/KOSIM.2022.39.3.311

초록보기

초록

본 연구는 도서관 정보서비스 환경에서 도서 이용자의 도서추천에 영향을 미치는 선호요인을 파악하기 위해 전 세계 도서 이용자의 참여로 이루어지는 사회적 목록 서비스인 Goodreads 리뷰 데이터를 대상으로 내용분석하였다. 이용자 선호의 내용을 보다 세부적인 관점에서 파악하기 위해 샘플 선정 과정에서 평점 그룹별, 도서별, 이용자별 하위 데이터 집합을 구성하였으며, 다양한 토픽을 고루 반영하기 위해 리뷰 텍스트의 토픽모델링 결과에 기반하여 층화 샘플링을 수행하였다. 그 결과, ‘내용’, ‘캐릭터’, ‘글쓰기’, ‘읽기’, ‘작가’, ‘스토리’, ‘형식’의 7개 범주에 속하는 총 90개 선호요인 관련 개념을 식별하는 한편, 평점에 따라 드러나는 일반적인 선호요인은 물론 호불호가 분명한 도서와 이용자에서 드러나는 선호요인의 양상을 파악하였다. 본 연구의 결과는 이용자 선호요인의 구체적 양상을 파악하여 향후 추천시스템 등에서 보다 정교한 추천에 기여할 수 있을 것으로 보인다.

Abstract

This study analyzed the contents of Goodreads review data, which is a social cataloging service with the participation of book users around the world, to identify the preference factors that affect book users’ book recommendations in the library information service environment. To understand user preferences from a more detailed point of view, sub-datasets for each rating group, each book, and each user were constructed in the sample selection process. Stratified sampling was also performed based on the result of topic modeling of review text data to include various topics. As a result, a total of 90 preference factors belonging to 7 categories(‘Content’, ‘Character’, ‘Writing’, ‘Reading’, ‘Author’, ‘Story’, ‘Form’) were identified. Also, the general preference factors revealed according to the ratings, as well as the patterns of preference factors revealed in books and users with clear likes and dislikes were identified. The results of this study are expected to contribute to more sophisticated recommendations in future recommendation systems by identifying specific aspects of user preference factors.

13

텍스트마이닝을 활용한 “잊힐 권리”의 토픽 분석

이소현(부산대학교 도서관) ; 구본진(부산대학교) 2022, Vol.39, No.2, pp.275-298 https://doi.org/10.3743/KOSIM.2022.39.2.275

초록보기

초록

본 연구는 잊힐 권리와 관련한 뉴스 기사와 학술지 게재 논문을 대상으로 텍스트마이닝 분석을 활용해 각 문서 내에 나타난 논점과 특성을 살펴보았다. 분석을 위해 ‘잊힐 권리’와 ‘잊혀질 권리’ 키워드를 검색어로 하여 2010년부터 2020년까지의 데이터를 수집하였다. 수집된 데이터를 대상으로 키워드 분석과 토픽모델링 분석을 수행한 결과, 지난 10년간 뉴스 기사와 학술지 논문에서 다루어진 쟁점은 크게 다르지 않으며, 접근 방법 또한 유사한 것으로 나타났다. 다만 뉴스 기사와 학술지 논문 간 비교를 통해 이들 간 공통적으로 나타나는 쟁점과 부분적인 쟁점의 차이가 있음을 확인하였다. 따라서 본 연구에서 도출된 쟁점을 중심으로 기록관리학 분야에서도 적극적인 논의가 이루어져야 할 필요가 있으며, 공통적인 쟁점들을 우선적으로 고려하되, 쟁점 상 이견이 존재하는 경우, 이를 다각적으로 논의하는 것이 필요하다고 볼 수 있다. 본 연구는 국내 기록관리학계에서 잊힐 권리와 관련된 논의가 이루어지고 있지 않은 현재의 상황에서 기록관리학 분야에서 잊힐 권리의 의미와 향후 발생할 수 있는 이슈를 도출해볼 수 있었다는데 의의가 있으며, 본 연구의 결과를 중심으로 기록관리학 분야에서 잊힐 권리에 대한 다양한 논의가 이루어지기를 기대한다.

Abstract

This study examined the issues and characteristics that appeared in news and journal articles related to the ‘right to be forgotten’ using text mining analysis. Data for analysis were collected from 2010 to 2020 with the keyword ‘right to be forgotten’. Keyword analysis and topic modeling analysis were performed on the collected data. As a result, in the last 10 years the issues about ‘right to be forgotten’ are not much different in news and journal articles and the approaches also are similar. However, it confirmed common issues and the partial difference between news and journal articles through comparison. Therefore in Archives and Records Management Studies, it is necessary to discuss derived in this study. In particular common issues are considered first but if there are differences in issues, it is needed to discuss them in various ways. This study is meaningful to understand the meaning and to draw issues that may arise in the future of the ‘right to be forgotten’. The results of this study will contribute to be variously discussed on the ‘right to be forgotten’ in Archives and Records Management Studies.

14

디지털 큐레이션 성숙도 모델 및 지표 개발에 관한 연구: 한국과학기술정보연구원 디지털큐레이션센터를 중심으로

김성훈(성균관대학교) ; 도슬기(성균관대학교 문헌정보학과) ; 한상은(카이스트 디지털인문사회과학센터) ; 김재훈(한국과학기술정보연구원) ; 임석종(한국과학기술정보연구원) ; 박진호(한성대학교) 2022, Vol.39, No.4, pp.269-306 https://doi.org/10.3743/KOSIM.2022.39.4.269

초록보기

초록

본 연구는 성숙도 모델 개념을 활용하여 디지털 전환 성과를 측정할 수 있는 지표 개발을 시도하였다. 디지털 전환을 위해서는 단순한 서비스 개선이 아니라 조직, 업무 변화까지를 고려할 필요가 있다. 여기서는 우리나라의 대표적인 과학기술정보서비스 기관인 KISTI의 디지털 전환 측정을 위한 모델 개발을 목표로 하였다. KSITI는 이미 디지털 전환을 위한 BPR 작업을 수행한 바 있으며, 성숙도 모델 개념을 차용하였다. 단, BPR에서는 해당 결과를 측정할 수 있는 방법은 존재하지 않는다. 본 논문에서는 성숙모 모델을 기반으로 디지털 전환을 측정할 수 있는 지표를 개발하였다. 지표개발은 모델 개발과 평가 두 가지 방법으로 수행하였다. 모델 구성을 위한 사례는 기존 KISTI에서 수행한 관련 연구, 다양한 국내․외 사례를 통해 이루어졌다. 검증 전 모델은 대분류를 기준으로 기술(37개), 데이터(45개), 전략(18개), 조직(인력)(36개), (사회적)영향력(14개)이었다. 검증 후에 최종 모델은 기술(20개/17개 지표 탈락), 데이터(36개/9개 지표 탈락), 전략(18개/유지), 조직(인력)(30개/6개 지표 탈락), (사회적)영향력(13개/1개 지표 탈락)으로 구성되었다.

Abstract

This study aimed to develop indicators that can measure the digital transformation performance of science and technology information construction and sharing systems by utilizing the Digital Curation Maturity Models. For digital transformation, it is necessary to consider not only simple service improvement but also organizational and business changes. In this study, we aimed to develop a model for measuring the digital transformation of KISTI, Korea’s representative science and technology information service organization. KISTI has already carried out BPR work for digital transformation and borrowed the concept of a maturity model. However, in BPR, there is no method to measure the result. Therefore, in this paper, we developed an index to measure digital transformation based on the maturity model. Indicator development was carried out in two ways: model development and evaluation. Cases for model construction were made through a comprehensive review of existing KISTI and various domestic and foreign cases. The models before verification were technology (37), data (45), strategy (18), organization (36), and (social)influence (14) based on the major categories. After verification using confirmatory factor analysis, the model is classified as technology (20 / 17 indicators dropped), data (36 / 9 indicators dropped), strategy (18 / maintenance), organization(30 / 6 indicators dropped), and (social) influence (13 indicators / 1 indicator dropped).

15

BERTopic을 활용한 불면증 소셜 데이터 토픽 모델링 및 불면증 경향 문헌 딥러닝 자동분류 모델 구축

고영수(연세대학교 문헌정보학과 석사과정) ; 이수빈(연세대학교 문헌정보학과 박사과정) ; 차민정(연세대학교 소셜오믹스 연구센터) ; 김성덕(연세대학교 문헌정보학과 석사과정) ; 이주희(연세대학교 문헌정보학과 석사과정) ; 한지영(연세대학교 문헌정보학과 석사과정) ; 송민(연세대학교 문헌정보학과) 2022, Vol.39, No.2, pp.111-129 https://doi.org/10.3743/KOSIM.2022.39.2.111

초록보기

초록

불면증은 최근 5년 새 환자가 20% 이상 증가하고 있는 현대 사회의 만성적인 질병이다. 수면이 부족할 경우 나타나는 개인 및 사회적 문제가 심각하고 불면증의 유발 요인이 복합적으로 작용하고 있어서 진단 및 치료가 중요한 질환이다. 본 연구는 자유롭게 의견을 표출하는 소셜 미디어 ‘Reddit’의 불면증 커뮤니티인 ‘insomnia’를 대상으로 5,699개의 데이터를 수집하였고 이를 국제수면장애분류 ICSD-3 기준과 정신의학과 전문의의 자문을 받은 가이드라인을 바탕으로 불면증 경향 문헌과 비경향 문헌으로 태깅하여 불면증 말뭉치를 구축하였다. 구축된 불면증 말뭉치를 학습데이터로 하여 5개의 딥러닝 언어모델(BERT, RoBERTa, ALBERT, ELECTRA, XLNet)을 훈련시켰고 성능 평가 결과 RoBERTa가 정확도, 정밀도, 재현율, F1점수에서 가장 높은 성능을 보였다. 불면증 소셜 데이터를 심층적으로 분석하기 위해 기존에 많이 사용되었던 LDA의 약점을 보완하며 새롭게 등장한 BERTopic 방법을 사용하여 토픽 모델링을 진행하였다. 계층적 클러스터링 분석 결과 8개의 주제군(‘부정적 감정’, ‘조언 및 도움과 감사’, ‘불면증 관련 질병’, ‘수면제’, ‘운동 및 식습관’, ‘신체적 특징’, ‘활동적 특징’, ‘환경적 특징’)을 확인할 수 있었다. 이용자들은 불면증 커뮤니티에서 부정 감정을 표현하고 도움과 조언을 구하는 모습을 보였다. 또한, 불면증과 관련된 질병들을 언급하고 수면제 사용에 대한 담론을 나누며 운동 및 식습관에 관한 관심을 표현하고 있었다. 발견된 불면증 관련 특징으로는 호흡, 임신, 심장 등의 신체적 특징과 좀비, 수면 경련, 그로기상태 등의 활동적 특징, 햇빛, 담요, 온도, 낮잠 등의 환경적 특징이 확인되었다.

Abstract

Insomnia is a chronic disease in modern society, with the number of new patients increasing by more than 20% in the last 5 years. Insomnia is a serious disease that requires diagnosis and treatment because the individual and social problems that occur when there is a lack of sleep are serious and the triggers of insomnia are complex. This study collected 5,699 data from ‘insomnia’, a community on ‘Reddit’, a social media that freely expresses opinions. Based on the International Classification of Sleep Disorders ICSD-3 standard and the guidelines with the help of experts, the insomnia corpus was constructed by tagging them as insomnia tendency documents and non-insomnia tendency documents. Five deep learning language models (BERT, RoBERTa, ALBERT, ELECTRA, XLNet) were trained using the constructed insomnia corpus as training data. As a result of performance evaluation, RoBERTa showed the highest performance with an accuracy of 81.33%. In order to in-depth analysis of insomnia social data, topic modeling was performed using the newly emerged BERTopic method by supplementing the weaknesses of LDA, which is widely used in the past. As a result of the analysis, 8 subject groups (‘Negative emotions’, ‘Advice and help and gratitude’, ‘Insomnia-related diseases’, ‘Sleeping pills’, ‘Exercise and eating habits’, ‘Physical characteristics’, ‘Activity characteristics’, ‘Environmental characteristics’) could be confirmed. Users expressed negative emotions and sought help and advice from the Reddit insomnia community. In addition, they mentioned diseases related to insomnia, shared discourse on the use of sleeping pills, and expressed interest in exercise and eating habits. As insomnia-related characteristics, we found physical characteristics such as breathing, pregnancy, and heart, active characteristics such as zombies, hypnic jerk, and groggy, and environmental characteristics such as sunlight, blankets, temperature, and naps.

16

도서추천 시스템 개선을 위한 도서이용 맥락 요소 탐색

심지영(연세대학교 대학도서관발전연구소) 2022, Vol.39, No.2, pp.299-324 https://doi.org/10.3743/KOSIM.2022.39.2.299

초록보기

초록

본 연구는 기존의 도서추천 시스템 연구에서 간과되어 온 도서이용의 맥락 요소를 파악하기 위해, 다양한 도서탐색 배경을 지닌 적극적인 도서 이용자 15명을 대상으로 6가지 도서탐색 상황에서 생성하는 내용을 사고구술(think-aloud) 프로토콜을 통해 수집하였다. 수집된 도서이용 내용은 내용분석 과정을 통해 독자자문 서비스의 이론적 개념인 ‘어필 요소(appeal factor)’를 토대로 도서이용에 영향을 미치는 내부 어필 요소와 외부 어필 요소를 각각 식별하였으며, 도서탐색에 사용하는 정보원과 탐색방법 관련 개념들을 또한 세분화하였다. 본 연구의 결과는 향후 도서추천 시스템 설계에 의미 있는 속성 데이터를 추출하고 반영하는 데 사용될 수 있을 것이다.

Abstract

In this study, in order to explore the contextual elements of book use that were overlooked in the existing book recommender system research, for 15 avid readers with various book search backgrounds, the contents generated in 6 book search situations were collected through the think-aloud protocol. By using content analysis from the collected book use contents, not only the internal and external appeal factors affecting book use, based on the ‘appeal factor’, the theoretical concept of the readers’ advisory service, but also information sources and search methods regarding book use were identified and categorized. The results of this study can be used to extract and reflect meaningful attribute data in the future book recommender system design process.

17

교사들의 교육지침 인식 개선을 위한 생활지도 정보요구와 교육지침 변화에 대한 인식 간 관계 연구

김진명(연세대학교 교육대학원 사서교육전공) ; 이지연(연세대학교) 2022, Vol.39, No.2, pp.131-157 https://doi.org/10.3743/KOSIM.2022.39.2.131

초록보기

초록

본 연구는 교사들의 생활교육에 대한 정보요구와 교육청으로부터 배부되는 교육지침에 대한 교사 인식 간 관계를 밝혀 이를 기반으로 한 학교도서관 정보서비스 제안에 목적을 두었다. 이론적 배경을 바탕으로 연구 모형을 제작하고 설계를 진행하였으며, 예비연구를 통한 심층면담을 분석하여 연구에서 고려할 요소들을 추출하고 설문 문항을 개발하였다. 경기도는 4개 권역으로 나누어지며, 각 권역별 3개 학교에 설문을 실시한 후 최종적으로 217부의 데이터를 최종 분석에 활용하였다. 연구 결과, 학교도서관의 생활교육 정보요구에 대한 관심이 궁극적으로 교사들의 학생생활교육 실태 개선에 영향을 미친다는 사실을 알 수 있었다. 연구 결과를 바탕으로 학교도서관에서 제공되는 정보서비스를 향상시키는 방법을 제시하였으며, 특히 교사들의 생활지도 정보요구를 학교 내부에서 해결하도록 학교도서관 서비스를 제안했다는 점에서 연구의 의의가 있다.

Abstract

This study aims to present the relationship between the teachers’ information needs regarding guidance and the perception of educational guidelines issued by the Office of Education. The research design was conducted based on the reviewing theoretical background studies, and questionnaires were found in in-depth semi-structured interviews in a pilot study. Gyeonggi Province is divided into four regions, which is the target of the survey. Teachers of three schools in each region were surveyed, and eventually 217 copies of the survey were used for the final analysis. The result shows that the role of school libraries in caring for teachers’ information needs ultimately influences the improvement of teachers’ guidance. Based on this result, the study suggests ways to improve information services provided by school libraries. In particular, the study is meaningful in that it has presented a potential service plan that can be performed by school libraries to help address teachers’ information needs about guidance in schools.

18

문헌정보학 분야의 리터러시 연구 동향 분석

장수현(중앙대학교 문헌정보학과) ; 남영준(중앙대학교) 2022, Vol.39, No.3, pp.263-292 https://doi.org/10.3743/KOSIM.2022.39.3.263

초록보기

초록

본 연구는 문헌정보학 현장인 도서관에서 제공되는 서비스인 이용자 교육의 관련 개념인 리터러시가 각종 문헌정보학 연구 분야에서 어떠한 연구 주제를 다루는지 확인하는 것을 목적으로 한다. 이를 위해 WoS와 KCI 데이터베이스에서 문헌정보학 분야 리터러시 관련 논문을 수집하여 키워드 분석 및 토픽 모델링 분석 기법을 상호보완적으로 사용해 분석하였다. 분석 결과, WoS와 KCI의 문헌정보학 분야 리티러시 관련 연구 동향은 저자 키워드, 주요 주제 등에서 차이가 있는 것으로 나타났으며, 토픽 모델링을 통해 KCI의 리터러시 관련 연구를 3개의 토픽으로 분류하였다. 또한, 연구에서 확인한 국내 문헌정보학 분야 리터러시 연구 동향은 전체 리터러시 관련 연구 동향과 연구량 급증 시기, 핵심 다빈출 키워드 차이가 있음을 분석하였다. 특히, 전체 분야 리터러시 연구는 ‘리터러시’, ‘교육’, ‘미디어’, ‘디지털’ 등의 단어가 다수 도출되었지만 문헌정보학 분야의 리터러시 연구는 ‘정보활용능력’, ‘학교도서관’ 등의 키워드가 다수 등장하였다. 이를 바탕으로 향후 국내에서도 정보가 급증하는 오늘날의 정보화 환경에 맞춰 정보에 대한 평가적인 안목을 기를 수 있는 능력에 관한 연구가 필요하다는 결론을 도출하였다.

Abstract

The purpose of this study is to identify the topics of research related to the concepts of literacy in the field of Library and Information Science which is related to user education in libraries. Data were collected from the WoS and KCI databases, and complementary keyword analysis and topic modeling analysis techniques were used to identify topics of literature-related research articles in the field of Library and Information Science. Findings presented that there was a difference in keywords and topics between the two databases. Literacy-related topics identified from the KCI database were classified into three groups through topic modeling. Also, it was analyzed that there is a difference between the overall literacy-related research trend, the timing of the surge in research volume, and key frequent keywords in the Library and Information Science field confirmed in the study. In particular, in the study of literacy in all fields, a number of words such as ‘literacy’, ‘education’, ‘media’, and ‘digital’ were derived. However, in literature research in the field of Library and Information Science, keywords such as ‘information utilization ability’ and ‘school library’ appeared. Based on this, it was concluded that research on the ability to develop an evaluative eye for information is needed in line with today’s information environment, where information is rapidly increasing in Korea in the future.

19

북한이탈주민의 정보빈곤에 관한 연구: Chatman의 정보빈곤이론을 기반으로

민수진(성균관대학교 문헌정보학과) ; 이용정(성균관대학교) 2022, Vol.39, No.3, pp.241-261 https://doi.org/10.3743/KOSIM.2022.39.3.241

초록보기

초록

본 연구는 Chatman(1996)의 정보빈곤이론(Theory of Information Poverty)을 바탕으로 정보 빈곤이 북한이탈주민의 한국사회적응에 미치는 영향을 알아보고자 하였다. 연구를 위해 정보빈곤이론을 기반으로 정보빈곤의 개념을 은폐(Secrecy), 기만(Deception), 위험감수(Risk-taking), 상황적 관련성(Situational relevance)에 따른 정보 수용이라는 네 가지 변인으로 구성하였고, 선행연구 분석 결과를 바탕으로 한국사회적응을 사회적 적응과 심리적 적응으로 구분하였다. 또한 생명윤리위원회(IRB)의 승인을 거쳐 2021년 8월 4일부터 8월 30일까지 북한이탈주민 지원 단체 <우리온>을 통해 국내 입국 후 최소 1년이 경과한 민법상 성년인 만 19세 이상의 북한이탈주민을 대상으로 설문조사를 실시하였다. 수집된 100개의 유효한 데이터를 빈도 분석, 신뢰도 분석, 상관관계 분석, 다중회귀분석을 통해 분석한 결과, 정보빈곤은 북한이탈주민의 사회적 적응과 심리적 적응에 유의한 영향을 미치는 것으로 나타났다. 특히, “기만” 변수는 북한이탈주민의 사회적 적응과 심리적 적응에 유의한 부(-)의 영향을 미치는 것으로 나타났다. 본 연구는 북한이탈주민을 정보빈곤층으로 정의하고, 그들의 한국사회적응을 Chatman의 정보빈곤이론을 기반으로 설명하였다는 점에서 학문적 의의가 있다. 무엇보다도, 질적 연구를 수행한 선행연구들과 달리 변수의 조작화를 통해 양적 연구를 시도하였다는 점에서 의미가 있다.

Abstract

The present study aims to investigate the effects of information poverty on North Korean refugees’ social adaptation to South Korea based on Chatman’s Theory of Information Poverty (1996). Based on the Theory of Information Poverty, information poverty consists of four variables: Secrecy, Deception, Risk-taking, and information acceptance in response to situational relevance. And based on the previous studies, adaptation to South Korean life is divided into social adaptation and psychological adaptation. From August 4 to August 30, 2021, after approval by the IRB through the North Korean refugee support organization <Urion>, surveys were conducted with North Korean refugees who had lived in South Korea for at least one year and were aged 19 or older. The 100 collected valid data were analyzed using frequency analysis, reliability analysis, correlation analysis, and multiple linear regression analysis. Findings of the study indicated that information poverty had significant effects on North Korean refugees’ social and psychological adaptation. In particular, the “deception” variable had negative effects on social and psychological adaptation. The study has theoretical implications that it explains North Korean refugees’ adaptation to South Korea based on Theory of Information Poverty by defining them as information poor. Above all, it attempts a quantitative approach through operationalization of key concepts unlike previous studies that were conducted with qualitative approaches.

20

LDA와 BERTopic을 이용한 토픽모델링의 증강과 확장 기법 연구

김선욱(경북대학교 사회과학대학 문헌정보학과) ; 양기덕(영남고문헌아카이브센터) 2022, Vol.39, No.3, pp.99-132 https://doi.org/10.3743/KOSIM.2022.39.3.099

초록보기

초록

본 연구의 목적은 LDA 토픽모델링 결과와 BERTopic 토픽모델링 결과를 합성하는 방법론인 Augmented and Extended Topics(AET)를 제안하고, 이를 사용해 문헌정보학 분야의 연구주제를 분석하는 데 있다. AET의 실제 적용결과를 확인하기 위해 2001년 1월부터 2021년 10월까지의 Web of Science 내 문헌정보학 학술지 85종에 게재된 학술논문 서지 데이터 55,442건을 분석하였다. AET는 서로 다른 토픽모델링 결과의 관계를 WORD2VEC 기반 코사인 유사도 매트릭스로 구축하고, 매트릭스 내 의미적 관계가 유효한 범위 내에서 매트릭스 재정렬 및 분할 과정을 반복해 증강토픽(Augmented Topics, 이하 AT)을 추출한 뒤, 나머지 영역에서 코사인 유사도 평균값 순위와 BERTopic 토픽 규모 순위에 대한 조화평균을 통해 확장토픽(Extended Topics, 이하 ET)을 결정한다. 최적 표준으로 도출된 LDA 토픽모델링 결과와 AET 결과를 비교한 결과, AT는 LDA 토픽모델링 토픽을 한층 더 구체화하고 세분화하였으며 ET는 유효한 토픽을 발견하였다. AT(Augmented Topics)의 성능은 LDA 이상이었으며 ET(Extended Topics)는 일부 경우를 제외하고 대부분 LDA와 유사한 수준의 성능을 나타내었다.

Abstract

The purpose of this study is to propose AET (Augmented and Extended Topics), a novel method of synthesizing both LDA and BERTopic results, and to analyze the recently published LIS articles as an experimental approach. To achieve the purpose of this study, 55,442 abstracts from 85 LIS journals within the WoS database, which spans from January 2001 to October 2021, were analyzed. AET first constructs a WORD2VEC-based cosine similarity matrix between LDA and BERTopic results, extracts AT (Augmented Topics) by repeating the matrix reordering and segmentation procedures as long as their semantic relations are still valid, and finally determines ET (Extended Topics) by removing any LDA related residual subtopics from the matrix and ordering the rest of them by (BERTopic topic size rank, Inverse cosine similarity rank). AET, by comparing with the baseline LDA result, shows that AT has effectively concretized the original LDA topic model and ET has discovered new meaningful topics that LDA didn’t. When it comes to the qualitative performance evaluation, AT performs better than LDA while ET shows similar performances except in a few cases.

바로가기메뉴

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

정보관리학회지