정보관리학회지, 한국정보관리학회

1

최상희(대구가톨릭대학교) 2010, Vol.27, No.2, pp.241-252 https://doi.org/10.3743/KOSIM.2010.27.2.241

초록보기

초록

Abstract

Titles have been regarded as having effective clustering features, but they sometimes fail to represent the topic of a document and result in poorly generated document clusters. This study aims to improve the performance of document clustering with titles by suggesting titles in the citation bibliography as a clustering feature. Titles of original literature, titles in the citation bibliography, and an aggregation of both titles were adapted to measure the performance of clustering. Each feature was combined with three hierarchical clustering methods, within group average linkage, complete linkage, and Ward's method in the clustering experiment. The best practice case of this experiment was clustering document with features from both titles by within-groups average method.

2

서지데이터 요소 채기 우선순에서 표제지의 기능성 연구

남태우(중앙대학교) 2004, Vol.21, No.1, pp.55-92 https://doi.org/10.3743/KOSIM.2004.21.1.055

초록보기

초록

본고에서는 서지데이터요소의 채기 과정에서 거의 신성권을 보장받았던 표제지의 기능성을 연구하고자 하였다. 그래서 우선적으로 서지통정상에서의 표제지의 출현배경과 그들의 개념정립을 고찰하였으며 편목과정에서 어떻게 취급하였는지도 규명하였다. 그리고 하이퍼텍스트환경에서의 표제지에 대한 탈-서지적 과정도 분석하였다.

Abstract

The title page of a book is a reliable source, since it, together with its verso, usually contains all bibliographically significant data. Generally, the title page is a page at the beginning of a book giving its title and the names of the author and publisher. Prescribing a source of information from which data elements should be derived is a way of specifying how an entity can represent itself. In simpler times, when bibliographic entities were for the most part books published in Western countries, the choice of source was obviously the title page, the "face of the book".

3

단행본 서명의 단어 임베딩에 따른 자동분류의 성능 비교

이용구(경북대학교 문헌정보학과) 2023, Vol.40, No.4, pp.307-327 https://doi.org/10.3743/KOSIM.2023.40.4.307

초록보기

초록

이 연구는 짧은 텍스트인 서명에 단어 임베딩이 미치는 영향을 분석하기 위해 Word2vec, GloVe, fastText 모형을 이용하여 단행본 서명을 임베딩 벡터로 생성하고, 이를 분류자질로 활용하여 자동분류에 적용하였다. 분류기는 k-최근접 이웃(kNN) 알고리즘을 사용하였고 자동분류의 범주는 도서관에서 도서에 부여한 DDC 300대 강목을 기준으로 하였다. 서명에 대한 단어 임베딩을 적용한 자동분류 실험 결과, Word2vec와 fastText의 Skip-gram 모형이 TF-IDF 자질보다 kNN 분류기의 자동분류 성능에서 더 우수한 결과를 보였다. 세 모형의 다양한 하이퍼파라미터 최적화 실험에서는 fastText의 Skip-gram 모형이 전반적으로 우수한 성능을 나타냈다. 특히, 이 모형의 하이퍼파라미터로는 계층적 소프트맥스와 더 큰 임베딩 차원을 사용할수록 성능이 향상되었다. 성능 측면에서 fastText는 n-gram 방식을 사용하여 하부문자열 또는 하위단어에 대한 임베딩을 생성할 수 있어 재현율을 높이는 것으로 나타났다. 반면에 Word2vec의 Skip-gram 모형은 주로 낮은 차원(크기 300)과 작은 네거티브 샘플링 크기(3이나 5)에서 우수한 성능을 보였다.

Abstract

To analyze the impact of word embedding on book titles, this study utilized word embedding models (Word2vec, GloVe, fastText) to generate embedding vectors from book titles. These vectors were then used as classification features for automatic classification. The classifier utilized the k-nearest neighbors (kNN) algorithm, with the categories for automatic classification based on the DDC (Dewey Decimal Classification) main class 300 assigned by libraries to books. In the automatic classification experiment applying word embeddings to book titles, the Skip-gram architectures of Word2vec and fastText showed better results in the automatic classification performance of the kNN classifier compared to the TF-IDF features. In the optimization of various hyperparameters across the three models, the Skip-gram architecture of the fastText model demonstrated overall good performance. Specifically, better performance was observed when using hierarchical softmax and larger embedding dimensions as hyperparameters in this model. From a performance perspective, fastText can generate embeddings for substrings or subwords using the n-gram method, which has been shown to increase recall. The Skip-gram architecture of the Word2vec model generally showed good performance at low dimensions(size 300) and with small sizes of negative sampling (3 or 5).

4

저자역할용어사전 구축 및 저작군집화에 관한 연구

윤재혁(성균관대학교 일반대학원 문헌정보학과) ; 도슬기(성균관대학교 일반대학원 문헌정보학과) ; 오삼균(성균관대학교 문헌정보학과) 2020, Vol.37, No.2, pp.197-223 https://doi.org/10.3743/KOSIM.2020.37.2.197

초록보기

초록

본 연구는 통합서지용 한국문헌자동화목록(KORMARC)으로 작성된 서지레코드를 FRBR의 저작(Work) 단위로 군집화 하는 과정에서 나타난 이슈사항들을 분석하고, 이에 대한 해결방안을 고안하였다. 특히 기존의 연구에서는 대표저작자를 식별하고 처리하는 기준이 명확하게 드러나지 않거나 파생저작 레코드의 대표저작자를 선정하는 방법에 대한 논의가 충분히 이루어지지 않았다. 따라서 본 연구는 저작을 창작하는 데 기여한 사람이 다수일 때 대표저작자를 명확하게 식별하기 위한 방법을 고안하는 데 초점을 맞추었다. 이를 위해 책임표시사항(245) 필드의 책임표시 태그(▼d, ▼e)에서 추출한 역할용어를 토대로 표준화된 저자역할용어사전을 개발하여 대표저작자 판별에 활용하는 방안을 마련하였다. 또한 저자명의 유사도와 표제의 유사도를 각각 계산하여 유사도가 일정 수준 이상인 경우 동일한 저작으로 군집화 하는 방법을 채택하였다. 각각의 유사도를 계산하여 동일 저작을 판단하므로 공백, 관제처리, 괄호제거와 같은 데이터 정제 조건을 조정하여 6가지 패턴에 따른 군집화의 정확도를 비교하였고, 저자명과 표제의 유사도가 모두 80퍼센트 이상일 때의 정확도가 가장 높게 나타났다. 본 연구는 대표저작자 선정을 위한 역할용어사전 개발, 대표저작자와 표제의 유사도를 별도로 측정하여 저작군집화를 시도한 실험연구이며 후속 연구에서는 표제 간 유사도 측정의 정확도를 향상시키는 방안과 FRBR 1그룹의 다른 개체(표현형, 구현형, 개별자료) 수준으로 확대하여 활용하는 방안, 국내에서 사용하고 있는 다른 형태의 MARC 데이터에 적용하는 방안을 고안할 예정이다.

Abstract

The purpose of this study is to analyze the issues resulted from the process of grouping KORMARC records using FRBR WORK concept and to suggest a new method. The previous studies did not sufficiently address the criteria or processes for identifying representative authors of records and their derivatives. Therefore, our study focused on devising a method of identifying the representative author when there are multiple contributors in a work. The study developed a method of identifying representative authors using an author role dictionary constructed by extracting role-terms from the statement of responsibility field (245). We also designed another way to group records as a work by calculating similarity measures of authors and titles. The accuracy rate of WORK grouping was the highest when blank spaces, parentheses, and controling processes were removed from titles and the measured similarity rates of authors and titles were higher than 80 percent. This was an experiment study where we developed an author-role dictionary that can be utilized in selecting a representative author and measured the similarity rate of authors and titles in order to achieve effective WORK grouping of KORMARC records. The future study will attempt to devise a way to improve the similarity measure of titles, incorporate FRBR Group 1 entities such as expression, manifestation and item data into the algorithm, and a method of improving the algorithm by utilizing other forms of MARC data that are widely used in Korea.

5

복수 자질에 의한 지적 구조의 계량정보학적 분석연구: 국내 대학도서관 분야 연구논문을 대상으로

최상희(대구가톨릭대학교) 2011, Vol.28, No.2, pp.65-78 https://doi.org/10.3743/KOSIM.2011.28.2.065

초록보기

초록

Abstract

The purpose of this study is to identify topic areas of academic library research using two informetric methods; word clustering and Pathfinder network. For the data analysis, 139 articles published in major library and information science journals from 2005 to 2009 were collected from the Korean Science Citation Index database. The keywords that represent research topics were gathered from two sections: an abstract and titles in references. Results showed that reference titles usefully represent topics in detail, and combining abstracts and reference titles can produce an expanded topic map.

6

ChatGPT가 자동 생성한 더블린 코어 메타데이터의 품질 평가: 국내 도서를 대상으로

김선욱(경북대학교 사회과학대학 문헌정보학과) ; 이혜경(경북대학교 문헌정보학과) ; 이용구(경북대학교) 2023, Vol.40, No.2, pp.183-209 https://doi.org/10.3743/KOSIM.2023.40.2.183

초록보기

초록

이 연구의 목적은 ChatGPT가 도서의 표지, 표제지, 판권기 데이터를 활용하여 생성한 더블린코어의 품질 평가를 통하여 ChatGPT의 메타데이터의 생성 능력과 그 가능성을 확인하는 데 있다. 이를 위하여 90건의 도서의 표지, 표제지와 판권기 데이터를 수집하여 ChatGPT에 입력하고 더블린 코어를 생성하게 하였으며, 산출물에 대해 완전성과 정확성 척도로 성능을 파악하였다. 그 결과, 전체 데이터에 있어 완전성은 0.87, 정확성은 0.71로 준수한 수준이었다. 요소별로 성능을 보면 Title, Creator, Publisher, Date, Identifier, Right, Language 요소가 다른 요소에 비해 상대적으로 높은 성능을 보였다. Subject와 Description 요소는 완전성과 정확성에 대해 다소 낮은 성능을 보였으나, 이들 요소에서 ChatGPT의 장점으로 알려진 생성 능력을 확인할 수 있었다. 한편, DDC 주류인 사회과학과 기술과학 분야에서 Contributor 요소의 정확성이 다소 낮았는데, 이는 ChatGPT의 책임표시사항 추출 오류 및 데이터 자체에서 메타데이터 요소용 서지 기술 내용의 누락, ChatGPT가 지닌 영어 위주의 학습데이터 구성 등에 따른 것으로 판단하였다.

Abstract

The purpose of this study is to evaluate the Dublin Core metadata generated by ChatGPT using book covers, title pages, and colophons from a collection of books. To achieve this, we collected book covers, title pages, and colophons from 90 books and inputted them into ChatGPT to generate Dublin Core metadata. The performance was evaluated in terms of completeness and accuracy. The overall results showed a satisfactory level of completeness at 0.87 and accuracy at 0.71. Among the individual elements, Title, Creator, Publisher, Date, Identifier, Rights, and Language exhibited higher performance. Subject and Description elements showed relatively lower performance in terms of completeness and accuracy, but it confirmed the generation capability known as the inherent strength of ChatGPT. On the other hand, books in the sections of social sciences and technology of DDC showed slightly lower accuracy in the Contributor element. This was attributed to ChatGPT’s attribution extraction errors, omissions in the original bibliographic description contents for metadata, and the language composition of the training data used by ChatGPT.

7

종합목록DB를 이용한 국내 대학도서관 서양서 소장 실태 분석

이지원(대구가톨릭대학교) ; 이재윤(명지대학교) 2018, Vol.35, No.1, pp.205-229 https://doi.org/10.3743/KOSIM.2018.35.1.205

초록보기

초록

본 연구는 국내 대학도서관 서양서 장서 개발의 변화를 살펴보기 위해 2003년과 2013년에 출판된 서양서 소장 실태를 KERIS 종합목록을 통해 분석하였다. 이를 위해 새로운 장서 지표로 소장 h-지수, 장서 고유성 지수, 그리고 공통장서 확보율을 제안하고 기본 지표인 종수 및 책수, 그리고 종당 책수와 함께 활용하였다. 분석 결과 2003년에 비해서 2013년에 출판된 서양서의 전체 소장 종수는 16.1% 감소하고 소장 책수는 42.2% 감소하여 소장 책수가 더 크게 감소하였다. 여러 도서관이 공통적으로 소장하는 공통 장서, 또는 기본 장서의 규모를 나타내는 공통장서 확보율은 줄어들었고, 장서고유성은 증가하였다. DDC 주류 중에서는 컴퓨터 관련 도서가 급감한 0XX(총류) 분야의 감소율이 가장 컸다. 도서관별 장서량 측면에서는 2003년에 비해서 2013년 출판도서의 경우에 상위 도서관이 더욱 과점하는 빈익빈 부익부 현상이 심화되었다.

Abstract

This study analyzed Korean university libraries’ holdings of Western language books published in 2003 and 2013 using the KERIS union catalog with a view to investigating the changes in collection development of Western language books in the libraries. To do that, new collection indexes - holding h-index, CUI (Collection Uniqueness Index), and CCHR (Common Collection Holding Ratio) - were suggested, and they were used with basic indexes such as the number of titles, the number of books, and the number of books per title. The analysis reveals that compared to those published in 2003, the number of titles was decreased by 16.1% with those published in 2013, and the number of books dropped more sharply, by 42.2%. Also, in 2013, CCHR was decreased while CUI was increased. In terms of subject, among DDC main classes, 0XX (Generalities) showed the greatest decrease rate in both the number of titles and books because of the radical reduction of computer-related books. In terms of each library’s holdings, the number of Western language books held by top libraries has been increased with those published in 2013.

8

유사문헌집단에서 적합/부적합정보의 유용성에 관한 연구

문성빈(연세대학교) 2015, Vol.32, No.3, pp.277-293 https://doi.org/10.3743/KOSIM.2015.32.3.277

초록보기

초록

본 논문에서는 문헌의 적합성수준을 적합성정도에 따라 4그룹(부적합한, 조금 적합한, 적합한, 매우 적합한)으로 나눈 후 서로 다른 심사자가 적합성 판정을 내린 4개의 적합성 판정세트(A, B, C, D)에서 “조금 적합한” 문헌을 부적합문헌으로 분류했을 때와 적합문헌으로 분류하였을 때에, 초록/표제 시스템과 전문검색시스템에서 적합성피드백으로 인한 검색효율성의 증진은 어느 쪽이 더 혜택을 받게 되는 지를 연구하였다. “조금 적합한” 문헌을 적합문헌으로 포함시켰을 때 초록/표제시스템이 전문검색시스템보다 모든 적합성판정세트에서 검색효율성의 증가율이 높았고, 반면에 전문검색시스템에서는 “조금 적합한” 문헌을 적합문헌그룹에서 제외시켰을 때 검색효율성의 증가율이 일관성 있게 높아지는 것을 발견하였다. 이는 전문검색시스템에서는 적합문헌으로 포함된 “조금 적합한” 문헌으로부터 얻어지는 적합성피드백 정보는 잡음의 역할을 하게 되어 검색효율성의 증진에 도움이 안 되고 있음을 암시하고 있다. 특히, 매우 동질적인 문헌을 색인 및 검색대상으로 하고 있는 전문검색시스템에서는 잡음에 의해 초래되는 낮은 정확률을 개선하는 정교한 검색기법에 대한 연구가 지속되어야만 한다.

Abstract

This study examined the relative retrieval effectiveness after relevance feedback between two systems (Title/Abstract and Full-text) using four different sets of relevance judgment. Four relevance levels (not relevant, marginally relevant, relevant, highly relevant) are also used, each of which is determined by referees giving a relevance score to documents. This study also investigated how much the average precision was improved after relevance feedback when “marginally relevant” documents are included in the relevant class with the Title/Abstract system, and with the Full-text retrieval system as well. It is found that the Title/Abstract system benefited from relevance feedback with the marginally relevant documents. In case of the Title/Abstract system, the higher percentage of improvement was consistently obtained when including the marginally relevant documents in the relevance class, however the result was vice versa in case of the Full-text retrieval system. It implied that the marginally relevant documents in the relevant class had caused noises in the Full-text retrieval system.

9

상호대차 활성화에 따른 대학도서관 이.공계열 외국학술지의 평가에 관한 연구

이창수(경북대학교) ; 김신영(숭의여자대학) 2002, Vol.19, No.1, pp.71-88 https://doi.org/10.3743/KOSIM.2002.19.1.071

초록보기

초록

본 연구는 상호대차활성화에 따른 이.공계열 외국학술지의 이용빈도를 조사해보고, 외국학술지 분담구독의 모델을 제시하는데 그 목적이 있다. 본 이용연구는 2000년도 E대학도서관의 459종의 구독학술지와 144종의 분담구독 학술지를 포함한 총 603종을 평가대상으로 하여 관내 이용통계와 상호대차통계를 분석한 것이다. 본 연구결과 439종의 학술지가 1회 이상 이용되었으며, 164종의 불용학술지가 조사되었다. 또한 E대학 이용자의 이용빈도와 JCR의 주제분야별 총인용빈도 상위학술지간에는 상관성이 있는 것으로 나타났다. 다양한 이용분석을 통하여 본 연구는 효과적인 학술지 관리방안을 제시하였다.

Abstract

The purpose of this use study is to evaluate foreign science and technology serials with reference to the result of interlibrary loan activity and to present the model of cooperative acquisition. This study was based upon an analysis of actual use data by library users and interlibrary loan from March 1, 2000 to February 28, 2001. 603 titles of foreign serials which were composed of 459 subscription titles of E University Library and 144 cooperative acquisition titles in the fields of science and technology were analyzed. The study reveals that only 439 serials(72.8%) were used even more once and 164 serials were not used at all during 1 year interval. A relationship was found between rankings of serials as measured by use and JCR citation ranking. Based upon various aspects of use analysis, this study suggested effective techniques for managing academic serials.

10

사회과학 분야 도서의 목차 텍스트에 대한 통계적 특성에 관한 연구

이용구(계명대학교) 2019, Vol.36, No.2, pp.255-273 https://doi.org/10.3743/KOSIM.2019.36.2.255

초록보기

초록

이 연구는 최근 접근 및 활용이 높아지고 있는 목차에 대해 품사 측면과 주제 측면에서 가지는 기술 통계와 비교 분석을 수행하였다. 이를 위해 대학 도서관의 수서 목록에서 사회과학분야 도서를 추출하고 해당하는 도서에 대해 종합목록으로부터 DDC 분류기호를, 인터넷 서점으로부터 목차 정보를 추출하였다. 서명과 목차를 대상으로 형태소 분석하여 명사 중심의 어휘에 대해 기술통계와 빈도 분석을 실시하였다. 그 결과 형태소 측면에서 서명과 목차는 명사가 대략 절반가량 차지하며, 서명과 비교하여 목차는 50배 정도 더 많은 명사를 가지며, 목차에 출현한 명사 중에 목차만이 고유하게 가지는 비율이 95.2%에 달하는 것으로 파악되었다. 또한 목차는 사회과학 학문분야에 따라 길이가 차이가 나는 것으로 나타났다.

Abstract

Recently, the table of contents (TOC) has been becoming increasingly accessible and utilized. The study conducted descriptive statistics and comparative analysis of the table of contents in terms of parts of speech and subject in text. For this purpose, this study chose the books of the social sciences field from acquisition lists of an academic library, obtained Dewey class numbers of target books from KERIS union catalog, and extracted TOC data from online bookstore. Morphological analysis was performed on each book titles and TOCs, and descriptive statistics and frequency analysis were carried out. As a result, nouns made up roughly half of the morphemes of titles or the TOCs. TOCs had about 50 times more nouns than titles. The percentage of unique nouns that appeared only in the table of contents is estimated to be 95.2% of the TOC’s total nouns. The table of contents also showed a differences in its lengths depending on the field of social science.

바로가기메뉴

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

정보관리학회지