정보관리학회지, 한국정보관리학회

1

문장 검색을 위한 색인시스템 구축 : 초 .중등 학생의 한국어 및 영어 문장을 중심으로

이태영(전북대학교) 2003, Vol.20, No.1, pp.145-163 https://doi.org/10.3743/KOSIM.2003.20.1.145

초록보기

초록

한국어 및 영어의 글쓰기를 도와주는 문장 및 문단 제공시스템을 구축하기 위하여 색인작성과 탐색시에 필요한 색인언어를 연구하였다. 색인언어로 명사어와 술어 및 부사어를 선정하였고 여러 가지 보조 색인기호들도 추가하였다. 접근점으로 주제명과 키워드를 사용하였고 키워드 검색은 1절, 2절, 3 절, 문맥첨가 탐색을 포함하였다. 검색의 만족도는 긍정적이었으며 데이터베이스의 양과 질을 충실히 보완한다면 문장이나 문단을 제공하여 주는 시스템은 효과적일 수 있다.

Abstract

An indexing language were studied to construct the sentences and paragraphs providing system aided to write a Korean or English composition. The indexing language includes the index terms like noun, predicate, and adverb. and also various index symbols. The subject name and the keyword Included the symbols, which Indicate the connectives between clauses in a sentence, is used as the access point. The search results show this system will be effective with large database and developed retrieval methods.

2

시맨틱 웹 환경에서 적합한 문장을 제공하는 이야기 쓰기 도우미에 관한 연구

이태영(전북대학교) 2009, Vol.26, No.4, pp.7-34 https://doi.org/10.3743/KOSIM.2009.26.4.007

초록보기

초록

이야기 쓰기를 돕는 본문 및 문장 검색시스템의 구축을 위해서 (1)이야기와 단락 및 문장의 구조를 분석하고 (2)색인작성과 탐색 질문에 적용되는 언어 추론을 연구하였다. 이야기 쓰기에 필요한 이야기, 단락, 그리고 문장으로 구성된 사항 데이터베이스와 필요한 추론규칙으로 이루어진 지식베이스와 온톨로지가 고안되었다. 추론의 기초인 실례(實例) 파일들은 시맨틱 웹 환경에서 작동될 마크업 언어 형식으로 만들어졌다. 시맨틱 웹 환경에서 실용적인 시스템이 되려면 단락과 문장을 정확히 대변하는 색인 방법론과 이를 정밀하게 지식베이스화 할 수 있는 마크업 언어의 창조가 필수적이라 사료된다.

Abstract

Structures of stories, paragraphs, and sentences and inferences applied to indexing and searching were studied to construct the full-text and sentence retrieval system for storytelling. The system designed the database of stories, paragraphs, and sentences and the knowledge-base of inference rules to aid to write the story. The Knowledge-base comprised the files of story frames, paragraph scripts, and sentence logics made by mark-up languages like SWRL etc. able to operate in semantic web. It is necessary to establish more precise indexing language represented the sentences and to create a mark-up languages able to construct more accurate inference rules.

3

자연어 질의 분석과 검색어 확장에 기반한 웹 정보 검색

윤성희(상명대학교) 2004, Vol.21, No.2, pp.235-248 https://doi.org/10.3743/KOSIM.2004.21.2.235

초록보기

초록

웹 문서 검색을 위해 키워드와 불리언 연산식을 사용하는 것에 비해 자연어 질의 문장을 입력하는 방법은 검색 시스템 사용자에게 훨씬 이상적인 인터페이스이다. 본 논문은 사용자가 입력하는 자연어 질의 문장을 구문 분석하고 그 구문 구조에 기반하여 검색어를 확장하는 다중 검색 기법을 제안한다. 구문 트리를 순회하여 구조적으로 연관된 복합 명사를 조합하거나 분할하는 과정을 거치고, 이형 표기 및 축약 표기 용어들에 대해 확장 다중 검색함으로써 웹 정보 검색 시스템의 재현율과 정확도를 높일 수 있다.

Abstract

For the users of information retrieval systems, natural language query is the more ideal interface, compared with keyword and boolean expressions. This paper proposes a retrieval technique with expanded keyword from syntactically-analyzed structures of natural language query as user input. Through the steps combining or splitting the compound nouns based on syntactic tree traversal of the query, and expanding the other-formed or shorten-formed into multiple keyword, it can enhance the precision and correctness of the retrieval system.

4

질의응답문서 검색에서 문서구조를 이용한 질의재생성에 관한 연구

최상희(대구가톨릭대학교) ; 서은경(한성대학교) 2006, Vol.23, No.2, pp.229-243 https://doi.org/10.3743/KOSIM.2006.23.2.229

초록보기

초록

질의응답문서는 이용자가 입력한 질의, 질의설명, 답을 아는 다른 이용자가 제시한 응답으로 구성된 구조화된 문서로서, 최근 웹 문서처럼 검색이 일반적으로 일어나고 있는 정보원이다. 이 연구에서는 질의응답문서의 구조적 특성을 기반으로 질의를 재생성하여 질의응답문서의 검색효율을 향상시키고자 하였다. 질의재생성 실험에서 성능이 비교된 문서구조는 질의와 응답내용이다. 질의를 기반으로 질의를 재생성하는 방식에서는 질의응답검색 시스템에 입력되어 있는 유사질의를 활용하여 클러스터링하는 기법이 적용되었다. 응답정보를 기반으로 질의를 재생성하는 방식에서는 가장 유사한 기존 질의에 대해 응답된 내용에서 단락검색으로 적합한 문장들을 선정하여 활용하는 기법이 적용되었다. 실험 결과 응답정보를 활용하여 질의를 재생성하는 방식이 정확률은 유지하면서 더 다양한 검색결과를 제공하는 것으로 나타났다.

Abstract

This study aims to suggest an effective way to enhance question-answer(QA) document retrieval performance by reconstructing queries based on the structural features in the QA documents. QA documents are a structured document which consists of three components: question from a questioner, short description on the question, answers chosen by the questioner. The study proposes the methods to reconstruct a new query using by two major structural parts, question and answer, and examines which component of a QA document could contribute to improve query performance. The major finding in this study is that to use answer document set is the most effective for reconstructing a new query. That is, queries reconstructed based on terms appeared on the answer document set provide the most relevant search results with reducing redundancy of retrieved documents.

5

초록의 구조와 내용에 관한 연구: 정보관리학회지를 중심으로

장우권(전남대학교) 2016, Vol.33, No.3, pp.107-131 https://doi.org/10.3743/KOSIM.2016.33.3.107

초록보기

초록

이 연구는 정보관리학회지의 초록의 현황에 대한 조사와 분석을 통해 초록의 특징을 진단하고자 한다. 이를 위해 학회지의 저자초록을 중심으로 초록의 구성요소, 초록의 유형 등을 분석하였다. 대상초록은 1984년부터 2015년까지 간행된 학회지 논문이다. 그 결과 수록논문은 1,168편, 지시적 초록은 96.6%, 통보적 초록은 3.4%, 국어와 영어 병기 초록은 99.5%였다. 연구방법에서 문헌사례 52.8%, 설문조사 21.1%, 실험이 26.1%였다. 문단과 문장에서 1문단이 92.1%, 2문단 이상이 7.9%, 5문장 이하가 79%, 6문장 이상이 21%, 1인칭 사용이 90.5%로 나타났다. 주제영역은 도서관/정보센터경영 19.4%, 정보서비스 17.3%, 정보공학 16.4%, 정보검색 15.1%, 계량정보 9.6% 등의 순으로 나타났다.

Abstract

This study aims to identify the characteristics of abstracts by analyzing the status of abstracts published in ‘Journal of the Korean Society for Information Management.’ To this end, the study analyzed the components of and the types of abstracts. Target abstracts were those published in the journal from 1984 to 2015. The journal published 1,168 articles with indicative abstracts accounting for 96.6%, informative abstracts 3.4%, and abstracts written in both English and Korean 99.5%. As for research methods, case study through literature review was 52.8%, surveys 21.1%, and experimentation 26.1%. The percentage of abstracts consisting of one paragraph was 92.1%, more than two paragraphs were 7.9%, fewer than 5 sentences were 79%, and 6 sentences or more were 21%. The use of the first person was 90.5%. In terms of topic areas, library and information center management was 19.4%, information services 17.3%, information technology 16.2%, information retrieval 15.1%, and informetrics 9.6%, etc.

6

과학기술 핵심개체 인식기술 통합에 관한 연구

최윤수(한국과학기술정보연구원) ; 정창후(한국과학기술정보연구원) ; 조현양(경기대학교) 2011, Vol.28, No.1, pp.89-104 https://doi.org/10.3743/KOSIM.2011.28.1.089

초록보기

초록

대용량 문서에서 정보를 추출하는 작업은 정보검색 분야뿐 아니라 질의응답과 요약 분야에서 매우 유용하다. 정보추출은 비정형 데이터로부터 정형화된 정보를 자동으로 추출하는 작업으로서 개체명 인식, 전문용어 인식, 대용어 참조해소, 관계 추출 작업 등으로 구성된다. 이들 각각의 기술들은 지금까지 독립적으로 연구되어왔기 때문에, 구조적으로 상이한 입출력 방식을 가지며, 하부모듈인 언어처리 엔진들은 특성에 따라 개발 환경이 매우 다양하여 통합 활용이 어렵다. 과학기술문헌의 경우 개체명과 전문용어가 혼재되어 있는 형태로 구성된 문서가 많으므로, 기존의 연구결과를 이용하여 접근한다면 결과물 통합과정의 불편함과 처리속도에 많은 제약이 따른다. 본 연구에서는 과학기술문헌을 분석하여 개체명과 전문용어를 통합 추출할 수 있는 기반 프레임워크를 개발한다. 이를 위하여, 문장자동분리, 품사태깅, 기저구인식 등과 같은 기반 언어 분석 모듈은 물론 이를 활용한 개체명 인식기, 전문용어 인식기를 개발하고 이들을 하나의 플랫폼으로 통합한 과학기술 핵심개체 인식 체계를 제안한다.

Abstract

Large-scaled information extraction plays an important role in advanced information retrieval as well as question answering and summarization. Information extraction can be defined as a process of converting unstructured documents into formalized, tabular information, which consists of named-entity recognition, terminology extraction, coreference resolution and relation extraction. Since all the elementary technologies have been studied independently so far, it is not trivial to integrate all the necessary processes of information extraction due to the diversity of their input/output formation approaches and operating environments. As a result, it is difficult to handle scientific documents to extract both named-entities and technical terms at once. In order to extract these entities automatically from scientific documents at once, we developed a framework for scientific core entity extraction which embraces all the pivotal language processors, named-entity recognizer and terminology extractor.

바로가기메뉴

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

정보관리학회지