정보관리학회지, 한국정보관리학회

1

이지숙(NHN㈜) ; 정영미(연세대학교) 2007, Vol.24, No.3, pp.201-218 https://doi.org/10.3743/KOSIM.2007.24.3.201

초록보기

초록

이 연구에서는 TREC이 제시한 토픽 검색의 정의에 따라 질의에 적합한 웹 사이트를 검색하는 효과적인 토픽 검색 알고리즘을 제안하고 실험을 통해 그 성능을 평가하였다. 이 연구의 토픽 검색 알고리즘은 먼저 질의에 대한 웹 페이지 검색 결과로부터 적합한 웹 사이트를 선정한 다음, 선정된 사이트의 구조를 이용하여 질의에 대한 적합성 점수를 산출한다. TREC의 .GOV 실험 문헌 집단과 TREC-2004 실험의 질의 및 적합문헌 리스트를 이용한 검색 실험 결과 이 토픽 검색 알고리즘은 상위 10위 안에 최소 2개 이상의 적합 사이트를 검색하여 비교적 높은 수준의 성능을 보였다. 또한 TREC-2004의 적합문헌 리스트 분석을 통해 적합문헌 선정에 토픽 검색의 정의가 엄격하게 적용되지 않은 경우가 있음을 확인하고, 수정된 적합문헌 리스트를 이용하여 토픽 검색 성능을 재평가한 결과 이 연구에서 제안한 토픽 검색 알고리즘의 성능이 월등히 향상되었다.

Abstract

This study proposes a topic distillation algorithm that ranks the relevant sites selected from retrieved web pages, and evaluates the performance of the algorithm. The algorithm calculates the topic score of a site using its hierarchical structure. The TREC .GOV test collection and a set of TREC-2004 queries for topic distillation task are used for the experiment. The experimental results showed the algorithm returned at least 2 relevant sites in top ten retrieval results. We performed an in-depth analysis of the relevant sites list provided by TREC-2004 to find out that the definition of topic distillation was not strictly applied in selecting relevant sites. When we re-evaluated the retrieved sites/sub-sites using the revised list of relevant sites, the performance of the proposed algorithm was improved significantly.

2

검색 포털들의 동영상 검색 서비스 분석 평가: 네이버와 구글을 중심으로

박소연(덕성여자대학교) 2014, Vol.31, No.3, pp.181-200 https://doi.org/10.3743/KOSIM.2014.31.3.181

초록보기

초록

본 연구에서는 주요 검색 포털들의 동영상 검색 서비스를 분석, 평가하였다. 이 연구에서는 네이버와 구글 코리아를 대상으로 동영상의 컬렉션별 분포, 작성 연도별 분포, 중복 동영상의 비중, 광고 동영상의 비중 및 특징, 검색 결과의 화질 등을 조사하고, 동영상의 적합도, 신뢰도, 최신성을 비교, 평가였다. 또한, 동영상의 적합도, 신뢰도에 영향을 미치는 요소들을 조사하였다. 마지막으로 동영상들 중 오류 동영상의 유형 및 특징도 조사하였다. 연구 결과, 구글이 네이버보다 동영상의 적합도가 높고, 네이버가 구글보다 동영상의 최신성이 다소 높은 것으로 나타났다. 동영상의 화질은 구글이 네이버보다 높은 것으로 나타났다. 또한 구글과 네이버 모두 중복되는 동영상의 비중이 높은 편이었으며, 광고 동영상은 네이버에서 구글보다 더 많이 노출되었다. 본 연구의 결과는 향후 포털들의 동영상 검색 서비스의 개선에 활용될 수 있을 것으로 기대된다.

Abstract

This study aims to analyze and evaluate video search services of major search portals, Naver and Google Korea. In particular, this study analyzed characteristics such as collection distribution, yearly distribution, the ratio of redundant search results, the ratio of advertising, and the quality of videos. This study also evaluated relevance, credibility, and currency of video search results, and investigated the factors that influence relevance and credibility. Finally, types and characteristics of error results were analyzed. The results of this study show that the relevance of Google’s video search results is higher than those of Naver, whereas currency of Naver’s search results is somewhat higher than those of Google. Google has more high resolution videos than Naver, and Naver has more advertising than Google. Both Google and Naver return many redundant videos in the search results. The results of this study can be implemented to the portal’s effective development of video search services.

3

온톨로지를 이용한 인터넷웹 검색에 관한 실험적 연구

김현희(명지대학교) ; 안태경(대외경제정책연구원) 2003, Vol.20, No.1, pp.417-455 https://doi.org/10.3743/KOSIM.2003.20.1.417

초록보기

초록

온톨로지는 웹자원을 지식화함으로써 정보의 효율적 검색, 통합, 재사용을 도모할 수 있는 새로운 기술인 시맨틱 웹의 구현을 위한 가장 핵심적인 요소 기술로 알려지고 있다. 온톨로지는 사람간에 그리고 서로 다른 응용 시스템간에 지식을 공유하고 재이용하는 방법을 제공하는 기술로서 특정 주제에 관한 지식 용어들의 집합으로서 이들 용어뿐만 아니라 용어간의 의미적 연결 관계와 간단한 추론 규칙을 포함한다. 본 연구에서는 인터넷 웹상에서 국제기구에 관한 정보를 체계적으로 관리하고 검색하기 위해서 국제기구 온톨로지를 설계하고 이 온톨로지에 기반 하여 검색 시스템을 구현해 보고, 이 시스템을 20개의 탐색 질문들을 이용하여 기존의 인터넷 검색엔진과 적합성과 탐색 시간이라는 두 가지 요인을 통해서 비교해 보았다. 실험 결과에 의하면 적합성 측정은 온톨로지 기반 시스템은 평균 4.53, 인터넷 검색엔진은 평균 2.51로 온톨로지 기반 시스템의 적합도가 1.80배 높은 것으로 나타났다. 또한 탐색시간은 온톨로지 기반 시스템은 평균 1.96분, 인터넷 검색엔진은 평균 4.74분으로 인터넷 검색엔진이 온톨로지 기반 시스템 보다 2.42배 정도 더 많은 탐색시간이 필요한 것으로 나타났다.

Abstract

Ontologies are formal theories that are suitable for implementing the semantic web, which is a new technology that attempts to achieve effective retrieval, integration, and reuse of web resources. Ontologies provide a way of sharing and reusing knowledge among people and heterogeneous applications systems. The role of ontologies is that of making explicit specified conceptualizations. In this context, domain and generic ontologies can be shared, reused, and integrated in the analysis and design stage of information and knowledge systems. This study aims to design an ontology for international organizations, and build an Internet web retrieval system based on the proposed ontology, and finally conduct an experiment to compare the system performance of the proposed system with that of Internet search engines focusing relevance and searching time. This study found that average relevance of ontology- based searching and Internet search engines are 4.53 and 2.51, and average searching time of ontology-based searching and Internet search engines are 1.96 minutes and 4.74 minutes.

4

웹기반 정보검색시스템의 검색관련 용어 표준에 관한 연구

남영준(중앙대학교) 2003, Vol.20, No.2, pp.199-217 https://doi.org/10.3743/KOSIM.2003.20.2.199

초록보기

초록

본 연구에서는 웹기반 정보검색시스템을 사용함에 있어 이용자 편의성을 최적화할 수 있는 검색 인터페이스 표준 용어를 제안하였다. 이를 위해 국립중앙도서관을 비롯하여 주요 전문 정보를 제공하고 있는 기관의 웹페이지를 조사. 분석하였다. 분석한 결과에 근거하여 웹기반 정보검색시스템에서 사용자 오류와 혼란을 최소화하고 검색 편의성을 극대화할 수 있는 표준 용어를 제안하였다. 제안의 기준은 해당 용어의 사용빈도와 의미를 활용하였다. 분석은 검색관련 기본 모듈을 비롯하여 검색범위설정 모듈, 이용자 지원 모듈에서 사용된 용어 가운대 최소 50%이상의 기관에서 제공하는 기능에 존재하는 용어만을 대상으로 하였다. 본 연구의 결과는 웹 기반 검색화면 설계 및 구축 전문가에게 검색 관련 용어선정을 위한 표준 자료로 활용될 것이다.

Abstract

This research suggesrs the method of standardizing terms for raising the dffectiveness of information retrival. Especiallly for web search, I propose the proper terms which they will use in retriveal by surveying and analysing the related terms abour information retrieval interface. The proper terms will solve the eqyivocaiton for user and increase the retrieval effectiveness. And I think the proposed terms will be used to standard data for designers who are construct the user interface systems.

5

자연어 질의 분석과 검색어 확장에 기반한 웹 정보 검색

윤성희(상명대학교) 2004, Vol.21, No.2, pp.235-248 https://doi.org/10.3743/KOSIM.2004.21.2.235

초록보기

초록

웹 문서 검색을 위해 키워드와 불리언 연산식을 사용하는 것에 비해 자연어 질의 문장을 입력하는 방법은 검색 시스템 사용자에게 훨씬 이상적인 인터페이스이다. 본 논문은 사용자가 입력하는 자연어 질의 문장을 구문 분석하고 그 구문 구조에 기반하여 검색어를 확장하는 다중 검색 기법을 제안한다. 구문 트리를 순회하여 구조적으로 연관된 복합 명사를 조합하거나 분할하는 과정을 거치고, 이형 표기 및 축약 표기 용어들에 대해 확장 다중 검색함으로써 웹 정보 검색 시스템의 재현율과 정확도를 높일 수 있다.

Abstract

For the users of information retrieval systems, natural language query is the more ideal interface, compared with keyword and boolean expressions. This paper proposes a retrieval technique with expanded keyword from syntactically-analyzed structures of natural language query as user input. Through the steps combining or splitting the compound nouns based on syntactic tree traversal of the query, and expanding the other-formed or shorten-formed into multiple keyword, it can enhance the precision and correctness of the retrieval system.

6

질의 로그 분석을 통한 네이버 이용자의 검색 형태 연구

이준호(숭실대학교) ; 권혁성(숭실대학교) ; 박소연() 2003, Vol.20, No.2, pp.27-41 https://doi.org/10.3743/KOSIM.2003.20.2.027

초록보기

초록

이용자와 검색 서비스 시스템의 모든 검색 과정을 기록한 질의 로그는 이용자의 실제 검색 행위를 사실적으로 반영한다. 따라서, 웹 검색 이용자들의 검색 행태를 이해하기위하여 웹 검색 서비스 시스템이 생성한 질의 로그를 분석하는 방법이 널리 사용되고 있다. 본 연구는 네이버 이용자의 웹 검색 행태를 파악하기 위하여 기존의 질의 로그 분석 방법론을 보완하여 제시한다. 또한, 본 연구는 통합 검색, 디텍토리 검색, 웹 문서 검색과 같은 다양한 검색 유형에 대하여 일주일 동안 생성된 질의 로그를 분석함으로써 네이버 웹 검색 이용자들의 전반적인 검색 행태를 파악하였다. 본 연구의 결과는 보다 효과적인 웹 검색 시스템 개발과 서비스 구축에 기여할 것으로 기대된다.

Abstract

Query logs are online records that capture user interactions with information retrieval systems and all the search processes. Query log analysis offers an advantage of providing reasonable and unobtrusive means of collecting search information from a large number of users. In this paper, query logs of NAVER, a major Korean Internet search service, were analyzed to investigate the information seeking behavior of NAVER users. The query logs were collected over one week from various collections such as comprehensive search, directory search and web document search. It is expected that this study could contribute to the development and implementation of more effective web search systems and services.

7

웹 검색 성능 최적화를 위한 융합적 방식

Yang, Kiduk(경북대학교) 2015, Vol.32, No.1, pp.7-22 https://doi.org/10.3743/KOSIM.2015.32.1.007

초록보기

초록

Abstract

This paper describes a Web search optimization study that investigates both static and dynamic tuning methods for optimizing system performance. We extended the conventional fusion approach by introducing the “dynamic tuning” process with which to optimize the fusion formula that combines the contributions of diverse sources of evidence on the Web. By engaging in iterative dynamic tuning process, where we successively fine-tuned the fusion parameters based on the cognitive analysis of immediate system feedback, we were able to significantly increase the retrieval performance.Our results show that exploiting the richness of Web search environment by combining multiple sources of evidence is an effective strategy.

8

검색 포털들의 검색어 추천 서비스 분석 평가: 네이버와 구글의 연관 검색어 서비스를 중심으로

박소연(덕성여자대학교) 2013, Vol.30, No.2, pp.297-315 https://doi.org/10.3743/KOSIM.2013.30.2.297

초록보기

초록

본 연구에서는 주요 검색 포털들의 검색어 추천 서비스를 분석, 평가하였다. 이 연구에서는 네이버와 구글 코리아를 대상으로 추천되는 연관 검색어의 적합도 및 최신성을 평가하고, 연관 검색어의 개수 및 분포, 연관 검색어가 제공되지 않는 질의의 특징을 조사하였다. 또한 연관 검색어의 유형을 질의와 연관 검색어의 관계 측면에서 분석하고, 연관 검색어들 중 유해 검색어의 유형 및 특징, 비표준어의 유형 및 특징도 조사하였다. 마지막으로, 한글 질의와 영어 질의, 대중적인 질의와 전문적인 질의의 연관 검색어의 특징을 비교하였다. 연구 결과, 네이버가 구글보다 연관 검색어의 적합도와 최신성이 다소 높은 것으로 나타났다. 또한 구글과 네이버 모두 새로운 연관 검색어를 제시하기보다는 질의에 단어를 추가 또는 삭제하거나, 질의와 동일한 검색어나 동의어 검색어를 제공하는 경우가 많은 것으로 나타났다. 본 연구의 결과는 향후 포털들의 검색어 추천 서비스의 개선에 활용될 수 있을 것으로 기대된다.

Abstract

This study aims to analyze and evaluate term suggestion services of major search portals, Naver and Google Korea. In particular, this study evaluated relevance and currency of related search terms provided, and analyzed characteristics such as number and distribution of terms, and queries that did not produce terms. This study also analyzed types of terms in terms of the relationship between queries and terms, and investigated types and characteristics of harmful terms and terms with grammatical errors. Finally, Korean queries and English queries, and popular queries and academic queries were compared in terms of the amount and relevance of search terms provided. The results of this study show that the relevance and currency of Naver's related search terms are somewhat higher than those of Google. Both Naver and Google tend to add terms to or delete terms from original queries, and provide identical search terms or synonym terms rather than providing entirely new search terms. The results of this study can be implemented to the portal's effective development of term suggestion services.

9

웹 환경에서의 학습 방법이 정보검색 및 정보종합 능력에 미치는 영향

함명식(서울맹학교) 2002, Vol.19, No.4, pp.5-34 https://doi.org/10.3743/KOSIM.2002.19.4.005

초록보기

초록

본 연구에서는 웹 환경에서의 학습 방법이 학생들의 정보검색 및 정보종합 능력에 어떠한 영향을 미치는가를 규명하고자 하였다. 본 연구의 결과는 다음과 같다. 첫째, 과제 중심형 학습 집단이 기법 중심형 학습 집단보다 정보검색 능력 중 정보성취도 검사점수가 높게 나타났으며, 통계적으로 유의미한 차이를 보였다 (t=3.59, p〈.05). 둘째, 네이버 국내 웹 1차 검색 (재현율 t=1.81, 정확율 t=.61)에서 과제 중심형 학습 집단과 기법 중심형 학습 집단간에 재현율과 정확율 모두 유의미한 차이가 없었다 (p〉.05). 그러나 2차 검색 (재현율 t=2.93, 정확율 t=2.45)과 3차 검색 (재현율 t=3.48, 정확율 t=2.50)에서는 과제중심형 학습 집단이 기법 중심형 학습 집단보다 재현율과 정확율이 높게 나타났으며, 통계적으로 유의미한 차이를 보였다 (p〈.05). 셋째, 과제 중심형 학습 집단과 기법 중심형 학습 집단은 정보종합 능력의 검사 점수 차이가 통계적으로 유의미하지 않았다 (t=1.95, p〉.05). 위 실험 결과를 종합해 보면, 인터넷에서 정보를 검색하는 경우에 과제에 대한 분석과 그에 알맞는 정보검색 기법을 적용하는 것이 중요하다. 기법에 의존하기보다는 과제를 분석하고 그에 알맞는 검색을 수행해야 한다. 또 정보 이용 교육이 정보검색 수준에서 머무르는 것이 아니라, 정보검색과 정보종합에 관한 교육이 정보 문제 해결의 맥락에서 이루어져야 할 것이다.

Abstract

The purpose of this study is to investigate the effects of learning methods on students'''' information retrieval and information synthesis capability in web. This is an experimental study comparing the two different learning methods as task-based learning and technic-based learning. The findings of this study were as follows: 1. The task-based learning was more effective than the technic-based learning in information achievements as information retrieval capability (t= 3.59, p〈.05). 2. In the 1st retrieval (recall ratio t=1.81 precision ratio t=.61) of Naver Korean Web Retrieval, there was no significant difference (p〉.05). In the 2nd retrieval (recall ratio t=2.93 precision ratio t=2.45) and 3rd retrieval (recall ratio t=3.48 precision ratio t= 2.50), the task-based group was more effective than the technic-based group (p〈.05). 3. There was no significant difference in students'''' information synthesis capability between the task-based learning and technic-based learning (t= 1.95, p〉.05). The findings of this study suggest that the task-based learning approach is more effective to improve students'''' information literacy, and that professionals should consider better instructional principles for the improvement of instructional quality.

10

로그분석을 통한 이용자의 웹 문서 검색 행태에 관한 연구

박소연(계명대학교) ; 이준호(숭실대학교) 2002, Vol.19, No.3, pp.111-122 https://doi.org/10.3743/KOSIM.2002.19.3.111

초록보기

초록

본 연구에서는 웹 검색 이용자들의 전반적인 검색 행태를 이해하기 위하여 국내에서 널리 사용되고 있는 웹 검색 서비스 네이버에서 생성된 검색 트랜잭션 로그를 분석하였다. 본 연구에서는 웹 검색 트랜잭션 로그 분석에 필요한 세션 정의 방법을 설명하고 로그 정제 및 질의 유형 분류방법을 제시하였으며, 한글 검색 트랜잭션 로그 분석에 필수절인 검색어 정의 방법을 제안하였다. 본 연구의 결과는 보다 효과적인 국내 웹 검색 시스템 개발과 서비스 구축에 기여할 것으로 기대된다.

Abstract

In order to investigate information seeking behavior of web search users, this study analyzes transaction logs posed by users of NAVER, a major Korean Internet search service. We present a session definition method for Web transaction log analysis, a way of cleaning original logs and a query classification method. We also propose a query term definition method that is necessary for Korean Web transaction log analysis. It is expected that this study could contribute to the development and implementation of more effective Web search systems and services.

바로가기메뉴

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

초록

Abstract

정보관리학회지