首页    期刊浏览 2024年11月24日 星期日
登录注册

文章基本信息

  • 标题:All-words Word Sense Disambiguation for Russian Using Automatically Generated Text Collection
  • 本地全文:下载
  • 作者:Bolshina Angelina ; Natalia Loukachevitch
  • 期刊名称:Cybernetics and Information Technologies
  • 印刷版ISSN:1311-9702
  • 电子版ISSN:1314-4081
  • 出版年度:2020
  • 卷号:20
  • 期号:4
  • 页码:90-107
  • DOI:10.2478/cait-2020-0049
  • 语种:English
  • 出版社:Bulgarian Academy of Science
  • 摘要:The limited amount of the sense annotated data is a big challenge for theword sense disambiguation task. As a solution to this problem, we propose analgorithm of automatic generation and labelling of the training collections based onthe monosemous relatives concept. In this article we explore the limits of thisalgorithm: we employ it to harvest training collections for all ambiguous nouns,verbs and adjectives presented in RuWordNet thesaurus and then evaluate the qualityof the obtained collections. We demonstrate that our approach can create high-quality labelled collections with almost full-coverage of the RuWordNet polysemouswords. Furthermore, we show that our method can be applied to the Word-in-Contexttask.
  • 关键词:Word Sense Disambiguation; Word-in-Context task; automatic annotation of training collections; monosemous relatives; Russian dataset;RuWordNet thesaurus.
国家哲学社会科学文献中心版权所有