首页    期刊浏览 2025年08月29日 星期五
登录注册

文章基本信息

  • 标题:Splitting Complex Sentences for Natural Language Processing Applications: Building a Simplified Spanish Corpus
  • 本地全文:下载
  • 作者:José Camacho Collados ; José Camacho Collados
  • 期刊名称:Procedia - Social and Behavioral Sciences
  • 印刷版ISSN:1877-0428
  • 出版年度:2013
  • 卷号:95
  • 页码:464-472
  • DOI:10.1016/j.sbspro.2013.10.670
  • 语种:English
  • 出版社:Elsevier
  • 摘要:AbstractThis paper presents a new Spanish parallel corpus of original and syntactically simplified texts. The simplification carried out basically consists of opportunistically splitting a complex original sentence into several simple ones. This parallel corpus is envisioned as a first step in order to create an automatic syntactic simplification system to be used as a preprocessing tool for other Natural Language Processing tasks such as Text Summarization, Information Extraction, parsing or Machine Translation. The corpus has been evaluated by human annotators regarding its grammaticality and preservation of meaning. The results suggest that the meaning of simplified and original sentences is almost identical.
  • 关键词:text simplification;syntactic simplification;parallel corpus;spanish;natural language processing.
国家哲学社会科学文献中心版权所有