首页    期刊浏览 2024年11月25日 星期一
登录注册

文章基本信息

  • 标题:Parsing Models for Identifying Multiword Expressions
  • 本地全文:下载
  • 作者:Spence Green ; Marie-Catherine de Marneffe ; Christopher D. Manning
  • 期刊名称:Computational Linguistics
  • 印刷版ISSN:0891-2017
  • 电子版ISSN:1530-9312
  • 出版年度:2013
  • 卷号:39
  • 期号:1
  • 页码:195-227
  • DOI:10.1162/COLI_a_00139
  • 语种:English
  • 出版社:MIT Press
  • 摘要:Multiword expressions lie at the syntax/semantics interface and have motivated alternative theories of syntax like Construction Grammar. Until now, however, syntactic analysis and multiword expression identification have been modeled separately in natural language processing. We develop two structured prediction models for joint parsing and multiword expression identification. The first is based on context-free grammars and the second uses tree substitution grammars, a formalism that can store larger syntactic fragments. Our experiments show that both models can identify multiword expressions with much higher accuracy than a state-of-the-art system based on word co-occurrence statistics. We experiment with Arabic and French, which both have pervasive multiword expressions. Relative to English, they also have richer morphology, which induces lexical sparsity in finite corpora. To combat this sparsity, we develop a simple factored lexical representation for the context-free parsing model. Morphological analyses are automatically transformed into rich feature tags that are scored jointly with lexical items. This technique, which we call a factored lexicon, improves both standard parsing and multiword expression identification accuracy.
国家哲学社会科学文献中心版权所有