首页    期刊浏览 2024年11月15日 星期五
登录注册

文章基本信息

  • 标题:The Czech Academic Corpus 2.0 Guide
  • 作者:Barbora Hladká ; Jan Hajič ; Jirka Hana
  • 期刊名称:The Prague Bulletin of Mathematical Linguistics
  • 印刷版ISSN:0032-6585
  • 电子版ISSN:1804-0462
  • 出版年度:2008
  • 卷号:89
  • 期号:1
  • 页码:41-96
  • DOI:10.2478/v10108-009-0003-9
  • 语种:English
  • 出版社:Walter de Gruyter GmbH
  • 摘要:The Czech Academic Corpus version 2.0 is a morphologically and syntactically annotated corpus of 650,000 words. The Czech Academic Corpus (CAC) was created by a team from the Institute of the Czech Language of the Academy of Sciences of the Czech Republic from 1971 to 1985. When the CAC project began there were only two computerized annotated corpora available since the 1960s - the Brown Corpus of American English and the LOB Corpus of British English. Both corpora became well known to corpus linguists, whereas the CAC remained hidden mainly because of the 1980s political regime in the Czech Republic. The idea of transferring the internal format and annotation scheme of the CAC into the Prague Dependency Treebank (PDT) concept emerged during the work on the PDT's second version. The main goal was to make the CAC and the PDT fully compatible and thus enable the integration of the CAC into the PDT. The currently released second version of the CAC presents the complete conversion of the internal format and morphological and syntactical annotation schemes. The Czech Academic Corpus v. 2.0 is being published by the Linguistic Data Consortium.
Loading...
联系我们|关于我们|网站声明
国家哲学社会科学文献中心版权所有