首页    期刊浏览 2024年09月02日 星期一
登录注册

文章基本信息

  • 标题:Computational Constancy Measures of Texts—Yule's K and Rényi's Entropy
  • 本地全文:下载
  • 作者:Kumiko Tanaka-Ishii ; Shunsuke Aihara
  • 期刊名称:Computational Linguistics
  • 印刷版ISSN:0891-2017
  • 电子版ISSN:1530-9312
  • 出版年度:2015
  • 卷号:41
  • 期号:3
  • 页码:481-502
  • DOI:10.1162/COLI_a_00228
  • 语种:English
  • 出版社:MIT Press
  • 摘要:This article presents a mathematical and empirical verification of computational constancy measures for natural language text. A constancy measure characterizes a given text by having an invariant value for any size larger than a certain amount. The study of such measures has a 70-year history dating back to Yule's K, with the original intended application of author identification. We examine various measures proposed since Yule and reconsider reports made so far, thus overviewing the study of constancy measures. We then explain how K is essentially equivalent to an approximation of the second-order Rényi entropy, thus indicating its signification within language science. We then empirically examine constancy measure candidates within this new, broader context. The approximated higher-order entropy exhibits stable convergence across different languages and kinds of text. We also show, however, that it cannot identify authors, contrary to Yule's intention. Lastly, we apply K to two unknown scripts, the Voynich manuscript and Rongorongo, and show how the results support previous hypotheses about these scripts.
国家哲学社会科学文献中心版权所有