首页    期刊浏览 2024年11月25日 星期一
登录注册

文章基本信息

  • 标题:Named Entity Recognition on Code-Mixed Cross-Script Social Media Content
  • 本地全文:下载
  • 作者:Somnath Banerjee ; Sudip Kumar Naskar ; Paolo Rosso
  • 期刊名称:Computación y Sistemas
  • 印刷版ISSN:1405-5546
  • 出版年度:2017
  • 卷号:21
  • 期号:4
  • 页码:681-692
  • 语种:English
  • 出版社:Instituto Politécnico Nacional
  • 其他摘要:Focusing on the current multilingual scenario in social media, this paper reports automatic extraction of named entities (NE) from code-mixed cross-script social media data. Our prime target is to extract NE for question answering. This paper also introduces a Bengali-English (Bn-En) code-mixed cross-script dataset for NE research and proposes domain specific taxonomies for NE. We used formal as well as informal language-specific features to prepare the classification models and employed four machine learning algorithms (Conditional Random Fields, Margin Infused Relaxed Algorithm, Support Vector Machine and Maximum Entropy Markov Model) for the NE recognition (NER) task. In this study, Bengali is considered as the native language while English is considered as the non-native language. However, the approach presented in this paper is generic in nature and could be used for any other code-mixed dataset. The classification models based on CRF and SVM performed well among the classifiers.
  • 其他关键词:Named entity recognition; code-mixed cross-script; Bengali-English social media content.
国家哲学社会科学文献中心版权所有