文章基本信息

标题：Text Classification Based on Convolutional Neural Networks and Word Embedding for Low-Resource Languages: Tigrinya
本地全文：下载
作者：Awet Fesseha ; Shengwu Xiong ; Eshete Derb Emiru 等
期刊名称：Information
电子版ISSN：2078-2489
出版年度：2021
卷号：12
期号：2
页码：52
DOI：10.3390/info12020052
出版社：MDPI Publishing
摘要：This article studies convolutional neural networks for Tigrinya (also referred to as Tigrigna), which is a family of Semitic languages spoken in Eritrea and northern Ethiopia. Tigrinya is a “low-resource” language and is notable in terms of the absence of comprehensive and free data. Furthermore, it is characterized as one of the most semantically and syntactically complex languages in the world, similar to other Semitic languages. To the best of our knowledge, no previous research has been conducted on the state-of-the-art embedding technique that is shown here. We investigate which word representation methods perform better in terms of learning for single-label text classification problems, which are common when dealing with morphologically rich and complex languages. Manually annotated datasets are used here, where one contains 30,000 Tigrinya news texts from various sources with six categories of “sport”, “agriculture”, “politics”, “religion”, “education”, and “health” and one unannotated corpus that contains more than six million words. In this paper, we explore pretrained word embedding architectures using various convolutional neural networks (CNNs) to predict class labels. We construct a CNN with a continuous bag-of-words (CBOW) method, a CNN with a skip-gram method, and CNNs with and without word2vec and FastText to evaluate Tigrinya news articles. We also compare the CNN results with traditional machine learning models and evaluate the results in terms of the accuracy, precision, recall, and F1 scoring techniques. The CBOW CNN with word2vec achieves the best accuracy with 93.41%, significantly improving the accuracy for Tigrinya news classification.
关键词：text classification; CNN; low-resource language; machine learning; word embedding; natural language processing text classification ; CNN ; low-resource language ; machine learning ; word embedding ; natural language processing