首页    期刊浏览 2024年11月08日 星期五
登录注册

文章基本信息

  • 标题:Generating Javanese Stopwords List using K-means Clustering Algorithm
  • 本地全文:下载
  • 作者:Aji Prasetya Wibawa ; Hidayah Kariima Fithri ; Ilham Ari Elbaith Zaeni
  • 期刊名称:Knowledge Engineering and Data Science
  • 印刷版ISSN:2597-4602
  • 电子版ISSN:2597-4637
  • 出版年度:2020
  • 卷号:3
  • 期号:2
  • 页码:106-111
  • DOI:10.17977/um018v3i22020p106-111
  • 出版社:Universitas Negeri Malang
  • 摘要:Stopword removal necessary in Information Retrieval. It can remove frequently appeared and general words to reduce memory storage. The algorithm eliminates each word that is precisely the same as the word in the stopword list. However, generating the list could be time-consuming. The words in a specific language and domain must be collected and validated by specialists. This research aims to develop a new way to generate a stop word list using the K-means Clustering method. The proposed approach groups words based on their frequency. The confusion matrix calculates the difference between the findings with a valid stopword list created by a Javanese linguist. The accuracy of the proposed method is 78.28% (K=7). The result shows that the generation of Javanese stopword lists using a clustering method is reliable.
  • 其他摘要:Stopword removal necessary in Information Retrieval. It can remove frequently appeared and general words to reduce memory storage. The algorithm eliminates each word that is precisely the same as the word in the stopword list. However, generating the list could be time-consuming. The words in a specific language and domain must be collected and validated by specialists. This research aims to develop a new way to generate a stop word list using the K-means Clustering method. The proposed approach groups words based on their frequency. The confusion matrix calculates the difference between the findings with a valid stopword list created by a Javanese linguist. The accuracy of the proposed method is 78.28% (K=7). The result shows that the generation of Javanese stopword lists using a clustering method is reliable.
  • 关键词:Stopwords Javanese language Clustering K-means
国家哲学社会科学文献中心版权所有