首页    期刊浏览 2025年07月02日 星期三
登录注册

文章基本信息

  • 标题:Qualitative Analysis of PLP in LSTM for Bangla Speech Recognition
  • 本地全文:下载
  • 作者:Nahyan Al Mahmud ; Shahfida Amjad Munni
  • 期刊名称:The International Journal of Multimedia & Its Applications (IJMA)
  • 印刷版ISSN:0975-5934
  • 电子版ISSN:0975-5578
  • 出版年度:2020
  • 卷号:12
  • 期号:5
  • 页码:1-8
  • DOI:10.5121/ijma.2020.12501
  • 出版社:Academy & Industry Research Collaboration Center (AIRCC)
  • 摘要:The performance of various acoustic feature extraction methods has been compared in this work using Long Short-Term Memory (LSTM) neural network in a Bangla speech recognition system. The acoustic features are a series of vectors that represents the speech signals. They can be classified in either words or sub word units such as phonemes. In this work, at first linear predictive coding (LPC) is used as acoustic vector extraction technique. LPC has been chosen due to its widespread popularity. Then other vector extraction techniques like Mel frequency cepstral coefficients (MFCC) and perceptual linear prediction (PLP) have also been used. These two methods closely resemble the human auditory system. These feature vectors are then trained using the LSTM neural network. Then the obtained models of different phonemes are compared with different statistical tools namely Bhattacharyya Distance and Mahalanobis Distance to investigate the nature of those acoustic features.
  • 关键词:LSTM; Perceptual linear prediction; Mel frequency cepstral coefficients; Bhattacharyya Distance; Mahalanobis Distance.
国家哲学社会科学文献中心版权所有