首页    期刊浏览 2024年10月07日 星期一
登录注册

文章基本信息

  • 标题:Fuzzy Clustering of Sequential Data
  • 本地全文:下载
  • 作者:B.K. Tripathy ; Rahul
  • 期刊名称:International Journal of Intelligent Systems and Applications
  • 印刷版ISSN:2074-904X
  • 电子版ISSN:2074-9058
  • 出版年度:2019
  • 卷号:11
  • 期号:1
  • 页码:43-54
  • DOI:10.5815/ijisa.2019.01.05
  • 出版社:MECS Publisher
  • 摘要:With the increase in popularity of the Internet and the advancement of technology in the fields like bioinformatics and other scientific communities the amount of sequential data is on the increase at a tremendous rate. With this increase, it has become inevitable to mine useful information from this vast amount of data. The mined information can be used in various spheres; from day to day web activities like the prediction of next web pages, serving better advertisements, to biological areas like genomic data analysis etc. A rough set based clustering of sequential data was proposed by Kumar et al recently. They defined and used a measure, called Sequence and Set Similarity Measure to determine similarity in data. However, we have observed that this measure does not reflect some important characteristics of sequential data. As a result, in this paper, we used the fuzzy set technique to introduce a similarity measure, which we termed as Kernel and Set Similarity Measure to find the similarity of sequential data and generate overlapping clusters. For this purpose, we used exponential string kernels and Jaccard's similarity index. The new similarity measure takes an account of the order of items in the sequence as well as the content of the sequential pattern. In order to compare our algorithm with that of Kumar et al, we used the MSNBC data set from the UCI repository, which was also used in their paper. As far as our knowledge goes, this is the first fuzzy clustering algorithm for sequential data.
  • 关键词:Clustering;Fuzzy Clustering;Sequence mining;Similarity measures;Pattern mining
国家哲学社会科学文献中心版权所有