文章基本信息

标题：Fuzzy Clustering of Sequential Data
本地全文：下载
作者：B.K. Tripathy ; Rahul
期刊名称：International Journal of Intelligent Systems and Applications
印刷版ISSN：2074-904X
电子版ISSN：2074-9058
出版年度：2019
卷号：11
期号：1
页码：43-54
DOI：10.5815/ijisa.2019.01.05
出版社：MECS Publisher
摘要：With the increase in popularity of the Internet and the advancement of technology in the fields like bioinformatics and other scientific communities the amount of sequential data is on the increase at a tremendous rate. With this increase, it has become inevitable to mine useful information from this vast amount of data. The mined information can be used in various spheres; from day to day web activities like the prediction of next web pages, serving better advertisements, to biological areas like genomic data analysis etc. A rough set based clustering of sequential data was proposed by Kumar et al recently. They defined and used a measure, called Sequence and Set Similarity Measure to determine similarity in data. However, we have observed that this measure does not reflect some important characteristics of sequential data. As a result, in this paper, we used the fuzzy set technique to introduce a similarity measure, which we termed as Kernel and Set Similarity Measure to find the similarity of sequential data and generate overlapping clusters. For this purpose, we used exponential string kernels and Jaccard's similarity index. The new similarity measure takes an account of the order of items in the sequence as well as the content of the sequential pattern. In order to compare our algorithm with that of Kumar et al, we used the MSNBC data set from the UCI repository, which was also used in their paper. As far as our knowledge goes, this is the first fuzzy clustering algorithm for sequential data.
关键词：Clustering;Fuzzy Clustering;Sequence mining;Similarity measures;Pattern mining