首页    期刊浏览 2024年10月06日 星期日
登录注册

文章基本信息

  • 标题:Lung Cancer Survival Prediction using Ensemble Data Mining on Seer Data
  • 本地全文:下载
  • 作者:Ankit Agrawal ; Sanchit Misra ; Ramanathan Narayanan
  • 期刊名称:Scientific Programming
  • 印刷版ISSN:1058-9244
  • 出版年度:2012
  • 卷号:20
  • 期号:1
  • 页码:29-42
  • DOI:10.1155/2012/920245
  • 出版社:Hindawi Publishing Corporation
  • 摘要:

    We analyze the lung cancer data available from the SEER program with the aim of developing accurate survival prediction models for lung cancer. Carefully designed preprocessing steps resulted in removal/modification/splitting of several attributes, and 2 of the 11 derived attributes were found to have significant predictive power. Several supervised classification methods were used on the preprocessed data along with various data mining optimizations and validations. In our experiments, ensemble voting of five decision tree based classifiers and meta-classifiers was found to result in the best prediction performance in terms of accuracy and area under the ROC curve. We have developed an on-line lung cancer outcome calculator for estimating the risk of mortality after 6 months, 9 months, 1 year, 2 year and 5 years of diagnosis, for which a smaller non-redundant subset of 13 attributes was carefully selected using attribute selection techniques, while trying to retain the predictive power of the original set of attributes. Further, ensemble voting models were also created for predicting conditional survival outcome for lung cancer (estimating risk of mortality after 5 years of diagnosis, given that the patient has already survived for a period of time), and included in the calculator. The on-line lung cancer outcome calculator developed as a result of this study is available at http://info.eecs.northwestern.edu:8080/LungCancerOutcomeCalculator/.

  • 关键词:Ensemble data mining; lung cancer; predictive modeling; outcome calculator
国家哲学社会科学文献中心版权所有