首页    期刊浏览 2024年11月25日 星期一
登录注册

文章基本信息

  • 标题:A document classifier for medicinal chemistry publications trained on the ChEMBL corpus
  • 本地全文:下载
  • 作者:George Papadatos ; Gerard JP van Westen ; Samuel Croset
  • 期刊名称:Journal of Cheminformatics
  • 印刷版ISSN:1758-2946
  • 电子版ISSN:1758-2946
  • 出版年度:2014
  • 卷号:6
  • 期号:1
  • 页码:40
  • DOI:10.1186/s13321-014-0040-8
  • 语种:English
  • 出版社:BioMed Central
  • 摘要:The large increase in the number of scientific publications has fuelled a need for semi- and fully automated text mining approaches in order to assist in the triage process, both for individual scientists and also for larger-scale data extraction and curation into public databases. Here, we introduce a document classifier, which is able to successfully distinguish between publications that are `ChEMBL-like’ (i.e. related to small molecule drug discovery and likely to contain quantitative bioactivity data) and those that are not. The unprecedented size of the medicinal chemistry literature collection, coupled with the advantage of manual curation and mapping to chemistry and biology make the ChEMBL corpus a unique resource for text mining. The method has been implemented as a data protocol/workflow for both Pipeline Pilot (version 8.5) and KNIME (version 2.9) respectively. Both workflows and models are freely available at: ftp://ftp.ebi.ac.uk/pub/databases/chembl/text-mining . These can be readily modified to include additional keyword constraints to further focus searches. Large-scale machine learning document classification was shown to be very robust and flexible for this particular application, as illustrated in four distinct text-mining-based use cases. The models are readily available on two data workflow platforms, which we believe will allow the majority of the scientific community to apply them to their own data.
  • 关键词:Machine learning ; Triage ; Curation ; Document classification
国家哲学社会科学文献中心版权所有