首页    期刊浏览 2024年11月25日 星期一
登录注册

文章基本信息

  • 标题:Large-scale virtual screening on public cloud resources with Apache Spark
  • 本地全文:下载
  • 作者:Marco Capuccini ; Laeeq Ahmed ; Wesley Schaal
  • 期刊名称:Journal of Cheminformatics
  • 印刷版ISSN:1758-2946
  • 电子版ISSN:1758-2946
  • 出版年度:2017
  • 卷号:9
  • 期号:1
  • 页码:15
  • DOI:10.1186/s13321-017-0204-4
  • 语种:English
  • 出版社:BioMed Central
  • 摘要:Structure-based virtual screening is an in-silico method to screen a target receptor against a virtual molecular library. Applying docking-based screening to large molecular libraries can be computationally expensive, however it constitutes a trivially parallelizable task. Most of the available parallel implementations are based on message passing interface, relying on low failure rate hardware and fast network connection. Google’s MapReduce revolutionized large-scale analysis, enabling the processing of massive datasets on commodity hardware and cloud resources, providing transparent scalability and fault tolerance at the software level. Open source implementations of MapReduce include Apache Hadoop and the more recent Apache Spark. We developed a method to run existing docking-based screening software on distributed cloud resources, utilizing the MapReduce approach. We benchmarked our method, which is implemented in Apache Spark, docking a publicly available target receptor against $$\sim $$ ∼ 2.2 M compounds. The performance experiments show a good parallel efficiency (87%) when running in a public cloud environment. Our method enables parallel Structure-based virtual screening on public cloud resources or commodity computer clusters. The degree of scalability that we achieve allows for trying out our method on relatively small libraries first and then to scale to larger libraries. Our implementation is named Spark-VS and it is freely available as open source from GitHub ( https://github.com/mcapuccini/spark-vs ). Graphical abstract .
  • 关键词:Virtual screening ; Docking ; Cloud computing ; Apache Spark
国家哲学社会科学文献中心版权所有