首页    期刊浏览 2024年12月01日 星期日
登录注册

文章基本信息

  • 标题:GDC 2: Compression of large collections of genomes
  • 本地全文:下载
  • 作者:Sebastian Deorowicz ; Agnieszka Danek ; Marcin Niemiec
  • 期刊名称:Scientific Reports
  • 电子版ISSN:2045-2322
  • 出版年度:2015
  • 卷号:5
  • 期号:1
  • DOI:10.1038/srep11565
  • 语种:English
  • 出版社:Springer Nature
  • 摘要:The fall of prices of the high-throughput genome sequencing changes the landscape of modern genomics. A number of large scale projects aimed at sequencing many human genomes are in progress. Genome sequencing also becomes an important aid in the personalized medicine. One of the significant side effects of this change is a necessity of storage and transfer of huge amounts of genomic data. In this paper we deal with the problem of compression of large collections of complete genomic sequences. We propose an algorithm that is able to compress the collection of 1092 human diploid genomes about 9,500 times. This result is about 4 times better than what is offered by the other existing compressors. Moreover, our algorithm is very fast as it processes the data with speed 200 MB/s on a modern workstation. In a consequence the proposed algorithm allows storing the complete genomic collections at low cost, e.g., the examined collection of 1092 human genomes needs only about 700 MB when compressed, what can be compared to about 6.7 TB of uncompressed FASTA files. The source code is available at http://sun.aei.polsl.pl/REFRESH/index.php?page=projects&project=gdc&subpage=about .
国家哲学社会科学文献中心版权所有