首页    期刊浏览 2025年02月20日 星期四
登录注册

文章基本信息

  • 标题:Available techniques in hadoop small file issue
  • 本地全文:下载
  • 作者:M. B. Masadeh ; M. S. Azmi ; S. S. S. Ahmad
  • 期刊名称:International Journal of Electrical and Computer Engineering
  • 电子版ISSN:2088-8708
  • 出版年度:2020
  • 卷号:10
  • 期号:2
  • 页码:2097-2101
  • DOI:10.11591/ijece.v10i2.pp2097-2101
  • 出版社:Institute of Advanced Engineering and Science (IAES)
  • 摘要:Hadoop is an optimal solution for big data processing and storing since being released in the late of 2006, hadoop data processing stands on master-slaves manner [1] that’s splits the large file job into several small files in order to process them separately, this technique was adopted instead of pushing one large file into a costly super machine to insights some useful information. Hadoop runs very good with large file of big data, but when it comes to big data in small files it could facing some problems in performance, processing slow down, data access delay, high latency and up to a completely cluster shutting down [2]. In this paper we will high light on one of hadoop’s limitations, that’s affects the data processing performance, one of these limits called “big data in small files” accrued when a massive number of small files pushed into a hadoop cluster which will rides the cluster to shut down totally. This paper also high light on some native and proposed solutions for big data in small files, how do they work to reduce the negative effects on hadoop cluster, and add extra performance on storing and accessing mechanism.
  • 关键词:big data in small files;big data;EHDFS;hadoop;HAR;HDFS;mapreduce;namenode;Nhar;sequance file;
国家哲学社会科学文献中心版权所有