首页    期刊浏览 2024年11月26日 星期二
登录注册

文章基本信息

  • 标题:Effective Job Execution in Hadoop Over Authorized Deduplicated Data
  • 本地全文:下载
  • 作者:Sachin Arun Thanekar ; K. Subrahmanyam ; A.B. Bagwan
  • 期刊名称:Webology
  • 印刷版ISSN:1735-188X
  • 出版年度:2020
  • 卷号:17
  • 期号:2
  • 页码:430-444
  • DOI:10.14704/WEB/V17I2/WEB17043
  • 出版社:University of Tehran
  • 摘要:Existing Hadoop treats every job as an independent job and destroys metadata of preceding jobs. As every job is independent, again and again it has to read data from all Data Nodes. Moreover relationships between specific jobs are also not getting checked. Lack of Specific user identities creation and forming groups, managing user credentials are the weaknesses of HDFS. Due to which overall performance of Hadoop becomes very poor. So there is a need to improve the Hadoop performance by reusing metadata, better space management, better task execution by checking deduplication and securing data with access rights specification. In our proposed system, task deduplication technique is used. It checks the similarity between jobs by checking block ids. Job metadata and data locality details are stored on Name Node which results in better execution of job. Metadata of executed jobs is preserved. Thus by preserving job metadata re computations time can be saved. Experimental results show that there is an improvement in job execution time, reduced storage space. Thus, improves Hadoop performance.
  • 关键词:Hadoop; H2Hadoop; Deduplication; HDFS; Storage;
国家哲学社会科学文献中心版权所有