首页 / 资料库 / 文献详情

MapReduce in the Cloud: Data-Location-Aware VM Scheduling

TungNguyenWeisongShi

2013Acta Scientiarum Naturalium Universitatis SunyatseniComputer Science被引 3

出版方页面 →

摘要

We have witnessed the fast-growing deployment of Hadoop,an open-source implementation of the MapReduce programming model,for purpose of data-intensive computing in the cloud.However,Hadoop was not originally designed to run transient jobs in which us ers need to move data back and forth between storage and computing facilities.As a result,Hadoop is inefficient and wastes resources when operating in the cloud.This paper discusses the inefficiency of MapReduce in the cloud.We study the causes of this inefficiency and propose a solution.Inefficiency mainly occurs during data movement.Transferring large data to computing nodes is very time-con suming and also violates the rationale of Hadoop,which is to move computation to the data.To address this issue,we developed a dis tributed cache system and virtual machine scheduler.We show that our prototype can improve performance significantly when run ning different applications.

引用本文(GB/T 7714)

Tung, Nguyen, Weisong, 等. MapReduce in the Cloud: Data-Location-Aware VM Scheduling[J]. Acta Scientiarum Naturalium Universitatis Sunyatseni, 2013.

引文网络

参考文献与被引分析加载中…

本站仅收录题录与摘要供学习参考,全文版权归属出版方;如有侵权请联系我们删除。