阅读背景:

具有大量数据的内存处理引擎有什么好处?

来源:互联网 

Spark performs the best if the dataset fits in memory, in case the dataset doesn't fit, it will use the disk and so it is as fast as hadoop. Let's assume that I m dealing with Tera/Peta bytes of data. with a small cluster. Obviously, there is no way to fit it in the memory. My observation is, in the big data era most of the dataset are in Giga bytes if not more. Spark performs the best if the dataset fits in




你的当前访问异常,请进行认证后继续阅读剩余内容。

分享到: