阅读背景:

将spark的MLLib例程与pandas数据帧一起使用

来源:互联网 

I have a pretty big data set (~20GB) stored on disk as Pandas/PyTables HDFStore, and I want to run random forests and boosted trees on it. Trying to do it on my local system takes forever, so I was thinking of farming it out to a spark cluster I have access to and instead using MLLib routines. I have a pretty big data set (~20GB) stored on




你的当前访问异常,请进行认证后继续阅读剩余内容。

分享到: