阅读背景:

CUDA如何获得网格,块,线程大小和parallalize非方矩阵计算

来源:互联网 

I am new to CUDA and need help understanding some things. I need help parallelizing these two for loops. Specifically how to setup the dimBlock and dimGrid to make this run faster. I know this looks like the vector add example in the sdk but that example is only for square matrices and when I try to modify that code for my 128 x 1024 matrix it doesn't work properly. I am new to CUDA and need help understanding so




你的当前访问异常,请进行认证后继续阅读剩余内容。

分享到: