Evolutive algorithms have an intrinsic stochastic nature, therefore they make large use of
random numbers generators, for instance the
C/C++ rand() function. Anyway, when GPGPU comes into play, using random numbers could be tricky.
The naive solution of create and pour them into GPU's global memory is a bad idea, because of the huge
bandwidth that would be wasted. Therefore, it takes a device-bound
generator. Nvidia's CURAND library makes it easy to
generate random numbers directly
inside CUDA kernels.
A bit of theory...
Obviously, it is impossible to generate
really random numbers on a deterministic machine. Any random function just applies some kind of transformation on another number, determining a succession that
looks like a random one. Anyway, two successions starting from the same number (the seed) will be completely identical. If our kernels start from the same seed, regardless the algorithm, they will produce the same results.
They must be different. Moreover, you have to
store the previous numbers (let's call them
global states) in order to produce the next ones during the following execution of
each kernel. Fortunately, CUDA allows the storage and update of our partials directly inside GPU's memory.
...and some sample code
I post the example code - comments will follow.
__global__ void setup_kernel ( curandState * state, unsigned long seed )
{
int id = threadIdx.x;
curand_init ( seed, id, 0, &state[id] );
}
__global__ void generate( curandState* globalState )
{
int ind = threadIdx.x;
curandState localState = globalState[ind];
float RANDOM = curand_uniform( &localState );
globalState[ind] = localState;
}
int main( int argc, char** argv)
{
dim3 tpb(N,1,1);
curandState* devStates;
cudaMalloc ( &devStates, N*sizeof( curandState ) );
// setup seeds
setup_kernel <<< 1, tpb >>> ( devStates, time(NULL) );
// generate random numbers
generate <<< 1, tpb >>> ( devStates );
return 0;
}
Obviously, you have to include not only CUDA's
includes and libraries, but CURAND kernel's as well (curand_kernel.h). Let's see what happens in the code, starting from main.
First, we create a
curandState pointer, that will point to our
global states.
Kernel setup_kernel will invoke
curand_init(), that takes some seeds (I used the seconds since the Epoch, but it's free to the user) and sets the global states.
Generate() kernel creates N threads. Each thread will have its own random floating point number, by using a local copy of its global state. In this case, the random number it's sampled from (0,1) with a
uniform probability, but CURAND gives the possibility to sample from a standard normal distribution as well.
Finally, we store the new seed into the global memory, and return.