Extend GPU multigrid backends and tests to 3D

147 views
Skip to first unread message

Andreas Iberl

unread,
Sep 11, 2026, 11:15:34 AMSep 11
to basil...@googlegroups.com
Dear all,

please find attached a patch adding multigrid GPU implementation in
three dimensions.

Several pieces of the GPU infrastructure have been hard-coded for
two-dimensional grid, such as point and coordinates types, data
indexing, kernel launch geometry, boundaries, reductions and
backend-specific data transfers. The current 2D implementations are kept
with the use of "#if dimension" conditions. Tests are added for the new
code with the OpenGL, CUDA and OpenCL backends.

It should be explicitly noted that I could not test the HIP backend
(neither at AMD or NVIDIA graphics cards), since my cluster doesn't have
the required libraries installed. Also note the convergence test failure
in the periodic.c test case on GPUs, which I attribute to the limitation
to single precision on GPUs (maybe the tolerance could be reduced at
GPUs to get rid of the warning, currently I just leave it as is).

Best regards,
Andreas
GPU3D

Andreas Iberl

unread,
Sep 17, 2026, 10:45:57 AMSep 17
to basilisk-fr
Dear all,

since my goal is to include also the 5x5 stencil (and the 5x5x5 stencil in 3D) boundary conditions, I rewrote a specific part of the GPU3D code, more specifically in "src/grid/gpu/grid.h". My initial implementation was somewhat redundant due to two branches for 2D and 3D. This patch unifies them, which makes it easier and shorter to implement the boundaries for the 5x5 stencil. Still, I couldn't test it with HIP, but the other tests passed (both 2D and 3D).

Best regards,
Andreas
UnifiedBCbranches

Stephane Popinet

unread,
Sep 18, 2026, 10:24:13 AMSep 18
to basil...@googlegroups.com
Hi Andreas,

Thanks a lot for the patches, following our review this morning I have
pushed all your patches, plus a few minor changes and results for the 3D
isotropic turbulence case, see:

https://basilisk.fr/src/?history

The 3D isotropic turbulence case in CUDA on the RTX4090 runs twice as
fast as 512 MPI cores on occigen.

https://basilisk.fr/src/examples/isotropic.c

well done!

Stephane

PS: Note also that the OpenGL version can run in 512^3 on the RTX4090,
unlike the CUDA version. I have put a note and explanation in the
documentation of isotropic.c. I am working on a patch.



Stephane Popinet

unread,
Sep 18, 2026, 1:35:42 PMSep 18
to basil...@googlegroups.com
> PS: Note also that the OpenGL version can run in 512^3 on the RTX4090,
> unlike the CUDA version. I have put a note and explanation in the
> documentation of isotropic.c. I am working on a patch.

I have just fixed this issue with this patch:

https://basilisk.fr/src/?changes=20260918172719


Andreas Iberl

unread,
Sep 22, 2026, 12:27:24 PMSep 22
to basilisk-fr
Dear all,

I have extended the multigrid GPU also to 1D (see the patch attached).

Keep in mind, that I still cannot test this with the HIP backend, and currently only two simple test cases are included in the "gpu1D.tests".

Best regards,

Andreas
GPU1D
Reply all
Reply to author
Forward
0 new messages