DBCSR-only GPU offload (OpenCL) on Intel PVC slower than CPU-only - expected?

26 views
Skip to first unread message

ganta.pra...@gmail.com

unread,
Aug 30, 2026, 12:21:19 PM (6 days ago) Aug 30
to cp2k
Dear all,

I am running CP2K with the OpenCL backend on Intel Data Center GPU Max 1550 (Ponte Vecchio) nodes on SuperMUC-NG Phase 2, offloading only the DBCSR sparse matrix-matrix multiplication library to the GPU while the rest of the workload remains on the CPU.

Benchmark results (below) show that the GPU-offloaded runs are consistently slower than the CPU-only runs. I assume that because only DBCSR is offloaded, the host-device data transfer and synchronization overhead outweighs the compute gain from the GPU 

Has this behavior of GPU offload underperforming CPU-only execution when only DBCSR is accelerated been reported before? I'd also appreciate pointers to published OpenCL benchmark results for DBCSR/CP2K on Intel GPUs.


Below are the results 
Screenshot From 2026-08-30 18-02-51.png

Have a nice day.

Thanks and Regards,
Prasanth.
Reply all
Reply to author
Forward
0 new messages