Large time difference between NEB images

15 views
Skip to first unread message

Sofia Scozziero

unread,
Sep 15, 2026, 5:32:16 PM (3 days ago) Sep 15
to cp2k
Hi everyone, 

I've recently started working with CP2K and I'm running into the following issue. When doing NEB calculations some images take significantly more time per OT step than others, I'm not talking about number of SCF steps to converge but wall time per OT step, which is up to 15x slower for some of the images (tipically I've found issues with maybe 2-3 images out of 12 per run). Now a 2x-3x difference is manageable but when it gets past the 5x its problematic. This has been happening running in a cluster on 3-4 nodes of 64 cpus each, and I'm running purely on MPI with this

mpirun -np $SLURM_NTASKS --bind-to core cp2k.psmp -i neb.inp -o neb.out

and open mp threads set explicitly to 1.

I've brought this up to my system admins and they haven't been able to detect any issues like over subscription, load balancing or any of the common culprits. I'm clueless on what could be happening. On the physics side everything seems ok... I'm not getting any unreasonable energies or gaps, the total charge and spin also seem ok.

I'm attaching my input (.inp file) and a shortened txt file with the lines I'm more concerned about. I tested the convergence parameters vs a plane wave calculation and with successive cp2k calculations so I believe they should be Ok.

Extra note: I've tried on 1 and 2 nodes with a smaller system and this problem didn't come up, granted being a different system it is not the same problem.

Any help is welcome, and thanks in advance,

Sofía
neb.inp
cp2k_neb.txt

Frederick Stein

unread,
Sep 16, 2026, 3:04:13 AM (2 days ago) Sep 16
to cp2k
Dear Sofia,
I am not familiar with NEB-calculations but it may help if you could post an output file/the output files of a calculation that actually finished (to have the CP2K timing report at the very end) to help us investigating your bottleneck.
Best,
Frederick

Sofia Scozziero

unread,
Sep 16, 2026, 10:33:39 AM (2 days ago) Sep 16
to cp2k
Hello Frederick, thanks for your reply,

I don't have an actual finished run for my production system, the time to converge is exceeding my cluster's time limits comfortably and even restarting stalled runs is taking too long. I can try to get one done and give you an update once its finished, or I could try to reproduce the issue on a smaller system (albeit it will not be the exact same set up). As of right now I can share the logs of the stalled runs (the NEB output log plus the images' SCF logs).

Let me know which option you think is more workable.
Best,
Sofía
Reply all
Reply to author
Forward
0 new messages