[slurm-dev] restarting checkpoint after "slurm_checkpoint_vacate" API call

2 views
Skip to first unread message

Manuel Rodríguez Pascual

unread,
Jan 29, 2015, 5:55:26 AM1/29/15
to slurm-dev
Good morning all,

I am facing a problem when using slurm.h API to manage checkpoints.

What I want to do is to checkpoint a running task, shut it down, and then restore it somewhere (in the same node or another one).

slurm.conf is configured with:
CheckpointType=checkpoint/blcr 
JobCheckpointDir=/home/slurm/ 

My code, after initial verifications goes like:

" int max_wait = 60;
if (slurm_checkpoint_vacate(opt.jobid, opt.stepid, max_wait, "/home/slurm/") != 0)
_show_error_and_exit();

//just in case it is still not stopped
slurm_kill_job(opt.jobid, 9,  KILL_JOB_ARRAY) ;

char* checkpoint_location = "/home/slurm";
if ( slurm_checkpoint_restart(opt.jobid, opt.stepid, 0,  checkpoint_location) != 0)
_show_error_and_exit();
"

The errno and error message I get is:  

"2011: Duplicate job id"

and this content in slurmctld:

"re-use active job_id 2570
slurmctld: _slurm_rpc_checkpoint restart 2570: Duplicate job id
"

if I do instead 

"if ( slurm_checkpoint_restart(opt.jobid +1 , opt.stepid, 0,  checkpoint_location) != 0)"

The errno and error message I get is:  

"2: No such file or directory"

and this content in slurmctld:

"No job ckpt file (/home/slurm//2571.ckpt) to read
slurmctld: _slurm_rpc_checkpoint restart 2570: No such file or directory
"
Which is right, the file does not exist, so of course it cannot start it. However if I specify "/home/slurm/2570/" as image_dir, the folder created by the checkpoint_vacate call, the result is the same. 

Besides that, it seems that the input parameter  " image_dir" is not read, only the default parameter. So if i set my "checkpoint_location" to "/a/b/c" for example, the output log returns the same error, showing that it is trying to find the image in "/home/slurm". 
   

So, this said, have you got any help or suggestion on how to deal with checkpoints with Slurm API? Am I doing something wrong? Is there any working example I can see? Should I be using other call instead of these ones?


Thanks for your help. Best regards,


Manuel

je...@schedmd.com

unread,
Jan 29, 2015, 11:53:23 AM1/29/15
to slurm-dev

The slurm_checkpoint_vacate() triggers a checkpoint operation, which
could take minutes to complete. You can't slurm_checkpoint_restart()
the job until the checkpoint operation and accounting are complete.
Adding some sleep/retry logic should do what you want.
--
Morris "Moe" Jette
CTO, SchedMD LLC
Commercial Slurm Development and Support

Manuel Rodríguez Pascual

unread,
Jan 30, 2015, 9:54:45 AM1/30/15
to slurm-dev
it works, thanks very much Moe :)


Just for the sake of completion (in case someone ends up googling for this), this is my complete code for a basic task checkpoint/restart. 


int max_wait = 60;
char* checkpoint_location = "/home/slurm";

if (slurm_checkpoint_vacate(opt.jobid, opt.stepid, max_wait, checkpoint_location) != 0){
slurm_perror("Error checkpointing: ");
exit(errno);
        }

int i = 0;
while ( slurm_checkpoint_restart(opt.jobid , opt.stepid, 0,  checkpoint_location) != 0) {
sleep (10);
i = i + 10;
slurm_perror("Error: ");
printf (". Still not posible to restart. Time: %i\n", i);

}

printf ("job has been restarted\n");



2015-01-29 17:53 GMT+01:00 <je...@schedmd.com>:

The slurm_checkpoint_vacate() triggers a checkpoint operation, which could take minutes to complete. You can't slurm_checkpoint_restart() the job until the checkpoint operation and accounting are complete. Adding some sleep/retry logic should do what you want.



--
Dr. Manuel Rodríguez-Pascual
skype: manuel.rodriguez.pascual
phone: (+34) 913466173 // (+34) 679925108
 
CIEMAT-Moncloa
Edificio 22, desp. 1.25
Avenida Complutense, 40 
28040- MADRID
SPAIN
Reply all
Reply to author
Forward
0 new messages