Leonardo: service restored – slurm controller issue resolved

  1. /
  2. HPC Center news
  3. /
  4. Leonardo: service restored –...

Dear Users,

The issue has been resolved, and it is now possible to submit jobs and query the queue status normally.

The problem was caused by instabilities on the local filesystem used by the Slurm service. We will continue monitoring the system closely over the coming hours to ensure stability.

Please note that when the controller was restored, an internal recovery procedure failed, causing all jobs that had started running on the Booster partitions to be terminated with the following message:

srun: PrologSlurmctld failed, job killed

This additional issue has now been fixed, and full production service has been restored across the entire cluster.

We apologize for the inconvenience and thank you for your understanding.

Best regards,

HPC User Support – CINECA