We have resolved the issue though we are continuing to monitoring our systems.
Posted Aug 25, 2026 - 08:54 EDT
Update
Discovery is now stable. We are waiting for jobs to finish on about 30 more nodes so we can reboot those and get them ready for new jobs. Until those 30 are reintegrated, you may see increased wait times in the queue for resources to be free.
Posted Aug 21, 2026 - 13:58 EDT
Monitoring
We experienced a large cluster issue while updating the Slurm database host to a newer version. This resulted in some jobs becoming stuck in the COMPLETING (CG) state and caused broader job instability.
We have since upgraded the cluster to the same Slurm version as the database and are rebooting affected nodes to clear jobs stuck in the CG state. The cluster is recovering, and stability is improving.
We appreciate your patience as we continue working to fully resolve the issue and return the cluster to a steady state.
Please check the status and results of your recent jobs. If you experienced job failures or lost compute time as a result of this outage, please reach out to research.computing@dartmouth.edu.
Where helpful, we can temporarily increase your available resources to help account for lost compute time.
Posted Aug 21, 2026 - 10:21 EDT
Update
Weve applied an update to the scheduler. All new jobs will begin as normal however as running jobs finish, we will need to reboot compute nodes.
Posted Aug 21, 2026 - 09:00 EDT
Update
HPC systems remain operational but degraded due to an ongoing issue. Troubleshooting is paused for tonight and will resume tomorrow morning. We appreciate your patience.
Posted Aug 20, 2026 - 19:12 EDT
Identified
The scheduler on Discovery(Slurm) is behaving erratically and we have a ticket open with the vendor. We will update when we know more.