Skip to main content Skip to secondary navigation

Slurm Primer

Main content start

What is Slurm? 

Slurm is a workload manager. Carina is a shared computing environment, and Slurm's role is to take in everyone's requests for resources and calculate the most efficient way to allocate them. 

The process of submitting a job to Slurm via sbatch on the login node.

Slurm Jobs

A job is a set of tasks that is submitted to Slurm, along with a request for the appropriate computing resources. A job is usually submitted in the form of a script, called an sbatch script. 

Requesting Resources

Before asking Slurm for resources, you need to determine how much compute power you need, and for how long. Slurm will only give you what you ask for, so if you ask for too little, your job may not finish before Slurm reclaims your resources. If you ask for too much, Slurm will hold your job in the queue until all requested resources are available, which could slow down your work.

Resource estimation can be tricky. We periodically offer a workshop on it, watch our Classes page for details.

Slurm has some reporting functions which can help you understand your usage.

Want to jump right in?

Try Slurm-O-Matic, the handy dandy tool that writes your sbatch script for you!

Carina Resources

Carina has three Slurm queues, also called partitions. The differences are whether or not GPUs can be requested, and the maximum runtime a job can have.

 Most of the time, you will use the  normal or gpu partitions, which would be called this way in your sbatch script: 

#!/bin/bash
#SBATCH --partition=normal 

The column marked -C is the value to use to request that specific hardware, using the --constraint or -C property. Specifying hardware may cause your job to wait in the queue longer than a job that accepts the next-available GPU.

#!/bin/bash
#SBATCH --partition=gpu
#SBATCH --constraint=GPU_SKU:A100_PCIE

Standard Queues

Queue NameCPUsMemoryNodesGPUsMax Time-C
normalDual Intel Xeon Gold 6130 (16C 2.1GHz)370GB2 2 Days 
gpuDual Intel Xeon Gold 6130 (16C 2.1GHz)350GB2Dual P1002 DaysGPU_SKU:P100_PCIE
gpuDual Intel Xeon Silver 4114 (10C 2.2GHz)250GB5Quad NVIDIA Tesla V1002 DaysGPU_SKU:V100_PCIE
gpuDual Intel Xeon Gold 6330 (56C 2GHz)250GB4Quad NVIDIA A100 with 14TB of local scratch2 DaysGPU_SKU:A100_PCIE

We have a special queue, gpu-long, that is set aside for jobs that need to run longer than two days. The maximum time is five days.

To use this queue, your sbatch file should include the partition name and the requested time. This would request five days on the gpu-long partition:

#!/bin/bash  
#SBATCH --partition=gpu-long  
#SBATCH --time=5-0

Specialty Queue

Queue NameCPUsMemoryNodesGPUsMax Time-C
gpu-longDual Intel Xeon Silver 4114 (20C 2.2GHz)250GB2Quad NVIDIA Tesla V1005 DaysGPU_SKU:V100_PCIE
gpu-longDual Intel Xeon Gold 6330 (56C 2GHz)250GB2Quad NVIDIA A100 with 14TB of local scratch5 DaysGPU_SKU:A100_PCIE

sbatch

At the most basic level, a sbatch script requests resources from Slurm and then tells it what to do with the resources. It's a shell or Bash script which follows a few rules. 

Any lines starting with #SBATCH are resource requests and other Slurm options. A complete list of options can be found in the sbatch documentation.

Job Submission

Once the submission script is written properly, you can submit it to the scheduler with the sbatch command. If the submission is successful, it will return a job id. If there are errors in the script, it will return error messages.

$ sbatch CatStats.sh
Submitted batch job 1377

If you have not specified a name for the job output using the --output option, you will find it in the working directory with the name slurm-<job id>.out

Sample Slurm sbatch Script

This script is overloaded; you may not need to use all of these settings. 


#!/bin/bash
# -------SLURM Parameters-------
#SBATCH --partition=gpu
#SBATCH --mem=1G
#SBATCH --gres=gpu:1
# Define how long the job will run 
#SBATCH --time=12:15:00 #d-hh:mm:ss
# Give your job a name
#SBATCH --job-name=CatStats
# Working directory
#SBATCH --chdir=/share/pi/drevil
# Name your output files
#SBATCH --output=CatStats.txt
#SBATCH --error=CatFail.txt
# -------Load Modules-------
 module load anaconda/2024.06
# -------Commands-------
# this loads a previously-created 
# conda environment
conda activate catdata
# run the python script
python catdata.py

Try Slurm-O-Matic to build better scripts, faster!

srun

Usually, you will submit a job via a sbatch script, which will both request your resources and do your computations. You can submit the job, close your laptop, and return later to see the results. 

Another way to interact with Carina is through an interactive command-line session using srun.  You can load modules, access your data, and more. It is especially good for debugging and benchmarking your script in preparation for submitting it via sbatch.

An interactive session with srun will end if you close your terminal window. (there is a small asterisk on that statement; there are workarounds but generally, do not trust a srun connection to remain open).

To start an interactive session in your terminal, type:

srun --pty bash

This will give you a compute node in the normal partition with one core for two hours (aka the default on Carina).

This command requests 2 GPUs on a compute node in the gpu partition:

srun --pty -p gpu --gres=gpu:2 bash

A quick breakdown of the command:

srun --pty -p gpu --gres=gpu:2 bash

Most sbatch commands are also used in srun; for example, you can request specific hardware using --constraint. This will request two Quad NVIDIA A100 GPUs, using the value from the partition list above.

srun --pty -p gpu --gres=gpu:2 --constraint=GPU_SKU:A100_PCIE bash

This is a very basic example of srun, to see more of what it can do, consult the official documentation.

If you finish your session before the requested time is up, you can either close your terminal or use scancel to end the job.

squeue

squeue is a powerful tool for monitoring your activity on Carina. To see the status of your jobs, use the command squeue -u <sunetid>. You will see something like this:

JOBIDPARTITIONNAMEUSERSTATETIMENODESNODELIST
80246normalbashsunetidR2:331secure-3

The usual progression of a job is:

If the job was submitted via sbatch, it will start in the PENDING (P) state until Slurm is ready to release it. Usually, if a job spends a long time in this state, it is because the requested resources are in high demand. Please be patient, and review your resource request to see if it can be reduced. 

When resources become available and the job has sufficient priority, an allocation is created for it and it moves to the RUNNING (R) state. This state can also indicate that you have an interactive session running.

When the job is finished, its state will be COMPLETED (C).  Output files will be available for viewing. 

If the job did not complete successfully, its state will be FAILED (F). Have you been staring at the screen for hours? Maybe get up and have a snack before you try again. It will work eventually.

These are the most common states on Carina, but there are more possible states in the Slurm documentation.

scancel

scancel is a useful utility for ending jobs. The typical use case is ending a job started with sbatch or an interactive session started with srun.

We used squeue -u to show our running job with the ID 80246. To end this job, the command would be

scancel 80246

It's also possible to close all of your own jobs using

scancel -u <sunetid>

If you are running an interactive session using srun the job will end and the resources will be released when you close your terminal window or your laptop drops the connection. If you are finished with the session but still connected to Carina, it's polite to use scancel to end the session.

Resource Utilization

Curious about your resource usage on Carina? Try the sreport function to see your cpu,memory,and gpu utilization statistics. This example returns information from December 1-31, 2024:

sreport cluster UserUtilizationByAccount -T GRES/gpu,cpu,Mem Start=2024-12-1T00:00:00 End=2024-12-31T23:59:59 user=<sunetid>

Real-time GPU Usage

This command works on a currently running job. Use squeue to find the job id; the STATE of the job must be R for this to work.

srun --jobid=<job id> --pty bash nvidia-smi