Running Jobs with SLURM
Submitting, monitoring, and troubleshooting interactive and batch SLURM jobs.
This page describes how workloads are scheduled on NowWhat and how to avoid the most common mistakes when submitting jobs.
Why SLURM Exists
SLURM is the workload manager of NowWhat. It decides where and when jobs run, enforces resource limits, and records execution metadata.
You do not run production workloads directly on the login node. You submit them to SLURM, and SLURM places them on the appropriate compute resources.
Core Terms
login node: the system where you connect, prepare files, and submit jobsjob: a scheduler-managed workloadbatch job: a non-interactive job submitted withsbatchinteractive job: an allocation started for an interactive shell or command, usually withsrunpartition: a queue of resourcesallocation: the CPUs, memory, GPUs, and time assigned to a jobjob ID: the numeric identifier returned by SLURM
Standard Workflow
- prepare scripts and inputs on the login node
- inspect partitions and node status
- submit the workload with
sbatchor request an interactive allocation withsrun - monitor the job
- inspect logs and accounting information
- adjust resource requests before the next run
Inspect the System
List partitions:
sinfoShow node-level information:
sinfo -N -lShow your jobs:
squeue -u your_usernameInspect one job in detail:
scontrol show job JOBIDInspect accounting data:
sacct -j JOBID --format=JobID,JobName,Partition,State,ExitCode,Elapsed,MaxRSSInteractive Jobs with srun
Interactive jobs are useful for:
- debugging
- checking the runtime environment on a compute node
- launching short exploratory sessions
Example CPU allocation:
srun --partition=cpu --time=00:30:00 --ntasks=1 --cpus-per-task=4 --mem=8G --pty bashExample GPU allocation:
srun --partition=gpu --time=00:30:00 --ntasks=1 --cpus-per-task=4 --mem=16G --gres=gpu:nvidia_h200_nvl_1g.35gb:1 --pty bashWhen the allocation starts, your shell is running on a compute node. If the command is still running on the login node, you are not where you think you are.
Batch Jobs with sbatch
Batch jobs are the normal way to run work on the cluster.
Submit a script:
sbatch my-job.shIf submission succeeds, SLURM prints a job ID:
Submitted batch job 12345Essential Directives
The most common #SBATCH directives are:
--job-name: job label--partition: target queue--time: wall-time limit--nodes: number of nodes--ntasks: total number of tasks--ntasks-per-node: tasks on each node--cpus-per-task: CPU cores per task--mem: memory request--output: standard output file--error: standard error file--gres=gpu:nvidia_h200_nvl_1g.35gb:N: number of H200 MIG instances requested
Request only what your application needs. Over-requesting resources increases queue time. Under-requesting memory or time causes jobs to fail in avoidable ways.
CPU Job Example
Save a file cpu-job.sbatch:
#!/bin/bash
#SBATCH --job-name=cpu-example
#SBATCH --partition=cpu
#SBATCH --time=00:10:00
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4
#SBATCH --mem=4G
#SBATCH --output=slurm-%j.out
#SBATCH --error=slurm-%j.err
set -euo pipefail
module purge
module load python/3.11
echo "Job ID: $SLURM_JOB_ID"
echo "Host: $(hostname)"
echo "Start time: $(date)"
python - <<'EOF'
from math import sqrt
values = [sqrt(i) for i in range(1, 11)]
print(values)
EOFSubmit it with:
sbatch cpu-job.sbatchGPU Job Example
#!/bin/bash
#SBATCH --job-name=gpu-example
#SBATCH --partition=gpu
#SBATCH --gres=gpu:nvidia_h200_nvl_1g.35gb:1
#SBATCH --time=00:15:00
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4
#SBATCH --mem=16G
#SBATCH --output=slurm-%j.out
#SBATCH --error=slurm-%j.err
set -euo pipefail
module purge
module load cuda
module load python/3.11
echo "Job ID: $SLURM_JOB_ID"
echo "Host: $(hostname)"
nvidia-smiMPI Job Pattern
MPI applications typically require a compiler and MPI stack that were built to work together.
Example pattern:
#!/bin/bash
#SBATCH --job-name=mpi-example
#SBATCH --partition=cpu
#SBATCH --nodes=2
#SBATCH --ntasks-per-node=16
#SBATCH --time=00:20:00
#SBATCH --mem=0
#SBATCH --output=slurm-%j.out
set -euo pipefail
module purge
module load gcc
module load openmpi
scontrol show hostnames "$SLURM_JOB_NODELIST"
mpirun -np "$SLURM_NTASKS" ./my_mpi_program input.datIf you change compiler or MPI toolchain, rebuild the application accordingly.
Monitor Running Jobs
Useful commands:
squeue -u your_username
tail -f slurm-JOBID.out
tail -f slurm-JOBID.errFor completed or failed jobs:
sacct -j JOBID --format=JobID,State,ExitCode,Elapsed,MaxRSSCommon Job States
PD: pendingR: runningCG: completingCD: completedF: failedTO: timed outCA: cancelled
If a job is pending, inspect the reason field in squeue. Read it before resubmitting the same mistake twenty times in a row.
Quick Troubleshooting
If a job is pending:
- confirm the partition is correct
- verify requested time, memory, CPUs, and GPUs
- inspect the reason field with
squeue - remove a job (or jobs) from the queue with
scancel <JOB_ID>
If a job fails:
- inspect
slurm-%j.outandslurm-%j.err - inspect accounting data with
sacct - verify input paths
- verify loaded modules
- verify that the application is compatible with the requested resources
Rules That Matter
- never run heavy workloads directly on the login node
- test new scripts with small inputs first
- keep output and error logs
- record the job ID for every important run
- use interactive allocations for debugging rather than improvising on shared infrastructure