Running Jobs with SLURM

Submitting, monitoring, and troubleshooting interactive and batch SLURM jobs.

This page describes how workloads are scheduled on NowWhat and how to avoid the most common mistakes when submitting jobs.

Why SLURM Exists

SLURM is the workload manager of NowWhat. It decides where and when jobs run, enforces resource limits, and records execution metadata.

You do not run production workloads directly on the login node. You submit them to SLURM, and SLURM places them on the appropriate compute resources.

Core Terms

  • login node: the system where you connect, prepare files, and submit jobs
  • job: a scheduler-managed workload
  • batch job: a non-interactive job submitted with sbatch
  • interactive job: an allocation started for an interactive shell or command, usually with srun
  • partition: a queue of resources
  • allocation: the CPUs, memory, GPUs, and time assigned to a job
  • job ID: the numeric identifier returned by SLURM

Standard Workflow

  1. prepare scripts and inputs on the login node
  2. inspect partitions and node status
  3. submit the workload with sbatch or request an interactive allocation with srun
  4. monitor the job
  5. inspect logs and accounting information
  6. adjust resource requests before the next run

Inspect the System

List partitions:

sinfo

Show node-level information:

sinfo -N -l

Show your jobs:

squeue -u your_username

Inspect one job in detail:

scontrol show job JOBID

Inspect accounting data:

sacct -j JOBID --format=JobID,JobName,Partition,State,ExitCode,Elapsed,MaxRSS

Interactive Jobs with srun

Interactive jobs are useful for:

  • debugging
  • checking the runtime environment on a compute node
  • launching short exploratory sessions

Example CPU allocation:

srun --partition=cpu --time=00:30:00 --ntasks=1 --cpus-per-task=4 --mem=8G --pty bash

Example GPU allocation:

srun --partition=gpu --time=00:30:00 --ntasks=1 --cpus-per-task=4 --mem=16G --gres=gpu:nvidia_h200_nvl_1g.35gb:1 --pty bash

When the allocation starts, your shell is running on a compute node. If the command is still running on the login node, you are not where you think you are.

Batch Jobs with sbatch

Batch jobs are the normal way to run work on the cluster.

Submit a script:

sbatch my-job.sh

If submission succeeds, SLURM prints a job ID:

Submitted batch job 12345

Essential Directives

The most common #SBATCH directives are:

  • --job-name: job label
  • --partition: target queue
  • --time: wall-time limit
  • --nodes: number of nodes
  • --ntasks: total number of tasks
  • --ntasks-per-node: tasks on each node
  • --cpus-per-task: CPU cores per task
  • --mem: memory request
  • --output: standard output file
  • --error: standard error file
  • --gres=gpu:nvidia_h200_nvl_1g.35gb:N: number of H200 MIG instances requested

Request only what your application needs. Over-requesting resources increases queue time. Under-requesting memory or time causes jobs to fail in avoidable ways.

CPU Job Example

Save a file cpu-job.sbatch:

#!/bin/bash
#SBATCH --job-name=cpu-example
#SBATCH --partition=cpu
#SBATCH --time=00:10:00
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4
#SBATCH --mem=4G
#SBATCH --output=slurm-%j.out
#SBATCH --error=slurm-%j.err
 
set -euo pipefail
 
module purge
module load python/3.11
 
echo "Job ID: $SLURM_JOB_ID"
echo "Host: $(hostname)"
echo "Start time: $(date)"
 
python - <<'EOF'
from math import sqrt
 
values = [sqrt(i) for i in range(1, 11)]
print(values)
EOF

Submit it with:

sbatch cpu-job.sbatch

GPU Job Example

#!/bin/bash
#SBATCH --job-name=gpu-example
#SBATCH --partition=gpu
#SBATCH --gres=gpu:nvidia_h200_nvl_1g.35gb:1
#SBATCH --time=00:15:00
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4
#SBATCH --mem=16G
#SBATCH --output=slurm-%j.out
#SBATCH --error=slurm-%j.err
 
set -euo pipefail
 
module purge
module load cuda
module load python/3.11
 
echo "Job ID: $SLURM_JOB_ID"
echo "Host: $(hostname)"
nvidia-smi

MPI Job Pattern

MPI applications typically require a compiler and MPI stack that were built to work together.

Example pattern:

#!/bin/bash
#SBATCH --job-name=mpi-example
#SBATCH --partition=cpu
#SBATCH --nodes=2
#SBATCH --ntasks-per-node=16
#SBATCH --time=00:20:00
#SBATCH --mem=0
#SBATCH --output=slurm-%j.out
 
set -euo pipefail
 
module purge
module load gcc
module load openmpi
 
scontrol show hostnames "$SLURM_JOB_NODELIST"
mpirun -np "$SLURM_NTASKS" ./my_mpi_program input.dat

If you change compiler or MPI toolchain, rebuild the application accordingly.

Monitor Running Jobs

Useful commands:

squeue -u your_username
tail -f slurm-JOBID.out
tail -f slurm-JOBID.err

For completed or failed jobs:

sacct -j JOBID --format=JobID,State,ExitCode,Elapsed,MaxRSS

Common Job States

  • PD: pending
  • R: running
  • CG: completing
  • CD: completed
  • F: failed
  • TO: timed out
  • CA: cancelled

If a job is pending, inspect the reason field in squeue. Read it before resubmitting the same mistake twenty times in a row.

Quick Troubleshooting

If a job is pending:

  • confirm the partition is correct
  • verify requested time, memory, CPUs, and GPUs
  • inspect the reason field with squeue
  • remove a job (or jobs) from the queue with scancel <JOB_ID>

If a job fails:

  • inspect slurm-%j.out and slurm-%j.err
  • inspect accounting data with sacct
  • verify input paths
  • verify loaded modules
  • verify that the application is compatible with the requested resources

Rules That Matter

  • never run heavy workloads directly on the login node
  • test new scripts with small inputs first
  • keep output and error logs
  • record the job ID for every important run
  • use interactive allocations for debugging rather than improvising on shared infrastructure