Calibrate Memory Requests for a SLURM Job

Estimate, measure, and calibrate the memory requested by an R job running on SLURM.

This tutorial shows how to estimate, measure, and calibrate the memory requested by a SLURM job. The source files are available in the dedicated tutorial folder.

The same R script is submitted three times:

  • run_pca_01.sbatch: initial estimate of 4G
  • run_pca_02.sbatch: abundant request of 16G
  • run_pca_03.sbatch: smaller request of 6G

The goal is not simply to make the job complete. A good SLURM request reserves enough resources for the workload without unnecessarily preventing other jobs from running.

On NowWhat, prefer reproducible batch scripts submitted with sbatch for computational jobs. Use interactive sessions only for small tests, debugging, and exploratory work. Batch jobs allow SLURM to schedule resources efficiently, keep the queue moving, and make completed workloads easier to inspect and reproduce.

R Data Types and Memory

R data typeBytes per elementNotes
raw1Stores raw bytes
logical4Stores TRUE, FALSE, or NA
integer4Stores whole numbers
double / numeric8Default type for numeric values
complex16Stores two doubles: real and imaginary parts
characterVariableA character vector stores pointers to strings; string contents and metadata require additional memory
factor4 + levelsStored as an integer vector with a character vector of levels

For example, an R numeric matrix with 150,000 rows and 1,000 columns contains doubles:

150,000 * 1,000 * 8 bytes = 1,200,000,000 bytes = 1,144.4 MiB

This calculation estimates only the payload of the matrix. It should be treated as a lower bound, not as an exact prediction of the job's peak memory.

Why Exact Memory Estimates Are Difficult

Predicting peak memory exactly is difficult for modern scientific software, especially for workloads written in R and Python. The source code describes the analysis, but the process running on the compute node includes several additional software layers and allocations:

  • the R or Python interpreter and its runtime state
  • object headers, attributes, reference tables, and allocator metadata
  • garbage-collected heaps and memory retained by the runtime for later reuse
  • temporary copies created by operations such as slicing, type conversion, centering, transposition, and vectorization
  • native libraries called by high-level packages, including BLAS, LAPACK, NumPy, and compiled R extensions
  • thread-local buffers, caches, imported packages, and dynamically loaded libraries
  • the memory behavior of the operating system, including resident memory, shared pages, and swap

The exact peak also depends on the runtime version, library implementation, input data, execution path, number of threads, and timing of garbage collection. Two executions of the same high-level script can therefore report different MaxRSS values even when they produce the same result.

For these reasons, static estimates from object dimensions remain useful, but they cannot usually account for every live allocation. In practice, the most reliable method is empirical and iterative:

estimate -> submit with a safe margin -> measure -> adjust -> repeat

The accounting data provided by NowWhat makes this process reproducible. After representative runs, use sacct and seff to compare requested resources with measured usage and converge on a robust configuration.

Get the Tutorial Files

Connect to NowWhat:

ssh your_username@nowwhat.stat.unipd.it

Clone the tutorials repository and enter the tutorial directory:

git clone http://nowwhat.stat.unipd.it/gitea/NowWhat/tutorials.git
cd tutorials/03-memory-cpu-estimates

If you already cloned the repository, update it instead:

cd tutorials
git pull
cd 03-memory-cpu-estimates

Run all commands in the rest of this tutorial from the 03-memory-cpu-estimates directory.

The tutorial contains:

03-memory-cpu-estimates/
├── README.md
├── scripts/
│   ├── 01_princomp.R
│   ├── run_pca_01.sbatch
│   ├── run_pca_02.sbatch
│   └── run_pca_03.sbatch
└── slurm_logs/

Accounting on NowWhat

NowWhat runs a SLURM accounting database that stores job statistics for 24 months. This makes it possible to inspect completed jobs long after they have left the active queue and use their measured CPU, memory, and elapsed-time statistics to improve future resource requests.

This accounting history is one reason to submit computational workloads with sbatch. A batch script records the requested resources and provides a reproducible workload that can be compared with the resources actually used. Interactive sessions remain useful for short tests, but long or repeatable computations should be converted into batch jobs so that the scheduler can manage the queue effectively.

Inspect Historical Jobs With sacct

sacct is the main command for querying the SLURM accounting database. Unlike squeue, which shows jobs currently waiting or running, sacct is a historical command and also reports completed, failed, cancelled, timed-out, and out-of-memory jobs.

Inspect one job by its job ID:

sacct -j JOBID \
  --format=JobID,JobName,State,ExitCode,AllocCPUS,ReqMem,MaxRSS,Elapsed

Inspect your jobs from a specific period:

sacct -u "$USER" --starttime 2026-06-01 --endtime 2026-06-09 \
  --format=JobID,JobName,State,Elapsed,AllocCPUS,ReqMem,MaxRSS

Useful fields include:

FieldMeaning
StateFinal or current state, such as COMPLETED, FAILED, or OUT_OF_MEMORY
ExitCodeExit status of the job or job step
AllocCPUSNumber of CPUs allocated by SLURM
ReqMemMemory requested when the job was submitted
MaxRSSMaximum resident memory measured for a job step
ElapsedWall-clock duration

Why sacct Reports Three Rows

A typical batch job appears as three related rows:

JobID           JobName      State ExitCode  AllocCPUS     ReqMem     MaxRSS    Elapsed
------------ ---------- ---------- -------- ---------- ---------- ---------- ----------
25904        pca_high_+  COMPLETED      0:0          2        16G              00:00:58
25904.batch       batch  COMPLETED      0:0          2              7528828K   00:00:58
25904.extern     extern  COMPLETED      0:0          2                         00:00:58

These rows represent the allocation and the steps that SLURM creates inside it:

RowMeaning
25904The top-level job allocation. It usually shows requested resources and the overall job state.
25904.batchThe batch step that executes the commands in the submitted sbatch script. This is usually the row containing the workload's MaxRSS.
25904.externThe external step used by SLURM to track processes associated with the job but outside the batch step, including job setup and cleanup activity.

The top-level row may have an empty MaxRSS because memory is measured on the job steps. For a simple sbatch job, inspect the .batch row when calibrating memory.

Jobs that launch additional commands with srun can contain more step rows. Job arrays also add array task identifiers, so a single submission can produce several groups of accounting rows.

Read an Efficiency Summary With seff

seff JOBID presents selected accounting data as a compact efficiency report:

seff JOBID

It compares:

  • CPU time used with the total CPU time made available by the allocation
  • maximum resident memory with the requested memory
  • the job's elapsed wall-clock time and allocation details

Low CPU efficiency can indicate that the job requested more CPUs than it used. Low memory efficiency can indicate an unnecessarily large memory request. Memory efficiency should not be pushed to 100%, however, because a robust request needs a safety margin for input variation and temporary allocations.

seff is convenient for a quick review, while sacct is better when you need specific fields, historical searches, job-step details, or output suitable for tables and scripts. Use both after representative jobs: first read the seff summary, then inspect the .batch row with sacct.

The R Workload

The script scripts/01_princomp.R creates a simulated gene-expression matrix and performs a PCA with pricomp function. The comments in the script calculate an initial memory estimate:

set.seed(100)
 
n_samples <- 1000
n_genes <- 150000
 
# X:
# - rows are genes
# - columns are samples
# - values are something
#
# Input size:
# n_samples * n_genes = 1000 * 150000 = 150,000,000 numbers
# R numeric matrices are double matrices: 8 bytes per number
# X size = 150,000,000 * 8 = 1,200,000,000 bytes = 1,144.4 MiB
X <- matrix(rnorm(n_samples * n_genes), nrow = n_genes, ncol = n_samples)
 
rownames(X) <- paste0("gene_", seq_len(n_genes))
colnames(X) <- paste0("sample_", seq_len(n_samples))
 
 
pca <- princomp(X, cor = FALSE)
# Operation performed:
# - princomp(..., cor = FALSE) computes PCA from the sample covariance matrix
#
# Memory estimate for princomp(X, cor = FALSE):
#
# 1. Input matrix already allocated:
#    X = 150000 * 1000 * 8 = 1,144.4 MiB
#
# 2. Internal working data:
#    princomp.default computes the covariance matrix with cov.wt().
#    cov.wt() centers the data before forming the covariance matrix, so it
#    needs a centered working matrix with the same size as X:
#    centered X = 150000 * 1000 * 8 = 1,144.4 MiB
#
# 3. Covariance matrix:
#    covariance = 1000 * 1000 * 8 = 7.6 MiB
#
# 4. Eigen decomposition of the covariance matrix:
#    eigenvectors = 1000 * 1000 * 8 = 7.6 MiB
#    eigenvalues  = 1000 doubles = 0.008 MiB
#
# 5. Scores kept in the PCA object:
#    scores = 150000 * 1000 * 8 = 1,144.4 MiB
#
# 6. Loadings kept in the PCA object:
#    loadings = 1000 * 1000 * 8 = 7.6 MiB
#
# Approximate subtotal:
# 1,144.4 + 1,144.4 + 7.6 + 7.6 + 0.008 + 1,144.4 + 7.6 = 3,456.0 MiB
 
dir.create("results", showWarnings = FALSE)
saveRDS(pca, "results/princomp.rds")

Rounding up, 4G initially appears to be a reasonable request.

1. Submit the Estimated Request

The first job requests:

#SBATCH --mem=4G

Submit it:

job_01=$(sbatch --parsable scripts/run_pca_01.sbatch)
echo "$job_01"

After it finishes, inspect it:

seff "$job_01"
sacct -j "$job_01" \
  --format=JobID,JobName,State,ExitCode,AllocCPUS,ReqMem,MaxRSS,Elapsed

Expected result:

Job ID: 25903
Cluster: e4_medooza_cluster
User/Group: rceccaroni/users
State: OUT_OF_MEMORY (exit code 0)
Nodes: 1
Cores per node: 2
CPU Utilized: 00:00:12
CPU Efficiency: 46.15% of 00:00:26 core-walltime
Job Wall-clock time: 00:00:13
Memory Utilized: 108.00 KB
Memory Efficiency: 0.00% of 4.00 GB (4.00 GB/node)
JobID           JobName      State ExitCode  AllocCPUS     ReqMem     MaxRSS    Elapsed
------------ ---------- ---------- -------- ---------- ---------- ---------- ----------
25903        pca_low_m+ OUT_OF_ME+    0:125          2         4G              00:00:13
25903.batch       batch OUT_OF_ME+    0:125          2                  108K   00:00:13
25903.extern     extern  COMPLETED      0:0          2                         00:00:13

The job fails with OUT_OF_MEMORY. The initial estimate considered the main objects required by the algorithm, but not all temporary allocations created internally by R.

For example, princomp(), cov.wt(), scale(), sweep(), and matrix multiplication may create additional full-size copies of the matrix. Several of these objects can exist simultaneously, making the real peak substantially higher than 3,456 MiB.

2. Measure With Abundant Memory

Instead of reconstructing every internal allocation, submit the same workload with an intentionally abundant request:

#SBATCH --mem=16G
job_02=$(sbatch --parsable scripts/run_pca_02.sbatch)
echo "$job_02"

After it finishes:

seff "$job_02"
sacct -j "$job_02" \
  --format=JobID,JobName,State,ExitCode,AllocCPUS,ReqMem,MaxRSS,Elapsed

Observed result:

Job ID: 25904
Cluster: e4_medooza_cluster
User/Group: rceccaroni/users
State: COMPLETED (exit code 0)
Nodes: 1
Cores per node: 2
CPU Utilized: 00:00:57
CPU Efficiency: 49.14% of 00:01:56 core-walltime
Job Wall-clock time: 00:00:58
Memory Utilized: 7.18 GB
Memory Efficiency: 44.88% of 16.00 GB (16.00 GB/node)
JobID           JobName      State ExitCode  AllocCPUS     ReqMem     MaxRSS    Elapsed
------------ ---------- ---------- -------- ---------- ---------- ---------- ----------
25904        pca_high_+  COMPLETED      0:0          2        16G              00:00:58
25904.batch       batch  COMPLETED      0:0          2              7528828K   00:00:58
25904.extern     extern  COMPLETED      0:0          2                         00:00:58

The job completes and reaches a MaxRSS of approximately 7.18G. Requesting 8G would therefore be a reasonable calibrated choice: it rounds the observed peak up and leaves some margin.

3. Try a Smaller Request

The third job requests only:

#SBATCH --mem=6G
job_03=$(sbatch --parsable scripts/run_pca_03.sbatch)
echo "$job_03"

After it finishes:

seff "$job_03"
sacct -j "$job_03" \
  --format=JobID,JobName,State,ExitCode,AllocCPUS,ReqMem,MaxRSS,Elapsed

Observed result:

Job ID: 25905
Cluster: e4_medooza_cluster
User/Group: rceccaroni/users
State: COMPLETED (exit code 0)
Nodes: 1
Cores per node: 2
CPU Utilized: 00:00:58
CPU Efficiency: 48.33% of 00:02:00 core-walltime
Job Wall-clock time: 00:01:00
Memory Utilized: 5.95 GB
Memory Efficiency: 99.16% of 6.00 GB (6.00 GB/node)
JobID           JobName      State ExitCode  AllocCPUS     ReqMem     MaxRSS    Elapsed
------------ ---------- ---------- -------- ---------- ---------- ---------- ----------
25905        pca_calib+  COMPLETED      0:0          2         6G              00:01:00
25905.batch       batch  COMPLETED      0:0          2              6238392K   00:01:00
25905.extern     extern  COMPLETED      0:0          2                         00:01:00

Surprisingly, this job also completes. Its reported resident-memory peak is approximately 5.95G, just below the 6G request.

There are two reasons why this can happen:

  1. This cluster permits an additional 20% of swap space. With --mem=6G, the job can use approximately:

    6G RAM + 20% swap = 6G + 1.2G = 7.2G total

    This is almost exactly the 7.17G peak observed during the 16G run. SLURM's MaxRSS and seff memory utilization report resident RAM, not the complete RAM-plus-swap peak.

    The 6G request therefore works on this cluster, but leaves almost no margin and depends on its swap configuration. For a robust request, 8G remains the better choice.

  2. Under memory pressure, R runs garbage collection more frequently and frees temporary objects sooner. With 16G available, it can retain allocated memory for longer and therefore reach a higher MaxRSS.

Takeaway

Estimating memory from R object sizes is a useful starting point, but it does not capture every temporary allocation made by the interpreter, runtime, packages, and native numerical libraries.

A reliable calibration workflow is:

estimate the main objects -> run with abundant memory -> inspect MaxRSS -> add a safety margin

For this workload, the initial 4G estimate is too small, 16G is unnecessarily abundant, and 8G is a robust calibrated request based on the observed peak. Repeating representative jobs while refining the request is not a failure of the estimation process: for complex R and Python workloads, it is often the only reliable way to identify an efficient configuration.

Apply the same process to production workloads on NowWhat:

  • submit repeatable computations with sbatch
  • reserve CPUs, memory, and time based on measurements rather than guesses
  • use seff for a quick efficiency summary
  • use the sacct history to inspect job steps and refine future requests
  • keep interactive sessions for small tests and debugging

Correctly sized batch jobs help SLURM schedule work fairly and efficiently, improving queue behavior for every NowWhat user.

Reimplementing a workload in a lower-level language such as C/Fortran can provide more explicit control over allocation, object lifetime, and memory layout. However, it also requires substantially more development effort, careful manual memory management, and additional validation. For most scientific workloads, calibrating a well-tested R or Python implementation is the more practical engineering choice.