Snakemake Single-Cell Spatial Demo

Run a small Snakemake workflow that combines Python preprocessing with R downstream analysis.

This tutorial shows how to run a small Snakemake workflow on NowWhat. The workflow uses Python for preprocessing and clustering, then R for downstream marker and plotting steps.

The source files for this tutorial are available in the dedicated tutorial folder.

What You Will Do

By the end of this tutorial, you will have:

  • cloned the tutorial workflow
  • activated the nw-snakemake environment
  • submitted workflow steps through SLURM
  • generated marker and pseudo-spatial plot outputs

Workflow Overview

The workflow uses PBMC3k, a real single-cell RNA-seq dataset available through Scanpy. The spatial coordinates are simulated, so they are intended only to demonstrate how spatial-style plotting can fit into a workflow.

Python performs the first part of the workflow:

  • downloads PBMC3k
  • runs simple QC filtering
  • normalizes and clusters the cells with Scanpy
  • exports cluster labels and a small expression matrix for R
  • creates simulated spatial coordinates

R performs the downstream steps:

  • calculates simple marker genes from the exported expression table
  • creates a PDF scatter plot using the simulated spatial coordinates

Get the Tutorial Files

Connect to NowWhat:

ssh your_username@nowwhat.stat.unipd.it

Clone the tutorials repository:

git clone https://nowwhat.stat.unipd.it/gitea/NowWhat/tutorials.git
cd tutorials/02-snakemake

If you already cloned it:

cd tutorials
git pull
cd 02-snakemake

Inspect the Workflow

The tutorial contains:

02-snakemake/
├── README.md
├── Snakefile
├── run_slurm.sh
├── envs/
│   └── scanpy.yaml
└── scripts/
    ├── 01_download_pbmc3k.py
    ├── 02_qc_filter.py
    ├── 03_cluster_scanpy.py
    ├── 04_make_spatial_coords.py
    ├── 05_markers_r.R
    └── 06_spatial_plot_r.R

The Snakefile defines the complete workflow and final targets:

results/markers.csv
plots/spatial_clusters.pdf

Snakemake creates the intermediate files under data/ and results/ as needed.

Load Snakemake

Load the micromamba module and activate the NowWhat Snakemake environment:

module load micromamba/latest
micromamba activate nw-snakemake

Verify that Snakemake is available:

snakemake --version

Prepare Conda for Snakemake

This workflow uses --use-conda, so the nw-snakemake environment must expose a conda command.

Check it:

micromamba run -n nw-snakemake conda info --json

If that command fails with conda: command not found, install conda once inside the environment:

micromamba install -n nw-snakemake -c conda-forge conda

The SLURM runner uses this command when it creates the per-rule conda environments.

Run the Workflow Through SLURM

From the 02-snakemake directory, run:

bash run_slurm.sh

The script configures strict conda channel priority, creates a slurm_logs directory, and submits Snakemake jobs with:

sbatch --parsable \
  --cpus-per-task=1 \
  --mem=2G \
  --time=00:10:00 \
  --output=slurm_logs/%x-%j.out \
  --error=slurm_logs/%x-%j.err

The Snakemake command uses the cluster-generic executor:

micromamba run -n nw-snakemake snakemake \
  --use-conda \
  --executor cluster-generic \
  --cluster-generic-submit-cmd "sbatch --parsable ..." \
  --jobs 2

Snakemake creates the required conda environment from envs/scanpy.yaml and submits workflow rules as SLURM jobs.

Follow Progress

Snakemake prints each rule as it runs. The main rule sequence is:

  1. download_pbmc3k
  2. qc_filter
  3. cluster_scanpy
  4. make_spatial_coords
  5. markers_r
  6. spatial_plot_r

If a step fails, read the error printed by Snakemake first. Then inspect the corresponding script under scripts/.

Check the Outputs

When the workflow finishes, check:

ls results
ls plots

Expected final outputs:

results/markers.csv
plots/spatial_clusters.pdf

You can inspect the marker table directly:

head results/markers.csv

Copy the PDF to your local machine if you want to open it locally:

scp your_username@nowwhat.stat.unipd.it:~/tutorials/02-snakemake/plots/spatial_clusters.pdf .

Adapt the remote path if you cloned the repository somewhere else.

Monitor SLURM Jobs

Check your jobs:

squeue -u $USER

Inspect SLURM logs:

ls slurm_logs
cat slurm_logs/*.out
cat slurm_logs/*.err

If a job fails, rerun Snakemake after fixing the issue. Completed outputs are reused automatically unless you remove them or force a rerun.

Clean and Rerun

To see what Snakemake would do without executing commands:

micromamba run -n nw-snakemake snakemake --dry-run

To rerun from scratch, remove generated outputs:

rm -r data results plots slurm_logs

Then run:

bash run_slurm.sh