Snakemake Single-Cell Spatial Demo
Run a small Snakemake workflow that combines Python preprocessing with R downstream analysis.
This tutorial shows how to run a small Snakemake workflow on NowWhat. The workflow uses Python for preprocessing and clustering, then R for downstream marker and plotting steps.
The source files for this tutorial are available in the dedicated tutorial folder.
What You Will Do
By the end of this tutorial, you will have:
- cloned the tutorial workflow
- activated the
nw-snakemakeenvironment - submitted workflow steps through SLURM
- generated marker and pseudo-spatial plot outputs
Workflow Overview
The workflow uses PBMC3k, a real single-cell RNA-seq dataset available through Scanpy. The spatial coordinates are simulated, so they are intended only to demonstrate how spatial-style plotting can fit into a workflow.
Python performs the first part of the workflow:
- downloads PBMC3k
- runs simple QC filtering
- normalizes and clusters the cells with Scanpy
- exports cluster labels and a small expression matrix for R
- creates simulated spatial coordinates
R performs the downstream steps:
- calculates simple marker genes from the exported expression table
- creates a PDF scatter plot using the simulated spatial coordinates
Get the Tutorial Files
Connect to NowWhat:
ssh your_username@nowwhat.stat.unipd.itClone the tutorials repository:
git clone https://nowwhat.stat.unipd.it/gitea/NowWhat/tutorials.git
cd tutorials/02-snakemakeIf you already cloned it:
cd tutorials
git pull
cd 02-snakemakeInspect the Workflow
The tutorial contains:
02-snakemake/
├── README.md
├── Snakefile
├── run_slurm.sh
├── envs/
│ └── scanpy.yaml
└── scripts/
├── 01_download_pbmc3k.py
├── 02_qc_filter.py
├── 03_cluster_scanpy.py
├── 04_make_spatial_coords.py
├── 05_markers_r.R
└── 06_spatial_plot_r.RThe Snakefile defines the complete workflow and final targets:
results/markers.csv
plots/spatial_clusters.pdfSnakemake creates the intermediate files under data/ and results/ as needed.
Load Snakemake
Load the micromamba module and activate the NowWhat Snakemake environment:
module load micromamba/latest
micromamba activate nw-snakemakeVerify that Snakemake is available:
snakemake --versionPrepare Conda for Snakemake
This workflow uses --use-conda, so the nw-snakemake environment must expose a conda command.
Check it:
micromamba run -n nw-snakemake conda info --jsonIf that command fails with conda: command not found, install conda once inside the environment:
micromamba install -n nw-snakemake -c conda-forge condaThe SLURM runner uses this command when it creates the per-rule conda environments.
Run the Workflow Through SLURM
From the 02-snakemake directory, run:
bash run_slurm.shThe script configures strict conda channel priority, creates a slurm_logs directory, and submits Snakemake jobs with:
sbatch --parsable \
--cpus-per-task=1 \
--mem=2G \
--time=00:10:00 \
--output=slurm_logs/%x-%j.out \
--error=slurm_logs/%x-%j.errThe Snakemake command uses the cluster-generic executor:
micromamba run -n nw-snakemake snakemake \
--use-conda \
--executor cluster-generic \
--cluster-generic-submit-cmd "sbatch --parsable ..." \
--jobs 2Snakemake creates the required conda environment from envs/scanpy.yaml and submits workflow rules as SLURM jobs.
Follow Progress
Snakemake prints each rule as it runs. The main rule sequence is:
download_pbmc3kqc_filtercluster_scanpymake_spatial_coordsmarkers_rspatial_plot_r
If a step fails, read the error printed by Snakemake first. Then inspect the corresponding script under scripts/.
Check the Outputs
When the workflow finishes, check:
ls results
ls plotsExpected final outputs:
results/markers.csv
plots/spatial_clusters.pdfYou can inspect the marker table directly:
head results/markers.csvCopy the PDF to your local machine if you want to open it locally:
scp your_username@nowwhat.stat.unipd.it:~/tutorials/02-snakemake/plots/spatial_clusters.pdf .Adapt the remote path if you cloned the repository somewhere else.
Monitor SLURM Jobs
Check your jobs:
squeue -u $USERInspect SLURM logs:
ls slurm_logs
cat slurm_logs/*.out
cat slurm_logs/*.errIf a job fails, rerun Snakemake after fixing the issue. Completed outputs are reused automatically unless you remove them or force a rerun.
Clean and Rerun
To see what Snakemake would do without executing commands:
micromamba run -n nw-snakemake snakemake --dry-runTo rerun from scratch, remove generated outputs:
rm -r data results plots slurm_logsThen run:
bash run_slurm.sh