Unified Pipeline Orchestrator (run_together)¶
Overview¶
run_together runs a complete CartoScope pipeline — ingest → FICTURE → cell decode → asset packaging → image import → (optional) publish — across many spatial platforms, for one or many samples, from a single command. It assembles the whole run as one Makefile (run_together.mk) and executes it, so the pipeline is resumable, parallel, and trains one joint FICTURE model across a multi-sample run.
The design has one guiding principle:
Convenience for the default case; full control in JSON.
A standard run needs only a few flags. Everything a run can express lives in a single canonical, list-based configuration that three layers assemble:
- Profile — per-platform defaults (auto-detected inputs, default analyses, image conventions).
- Tier-1 CLI — the common knobs (
--in-dir/--samples,--width/--n-factoror--project-models,--out-dir). - Tier-2 JSON (
--config) — augments or fully specifies anything.
The layers compose: set the base run on the CLI and add only the extra analyses/images in JSON.
Read next¶
| To… | See |
|---|---|
| Specify inputs (single CLI, sample sheet, or config JSON) | Specifying Inputs |
| Attach DAPI / H&E / protein images | Image Modalities |
| See exact inputs + examples for your platform | Platforms: Xenium · Visium HD · CosMx SMI · MERSCOPE |
A minimal run:
1 2 | |
Requirements¶
makeon thePATH.- Tools used by the stages:
spatula,punkst(FICTURE2),pigz,sort,python3,go-pmtiles,gdal,tippecanoe,parquet-tools,jq(Visium HD H&E), andaws(publish).run_togetherdelegates to the CartLoader modules, which findspatula/punkstas built repo submodules or onPATH.
Stages, selection & resume¶
The pipeline stages are ingest, ficture, cells, cartload, images, anno, upload.
- Resume: re-run the same command (or
make -f OUT/run_together.mk -j N); completed stages are skipped via flag files. --only ingest,ficture/--skip images— run a subset; excluded upstream stages are assumed done (prereqs are pruned somakewon't error).--restart— rebuild everything (make -B).--dry-run— write the Makefile and print commands (make -n) without executing.
Packaging bundles everything produced — every FICTURE pixel decode plus every cell analysis that ran.
Publishing¶
Publishing is opt-in and entirely CLI-driven (no config block). Two independent actions — enable either or both:
--anno— AI-annotate each packaged sample directory (viaanno_cartload_folder). Requires--tissueand--organism(no defaults);--anno-api-type(defaultumgpt),--anno-model(defaultclaude-opus-4-7), and--anno-threads(default10) are overridable. For a joint run the shared factors are annotated once at thecartl/root and reused into every sample.--s3-upload— upload each self-contained sample directory to S3.
When both run, upload waits on anno (which edits catalog.yaml).
S3 destination: <s3-prefix>/batch=<batch>/<collection>/<dir-id>/, where <dir-id> is the sample's output directory name (<sample_id> single, <multi_id>-<sample_id> joint). For a joint run, multi-catalog.yaml and the shared factor files are additionally uploaded to the parent <s3-prefix>/batch=<batch>/<collection>/ so relative pointers resolve.
--s3-prefix— defaults3://cartostore/data·--batch— default currentYYYY_MM·--collection— default out-dir basename--aws-profile(defaultcartostore) ·--aws(binary path) ·--s3-jobs(parallel copies, default4)
1 2 3 4 | |
How samples are packaged¶
- Single sample → one
run_cartload2call, output atcartl/<id>/. - Joint multi-sample run (samples sharing one
--out-dir) → a singlerun_cartload2_multicall that packages every sample in parallel. Each sample is a self-containedcartl/<multi_id>-<sample_id>/directory, and acartl/multi-catalog.yamllinks every per-samplecatalog.yamland copies the shared factor files (post/rgb/de/info/umap) into thecartl/root — socartl/uploads to S3 as one deployable unit.
This mirrors the FICTURE manifests: run_ficture2_multi writes a shared ficture.multi.params.json that run_cartload2_multi reads to discover samples and shared assets.
Command-line parameters¶
Run: --dry-run, --restart, -j/--n-jobs, --threads, --makefn, --only, --skip.
Input/output: --platform, --in-dir, --samples, --out-dir, --out-root, --id, --config, --platform-json (external profile override). Single-sample file inputs and column-name overrides are documented in Specifying Inputs.
FICTURE mode: --width, --n-factor (de-novo); --project-models (projection-only); --no-ficture (tiling only — package with no factor layers); --segment-10x / --segment-width-10x (also export the hexagon files as 10x MEX directories, in any mode); --tile-size / --tile-buffer (punkst tiling knobs, default 500 µm / 1000 lines; a wide hexagon such as --width 100 needs about 50× the width). See Specifying Inputs → FICTURE mode.
Common decode overrides: --exclude-feature-regex, --min-ct-per-unit-hexagon (default 50), --always-single-molecule / --never-single-molecule (default: single-molecule ON for pixel FICTURE, OFF for cell decode). An explicit CLI flag wins over a --config/profile value, which wins over the built-in default.
Packaging overrides: --bin-count — number of gene bins for the point PMTiles layers (profile default 500). Equivalent to {"cartload": {"bin_count": N}} in a --config; the flag wins.
Publish: --anno, --s3-upload, --tissue, --organism, --anno-api-type, --anno-model, --anno-threads, --collection, --batch, --s3-prefix, --aws-profile, --aws, --s3-jobs (see Publishing).
Output¶
1 2 3 4 5 6 7 8 9 10 11 12 | |
For a joint run, upload the whole cartl/ directory. For --out-root batches (independent per-sample models), this layout is created under <out_root>/<id>/, and the master Makefile is written at <out_root>/.
See also¶
- Specifying Inputs · Image Modalities · Supported Platforms
- Tutorials: single-sample · multi-sample