End-to-End with run_together (multi-sample)¶
When you have several sections that should share one FICTURE model, list them in a sample sheet and run_together will:
- ingest every sample,
- train a single joint model on all of them together,
- package each sample independently against that shared model.
The dependency graph fans in on the joint model, then fans back out per sample — so make -j ingests all samples in parallel, blocks once on the shared model, then packages each sample in parallel:
1 2 3 | |
When to use a joint model
Sharing one model across sections makes factors directly comparable between samples and is the recommended setup for replicates or a cohort. For unrelated samples, run them independently (see batches). Background: Why multi-sample?
1. List the samples in a TSV sheet¶
A multi-sample run is just a single-sample run with --samples samples.tsv in place of --in-dir. The sheet is a wide table of input roles; for the common case you only need id + in_dir:
1 2 3 4 | |
2. Run¶
1 2 3 | |
All rows share one --out-dir, so they train one joint model. Everything else (FICTURE mode, images, cell analyses) is assembled exactly as in the single-sample run and applied to every sample. Add --dry-run to inspect run_together.mk and run_together.resolved.json first.
Output layout
Packaging is delegated to run_cartload2_multi: each sample is written to a self-contained cartl/<multi_id>-<sample_id>/ directory (with <multi_id> defaulting to the --out-dir basename), plus a cartl/multi-catalog.yaml that links every per-sample catalog.yaml and copies the shared factor files (post/rgb/de/info/umap) into the cartl/ root as a unified factors: map. Upload the whole cartl/ directory to S3 as one unit.
Extra role columns
The sheet can also carry per-sample inputs — raw_transcript (a raw CSV to ingest), transcript (pre-converted TSV, skips ingest), xy, boundaries, clusters, mex, cellxgene (a cell×gene matrix CSV, e.g. MERSCOPE cell_by_gene.csv):
1 2 3 4 | |
When to use JSON instead¶
Escalate to a --config JSON when samples need different settings, or to add analyses/images beyond the profile defaults. JSON augments the CLI (append-by-default); a samples block gives full per-sample control:
1 2 3 4 5 6 7 | |
See the configuration reference for every key.
Visium HD example¶
The same shape works for Visium HD; the 10x_visium_hd profile adds the square/cell imports and (with a per-sample hne column) the H&E layer automatically. Add hne as a sheet column:
1 2 3 | |
1 2 | |
Publishing (opt-in)¶
Publishing is CLI-driven — no config block. Enable annotation (--anno, needs --tissue/--organism) and/or S3 upload (--s3-upload) for every sample:
1 2 3 4 | |
Each sample uploads to s3://cartostore/data/batch=<YYYY_MM>/<collection>/<multi_id>-<sample_id>/, where <YYYY_MM> defaults to the current date (override --batch) and <collection> defaults to the out-dir basename (override --collection). See the reference → Publish for all flags (--s3-prefix, --aws-profile, annotation model/threads, …).
Batches of independent samples¶
For many unrelated samples (each its own model), use --out-root instead of --out-dir:
1 2 | |
Each row becomes its own <out_root>/<id>/ with an independent model. See Specifying Inputs → Multi-sample.
Next steps¶
- Full option and configuration reference:
run_togetherreference. - Supported platforms and expected inputs: Supported Platforms & Inputs.
- A hand-built multi-sample walkthrough (lower-level modules): Human Cortex Multi-sample Tutorial.