Specifying Inputs to run_together¶
run_together accepts inputs in three ways, from most convenient to most expressive. They compose — you can set the base run on the CLI and add only the extras in a config JSON.
| Mode | Flag | Use when |
|---|---|---|
| 1. Single sample (direct CLI) | --in-dir or explicit --in-* file flags |
One sample; a standard platform directory, or a few individually-named files. |
| 2. Multi-sample (sample sheet) | --samples samples.tsv |
Several samples (a joint model by default). A wide TSV, one row per sample. |
| 3. Full config (JSON) | --config run.json |
Per-sample settings differ, or you need analyses/images beyond the profile defaults. |
Whichever you use, the platform profile (--platform <name>) supplies the defaults: which files to auto-detect, the default FICTURE analyses, and the image conventions. See Supported Platforms for what each profile expects and per-platform pages for worked examples of all three modes.
Mode 1 — Single sample (direct CLI)¶
Two shapes, depending on how your files are laid out.
A standard platform directory — point --in-dir at it and the profile auto-detects every role inside:
1 2 3 | |
Individually-named files (not in a standard directory) — name each explicitly. Each flag is the CLI equivalent of a sample-sheet column:
1 2 3 4 5 | |
| Flag | Sample-sheet column | Meaning |
|---|---|---|
--in-dir |
in_dir |
Raw platform directory → every role auto-detected inside it. A platform may instead consume the directory whole (Seq-Scope: it is the MEX directory) |
--in-prefix |
in_prefix |
Raw platform path prefix → roles auto-detected by suffix (Stereo-seq) |
--raw-transcript |
raw_transcript |
Raw transcript CSV/TSV (or .parquet, converted first) to ingest through sge_convert |
--in-transcript |
transcript |
An already-ingested transcript TSV; skips ingest |
--in-cell-xy |
xy |
Cell centroids / metadata file |
--in-cell-boundary |
boundaries |
Cell boundary polygons |
--in-cellxgene |
cellxgene |
Cell×gene matrix CSV → converted to a MEX (drives cell clustering) |
--id |
id |
Sample id. With --out-dir, defaults to rep1 (the packaged dir/catalog id becomes <out-dir basename>-rep1); with --out-root, defaults to the in_dir/in_prefix basename |
--in-dir is just a one-row sample sheet with only in_dir; the explicit --in-* flags are a one-row sheet with those columns.
Mode 2 — Multi-sample (sample sheet)¶
A joint multi-sample run is a single-sample run with --samples samples.tsv in place of --in-dir. All rows share one --out-dir, so they train one joint FICTURE model. For the common case you only need id + in_dir:
1 2 | |
1 2 3 4 | |
For independent per-sample models (each its own model), use --out-root instead of --out-dir; each row becomes its own <out_root>/<id>/.
The sample sheet is a wide table of input roles¶
Columns map to per-sample input roles. Every column is optional except that each sample needs a transcript source (in_dir, raw_transcript, or transcript). Column names are case-sensitive; the listed aliases are accepted interchangeably. The transcript source is checked before anything is planned: a sample with none, or one whose file is missing (including an in_dir without the platform's expected input), stops the run with an error naming each affected sample. Because an unrecognized column is otherwise ignored, that error also lists any unrecognized columns (e.g. transcripts for transcript). The check is skipped when the ingest stage is not run (--only / --skip ingest), since a resumed run reuses the already-converted transcripts.
| Column (aliases) | Role | Meaning |
|---|---|---|
id |
— | Sample identifier. Optional: defaults to the in_dir/in_prefix basename, or to rep1 for a lone --out-dir sample with no named input. |
in_dir |
— | Raw platform directory → the profile auto-detects every role inside it. |
in_prefix |
— | Raw platform path prefix → the profile auto-detects roles by suffix ({prefix}.tissue.gef, …). For platforms whose files share a name rather than a directory; see BGI Stereo-seq. |
raw_transcript |
— | Explicit path to a raw transcript file to ingest (e.g. MERSCOPE detected_transcripts.csv, Xenium transcripts.csv.gz). Runs through sge_convert. Sheet equivalent of --raw-transcript. |
transcript (tsv) |
transcript | A pre-converted transcripts.tsv.gz (the output of sge_convert) → skips ingest. Not for raw CSVs. |
xy (cell_xy) |
xy | Cell centroids file. |
boundaries (cell_boundary, cell_boundaries) |
boundaries | Cell boundary polygons. |
clusters |
clusters | External cluster labels. |
mex (mex_dir) |
mex | MEX directory. Or give the explicit triple mex_bcd / mex_ftr / mex_mtx when one directory does not apply. |
cellxgene (cell_by_gene) |
cellxgene | Cell×gene matrix CSV (e.g. MERSCOPE cell_by_gene.csv) → converted to a MEX that drives cell clustering (works without boundaries). |
gef / cellbin_gef |
gef, cellbin_gef | Stereo-seq binary GEFs, when they do not match in_prefix + the standard suffix. |
cell_tsv |
cell_tsv | A standalone pixel TSV (X, Y, gene, count, cell_id) that supplies cell counts on its own, for platforms whose cell assignment cannot be carried on the transcript (Stereo-seq cell bins). |
dapi |
— | A single DAPI image (.ome.tif/.tif/.png) → a colorized dapi layer. Xenium's morphology.ome.tif z-stack is recognized by name (any prefix, e.g. GSM123_morphology.ome.tif) and imported with --use-middle-page --high-memory. |
hne |
— | H&E image (Visium HD adds the layer automatically). |
Other images are not sample-sheet columns (only dapi and hne) — they are a separate concern; see Image Modalities.
Parquet inputs. raw_transcript, xy, boundaries and clusters may point at a .parquet file (Xenium ships transcripts.parquet, cells.parquet, cell_boundaries.parquet beside the .csv.gz files). Their consumers read CSV only, so each is converted to .csv.gz under OUT/tsv/<id>/parquet2csv/ with cartloader parquet_to_csv_rapid as the first ingest step — announced with a NOTE: at planning time — and every later stage uses the converted file. A failed or empty conversion aborts the run with an explicit error.
Rules: an explicit column overrides auto-detection for that role; the cell values ` (empty),-,., andNAall mean *unset*. Role columns are resolved relative toin_dirwhen relative, butraw_transcript` is not — give it an absolute path (or one relative to the working directory).
The three transcript sources¶
Each sample gets its transcripts from exactly one of:
in_dir— the standard path: point at the raw platform folder and the profile findsdetected_transcripts.csv[.gz](and every other role) inside it.raw_transcript— a raw CSV/TSV that still needs ingesting, when it is arbitrarily named or not laid out as a standardin_dir.transcript— an already-ingested TSV (transcripts.tsv.gz); ingest is skipped and it feeds FICTURE directly.
1 2 3 4 5 | |
Column-name overrides¶
When input columns are non-standard, name them (CLI or the profile/--config):
--colname-transcript-x/-y/-feature/-count— the raw transcript's coordinate/gene/count columns.--colname-xy-cell/-x/-y— the xy file's columns (use--colname-xy-cell ''for an unnamed index column).--colname-boundary-cell/-x/-y— the boundary file's columns.
Transcript that already carries a cell_id column (rare)
If your raw transcript CSV already has a per-molecule cell-id column, name it with
--colname-transcript-cell <name> (or ingest.csv_colname_cell in --config).
The column is carried through ingest to transcript column 5 and used directly for
cell analysis, skipping spatula tsv-add-cell-id. A cellxgene MEX, if also
present, still drives the clustering.
FICTURE mode (de-novo vs. projection)¶
Independent of how inputs are specified, Tier-1 selects the base FICTURE work. Exactly one mode is the base (a config JSON can add more analyses on top).
Train new LDA models. --width and --n-factor accept comma lists → the cross-product is trained (multiple widths supported).
1 2 | |
Reuse already-trained models — no LDA training runs. Point --project-models at one or more existing FICTURE directories; run_together reads each ficture.params.json and re-projects every model it lists onto the current data.
1 2 | |
Ingest still runs (the data is current); only training is skipped.
Host a dataset with no FICTURE at all — transcripts, the SGE raster, and histology, with no factor layers. --no-ficture runs only FICTURE's tiling step (run_ficture2_multi --prepare-only), and packaging reads the resulting tiled TSV directly.
1 | |
Applies to every platform, not just Seq-Scope. Cell analyses are skipped too (they decode against a model). Cannot be combined with --n-factor or --project-models. The hexagon files are still built alongside the tiles, so adding factors later re-uses them instead of re-tiling — rerun without --no-ficture in the same --out-dir.
Hexagon MEX export (any mode): --segment-10x additionally writes each sample's hexagon file as a 10x MEX directory, fic/samples/<id>/<id>.hex_<width>.mex/ (barcodes.tsv.gz with hexagon centers as x:y, features.tsv.gz, matrix.mtx.gz). It converts the very hexagon files the factor analysis uses (same --min-ct-per-unit-hexagon filter), so the two are consistent. By default every analysis' --width is exported; --segment-width-10x 12,24 picks the widths instead (extra widths get their hexagon files built too). Works with --no-ficture — the way to get hexagon MEX files without any factor analysis, e.g. for Seq-Scope:
1 2 | |
Config equivalent: top-level "segment_10x": true or { "widths": "12,24" } (a CLI flag wins). The directories are recorded under mex in the per-sample ficture.params.json and in ficture.multi.params.json; see run_ficture2_multi --segment-10x.
Tiling (run-wide, rarely needed): --tile-size and --tile-buffer are forwarded to run_ficture2_multi (punkst multisample-prepare) and otherwise keep its defaults of 500 µm and 1000 lines. punkst rejects a tile size under 20× the hexagon side length (width / √3) and recommends 50–100×, so only a wide hexagon needs this — e.g. --width 100 wants roughly --tile-size 5000. The buffer is a per-tile line buffer, not a spatial size. Config equivalents: top-level "tile_size" / "tile_buffer" (a CLI flag wins).
Common decode overrides (else profile / built-in default): --exclude-feature-regex, --include-feature-list / --exclude-feature-list, the --ingest-*-feature-* family, --min-ct-per-unit-hexagon (default 50), --min-ct-per-unit-train, --cell-min-cell-count / --cell-min-feature-count, and --always-single-molecule / --never-single-molecule (default: single-molecule ON for pixel FICTURE, OFF for cell decode). An explicit CLI flag wins over a --config/profile value, which wins over the built-in default.
Feature filtering: two independent layers¶
Feature filters come in two flavours that do not interact. Each takes a regex and/or a plain text file of feature names (one per line), and each has a default that applies to every platform:
| default regex | rationale | |
|---|---|---|
| ingest | ^(Unassigned\|Neg\|BLANK\|Blank\|Intergenic\|Deprecated\|System\|NCS-\|NCP-) |
technical artifacts — not genes, so they never enter the data |
| analysis | ^(Gm[0-9]\|MT-\|mt-\|Rps\|Rpl) |
real genes that distort a factorization — kept in the data, kept out of the models |
Pass '' to either flag to disable its default; a profile or --config value overrides it.
1. Analysis (FICTURE) filters — --include-feature-list / --exclude-feature-list / --exclude-feature-regex. These restrict the factor model only (lda4hex --features); the data — counts, pseudobulk, DE — keep every gene. --include-feature-list is the natural place for a highly-variable-gene set: the model is built on those genes, but all genes are still emitted afterwards.
| stage | affected |
|---|---|
| ingest, tiled transcripts, feature list, packaged tiles | no — every gene is kept |
| pixel FICTURE: the LDA/projection model (and hence pixel decoding, which reads that model) | yes |
| cell clustering: the LDA is a projection onto the pixel model, so it inherits the same restriction | yes (via the model) |
| cell counts, pseudobulk, DE | no — excluded/non-HVG genes stay in and reappear here |
A list and the regex combine: the regex narrows what the list leaves. The features the pixel model was trained on are written to fic/multi.selected_features.tsv and recorded per model as feature in ficture.params.json.
Because the restriction lives in the model (not the counts), the cell pseudobulk and DE report every gene, and the pixel decode — which reads the restricted model — reports only the model's genes. That asymmetry is intentional: the pixel decode is the model, the cell pseudobulk is an independent aggregate of the raw counts.
2. Ingest filters — --ingest-include-feature-list / --ingest-exclude-feature-list / --ingest-include-feature-regex / --ingest-exclude-feature-regex. These drop features from the transcript TSV as it is written, so a filtered feature is gone from everything downstream: the feature list, the packaged tiles, the browser's gene list and every analysis. Use them for features that should not be part of the dataset at all (negative-control probes, blanks); use the analysis filters for features that should be visible but not drive the factorization.
The ingest regex always replaces sge_convert's per-platform default (e.g. Xenium's negative-probe pattern), so one pattern governs every platform.
- Applied by whichever step writes the transcript:
sge_convert,reformat_cosmx(CosMx), and the Stereo-seq cell-bin conversion. - Also applied to the cell count matrices (
mex2sptsv/pixel2sptsv), so a MEX-derived cell source (acell_by_genematrix or an externalmexrole) — which never passes throughsge_convert— drops the same technical artifacts. This is the only feature filter on the cell counts; the FICTURE filters deliberately are not applied there (see above), so the cell pseudobulk keeps every real gene. - Ignored, with a warning, for a sample that supplies an already-ingested
transcript— that file is used as given. Filter it beforehand, or supply the raw input. - On the MEX-based platforms (
10x_visium_hd,seqscope,illumina) filtering happens insidespatula convert-sge, which accepts only one include-type and one exclude-type filter; a list and a regex of the same polarity is an error there. Resolve them into one list withcartloader feature_select.
Count thresholds are applied before the analysis filter
--min-ct-per-unit-hexagon (hexagons) and the cell analysis's minimum cell count are applied over all genes, while the model is fit over the restricted set. Restricting to a small panel therefore leaves units/cells whose surviving counts are low; lower --min-ct-per-unit-train and --cell-min-cell-count accordingly. Ingest filters do not have this problem — they run before any counting.
Coordinate units at ingest¶
--units-per-um <n> tells ingest how many coordinate units of the raw input make one micron — the factor sge_convert divides the input X/Y by. 1 means the input is already in microns, 1000 means nanometers, 2 is Stereo-seq's 500 nm bins. Everything downstream of ingest is in microns, so this is the one place the input's unit convention is declared.
Normally you never set it: each platform's ingest preset knows its own convention. Set it when a dataset departs from that convention — notably Illumina StrataMap, which ships two barcode formats:
| Barcode | Units | --units-per-um |
|---|---|---|
SBC:433503:2393851 (older) |
nanometers | 1000 |
SBC:686.951:4668.15 (current) |
microns | 1 |
Ingest detects which of the two a StrataMap barcode file uses (a fractional coordinate means microns; otherwise the coordinate magnitude decides) and reports the value it picked, so both formats run correctly with no flag. --units-per-um overrides the detection; if the value you give contradicts the file, the run warns and uses yours.
- Applied by
sge_convert, in both of its routes (MEX viaspatula convert-sge, CSV viasge_format_generic), so it covers every platform whose ingest issge_convert. Onstereoseqit also scales the cell-bin TSV, keeping the pixel and cell coordinates on one system. - Not supported on
cosmx_smi, whose ingest (reformat_cosmx) writes the transcript TSV itself. Asking for it there is an error. - Ignored, with a warning, for a sample that supplies an already-ingested
transcript— that file is used as given, in microns. - Boundary/centroid files are not rescaled by this flag: they must already be in microns (see the platform pages).
1 2 | |
The config equivalent lives in the ingest block (a CLI --units-per-um wins over it):
1 | |
Coordinate jitter at ingest¶
--jitter-xy <um> adds a uniform random offset in [-um, +um] to each transcript's X and Y — drawn independently per transcript and per axis — as the transcript TSV is written. It is off by default (0).
Use it on coarse-resolution platforms whose molecules sit on a lattice rather than at measured positions. On 10x Visium HD, for example, every transcript in a 2 µm bin carries that bin's single coordinate, so the transcript cloud is a grid of stacked points; hexagon binning, pixel decoding and the rendered tiles all see the lattice. Jittering by roughly half the bin pitch (--jitter-xy 0.8 for 2 µm bins) spreads each bin's molecules across the area they came from.
Like the ingest feature filters, this rewrites the data itself: the jittered coordinates are what every later stage reads — hexagons, FICTURE, the packaged tiles, the browser.
- Applied by
sge_convert, in both of its routes:spatula convert-sgefor MEX input (10x_visium_hd,seqscope,illumina) andsge_format_genericfor CSV input (10x_xenium,merfish,stereoseq's bin1 GEM,generic). Jitter is in microns, applied after the input's units are converted. - Not supported on
cosmx_smi, whose ingest (reformat_cosmx) writes the transcript TSV itself. Asking for it there is an error, not a silently dropped flag. - On
stereoseqit covers the bin1 pixel transcript only; the cell-bin TSV (convert_stereoseq_cellbin) has no jitter option and stays on the original grid. The run warns when you ask. - Ignored, with a warning, for a sample that supplies an already-ingested
transcript— that file is used as given.
1 | |
The config equivalent lives in the ingest block (a CLI --jitter-xy wins over it):
1 | |
Mode 3 — Full config (JSON)¶
Escalate to --config run.json when samples need different settings, or to add analyses/images beyond the profile defaults. Everything a run can express reduces to one canonical, list-based configuration that the three layers (profile → CLI → JSON) assemble:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 | |
Publishing (annotation + S3 upload) is not part of this config — it is driven entirely by CLI flags (see Overview → Publishing).
List assembly rule (append-by-default, keyed by id)¶
For ficture, cell_analyses, and images, JSON entries are merged into the profile/CLI-derived list:
- entry with a new
id→ appended; - entry reusing an existing
id→ deep-merged (override); - to discard the base list entirely, write the section as
{ "replace": [ ... ] }.
This is what lets you set the base on the CLI and add only the extras in JSON:
1 2 | |
1 2 3 4 5 6 | |
ficture analyses¶
Each entry is either de-novo or a projection:
1 2 | |
ficture_defaults (per-analysis decode params like decode_scale, min_ct_per_unit_hexagon, min_ct_per_unit_train) apply to every analysis, including projections; per-entry keys win. cell_defaults does the same for cell_analyses entries (min_cell_count, min_feature_count).
cell_analyses¶
Cell-level decode is platform-default and automatic: an analysis runs whenever a sample provides its required roles. A sample contributes when it has every role in uses and (if present) at least one role in any_uses; roles in optional_uses are added to the decode when available. model_id picks which FICTURE model decodes the cells (default: the largest-factor model).
1 | |
any_uses lets an analysis accept alternative cell-count sources — e.g. MERSCOPE runs cell analysis from either boundaries or a cellxgene MEX, and a mixed joint run (some samples with boundaries, some with a MEX) is resolved per sample and packaged from one call. See the MERSCOPE page for the source-precedence rules.
An analysis may carry a multi_import command (e.g. Xenium's xeniumranger → import_xenium_cell). Such an analysis relies on sample-specific cluster labels: on a single-sample run it goes through run_ficture2_multi_cells as usual; on a joint run it is imported per sample instead (sheet-provided xy/boundaries/clusters paths are forwarded as --csv-* overrides). See the Xenium page.
extra_flags is a list of raw flags appended to this analysis's run_ficture2_multi_cells call, for options run_together does not model (e.g. ["--zero-based-clust-id"]).
TSNE manifolds are not generated by default (only UMAP is; nothing downstream reads TSNE, and it is slow on large datasets). To add them for an analysis, pass "extra_flags": ["--tsne"].
name sets the factor's human-readable name: in catalog.yaml and multi-catalog.yaml — the label shown for the layer. The factor id is unchanged (it still names every file and asset key), so this is purely cosmetic:
1 | |
Without it, a cell analysis displays its bare id (published), and the multi-catalog's cell factors carry no name at all. The name is written after packaging, into the shared multi-catalog and every contributing sample's catalog. Plain text only — quotes, $, backticks and backslashes are rejected up front, since the name travels through a generated make recipe. Unrelated to alias, which points a factor at a companion factor-label TSV.
alias names a companion cluster-label file for the analysis — one label per cluster id:
1 2 | |
1 2 3 4 | |
Two columns, cluster id then label, with an optional header; tab-, comma- or whitespace-separated (a label may contain spaces unless the file is comma-separated). The ids use the same numbering as the analysis's cluster files: 1-based for supplied cluster labels, unless the analysis passes --zero-based-clust-id in extra_flags, in which case they are 0-based. A Leiden analysis clusters on demand, so its alias ids are 0-based, matching the packaged output. At planning time the file is validated (integer ids at or above the base, no duplicates, non-empty labels) and rewritten to the packaged layout (index<TAB>alias, 0-based) as <out_dir>/tsv/alias.<analysis_id>.tsv. After packaging it is deployed as <analysis_id>-alias.tsv and recorded under the factor's alias key in the multi-catalog and every contributing sample's catalog. A factor that carries alias is skipped by the AI annotation stage (--anno / --anno-deep), so curated labels are not overwritten.
Supplying your own --list-* files¶
By default run_together derives each --list-* file that run_ficture2_multi_cells consumes, writing <out_dir>/tsv/in_<role>.<analysis_id>.tsv from the samples' resolved roles. To supply one yourself instead — most often externally assigned cell clusters, since without --list-cluster the cells stage computes Leiden clusters on demand — name it per role, either run-wide or per analysis:
| role | flag it feeds | CLI flag | line format |
|---|---|---|---|
clusters |
--list-cluster |
--list-cluster |
SAMPLE_ID<TAB>CLUSTER_FILE |
xy |
--list-xy |
--list-xy |
SAMPLE_ID<TAB>XY_FILE |
boundaries |
--list-boundaries |
--list-boundaries |
SAMPLE_ID<TAB>BOUNDARY_FILE |
mex |
--mex-list |
--list-mex |
SAMPLE_ID<TAB>MEX_DIR (or a bcd/ftr/mtx triple) |
cell_tsv |
--tsv-list |
--list-cell-tsv |
SAMPLE_ID<TAB>CELL_TSV |
1 2 3 | |
1 2 3 4 5 6 7 | |
A named list is passed verbatim (no file is generated for that role) and:
- satisfies that role's gating — no sample has to carry the role on disk, and the role need not appear in the analysis's
usesat all, so a cluster list can be attached to an analysis that would otherwise cluster on demand; - is validated at plan time — unknown role name, missing file, empty file, and a first column matching none of the run's sample ids are hard errors; partial coverage and ids outside the run are warnings;
- does not change
multi_importrouting — that path is per-sample by nature, so on a joint Xenium run attach the list to the jointly decodedcartloaderanalysis, andxeniumrangerkeeps its per-sample import.
Cluster ids are read as 1-based and decremented; for 0-based labels add "extra_flags": ["--zero-based-clust-id"].
See also¶
run_togetherOverview — the pipeline, stages/resume, publishing, output.- Image Modalities — how to attach DAPI/H&E/protein images.
- Supported Platforms and the per-platform pages.
- Tutorials: single-sample, multi-sample.