Skip to content

Take the input samplesheet as a validated row channel for pipeline composition - #1966

Draft
pinin4fjords wants to merge 35 commits into
output-recordsfrom
feat/composable-input
Draft

pinin4fjords wants to merge 35 commits into
output-recordsfrom
feat/composable-input

Conversation

@pinin4fjords

@pinin4fjords pinin4fjords commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Proof of concept for including rnaseq in another pipeline with nextflow#7213 (pipeline composition), without weakening input validation. Targets output-records (#1945) and contains only what composition needs; #1945 stays the modernisation (records, typing, output block). This is not a statement that nf-core/rnaseq supports composition: the standards below need agreeing first.

What changes

  • Input is a validated row channel. input is Channel<SampleRow>: loaded from the CSV on the command line, or given as a channel by an including pipeline. Rows are validated against assets/schema_input.json with nf-schema validate() and grouped per sample. The auto-strandedness flag, MultiQC paired-end ids and the SummarizedExperiment colData are derived from the rows. Rows are sorted by path before the runs of a sample are merged, because the channel has no fixed order.
  • Params are passed in. NFCORE_RNASEQ, RNASEQ, PIPELINE_INITIALISATION and PIPELINE_COMPLETION take params as an input, params that had defaults only in nextflow.config get them in the params block, assets and the schema are found through moduleDir, and conf/ is split into params.config (config-only params) and modules.config, with alias-independent process selectors.
  • The GTF and the primary quantification are returned to an including pipeline.
  • Tool arguments travel on the records (the part that needs the nf-core discussion). Process config closures read global params, which an including pipeline does not share. Arguments that follow from a param are built in the workflow (tool_args.nf), collected in one ToolArgs record and attached as args to each module's input record; ext.args replaces them. The affected modules get their own input record types.
  • nf-schema and CI. Pins a patched 2.7.2-channel.5 (Validate and print the value of Channel and Value params instead of hanging nextflow-io/nf-schema#230, Major overhaul of Salmon requirements #232), because validateParameters blocks on Channel params. CI sets NXF_PLUGINS_TEST_REPOSITORY and builds Nextflow from a master commit with #7213.

Status

Notes

  • HISAT2_BUILD takes its memory threshold as whole GB, converted in the workflow.
  • conf/test.config seeds only the rRNA removal alignment.

🤖 Generated with Claude Code

pinin4fjords and others added 5 commits October 1, 2026 10:35
input is now a Channel<SampleRow>: the rows are collected, validated against
assets/schema_input.json, and grouped per sample. The auto-strandedness flag,
the MultiQC paired-end ids and the SummarizedExperiment samplesheet are
derived from the rows instead of reading the file. Uses a patched nf-schema
(validateParameters blocks on Channel params) until the fix is released.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…tion tests [skip ci]

Params that had a default only in nextflow.config (gencode, prokaryotic,
featurecounts_group_type, with_umi, aligner, bam_csi_index, save_unaligned,
contaminant_screening_input, monochrome_logs) now have it in the params block,
since an including pipeline does not load nextflow.config and would otherwise
see them as required.

Adds function tests that feed rows to readSamplesheet (the channel path) and an
invalid samplesheet fixture for the command-line path.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…e nested dataflow params [skip ci]

The gated SALMON_INDEX input uses named combine arguments, as the other
branch does, because positional combine rejects a null-valued channel. The
any_auto_strandedness test inputs are dataflow values. The params dump in
utils_nextflow_pipeline replaces dataflow values at any nesting depth, since
the params of an included pipeline are a nested record. nf-schema is pinned
to 2.7.2-channel.3.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…uded in another pipeline

In an including pipeline, params read outside the entry workflow resolve to the
including pipeline's params, and projectDir to its project directory. NFCORE_RNASEQ,
RNASEQ, PIPELINE_INITIALISATION and PIPELINE_COMPLETION now take the pipeline's
params as an input named params, so their bodies are unchanged; the functions that
read params take them as an argument. ALIGN_BOWTIE2 takes save_unaligned as an
input. Bundled assets and the samplesheet schema are located relative to moduleDir.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

Warning

Newer version of the nf-core template is available.

Your pipeline is using an old version of the nf-core template: 4.0.3.
Please update your pipeline to the latest version.

For more documentation on how to update your pipeline, please see the Synchronisation documentation.

…e patched nf-schema; document including the pipeline [skip ci]

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

❌ nf-test failed with latest Nextflow version

Note

Tests with Nextflow's latest version failed but it will not cause a CI workflow failure.
Please check if the failure is expected with newer (edge-)releases of Nextflow or if it needs fixing.

  • ❌ docker | latest-everything | Shard 18/20

See the full run for details.

pinin4fjords and others added 4 commits October 1, 2026 11:45
…ncluding pipeline [skip ci]

gtf and gene_quant are declared in the output block with enabled false, so they
are returned to a pipeline that includes this one but are not published.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…oadable files, make selectors alias-independent [skip ci]

nextflow.config included params defaults and the process config inline, so a pipeline that
includes this one (and does not load nextflow.config) had no way to get them. They are now
conf/params.config and conf/process.config, included from nextflow.config at the same place;
the resolved config of a standalone run is unchanged.

Eight process selectors began with NFCORE_RNASEQ:RNASEQ:, which cannot match when the
pipeline is included under an alias (the qualified name then has the alias in front); they
now accept a prefix.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…bin/ [skip ci]

deseq2_qc.r was called by name from the pipeline's bin/, which is not on a process's PATH when the
pipeline is included in another one. It is now the module's template, with the options set the way
the other R templates do (opt defaults overridden from task.ext.args); the MultiQC header handling
that followed it in the process script is done in R in the same template, and versions.yml is written
by the template. The quoted empty --sample_suffix is dropped from ext.args since it cannot sit in an R
string and is the default.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
bin/mqc_features_stat.py was an older copy of the template that
custom/multiqccustombiotype uses, and nothing referenced it. bin/deseq2_qc.r moved to its module's
templates. The path filters and ignore entries for bin/ go with it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
pinin4fjords and others added 17 commits October 1, 2026 14:07
…hrough the process input [skip ci]

FQ_LINT, UMITOOLS_EXTRACT, FASTP and TRIMGALORE take the arguments that follow from the
pipeline params in an args field of their input record (FqLintInput, UmitoolsExtractInput,
FastpInput, TrimgaloreInput), built in the workflow by functions in tool_args.nf, and join them
with task.ext.args, which stays the hook for users and including pipelines. The closures that read
params are removed from config. The typed params block declares the params that were config params
only, so that the params record of an including pipeline controls them.

Changed signatures: FASTQ_QC_TRIM_FILTER_SETSTRANDEDNESS (fq_lint_args, umi_extract_args, fastp_args,
trimgalore_args), FASTQ_FASTQC_UMITOOLS_FASTP and FASTQ_FASTQC_UMITOOLS_TRIMGALORE (umi_extract_args and
the trimmer arguments).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…t [skip ci]

STAR_ALIGN, SENTIEON_STAR_ALIGN, PARABRICKS_RNA_FQ2BAM, HISAT2_ALIGN and BOWTIE2_ALIGN take their
arguments (defaults, extra arguments and read groups) in an args field of their input record
(StarAlignInput, SentieonStaralignInput, ParabricksRnafq2bamInput, Hisat2AlignInput, Bowtie2AlignInput).
The policy that was in the config closures is in starAlignArgs, hisat2AlignArgs and bowtie2AlignArgs
in tool_args.nf, shared by the three STAR processes. The align_star, align_hisat2 and align_bowtie2
config files are removed, and the align_star tests give the arguments in the input records.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ough the process input [skip ci]

SAMTOOLS_INDEX, UMITOOLS_DEDUP and UMICOLLAPSE take their arguments in an args field of their input
record (SamtoolsIndexInput, UmitoolsDedupInput, UmicollapseInput). The BAM sort, markduplicates and UMI
dedup subworkflows, and the aligner subworkflows that call them, take the index arguments and the UMI
grouping method and separator. The transcriptome BAM index of the UMI dedup now also uses --bam_csi_index,
like every other index in the run.

Changed signatures: BAM_SORT_STATS_SAMTOOLS, BAM_MARKDUPLICATES_PICARD, BAM_DEDUP_UMI,
BAM_DEDUP_STATS_SAMTOOLS_UMITOOLS, BAM_DEDUP_STATS_SAMTOOLS_UMICOLLAPSE, ALIGN_STAR, ALIGN_BOWTIE2,
FASTQ_ALIGN_HISAT2 (index_args and, for the UMI ones, the grouping method and separator).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ough the process input [skip ci]

SALMON_QUANT, KALLISTO_QUANT, DESEQ2_QC and SUMMARIZEDEXPERIMENT_SUMMARIZEDEXPERIMENT take their
arguments, label or file prefix in their input record (SalmonQuantInput, KallistoQuantInput,
Deseq2QcInput, SummarizedexperimentInput). The Salmon library type and Kallisto strandedness policy is
in the pseudo-alignment subworkflow, DESeq2 QC arguments in tool_args.nf. The configs of the pseudo-alignment
subworkflow and of DESeq2 QC are removed.

Changed signatures: QUANTIFY_PSEUDO_ALIGNMENT (salmon_libtype, extra_salmon_args, extra_kallisto_args,
se_prefix), QUANT_TXIMPORT_SUMMARIZEDEXPERIMENT (se_prefix).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…t [skip ci]

SUBREAD_FEATURECOUNTS, RUSTQC, BRACKEN, MULTIQC and RIBODETECTOR take their arguments, or whether they run
on a GPU, in their input record (FeaturecountsInput, RustqcInput, BrackenInput, MultiqcInput,
RibodetectorInput). The featurecounts, rustqc and bracken config files are removed.

Changed signatures: BAM_QC_RNASEQ (featurecounts_feature_type), FASTQ_REMOVE_RRNA and
FASTQ_QC_TRIM_FILTER_SETSTRANDEDNESS (use_gpu_ribodetector), MULTIQC_RNASEQ (multiqc_args).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…rocess input [skip ci]

STAR_GENOMEGENERATE, KALLISTO_INDEX, SALMON_INDEX, CUSTOM_GTFFILTER and CUSTOM_CATADDITIONALFASTA take their
arguments or file prefix in their input record (StarGenomegenerateInput, KallistoIndexInput,
SalmonIndexInput, CustomGtffilterInput, CustomCatadditionalfastaInput). Their config closures are removed;
no process config reads a pipeline param any more.

Changed signatures: PREPARE_GENOME_REFERENCES (skip_gtf_transcript_filter, genome), PREPARE_GENOME_INDICES
(prokaryotic, salmon_index_args, kallisto_index_args), FASTQ_SUBSAMPLE_FQ_SALMON (index_args),
FASTQ_QC_TRIM_FILTER_SETSTRANDEDNESS (salmon_index_args).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…nd document the convention [skip ci]

Every pipeline param is declared, with its default, in the params block of main.nf. conf/params.config
holds the params that only affect config settings. CONTRIBUTING and the usage docs describe that process
config does not read pipeline params and what an including pipeline needs.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…eline params [skip ci]

The samplesheet output read the global aligner and skip_quantification_merge params, which are those of
an including pipeline, so the BAM paths of the rows were wrong and the run warned about undefined params.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ries the pipeline's own schema [skip ci]

conf/modules.config is the name the pipeline composition ADR and nf-core use for the process config an
including pipeline reuses. The parameter summary, the completion summary and the MultiQC workflow summary
used a bare nextflow_schema.json, which is looked up in the project that runs (the including project
when this pipeline is included), so they use the pipeline's own. The banner and the methods description
no longer fail when the manifest has no doi. nf-schema 2.7.2-channel.4.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…en it is set [skip ci]

A user or an including pipeline that sets ext.args for a tool replaces the arguments the pipeline
builds for it, as when the pipeline runs directly, instead of having them added after.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…enome indices and MultiQC take it [skip ci]

ToolArgs holds the arguments and options that follow from the params. NFCORE_RNASEQ builds it once and
hands it to RNASEQ and PREPARE_GENOME_INDICES, which replaces their per-tool arguments with one tool_args
entry that declares only the fields it reads (as MULTIQC_RNASEQ does).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… ci]

FASTQ_QC_TRIM_FILTER_SETSTRANDEDNESS replaces six per-tool arguments with tool_args, which it reads for fq lint
and forwards; the trimming, fastp, rRNA removal and strandedness-inference subworkflows each declare the
fields they read.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…h the BAM and aligner subworkflows [skip ci]

The nine subworkflows that forwarded index_args to the SAMTOOLS_INDEX call take tool_args instead; the
four that run the index declare the one field they read and attach only that string to the process input.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…lds they read [skip ci]

ALIGN_STAR, ALIGN_BOWTIE2, FASTQ_ALIGN_HISAT2, BAM_DEDUP_UMI, FASTQ_QC_TRIM_FILTER_SETSTRANDEDNESS and
QUANTIFY_PSEUDO_ALIGNMENT took tool_args as an untyped Record. Each now declares a record with the fields
that it and the subworkflows it forwards to read, so a missing field is a lint error.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ass [skip ci]

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
pinin4fjords and others added 5 commits October 1, 2026 21:13
…erted by a helper [skip ci]

The module compared a String? cast to a memory unit, which the runtime type check cannot resolve.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…group test args [skip ci]

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…on tests for params and sample inputs [skip ci]

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…t rows by path before merging runs; align indices test setup inputs [skip ci]

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…lace; scope rRNA seeds to FASTQ_REMOVE_RRNA [skip ci]

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
pinin4fjords and others added 3 commits October 2, 2026 16:21
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…nd of deseq2_qc.r

nextflow config resolves the plugins block, and the patched nf-schema is not in
the public registry.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@bentsherman

bentsherman commented Oct 7, 2026 •

Copy link
Copy Markdown

Comparing this PR to my approach with methylseq (nf-core/methylseq#626) and what I plan to do, since this PR goes beyond my methylseq PR in some ways.

Params and input. Taking params.input as a Channel<SampleRow> looks right. Passing params into NFCORE_RNASEQ looks right. Consider collapsing NFCORE_RNASEQ into the entry workflow to reduce boilerplate. I believe the PIPELINE_* workflows are absorbed by the nf-core-utils plugin.

Tool args. We used the same pattern: add ext settings to process inputs, allow runtime override with task.ext.args ?: <args input> ?: ''. We both extracted the nastier tool args to helper functions to avoid bloating the workflow logic. I first build an options map and then render it to a string with a small cli() helper, which avoids string parsing. I don't have a strong opinion on the ToolArgs record vs passing tool args separately.

Module config. conf/modules/ keeps the settings that don't depend on params. I went ahead and moved everything into pipeline code so that I could delete conf/modules/ entirely. I recommend you do the same, since it makes the pipeline easier to compose.

Smaller things:

  • Params block should go after includes
  • Boolean params default to false so you don't need Boolean = false
  • The params.config seems like an unnecessary level of indirection. Nevermind, I see now it is used by the meta-pipeline
  • Wrapping process inputs with channel.value() is unnecessary

@pinin4fjords

Copy link
Copy Markdown
Member Author

Thanks for that @bentsherman ! The reason for the similarity is that I shamelessly stole some of your work once I realised the issues with params scopes etc and and my existing arg handling.

That said, the args handling will be a LOT more contentious in nf-core, I'm a bit more nervous of this one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants