Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
461 changes: 461 additions & 0 deletions adr/20260608-pipeline-composition.md

Large diffs are not rendered by default.

6 changes: 1 addition & 5 deletions docs/modules/using-modules.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -134,11 +134,7 @@ See [module run][cli-module-run] for the full command reference.

Parameters are inferred from the module's declared inputs: the `input:` section for a process, or the `take:` section for a named workflow.

Type conversions are handled the same way as [typed parameters][typed-params]. Workflow modules use the following additional rules for dataflow types:

- A `Channel<E>` input accepts a samplesheet path, which Nextflow loads as a channel of records. The samplesheet file can be CSV, JSON, or YAML. The element type must be `Map`, `Record`, or a record type. Each row is validated against and converted to the declared type.

- A `Value<V>` input accepts a value of type `V`, which Nextflow wraps in a value channel.
Type conversions are handled the same way as [typed parameters][typed-params], including the rules for dataflow types (`Channel` and `Value`).

Consider the following workflow:

Expand Down
16 changes: 15 additions & 1 deletion docs/reference/syntax.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -88,7 +88,7 @@ include { hello as sayHello } from './some/module'
The include source should be a string literal. Each include clause should specify a name, and may also specify an *alias*. In the above example, `hello` is included under the alias `sayHello`.

:::note
[Enum](#enum-type) and [record](#record-type) types cannot be aliased. They must be included under their original name.
[Enum](#enum-type) and [record](#record-type) types cannot be aliased. They must be included under their original name. The `params` block of an included pipeline is the exception -- it is a type created for the include, so it must be aliased.
:::

Include clauses can be separated by semi-colons or newlines:
Expand All @@ -111,6 +111,19 @@ include { hello } from './some/module'
include { bye as goodbye } from './some/module'
```

<AddedInVersion version="26.10" />

The `workflow` name refers to the entire pipeline defined by the included script, i.e., its params block, entry workflow, and output block. The `params` name refers to the params block as a record type. Each of these names must be given an alias:

```nextflow
include {
params as RnaseqParams ;
workflow as RNASEQ
} from './pipelines/rnaseq.nf'
```

See [pipeline composition][pipeline-composition] for more information.

### Params block

The params block consists of one or more *parameter declarations*. A parameter declaration consists of a name, type, and an optional default value:
Expand Down Expand Up @@ -1015,3 +1028,4 @@ See [strict syntax][strict-syntax-page] for more information.
[strict-syntax-page]: ../strict-syntax
[workflow-output-def]: ../workflow#outputs
[workflow-typed-page]: ../workflow-typed
[pipeline-composition]: ../workflow-typed#pipeline-composition
38 changes: 37 additions & 1 deletion docs/typed-parameters.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,42 @@ The `params` block does not require the `nextflow.enable.types` feature flag. Fo

## Supported types

Parameters can use any [standard type][stdlib-types] except the dataflow types (`Channel` and `Value`). This includes primitive types such as `Path`, `String`, `Integer`, and `Boolean`, as well as collections and records.
Parameters can use any [standard type][stdlib-types]. This includes primitive types such as `Path`, `String`, `Integer`, and `Boolean`, as well as collections, records, and the dataflow types (`Channel` and `Value`).

<AddedInVersion version="26.10" />

A parameter can be declared with a dataflow type:

- A `Channel<E>` parameter accepts a samplesheet path, which Nextflow loads as a channel of records. The samplesheet file can be CSV, JSON, or YAML. The element type must be `Map`, `Record`, or a record type. Each row is validated against and converted to the declared type.

- A `Value<V>` parameter accepts a value of type `V`, which Nextflow wraps in a value channel.

For example:

```nextflow
params {
samples: Channel<Sample>
index: Value<Path>
}

workflow {
RNASEQ(params.samples, params.index)
}

record Sample {
id: String
fastq_1: Path
fastq_2: Path
}
```

The pipeline can be run with a samplesheet and an index file:

```console
$ nextflow run main.nf --samples samples.csv --index genome.fa
```

Dataflow types are useful when a pipeline is [included by another pipeline][pipeline-composition], because they allow the calling pipeline to provide a parameter from dataflow logic.

## Default and required parameters

Expand Down Expand Up @@ -66,6 +101,7 @@ The language server validates each parameter reference against its declared type
- [Pipeline parameters][cli-params]: How parameter values are resolved across sources.

[cli-params]: ./cli#pipeline-parameters
[pipeline-composition]: ./workflow-typed#pipeline-composition
[static-typing-page]: ./static-typing
[stdlib-types]: ./reference/stdlib-types
[syntax-params-block]: ./reference/syntax#params-block
124 changes: 124 additions & 0 deletions docs/workflow-typed.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -62,6 +62,130 @@ The operator library supports static typing and records. All operators work in b

For more information about best practices when migrating existing code, see [Using operators with static typing][migrating-static-types-operators].

## Pipeline composition

<AddedInVersion version="26.10" />

:::warning{title="Experimental: may change in a future release."}
:::

An entire pipeline -- the `params` block, entry workflow, and `output` block of a script -- can be included as a named workflow and called like any other workflow. The params block acts as the `take:` section and the output block acts as the `emit:` section.

Given the following pipeline:

```nextflow
// pipelines/rnaseq.nf
nextflow.enable.types = true

params {
input: Channel<Sample>
aligner: String = 'star_salmon'
fasta: Path
}

workflow {
main:
// ...

publish:
bams = ch_bams
multiqc = val_multiqc
}

output {
bams: Channel<Path> { path 'bams' }
multiqc: Path { path 'multiqc' }
}
```

It can be included and called as follows:

```nextflow
include { workflow as RNASEQ } from './pipelines/rnaseq.nf'

workflow {
main:
rnaseq = RNASEQ(record(
input: samples,
fasta: file('index.fasta')
))
rnaseq.bams.view() // Channel<Path>
rnaseq.multiqc.view() // Value<Path>
}
```

The included pipeline must be aliased to a specific name (`RNASEQ`), which is also used to scope its processes in the config:

```nextflow
process {
withName: 'RNASEQ:STAR_ALIGN' {
cpus = 12
memory = 72.GB
}
}
```

Because the included pipeline is part of the calling pipeline's dataflow graph, it can consume a channel produced by another pipeline, and begin working on each item as soon as it is emitted.

Note the following:

- Both the calling script and the included pipeline must enable static typing (`nextflow.enable.types = true`).

- The pipeline is called with a single record, with one field for each param. Params with a default value can be omitted.

- The outputs published by the included pipeline are emitted to the calling workflow instead of being published. Only the calling pipeline decides what is published, by declaring its own `output` block.

- Params and outputs are not inherited. The calling pipeline declares its own params and outputs and passes them explicitly to the included pipeline.

### Importing the params block

Redeclaring every param of an included pipeline gets tedious. The `params` block can be imported as a record type instead:

```nextflow
include {
params as RnaseqParams ;
workflow as NFCORE_RNASEQ
} from './pipelines/rnaseq.nf'

params {
strandedness: String = 'auto' // unique to this pipeline
rnaseq: RnaseqParams // input, aligner, fasta
}

workflow {
main:
rnaseq = NFCORE_RNASEQ( params.rnaseq + record(input: samples) )
}
```

`RnaseqParams` is a *partial* record type: every field is nullable, so a user can provide any rnaseq param as `--rnaseq.<name>`, the calling pipeline can override specific params with `params.rnaseq + record(input: samples)`, and the `NFCORE_RNASEQ()` call reports any param that is still missing.

Note the following:

- A param that the pipeline defaults (`aligner`) can be omitted from the record. The pipeline applies its own default when it is called.

- `rnaseq.input` is supplied by the dataflow, which overrides any value given by the user.

- `rnaseq.fasta` must still be provided, but the error surfaces at the `NFCORE_RNASEQ()` call rather than at launch.

A pipeline receives its params when it is called, so `params` refers to a single execution of the pipeline. A pipeline can be included under any number of aliases, and each alias resolves its own params. As with a named workflow, a pipeline can be called only once per alias -- include it again under a different alias to call it again.

### Best practices

Pipeline inclusion only captures the pipeline script and the modules it includes. It does not capture external context such as the config or the `lib` directory. As a result, an included pipeline should be written so that it works when included by another pipeline:

- Pipeline parameters should be declared in the `params` block. The config should only declare *config params*, i.e. params that only affect config settings.

- Project-level assets (`projectDir`, `bin`, `lib`) should not be used, since the calling pipeline has a different project root. Module-level assets can be safely used through the module `resources/` bundle and `moduleDir`.

- Params should be referred to only in the entry workflow and output block. A process or workflow that reads `params` directly should declare an explicit input instead.

- Process configuration (`container`, `conda`, `ext`) should be specified in the process definition or avoided in favor of process inputs.

- Workflow outputs should be published using the `output` block, not `publishDir`.

None of these constraints are absolute. Each of them can be circumvented by replicating the external context in the calling pipeline. Following them simply makes it easier to include a pipeline with minimal extra work.

## Validation

The language server validates each workflow input and output against its declared type. Calling a workflow with an argument whose type does not match its `take:` declaration, or emitting a value that does not match its `emit:` declaration, is reported as an error before the pipeline runs.
Expand Down
4 changes: 4 additions & 0 deletions examples/pipeline-composition/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
# Nextflow run artifacts
.nextflow*
work/
results/
Loading
Loading