Repository navigation
Pipeline composition - #7213
Pipeline composition#7213
Conversation
This comment was marked as outdated.
This comment was marked as outdated.
|
Great write up, thanks for this Ben! As you might expect, I'm most concerned about the params. You characterise it as a one-off cost which is mitigated by LLMs, however that doesn't take into account updates to included pipelines (a core functionality with included modules). The I'd still love to look into how we could bulk import nested config and apply it at root level. Even if it is a separate import + apply mechanism (eg. like config profiles in a sense?). I think without it, the use of the meta pipeline functionality is substantially limited. |
This comment was marked as resolved.
This comment was marked as resolved.
This comment was marked as resolved.
This comment was marked as resolved.
This comment was marked as resolved.
This comment was marked as resolved.
This comment was marked as resolved.
This comment was marked as resolved.
This comment was marked as resolved.
This comment was marked as resolved.
@adamrtalbot agreed, I never said global. I would love it if the pipeline config is imported within a dedicated scope and treated as a baseline default. Then the import-ing pipeline can override anything, but doesn't need to duplicate config that isn't being changed. Doing this would not be trivial. The only way I can think of is to do something fairly radical like rendering the config at import time and saving that to a locked config file somewhere. Or some other crazy mechanism. |
Config or params? In my mind they are very different concepts, I was referring to parameters here. |
This comment was marked as resolved.
This comment was marked as resolved.
Ideally params, but might need to be config for all the
Yeah as it stands I think this basically boils down to the functionality we already have with |
Can the nf-core tooling install a workflow from a pipeline repo? e.g. NFCORE_RNASEQ from nf-core/rnaseq? I think that is the main thing that this ADR adds |
This comment was marked as outdated.
This comment was marked as outdated.
edmundmiller
left a comment
There was a problem hiding this comment.
left a few thoughts on scope/clarity — overall the direction makes sense to me.
This comment was marked as resolved.
This comment was marked as resolved.
|
Note, my previous comment is based on the assumption is focused on sub-workflow only, not plain pipelines. |
1844652 to
4e9f0d2
Compare
ce77297 to
bcad1f6
Compare
|
Marking as draft until I finish my deep clean. Example is still valid |
cfe83b0 to
150e0bc
Compare
|
Thanks Ben, I went through the changes since my last review (07-01). The ADR reads much better framed as pipeline composition, and it's great to see it running end to end with the example. Most of my earlier points are addressed or out of scope now: the spec/lockfile question went away with remote inclusion, and the binding rules are concrete in Here's what I found: 1. Including a pipeline from two scripts crashes when the path contains a symlink. 2. Renaming 3. Stale docs. 4. 5. Unrelated behaviour change: 6. Minor
I'll follow up with a separate review focused on the composition semantics. |
|
Thanks Paolo. Going through your points: 1. Symlinked paths: already fixed 2. 3. Stale docs: fixed 4. Overridden params: agreed that an error would be better than silently dropping the value. I tried it, and it isn't a quick fix. I added my findings to the ADR as an open question, so we can look into it later if it becomes a common issue. 5. 6. |
pditommaso
left a comment
There was a problem hiding this comment.
Thanks Ben, all my points are addressed. I checked the changes on macOS: the composition tests pass, the example runs end to end, and an output index CSV now loads fine as a Channel param.
A few small things left, none of them blocking:
- Merge order with the language server: nextflow-io/language-server#183 needs to land together with the nf-lang bump, otherwise go-to-definition on named arguments would silently stop working.
- README: "
--rnaseq.inputis ignored" is only half true, because an invalid value still fails the run. Maybe point to the ADR open question instead. - Standalone runs: running
./pipelines/nf-core/rnaseqfrom the example root also loads the meta-pipeline'snextflow.config(the run shows asfetchngs-rnaseq). It's normal Nextflow behaviour, but it may be worth a note next to "it can still be run on its own". - CSV index (pre-existing):
CsvWriterquotes every value but doesn't escape an embedded", so such values won't load back. Could be a follow-up.
Thanks again for all the work on this!
Add an ADR for pipeline composition and implement it with an end-to-end example. Signed-off-by: Ben Sherman <bentshermann@gmail.com>
7496ecb to
91766b8
Compare
…atched nf-schema, deploy output via outputDir [skip ci] The typed pipeline needs nextflow-io/nextflow#7213, #7646 and #7674, which are not in a release yet, so the nf-test jobs run a launcher built from a pinned upstream commit, and the patched nf-schema comes from a plugin repository. The AWS test workflows set the output location with the outputDir config setting, since the outdir param is gone. nf-test no longer lists bin/ as a trigger and lists conf/. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This PR adds an ADR for remote pipeline inclusion, aka "meta-pipelines".
It describes an approach for including remote pipelines into a meta-pipeline in a way that preserves dataflow concurrency between pipeline inputs/outputs.
It discusses alternative approaches such as pipeline chaining / nf-cascade and why they don't satisfy certain use cases (preserving dataflow concurrency).
It also walks through a basic example of fetchngs -> rnaseq.