Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
109 commits
Select commit Hold shift + click to select a range
ff7289b
WIP add custom RefCDS
FerriolCalvet Jan 4, 2025
d9cb420
WIP
FerriolCalvet Jan 5, 2025
c2c91bd
update biomart query
FerriolCalvet Jan 6, 2025
3bb1182
partially tested updates
FerriolCalvet Jan 8, 2025
1798fc2
add old branch + new python script
FerriolCalvet Apr 10, 2026
3ba2385
dNdScv dynamic reference working
FerriolCalvet Apr 11, 2026
10de279
fix review comments
FerriolCalvet Apr 11, 2026
c74c70b
Merge branch 'dev' into feat/dynamic-refcds
FerriolCalvet Apr 11, 2026
0e98dcf
update biomart querying by batches
FerriolCalvet Apr 12, 2026
3537285
apply review comments
FerriolCalvet Apr 14, 2026
660131a
update biomart query (may not work)
FerriolCalvet Apr 14, 2026
148a486
temp update to latest version
FerriolCalvet Apr 15, 2026
26ec2a6
update query to filter canonical
FerriolCalvet Apr 16, 2026
55b72e2
fix file storage in biomart querying
FerriolCalvet Apr 16, 2026
b34a5bb
reduce batch size
FerriolCalvet Apr 16, 2026
7bd6c2d
update waiting time
FerriolCalvet Apr 16, 2026
54430ff
fix bug in dna2protein
FerriolCalvet Apr 21, 2026
c6f7503
Add depth-vs-metric QC plots and summaries for mutdensity/omega outpu…
Copilot Apr 22, 2026
32edc25
Fix dna2protein querying (#446)
FerriolCalvet Apr 26, 2026
3d120ae
Add configurable VAF distortion threshold (#447)
Copilot Apr 29, 2026
eae311e
fix bug in storing panel
FerriolCalvet May 1, 2026
6fdfe17
remove internal query
FerriolCalvet May 1, 2026
abc50b0
add biomart reference creation information
FerriolCalvet May 1, 2026
46b495e
fix bug in depths_comparison script
FerriolCalvet May 1, 2026
fc84613
working version
FerriolCalvet May 1, 2026
f8de01c
update outputs
FerriolCalvet May 1, 2026
a45688b
remove unnecessary code
FerriolCalvet May 1, 2026
36d7aef
Merge pull request #444 from bbglab/feat/dynamic-refcds
FerriolCalvet May 4, 2026
1e0a122
fix bug in selected_genes processing
FerriolCalvet May 5, 2026
2734934
add dn/ds proxy (#449)
FerriolCalvet May 6, 2026
498559a
Apply multiple-testing correction to omega outputs (#452)
Copilot May 20, 2026
f1cc220
Update deepCSA documentation (#457)
m-huertasp May 27, 2026
7fe4437
fix minor bugs in documentation
FerriolCalvet May 27, 2026
51bfe75
fix flagged in plot selection summary
FerriolCalvet May 27, 2026
86d49b9
add pvalue_adj to omega_flagged
FerriolCalvet May 27, 2026
5f7ff6e
full fix for the bug
FerriolCalvet May 27, 2026
a06c90a
Expand deepCSA documentation coverage and output interpretation (#448)
Copilot May 27, 2026
1d41152
bug fixed
FerriolCalvet May 27, 2026
79b624a
several bug fixes and updates (#465)
FerriolCalvet May 31, 2026
e47ab1c
Add scripts to assess the panel (#467)
FerriolCalvet May 31, 2026
d012b74
Apply per-gene multiple testing correction for site selection metrics…
Copilot Jun 1, 2026
f199f73
feat: add somatic/germline mask to utils_filter
m-huertasp Jun 2, 2026
30b8521
test: add unit test for all utils filter functions
m-huertasp Jun 2, 2026
39e5bb7
refactor: use shared helper in filter_cohort
m-huertasp Jun 2, 2026
8d1ce9a
refactor: use helper in check_contamination
m-huertasp Jun 2, 2026
07f4e27
docs: add docstrings to check contamination
m-huertasp Jun 2, 2026
d39c828
docs: fix typo in docstring
m-huertasp Jun 2, 2026
38bc0c1
refactor: use somatic threshold parameter in check contamination
m-huertasp Jun 2, 2026
f207ec4
style_ remove unused import and f-string prefixes
m-huertasp Jun 2, 2026
7a448f6
fix: added loc for df division
m-huertasp Jun 2, 2026
66c8936
Add VAF pseudocount exploration (#470)
FerriolCalvet Jun 3, 2026
1e5288c
chore: remove dead SUBSETPILEUP and READSPOSBED
m-huertasp Jun 4, 2026
a4a0131
chore: remove dead withName selectors
m-huertasp Jun 4, 2026
f53ff69
chore: remove dead SUBSETPANEL and MUTRATE from exome.config
m-huertasp Jun 5, 2026
bee9b6e
chore: remove dead CREATESAMPLEPANELS
m-huertasp Jun 5, 2026
66902a1
chore: remove orphan CREATESAMPLEPANELS module
m-huertasp Jun 5, 2026
173d2cc
chore: strip names prefix on elegible selectors
m-huertasp Jun 5, 2026
495e29c
docs: add comment for selectors raising warnings
m-huertasp Jun 5, 2026
721185f
chore: remove redundant .first() on already-pinned channels
m-huertasp Jun 5, 2026
a691c9d
chore: pin singleton channels at source and drop redundant .first()
m-huertasp Jun 5, 2026
0ca40e3
fix: omega preprocessing and estimator selector warning
m-huertasp Jun 8, 2026
ddfae4c
docs: add total path for omega selectors
m-huertasp Jun 8, 2026
094e2ac
fix: remove selector warning sigprofilerassignment
m-huertasp Jun 8, 2026
95e3cf3
refactor: simplify and remove selector warning for expand panels
m-huertasp Jun 8, 2026
44d0381
refactor: remove selector warnings for regressions
m-huertasp Jun 8, 2026
5a266fe
chore: remove ".first()" warning in mutatiopreprocessing
m-huertasp Jun 8, 2026
0a04146
fix: remove stray closing paren in exome.config withName selector
m-huertasp Jun 8, 2026
04e98a8
docs: remove non used comment
m-huertasp Jun 8, 2026
91100a8
remove: _plot_selectionfeatures.py (unused)
m-huertasp Jun 9, 2026
5480178
chore: update reference datasets paths
m-huertasp Jun 12, 2026
dcd1bec
update datasets defaults
FerriolCalvet Jun 15, 2026
f9609fc
Fix null values in plot saturation (#476)
FerriolCalvet Jun 17, 2026
0c6ea74
fix error in compute proportions when missing panel
FerriolCalvet Jun 17, 2026
48d885a
update scripts
FerriolCalvet Jun 21, 2026
2dbb2b0
working mutdensity plots
FerriolCalvet Jun 22, 2026
4b6653b
fix concat profile error when single sample
FerriolCalvet Jun 27, 2026
82b6999
tests: check contamination initial test
m-huertasp Jul 7, 2026
a68ad3f
refactor: move heatmap creation to parametrized function
m-huertasp Jul 10, 2026
1f0c806
refactor: extract prepare datasets
m-huertasp Jul 10, 2026
b93e547
refactor: extract somatic vs germline
m-huertasp Jul 10, 2026
88354a7
refactor: common code to single function
m-huertasp Jul 13, 2026
d2196a1
refactor: apply two_way_comparison to all possible
m-huertasp Jul 13, 2026
e91717b
refactor: exploration of contaminated samples
m-huertasp Jul 13, 2026
19dee49
refactor: contamination detection snps
m-huertasp Jul 13, 2026
8e9d95f
fix: not plotting x-axis labels
m-huertasp Jul 22, 2026
ec39c36
Merge pull request #469 from bbglab/tests/418-uniformize-somatic-vari…
m-huertasp Jul 22, 2026
11039c9
Compute selection for known hotspots analysis (#474)
Copilot Jul 24, 2026
96be6c9
update oncodrive3d documentation
FerriolCalvet Jul 24, 2026
3e4f851
refactor: remove duplicated selector for MUTATEDGENOMESFROMVAFAM
m-huertasp Jul 28, 2026
3433992
Merge pull request #472 from bbglab/chore/clean-nextflow-warnings
FerriolCalvet Jul 28, 2026
a3fcac9
add VAF_smoothing
FerriolCalvet Jul 29, 2026
0a132f7
add exploratory notebook
FerriolCalvet Jul 29, 2026
a0deaa8
working version also with plotting
FerriolCalvet Jul 29, 2026
616c6e1
fix output format
FerriolCalvet Jul 29, 2026
c538edc
handle errors in plotting
FerriolCalvet Jul 30, 2026
7192c2f
add filter of AM_not_supported in default filters
FerriolCalvet Jul 30, 2026
c6104ac
complete analysis and output reorganization
FerriolCalvet Jul 30, 2026
272ef6f
QC plot on omega vs dndscv values comparison (#479)
FerriolCalvet Jul 31, 2026
75d6386
working version of whole-genome scaled mutation rate
FerriolCalvet Jul 31, 2026
2415947
complete version
FerriolCalvet Aug 1, 2026
a10608e
compile and organize outputs
FerriolCalvet Aug 5, 2026
630161f
remove exploration notebook
FerriolCalvet Aug 5, 2026
6910e33
ignore notebooks
FerriolCalvet Aug 5, 2026
f093363
define panel version properly
FerriolCalvet Aug 5, 2026
5ed0d7a
fix hotspots table matching
FerriolCalvet Aug 6, 2026
496011b
update batch size in gene coverage summary
FerriolCalvet Aug 7, 2026
9e34bdf
update VAF smoothing computations and plotting
FerriolCalvet Aug 7, 2026
6e3f8c7
fix outputs in stub
FerriolCalvet Aug 7, 2026
3b16560
Merge pull request #483 from bbglab/vaf-smoothingv2
FerriolCalvet Aug 7, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -15,3 +15,4 @@ scratch/
scratchhhh/
tests/test_data/all_samples.somatic.mutations.maf
tests/test_data/all_samples_indv.depths.tsv.gz
assets/useful_scripts/*.ipynb
60 changes: 54 additions & 6 deletions CITATIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,10 @@

## Sources of data and tools

- **Ensembl VEP**

> McLaren W, Gil L, Hunt SE, et al. The Ensembl Variant Effect Predictor. Genome Biol. 2016;17:122. doi: 10.1186/s13059-016-0974-4.

- Nanoseq masks

> Abascal, F., Harvey, L.M.R., Mitchell, E. et al. Somatic mutation landscapes at single-molecule resolution. Nature 593, 405–410 (2021). https://doi.org/10.1038/s41586-021-03477-4
Expand All @@ -22,7 +26,7 @@

> https://cancer.sanger.ac.uk/signatures/sbs

- **dNdScv covariates**
- **dNdScv method + covariates**

> Martincorena I, et al. (2017) Universal Patterns of Selection in Cancer and Somatic Tissues. Cell. http://www.cell.com/cell/fulltext/S0092-8674(17)31136-4

Expand All @@ -38,11 +42,55 @@

> Ewels P, Magnusson M, Lundin S, Käller M. MultiQC: summarize analysis results for multiple tools and samples in a single report. Bioinformatics. 2016 Oct 1;32(19):3047-8. doi: 10.1093/bioinformatics/btw354. Epub 2016 Jun 16. PubMed PMID: 27312411; PubMed Central PMCID: PMC5039924.

- Python
- SigProfilerAssignment, MatrixGenerator
- HDP
- OncodriveFML
- OncodriveCLUSTL
- **bgreference / bgdata**

> Repository: https://github.com/bbglab/bgreference
> Repository: https://github.com/bbglab/bgdata

- **SAMtools**

> Li H, Handsaker B, Wysoker A, et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25(16):2078-9. doi: 10.1093/bioinformatics/btp352.

- **BEDTools**

> Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010;26(6):841-842. doi: 10.1093/bioinformatics/btq033.

- **HTSlib / Tabix**

> Li H. Tabix: fast retrieval of sequence features from generic TAB-delimited files. Bioinformatics. 2011;27(5):718-719. doi: 10.1093/bioinformatics/btq671.

- **Python**

> Python Software Foundation. Python Language Reference, version 3.x. https://www.python.org/

- **SigProfilerAssignment, SigProfilerMatrixGenerator**

> Alexandrov LB, et al. The repertoire of mutational signatures in human cancer. Nature 578, 94–101 (2020). doi:10.1038/s41586-020-1943-3.
> SigProfiler tools (SigProfilerMatrixGenerator, SigProfilerAssignment). Alexandrov Lab. https://github.com/AlexandrovLab
> Repository: https://github.com/AlexandrovLab/SigProfilerAssignment
> Repository: https://github.com/AlexandrovLab/SigProfilerMatrixGenerator


- **HDP (Hierarchical Dirichlet Processes)**

> https://github.com/nicolaroberts/hdp
> Roberts, N. D. (2018). Patterns of somatic genome rearrangement in human cancer. https://doi.org/10.17863/CAM.22674

- **OncodriveFML**

> Repository: https://github.com/bbglab/oncodrivefml (see repository for citation details)
> Mularoni L, Sabarinathan R, Deu-Pons J, Gonzalez-Perez A, Lopez-Bigas N. OncodriveFML: a general framework to identify coding and non-coding regions with cancer driver mutations. Genome Biology. 2016;17:128. doi:10.1186/s13059-016-0994-0. https://github.com/bbglab/oncodrivefml

- **OncodriveCLUSTL**

> Repository: https://github.com/bbglab/oncodriveclustl (see repository for citation details)
> Claudia Arnedo-Pac, Loris Mularoni, Ferran Muiños, Abel Gonzalez-Perez, Nuria Lopez-Bigas, OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers, Bioinformatics, Volume 35, Issue 22, November 2019, Pages 4788–4790, https://doi.org/10.1093/bioinformatics/btz501; https://github.com/bbglab/oncodriveclustl

- **Omega (dN/dS)**

> Repository: https://github.com/bbglab/omega (see repository for citation details)



## Software packaging/containerisation tools

Expand Down
58 changes: 30 additions & 28 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,34 @@

![deepCSA workflow overview](docs/images/deepCSA.png)

## Documentation

Find the documentation ([link to docs](https://github.com/bbglab/deepCSA/tree/main/docs)).

We are working to provide the biggest possible detail on the [usage](docs/usage.md) and explanation of the rationale and [tools](docs/tools.md).

You can also find an explanation with examples on the [output](docs/output.md).

For more examples and description of the entire process, you can check the publications listed below.

## Publications

> **Sex and smoking bias in the selection of somatic mutations in human bladder**
>
> Ferriol Calvet*, Raquel Blanco Martinez-Illescas*, Ferran Muiños, Maria Tretiakova, Elena S. Latorre-Esteves, Jeanne Fredrickson, Maria Andrianova, Stefano Pellegrini, Axel Rosendahl Huber, Joan Enric Ramis-Zaldivar, Shuyi Charlotte An, Elana Thieme, Brendan F. Kohrn, Miguel L. Grau, Abel Gonzalez-Perez, Nuria Lopez-Bigas & Rosa Ana Risques
>
>_Nature_ (2025) doi:[10.1038/s41586-025-09521-x](https://doi.org/10.1038/s41586-025-09521-x)
>
> *these authors contributed equally and the order was decided randomly

&

> **DeepClone, an end-to-end protocol to study somatic mutagenesis and selection at high resolution**
>
> Ferriol Calvet, Morena Pinheiro-Santin, Erika Lopez, Raquel Blanco Martinez-Illescas, Núria Samper, Miguel L. Grau, Ferran Muiños, Rocío Chamorro González, Maria Andrianova, Federica Brando, Stefano Pellegrini, Marta Huertas, Elisabet Figuerola-Bou, Coohleen Coombes, Brendan F. Kohrn, Jeanne Fredrickson, Rosa Ana Risques, Nuria Lopez-Bigas, Abel Gonzalez-Perez
>
> _protocols.io_ (2026) doi:[10.17504/protocols.io.dm6gp1jodgzp/v2](https://dx.doi.org/10.17504/protocols.io.dm6gp1jodgzp/v2)

## Usage

You can find a detailed documentation in the [docs section](docs/README.md), but here there is a minimal summary on how to prepare the inputs. Still for your first runs if you need to make the complete set up you have to check the deeper documentation.
Expand All @@ -22,6 +50,8 @@ sample2,sample2.filtered.vcf,sample2.sorted.bam

Each row represents a single sample with a single-sample VCF containing the mutations called in that sample and the BAM file that was used for getting those variant calls. The mutations will be obtained from the VCF and the BAM file will be used for computing the sequencing depth at each position and using this for the downstream analysis.

Two alternative input modes are also supported: a samplesheet with only `sample,vcf` columns combined with a precomputed depths table, or a single cohort-level MAF file passed via `--input_maf` together with a precomputed depths table. See [Input scenarios](docs/input_scenarios.md) for details.

**Make sure that you do not use any '.' in your sample names, and also use text-like names for the samples, try to avoid having only numbers.** This second case should be handled properly but using string-like names will ensure consistency.

**There are specific datasets that need to be prepared before running deepCSA. You can find a list of those, and instructions for downloading them in [the documentation section of the repo](docs/usage.md#mandatory-parameter-configuration).**
Expand Down Expand Up @@ -57,31 +87,3 @@ We thank the following people for their extensive assistance in the development
An extensive list of references for the tools used by the pipeline can be found in the [`CITATIONS.md`](CITATIONS.md) file.

This pipeline uses code and infrastructure developed and maintained by the [nf-core](https://nf-co.re) community, reused here under the [MIT license](https://github.com/nf-core/tools/blob/master/LICENSE).

> **The nf-core framework for community-curated bioinformatics pipelines.**
>
> Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso & Sven Nahnsen.
>
> _Nat Biotechnol._ 2020 Feb 13. doi: [10.1038/s41587-020-0439-x](https://dx.doi.org/10.1038/s41587-020-0439-x).

## Documentation

Find the documentation ([link to docs](https://github.com/bbglab/deepCSA/tree/main/docs)).

We are working to provide the biggest possible detail on the [usage](docs/usage.md) and explanation of the rationale and [tools](docs/tools.md), but this is still in progress.

## Publications

> **Sex and smoking bias in the selection of somatic mutations in human bladder**
>
> Ferriol Calvet*, Raquel Blanco Martinez-Illescas*, Ferran Muiños, Maria Tretiakova, Elena S. Latorre-Esteves, Jeanne Fredrickson, Maria Andrianova, Stefano Pellegrini, Axel Rosendahl Huber, Joan Enric Ramis-Zaldivar, Shuyi Charlotte An, Elana Thieme, Brendan F. Kohrn, Miguel L. Grau, Abel Gonzalez-Perez, Nuria Lopez-Bigas & Rosa Ana Risques
>
>_Nature_ (2025) doi:[10.1038/s41586-025-09521-x](https://doi.org/10.1038/s41586-025-09521-x)
>
> *these authors contributed equally and the order was decided randomly

> **DeepClone, an end-to-end protocol to study somatic mutagenesis and selection at high resolution**
>
> Ferriol Calvet, Morena Pinheiro-Santin, Erika Lopez, Raquel Blanco Martinez-Illescas, Núria Samper, Miguel L. Grau, Ferran Muiños, Rocío Chamorro González, Maria Andrianova, Federica Brando, Stefano Pellegrini, Marta Huertas, Elisabet Figuerola-Bou, Coohleen Coombes, Brendan F. Kohrn, Jeanne Fredrickson, Rosa Ana Risques, Nuria Lopez-Bigas, Abel Gonzalez-Perez
>
> _protocols.io_ (2026) doi:[10.17504/protocols.io.dm6gp1jodgzp/v2](https://dx.doi.org/10.17504/protocols.io.dm6gp1jodgzp/v2)
44 changes: 44 additions & 0 deletions assets/assess_panel/assess_panel.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
# Experimental design principles

## Definition of relationship between number of mutations detected and sequencing depth

$$
\text{Expected number of mutations} = (\text{ average depth * region of interest (bp) }) \times (\text{expected mutation density (muts/bp)})
$$

## Use cases

### Predict average depth required

Given a fixed panel size based on which are the regions of interest, a desired number of mutations (enough to get good selection/mutagenesis metrics) and a known estimation of the expected mutation density of the sample, you can estimate which is the sequencing depth required using this formula:

$$
\text{average depth} = \frac{\text{minimum desired mutations}}{(\text{region of interest (bp)}) \times (\text{expected mutation density (muts/bp)})}
$$

To get an estimation of the expected mutation density you can use this formula even if it is with some approximate values:

$$
\text{expected mutation density (muts/bp)} = \text{mutation rate (mutations} \times \text{bp} \times \text{year)} \times \text{time (year)}
$$

For the mutation rate this could be a possible formula; but you may also rely on previously published data.

$$
\text{mutation rate (} \frac{\text{mutations}}{\text{bp * year}} \text{)} = \frac{\text{observed mutations}}{ \text{genome size} \times \text{age (year)} }
$$

In this example, we are assuming that the units of the mutation rate are mutations x genome x year, but this could be in any other units
Particularly the time units could be different from year.

In case that the sequencing depth is restricted by the availability of DNA or any other reason, you can decide to adjust the panel so that by increasing the depth you can still reach a stable measurement of the mutation density and other variables.

An additional scenario to consider is the trinucleotide mutation probabilities. In a given sample, the different sites may have a very different mutation probablility depending on which are the active mutational processes. If these are known, you can factor in this information to adjust the requirements of the experimental design.

### Expected number of mutations given an average depth

The average depth represents the number of times that you sequence your region of interest, then it can be seen as:

$$
\text{Expected number of mutations} = (\text{ average depth * region of interest (bp) }) \times (\text{expected mutation density (muts/bp)})}
$$
101 changes: 101 additions & 0 deletions assets/assess_panel/assess_panel.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,101 @@
#!/usr/bin/env python


import sys
import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt
import click


sys.path.append("../../bin")

from utils_context import triplet_context_iterator

# we need a stable background mutation density estimation,
# for this we need at least 3 mutations

def mutation_based_assessment(panel_tsv, expected_mutation_density_in_mb, samples, depth, mutational_profile = None):
if mutational_profile is None:
mutational_profile = pd.DataFrame(np.ones(96) / 96)
mutational_profile.index = triplet_context_iterator()
mutational_profile = mutational_profile / mutational_profile.mean()
mutational_profile.columns = ["Probability"]

size_per_consequence = panel_tsv.groupby(by = ["GENE", "IMPACT", "CONTEXT_MUT"]).size().reset_index(name = "DEPTH")
size_per_consequence["DEPTH"] *= depth/3

size_per_consequence = size_per_consequence.merge(mutational_profile,
left_on = "CONTEXT_MUT",
right_index = True,
how = "left")
size_per_consequence["ADJUSTED_DEPTH"] = size_per_consequence["DEPTH"] * size_per_consequence["Probability"]

total_adjusted_size = size_per_consequence.groupby(by = ["GENE", "IMPACT"])["ADJUSTED_DEPTH"].sum() / 3
expected_mutations_per_cnsq = (expected_mutation_density_in_mb * total_adjusted_size * samples) / 1e6

return expected_mutations_per_cnsq



@click.command()
@click.option('--panel', '-p', 'panel_path', required=True, type=click.Path(exists=True, dir_okay=False),
help='Path to panel TSV file')
@click.option('--expected-mutation-density', '-e', 'expected_mutation_density', default=1.2, show_default=True,
type=float, help='Expected mutation density (mutations per Mb)')
@click.option('--samples', '-s', default=100, show_default=True, type=int,
help='Number of samples to assume')
@click.option('--depth', '-d', default=5000, show_default=True, type=int,
help='Sequencing depth (used in size adjustment)')
@click.option('--mutational-profile', '-m', 'mutational_profile_path', default=None, type=click.Path(exists=True, dir_okay=False),
help='Optional mutational profile TSV with CONTEXT_MUT and Probability (or column named like all_samples.all)')
@click.option('--output-prefix', '-o', default=None, help='Optional prefix to write TSV outputs (will create <prefix>*.tsv)')
def main(panel_path, expected_mutation_density, samples, depth, mutational_profile_path, output_prefix):
"""CLI entrypoint for assess_panel.

Loads the panel TSV and optional mutational profile, runs the mutation-based assessment
(with and without the provided profile) and optionally writes results to TSV files.
"""
panel = pd.read_table(panel_path)

mutational_prof = None
if mutational_profile_path:
mutational_prof = pd.read_table(mutational_profile_path)
# Try to normalize common formats: expect CONTEXT_MUT as index and a column named Probability
if 'CONTEXT_MUT' in mutational_prof.columns:
# If the profile already has a Probability column, keep it; otherwise try to rename
if 'Probability' in mutational_prof.columns:
mutational_prof = mutational_prof.set_index('CONTEXT_MUT')[['Probability']]
elif 'all_samples.all' in mutational_prof.columns:
mutational_prof = mutational_prof.set_index('CONTEXT_MUT').rename(columns={'all_samples.all': 'Probability'})[['Probability']]
else:
# Fallback: use all other numeric column(s) and take the first as Probability
numeric_cols = mutational_prof.select_dtypes(include=[np.number]).columns.tolist()
if numeric_cols:
mutational_prof = mutational_prof.set_index('CONTEXT_MUT')[[numeric_cols[0]]].rename(columns={numeric_cols[0]: 'Probability'})
else:
# If no numeric columns, create uniform profile (will be normalized in function)
mutational_prof = None

# Run assessments
mutations_per_gene_consequence = mutation_based_assessment(panel, expected_mutation_density, samples, depth)
mutations_per_gene_consequence_mut_prof = mutation_based_assessment(panel, expected_mutation_density, samples, depth, mutational_prof)


signatures_based_assessment(panel_tsv, expected_mutation_density_in_mb, samples, depth, mutational_profile)


# Print a short summary to stdout
click.echo("Mutation-based assessment (uniform profile) - top rows:")
click.echo(mutations_per_gene_consequence.head().to_string())
click.echo("\nMutation-based assessment (provided mutational profile) - top rows:")
click.echo(mutations_per_gene_consequence_mut_prof.head().to_string())

# Optionally write outputs
if output_prefix:
out1 = f"{output_prefix}.mutations_per_gene_consequence.tsv"
out2 = f"{output_prefix}.mutations_per_gene_consequence_mutprof.tsv"
mutations_per_gene_consequence.to_csv(out1, sep='\t')
mutations_per_gene_consequence_mut_prof.to_csv(out2, sep='\t')
click.echo(f"Wrote results to: {out1} and {out2}")
Loading