diff --git a/.gitignore b/.gitignore index a20ec1c1..33b11cd7 100644 --- a/.gitignore +++ b/.gitignore @@ -15,3 +15,4 @@ scratch/ scratchhhh/ tests/test_data/all_samples.somatic.mutations.maf tests/test_data/all_samples_indv.depths.tsv.gz +assets/useful_scripts/*.ipynb diff --git a/CITATIONS.md b/CITATIONS.md index 4aaab46f..79a747d5 100644 --- a/CITATIONS.md +++ b/CITATIONS.md @@ -10,6 +10,10 @@ ## Sources of data and tools +- **Ensembl VEP** + + > McLaren W, Gil L, Hunt SE, et al. The Ensembl Variant Effect Predictor. Genome Biol. 2016;17:122. doi: 10.1186/s13059-016-0974-4. + - Nanoseq masks > Abascal, F., Harvey, L.M.R., Mitchell, E. et al. Somatic mutation landscapes at single-molecule resolution. Nature 593, 405–410 (2021). https://doi.org/10.1038/s41586-021-03477-4 @@ -22,7 +26,7 @@ > https://cancer.sanger.ac.uk/signatures/sbs -- **dNdScv covariates** +- **dNdScv method + covariates** > Martincorena I, et al. (2017) Universal Patterns of Selection in Cancer and Somatic Tissues. Cell. http://www.cell.com/cell/fulltext/S0092-8674(17)31136-4 @@ -38,11 +42,55 @@ > Ewels P, Magnusson M, Lundin S, Käller M. MultiQC: summarize analysis results for multiple tools and samples in a single report. Bioinformatics. 2016 Oct 1;32(19):3047-8. doi: 10.1093/bioinformatics/btw354. Epub 2016 Jun 16. PubMed PMID: 27312411; PubMed Central PMCID: PMC5039924. -- Python -- SigProfilerAssignment, MatrixGenerator -- HDP -- OncodriveFML -- OncodriveCLUSTL +- **bgreference / bgdata** + + > Repository: https://github.com/bbglab/bgreference + > Repository: https://github.com/bbglab/bgdata + +- **SAMtools** + + > Li H, Handsaker B, Wysoker A, et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25(16):2078-9. doi: 10.1093/bioinformatics/btp352. + +- **BEDTools** + + > Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010;26(6):841-842. doi: 10.1093/bioinformatics/btq033. + +- **HTSlib / Tabix** + + > Li H. Tabix: fast retrieval of sequence features from generic TAB-delimited files. Bioinformatics. 2011;27(5):718-719. doi: 10.1093/bioinformatics/btq671. + +- **Python** + + > Python Software Foundation. Python Language Reference, version 3.x. https://www.python.org/ + +- **SigProfilerAssignment, SigProfilerMatrixGenerator** + + > Alexandrov LB, et al. The repertoire of mutational signatures in human cancer. Nature 578, 94–101 (2020). doi:10.1038/s41586-020-1943-3. + > SigProfiler tools (SigProfilerMatrixGenerator, SigProfilerAssignment). Alexandrov Lab. https://github.com/AlexandrovLab + > Repository: https://github.com/AlexandrovLab/SigProfilerAssignment + > Repository: https://github.com/AlexandrovLab/SigProfilerMatrixGenerator + + +- **HDP (Hierarchical Dirichlet Processes)** + + > https://github.com/nicolaroberts/hdp + > Roberts, N. D. (2018). Patterns of somatic genome rearrangement in human cancer. https://doi.org/10.17863/CAM.22674 + +- **OncodriveFML** + + > Repository: https://github.com/bbglab/oncodrivefml (see repository for citation details) + > Mularoni L, Sabarinathan R, Deu-Pons J, Gonzalez-Perez A, Lopez-Bigas N. OncodriveFML: a general framework to identify coding and non-coding regions with cancer driver mutations. Genome Biology. 2016;17:128. doi:10.1186/s13059-016-0994-0. https://github.com/bbglab/oncodrivefml + +- **OncodriveCLUSTL** + + > Repository: https://github.com/bbglab/oncodriveclustl (see repository for citation details) + > Claudia Arnedo-Pac, Loris Mularoni, Ferran Muiños, Abel Gonzalez-Perez, Nuria Lopez-Bigas, OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers, Bioinformatics, Volume 35, Issue 22, November 2019, Pages 4788–4790, https://doi.org/10.1093/bioinformatics/btz501; https://github.com/bbglab/oncodriveclustl + +- **Omega (dN/dS)** + + > Repository: https://github.com/bbglab/omega (see repository for citation details) + + ## Software packaging/containerisation tools diff --git a/README.md b/README.md index a254570a..d673dddb 100644 --- a/README.md +++ b/README.md @@ -6,6 +6,34 @@ ![deepCSA workflow overview](docs/images/deepCSA.png) +## Documentation + +Find the documentation ([link to docs](https://github.com/bbglab/deepCSA/tree/main/docs)). + +We are working to provide the biggest possible detail on the [usage](docs/usage.md) and explanation of the rationale and [tools](docs/tools.md). + +You can also find an explanation with examples on the [output](docs/output.md). + +For more examples and description of the entire process, you can check the publications listed below. + +## Publications + +> **Sex and smoking bias in the selection of somatic mutations in human bladder** +> +> Ferriol Calvet*, Raquel Blanco Martinez-Illescas*, Ferran Muiños, Maria Tretiakova, Elena S. Latorre-Esteves, Jeanne Fredrickson, Maria Andrianova, Stefano Pellegrini, Axel Rosendahl Huber, Joan Enric Ramis-Zaldivar, Shuyi Charlotte An, Elana Thieme, Brendan F. Kohrn, Miguel L. Grau, Abel Gonzalez-Perez, Nuria Lopez-Bigas & Rosa Ana Risques +> +>_Nature_ (2025) doi:[10.1038/s41586-025-09521-x](https://doi.org/10.1038/s41586-025-09521-x) +> +> *these authors contributed equally and the order was decided randomly + +& + +> **DeepClone, an end-to-end protocol to study somatic mutagenesis and selection at high resolution** +> +> Ferriol Calvet, Morena Pinheiro-Santin, Erika Lopez, Raquel Blanco Martinez-Illescas, Núria Samper, Miguel L. Grau, Ferran Muiños, Rocío Chamorro González, Maria Andrianova, Federica Brando, Stefano Pellegrini, Marta Huertas, Elisabet Figuerola-Bou, Coohleen Coombes, Brendan F. Kohrn, Jeanne Fredrickson, Rosa Ana Risques, Nuria Lopez-Bigas, Abel Gonzalez-Perez +> +> _protocols.io_ (2026) doi:[10.17504/protocols.io.dm6gp1jodgzp/v2](https://dx.doi.org/10.17504/protocols.io.dm6gp1jodgzp/v2) + ## Usage You can find a detailed documentation in the [docs section](docs/README.md), but here there is a minimal summary on how to prepare the inputs. Still for your first runs if you need to make the complete set up you have to check the deeper documentation. @@ -22,6 +50,8 @@ sample2,sample2.filtered.vcf,sample2.sorted.bam Each row represents a single sample with a single-sample VCF containing the mutations called in that sample and the BAM file that was used for getting those variant calls. The mutations will be obtained from the VCF and the BAM file will be used for computing the sequencing depth at each position and using this for the downstream analysis. +Two alternative input modes are also supported: a samplesheet with only `sample,vcf` columns combined with a precomputed depths table, or a single cohort-level MAF file passed via `--input_maf` together with a precomputed depths table. See [Input scenarios](docs/input_scenarios.md) for details. + **Make sure that you do not use any '.' in your sample names, and also use text-like names for the samples, try to avoid having only numbers.** This second case should be handled properly but using string-like names will ensure consistency. **There are specific datasets that need to be prepared before running deepCSA. You can find a list of those, and instructions for downloading them in [the documentation section of the repo](docs/usage.md#mandatory-parameter-configuration).** @@ -57,31 +87,3 @@ We thank the following people for their extensive assistance in the development An extensive list of references for the tools used by the pipeline can be found in the [`CITATIONS.md`](CITATIONS.md) file. This pipeline uses code and infrastructure developed and maintained by the [nf-core](https://nf-co.re) community, reused here under the [MIT license](https://github.com/nf-core/tools/blob/master/LICENSE). - -> **The nf-core framework for community-curated bioinformatics pipelines.** -> -> Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso & Sven Nahnsen. -> -> _Nat Biotechnol._ 2020 Feb 13. doi: [10.1038/s41587-020-0439-x](https://dx.doi.org/10.1038/s41587-020-0439-x). - -## Documentation - -Find the documentation ([link to docs](https://github.com/bbglab/deepCSA/tree/main/docs)). - -We are working to provide the biggest possible detail on the [usage](docs/usage.md) and explanation of the rationale and [tools](docs/tools.md), but this is still in progress. - -## Publications - -> **Sex and smoking bias in the selection of somatic mutations in human bladder** -> -> Ferriol Calvet*, Raquel Blanco Martinez-Illescas*, Ferran Muiños, Maria Tretiakova, Elena S. Latorre-Esteves, Jeanne Fredrickson, Maria Andrianova, Stefano Pellegrini, Axel Rosendahl Huber, Joan Enric Ramis-Zaldivar, Shuyi Charlotte An, Elana Thieme, Brendan F. Kohrn, Miguel L. Grau, Abel Gonzalez-Perez, Nuria Lopez-Bigas & Rosa Ana Risques -> ->_Nature_ (2025) doi:[10.1038/s41586-025-09521-x](https://doi.org/10.1038/s41586-025-09521-x) -> -> *these authors contributed equally and the order was decided randomly - -> **DeepClone, an end-to-end protocol to study somatic mutagenesis and selection at high resolution** -> -> Ferriol Calvet, Morena Pinheiro-Santin, Erika Lopez, Raquel Blanco Martinez-Illescas, Núria Samper, Miguel L. Grau, Ferran Muiños, Rocío Chamorro González, Maria Andrianova, Federica Brando, Stefano Pellegrini, Marta Huertas, Elisabet Figuerola-Bou, Coohleen Coombes, Brendan F. Kohrn, Jeanne Fredrickson, Rosa Ana Risques, Nuria Lopez-Bigas, Abel Gonzalez-Perez -> -> _protocols.io_ (2026) doi:[10.17504/protocols.io.dm6gp1jodgzp/v2](https://dx.doi.org/10.17504/protocols.io.dm6gp1jodgzp/v2) diff --git a/assets/assess_panel/assess_panel.md b/assets/assess_panel/assess_panel.md new file mode 100644 index 00000000..b5ae8b17 --- /dev/null +++ b/assets/assess_panel/assess_panel.md @@ -0,0 +1,44 @@ +# Experimental design principles + +## Definition of relationship between number of mutations detected and sequencing depth + +$$ +\text{Expected number of mutations} = (\text{ average depth * region of interest (bp) }) \times (\text{expected mutation density (muts/bp)}) +$$ + +## Use cases + +### Predict average depth required + +Given a fixed panel size based on which are the regions of interest, a desired number of mutations (enough to get good selection/mutagenesis metrics) and a known estimation of the expected mutation density of the sample, you can estimate which is the sequencing depth required using this formula: + +$$ +\text{average depth} = \frac{\text{minimum desired mutations}}{(\text{region of interest (bp)}) \times (\text{expected mutation density (muts/bp)})} +$$ + +To get an estimation of the expected mutation density you can use this formula even if it is with some approximate values: + +$$ +\text{expected mutation density (muts/bp)} = \text{mutation rate (mutations} \times \text{bp} \times \text{year)} \times \text{time (year)} +$$ + +For the mutation rate this could be a possible formula; but you may also rely on previously published data. + +$$ +\text{mutation rate (} \frac{\text{mutations}}{\text{bp * year}} \text{)} = \frac{\text{observed mutations}}{ \text{genome size} \times \text{age (year)} } +$$ + +In this example, we are assuming that the units of the mutation rate are mutations x genome x year, but this could be in any other units +Particularly the time units could be different from year. + +In case that the sequencing depth is restricted by the availability of DNA or any other reason, you can decide to adjust the panel so that by increasing the depth you can still reach a stable measurement of the mutation density and other variables. + +An additional scenario to consider is the trinucleotide mutation probabilities. In a given sample, the different sites may have a very different mutation probablility depending on which are the active mutational processes. If these are known, you can factor in this information to adjust the requirements of the experimental design. + +### Expected number of mutations given an average depth + +The average depth represents the number of times that you sequence your region of interest, then it can be seen as: + +$$ +\text{Expected number of mutations} = (\text{ average depth * region of interest (bp) }) \times (\text{expected mutation density (muts/bp)}) +$$ diff --git a/assets/assess_panel/assess_panel.py b/assets/assess_panel/assess_panel.py new file mode 100644 index 00000000..eb6aae5e --- /dev/null +++ b/assets/assess_panel/assess_panel.py @@ -0,0 +1,114 @@ +#!/usr/bin/env python + + +import sys +import pandas as pd +import numpy as np +import seaborn as sns +import matplotlib.pyplot as plt +import click + + +sys.path.append("../../bin") + +from utils_context import triplet_context_iterator + +# we need a stable background mutation density estimation, +# for this we need at least 3 mutations + +def mutation_based_assessment(panel_tsv, expected_mutation_density_in_mb, samples, depth, mutational_profile = None): + if mutational_profile is None: + mutational_profile = pd.DataFrame(np.ones(96) / 96) + mutational_profile.index = triplet_context_iterator() + mutational_profile = mutational_profile / mutational_profile.mean() + mutational_profile.columns = ["Probability"] + + size_per_consequence = panel_tsv.groupby(by = ["GENE", "IMPACT", "CONTEXT_MUT"]).size().reset_index(name = "DEPTH") + size_per_consequence["DEPTH"] *= depth/3 + + size_per_consequence = size_per_consequence.merge(mutational_profile, + left_on = "CONTEXT_MUT", + right_index = True, + how = "left") + size_per_consequence["ADJUSTED_DEPTH"] = size_per_consequence["DEPTH"] * size_per_consequence["Probability"] + + total_adjusted_size = size_per_consequence.groupby(by = ["GENE", "IMPACT"])["ADJUSTED_DEPTH"].sum() / 3 + expected_mutations_per_cnsq = (expected_mutation_density_in_mb * total_adjusted_size * samples) / 1e6 + expected_mutations_per_cnsq.name = "EXPECTED_MUTATIONS" + expected_mutations_per_cnsq = expected_mutations_per_cnsq.round(5) + + return expected_mutations_per_cnsq + + + +@click.command() +@click.option('--panel', '-p', 'panel_path', required=True, type=click.Path(exists=True, dir_okay=False), + help='Path to panel TSV file') +@click.option('--expected-mutation-density', '-e', 'expected_mutation_density', default=1.2, show_default=True, + type=float, help='Expected mutation density (mutations per Mb)') +@click.option('--samples', '-s', default=100, show_default=True, type=int, + help='Number of samples to assume') +@click.option('--depth', '-d', default=5000, show_default=True, type=int, + help='Sequencing depth (used in size adjustment)') +@click.option('--mutational-profile', '-m', 'mutational_profile_path', default=None, type=click.Path(exists=True, dir_okay=False), + help='Optional mutational profile TSV with CONTEXT_MUT and Probability (or column named like all_samples.all)') +@click.option('--output-prefix', '-o', default=None, help='Optional prefix to write TSV outputs (will create *.tsv)') +@click.option('--genes-list', '-g', 'genes_list', default=None, type=str, + help='Optional list of genes to consider') +def main(panel_path, expected_mutation_density, samples, depth, mutational_profile_path, output_prefix, genes_list): + """CLI entrypoint for assess_panel. + + Loads the panel TSV and optional mutational profile, runs the mutation-based assessment + (with and without the provided profile) and optionally writes results to TSV files. + """ + panel = pd.read_table(panel_path) + if genes_list: + genes_list = [gene.strip() for gene in genes_list.split(",")] + panel = panel[panel["GENE"].isin(genes_list)] + + mutational_prof = None + if mutational_profile_path: + mutational_prof = pd.read_table(mutational_profile_path) + # Try to normalize common formats: expect CONTEXT_MUT as index and a column named Probability + if 'CONTEXT_MUT' in mutational_prof.columns: + # If the profile already has a Probability column, keep it; otherwise try to rename + if 'Probability' in mutational_prof.columns: + mutational_prof = mutational_prof.set_index('CONTEXT_MUT')[['Probability']] + elif 'all_samples.all' in mutational_prof.columns: + mutational_prof = mutational_prof.set_index('CONTEXT_MUT').rename(columns={'all_samples.all': 'Probability'})[['Probability']] + else: + # Fallback: use all other numeric column(s) and take the first as Probability + numeric_cols = mutational_prof.select_dtypes(include=[np.number]).columns.tolist() + if numeric_cols: + mutational_prof = mutational_prof.set_index('CONTEXT_MUT')[[numeric_cols[0]]].rename(columns={numeric_cols[0]: 'Probability'}) + else: + # If no numeric columns, create uniform profile (will be normalized in function) + mutational_prof = None + + # Run assessments + print(f"Running mutation-based assessment with expected mutation density: {expected_mutation_density}, samples: {samples}, depth: {depth}") + mutations_per_gene_consequence = mutation_based_assessment(panel, expected_mutation_density, samples, depth) + + # Print a short summary to stdout + click.echo("Mutation-based assessment (uniform profile) - top rows:") + click.echo(mutations_per_gene_consequence.head().to_string()) + + if mutational_prof is not None: + print(f"Running mutation-based assessment with provided mutational profile: {mutational_profile_path}") + mutations_per_gene_consequence_mut_prof = mutation_based_assessment(panel, expected_mutation_density, samples, depth, mutational_prof) + + click.echo("\nMutation-based assessment (provided mutational profile) - top rows:") + click.echo(mutations_per_gene_consequence_mut_prof.head().to_string()) + + # Optionally write outputs + if output_prefix: + out1 = f"{output_prefix}.mutations_per_gene_consequence.tsv" + mutations_per_gene_consequence.to_csv(out1, sep='\t') + click.echo(f"Wrote results to: {out1}") + if mutational_prof is not None: + out2 = f"{output_prefix}.mutations_per_gene_consequence_mutprof.tsv" + mutations_per_gene_consequence_mut_prof.to_csv(out2, sep='\t') + click.echo(f"Wrote results to: {out1} and {out2}") + +if __name__ == "__main__": + main() \ No newline at end of file diff --git a/assets/assess_panel/panel.ipynb b/assets/assess_panel/panel.ipynb new file mode 100644 index 00000000..90491a88 --- /dev/null +++ b/assets/assess_panel/panel.ipynb @@ -0,0 +1,587 @@ +{ + "cells": [ + { + "cell_type": "code", + "execution_count": 8, + "id": "a87db968", + "metadata": {}, + "outputs": [], + "source": [ + "import pandas as pd\n", + "import numpy as np\n", + "import seaborn as sns\n", + "import matplotlib.pyplot as plt\n" + ] + }, + { + "cell_type": "markdown", + "id": "74d60b27", + "metadata": {}, + "source": [ + "## deepCSA path" + ] + }, + { + "cell_type": "code", + "execution_count": 9, + "id": "ffc70136", + "metadata": {}, + "outputs": [], + "source": [ + "deepCSA_run_dir = \"/data/bbg/nobackup/lung_duplex/analysis/fullcohortnormal/asxs10/2026-08-11_PEACE_TRACERx\"" + ] + }, + { + "cell_type": "code", + "execution_count": 10, + "id": "d01074db", + "metadata": {}, + "outputs": [], + "source": [ + "all_mutdensities = pd.read_table(f\"{deepCSA_run_dir}/mutdensity/all_mutdensities.tsv\", sep = \"\\t\")" + ] + }, + { + "cell_type": "code", + "execution_count": 11, + "id": "7d0b7bfa", + "metadata": {}, + "outputs": [ + { + "data": { + "text/html": [ + "
\n", + "\n", + "\n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + "
SAMPLE_IDMUTDENSITY_MB
19653L8830.297713
33629L8650.750569
34889L9350.488972
38461SmokingstatusunifNever0.263338
42033L9050.399045
.........
893385L3930.497468
899269L9280.614416
900529L3940.436490
901789L8750.296914
904205L8770.670337
\n", + "

192 rows × 2 columns

\n", + "
" + ], + "text/plain": [ + " SAMPLE_ID MUTDENSITY_MB\n", + "19653 L883 0.297713\n", + "33629 L865 0.750569\n", + "34889 L935 0.488972\n", + "38461 SmokingstatusunifNever 0.263338\n", + "42033 L905 0.399045\n", + "... ... ...\n", + "893385 L393 0.497468\n", + "899269 L928 0.614416\n", + "900529 L394 0.436490\n", + "901789 L875 0.296914\n", + "904205 L877 0.670337\n", + "\n", + "[192 rows x 2 columns]" + ] + }, + "execution_count": 11, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "samples_and_mutdensity = all_mutdensities[(all_mutdensities[\"GENE\"] == 'ALL_GENES')\n", + " & (all_mutdensities[\"MUTTYPES\"] == 'SNV')\n", + " & (all_mutdensities[\"REGIONS\"] == \"all\")\n", + " ] [[\"SAMPLE_ID\", \"MUTDENSITY_MB\"]]\n", + "samples_and_mutdensity" + ] + }, + { + "cell_type": "code", + "execution_count": 12, + "id": "f7224ea0", + "metadata": {}, + "outputs": [ + { + "data": { + "text/html": [ + "
\n", + "\n", + "\n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + "
SAMPLE_IDMUTDENSITY_MB
247337all_samples0.379713
201317SmokingstatusunifUnknown0.216706
38461SmokingstatusunifNever0.263338
817425SmokingstatusunifFormer0.444664
622421SmokingstatusunifCurrent0.536064
98353CohortTracerx_SmokingstatusunifNever0.281333
279589CohortTracerx_SmokingstatusunifFormer0.507360
718449CohortTracerx_SmokingstatusunifCurrent0.540464
739569CohortTracerx_SexM0.469281
596677CohortTracerx_SexF0.446257
221281CohortTracerx_LungcancerYes0.456712
503273CohortTracerx_BinarysmokingstatusunifNever0.281333
499701CohortTracerx_BinarysmokingstatusunifEver0.521992
125669CohortTracerx0.456712
813853CohortPeace_SmokingstatusunifUnknown0.216706
717189CohortPeace_SmokingstatusunifNever0.255042
762053CohortPeace_SmokingstatusunifFormer0.412204
694601CohortPeace_SmokingstatusunifCurrent0.446315
551813CohortPeace_SexM0.338563
688509CohortPeace_SexF0.321420
461773CohortPeace_LungcancerYes0.378748
94677CohortPeace_LungcancerNo0.314176
670753CohortPeace_BinarysmokingstatusunifUnknown0.216706
429833CohortPeace_BinarysmokingstatusunifNever0.255042
285577CohortPeace_BinarysmokingstatusunifEver0.412876
266561CohortPeace0.330824
136281BinarysmokingstatusunifUnknown0.216706
274861BinarysmokingstatusunifNever0.263338
169585BinarysmokingstatusunifEver0.464849
\n", + "
" + ], + "text/plain": [ + " SAMPLE_ID MUTDENSITY_MB\n", + "247337 all_samples 0.379713\n", + "201317 SmokingstatusunifUnknown 0.216706\n", + "38461 SmokingstatusunifNever 0.263338\n", + "817425 SmokingstatusunifFormer 0.444664\n", + "622421 SmokingstatusunifCurrent 0.536064\n", + "98353 CohortTracerx_SmokingstatusunifNever 0.281333\n", + "279589 CohortTracerx_SmokingstatusunifFormer 0.507360\n", + "718449 CohortTracerx_SmokingstatusunifCurrent 0.540464\n", + "739569 CohortTracerx_SexM 0.469281\n", + "596677 CohortTracerx_SexF 0.446257\n", + "221281 CohortTracerx_LungcancerYes 0.456712\n", + "503273 CohortTracerx_BinarysmokingstatusunifNever 0.281333\n", + "499701 CohortTracerx_BinarysmokingstatusunifEver 0.521992\n", + "125669 CohortTracerx 0.456712\n", + "813853 CohortPeace_SmokingstatusunifUnknown 0.216706\n", + "717189 CohortPeace_SmokingstatusunifNever 0.255042\n", + "762053 CohortPeace_SmokingstatusunifFormer 0.412204\n", + "694601 CohortPeace_SmokingstatusunifCurrent 0.446315\n", + "551813 CohortPeace_SexM 0.338563\n", + "688509 CohortPeace_SexF 0.321420\n", + "461773 CohortPeace_LungcancerYes 0.378748\n", + "94677 CohortPeace_LungcancerNo 0.314176\n", + "670753 CohortPeace_BinarysmokingstatusunifUnknown 0.216706\n", + "429833 CohortPeace_BinarysmokingstatusunifNever 0.255042\n", + "285577 CohortPeace_BinarysmokingstatusunifEver 0.412876\n", + "266561 CohortPeace 0.330824\n", + "136281 BinarysmokingstatusunifUnknown 0.216706\n", + "274861 BinarysmokingstatusunifNever 0.263338\n", + "169585 BinarysmokingstatusunifEver 0.464849" + ] + }, + "execution_count": 12, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "samples_and_mutdensity[~(samples_and_mutdensity[\"SAMPLE_ID\"].str.startswith(\"L\"))].sort_values(by = \"SAMPLE_ID\", ascending = False)" + ] + }, + { + "cell_type": "code", + "execution_count": null, + "id": "9a6cff4e", + "metadata": {}, + "outputs": [], + "source": [] + }, + { + "cell_type": "markdown", + "id": "9ffc1877", + "metadata": {}, + "source": [ + "## Load data and plot results" + ] + }, + { + "cell_type": "code", + "execution_count": 17, + "id": "f631f449", + "metadata": {}, + "outputs": [], + "source": [ + "output_path = '/data/bbg/datasets/pipelines/deepCSA/files_for_exploration'" + ] + }, + { + "cell_type": "code", + "execution_count": 23, + "id": "4002d478", + "metadata": {}, + "outputs": [ + { + "data": { + "text/html": [ + "
\n", + "\n", + "\n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + " \n", + "
EXPECTED_MUTATIONS
GENEIMPACT
ALKessential_splice0.18797
missense4.76482
nonsense0.24592
splice_region_variant0.78864
synonymous1.52028
\n", + "
" + ], + "text/plain": [ + " EXPECTED_MUTATIONS\n", + "GENE IMPACT \n", + "ALK essential_splice 0.18797\n", + " missense 4.76482\n", + " nonsense 0.24592\n", + " splice_region_variant 0.78864\n", + " synonymous 1.52028" + ] + }, + "execution_count": 23, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "mutations_per_gene_consequence = pd.read_table(f\"{output_path}/exome_lung_panel_genes_smokers.8x1000.mutations_per_gene_consequence.tsv\",\n", + " sep = \"\\t\")\n", + "mutations_per_gene_consequence = mutations_per_gene_consequence.set_index([\"GENE\", \"IMPACT\"])\n", + "mutations_per_gene_consequence.head()" + ] + }, + { + "cell_type": "code", + "execution_count": 24, + "id": "3a9e7a84", + "metadata": {}, + "outputs": [], + "source": [ + "consequences_to_plot = ['synonymous', 'missense', 'nonsense',\n", + " 'essential_splice', 'splice_region_variant', \n", + " # 'non_genic_variant', 'intron_variant',\n", + " # 'non_coding_exon_region', 'non_coding_transcript_variant'\n", + " ]" + ] + }, + { + "cell_type": "code", + "execution_count": 31, + "id": "07588745", + "metadata": {}, + "outputs": [ + { + "data": { + "text/plain": [ + "Index(['ALK', 'ARID1A', 'ARID2', 'ATM', 'BRAF', 'CDKN2A', 'CHEK2', 'CTNNB1',\n", + " 'EGFR', 'ERBB2', 'ERBB4', 'FAT1', 'KEAP1', 'KMT2D', 'KRAS', 'MAX',\n", + " 'NF1', 'NFE2L2', 'NOTCH1', 'PIK3CA', 'PTEN', 'RB1', 'RET', 'ROS1',\n", + " 'SETBP1', 'SETD2', 'SMAD4', 'SMARCA4', 'STK11', 'TIAM2', 'TP53'],\n", + " dtype='object', name='GENE')" + ] + }, + "execution_count": 31, + "metadata": {}, + "output_type": "execute_result" + } + ], + "source": [ + "mutations_per_gene_consequence.unstack(level=0).loc[consequences_to_plot].columns.droplevel(0)" + ] + }, + { + "cell_type": "code", + "execution_count": 34, + "id": "8d07e817", + "metadata": {}, + "outputs": [ + { + "data": { + "image/png": "iVBORw0KGgoAAAANSUhEUgAAAoMAAAPGCAYAAACYjJDlAAAAOnRFWHRTb2Z0d2FyZQBNYXRwbG90bGliIHZlcnNpb24zLjEwLjAsIGh0dHBzOi8vbWF0cGxvdGxpYi5vcmcvlHJYcgAAAAlwSFlzAAAPYQAAD2EBqD+naQABAABJREFUeJzs3XdYFNcaBvB3WWBtdMTFCgKCUpQYFRELRhEVI0ViQ4MtsWEUFcWCLbhqNPYSlaIREAuWeI09VqwREBVRo8YYFpUuve39g7jJ0FwUmN3Z73efeZ47Z8/MvHsyLmfPmZnlSSQSCQghhBBCiFJSYTsAIYQQQghhD3UGCSGEEEKUGHUGCSGEEEKUGHUGCSGEEEKUGHUGCSGEEEKUGHUGCSGEEEKUGHUGCSGEEEKUGHUGCSGEEEKUmCrbAYjsJOInbEfgpNJbZ9mOwFn8L0awHYG7eDy2E3BXYT7bCbhLx7DeDjWZp1lvx6qJHZIstiNUQCODhBBCCCFKjDqDhBBCCCFKjKaJCSGEEMI5NNolO2orQgghhBAlRp1BQgghhBAlRtPEhBBCCOEcFbrjXmY0Miij6Oho8Pl8ODs7M8pfvHgBHo+H2NjYSrcLDQ2FtrY2oywhIQEtW7aEu7s7CgoK6ihx/fsp7ACGfTsLnw30hL3raExb+D2evXzFdiyFt+XMDXTw28hYei7fxXYsTgk7cBh9h7jDuntvuI/2xp2YWLYjcUbYgUPo6+IGa7tecB/1Ne7cjWU7ksK7HROHybP94eDiAXO7Pjh36QrbkYiCo86gjIKDg+Hj44OrV6/i5cuXH72f27dvo2fPnhgwYAAOHjwIgUBQiynZdTv2Pka5DkbktrUIXrsCxSUlmDh3MXLz6Jldn8q0mR4uLZ4oXY75jmY7EmecPHMOonUbMGW8N46G70Fn246Y5OOLJHEy29EU3snTZyFauwFTJrxv206Y5DOL2vYT5eblw9zMBAGzv2M7CuEI6gzKICcnBwcOHMCUKVPg4uKC0NDQj9rPhQsX0LdvX4wbNw5BQUHg8/m1G5Rlu39YDveB/WBm3AYWpm0hmj8TSa/f4sHjp2xHU3h8FR6aajSWLrpNGrEdiTNC9kXAY+gQeLp9CRNjIyycMwvCZgaIOBTFdjSFFxIWAQ/XIfB0GwqTtsZYOJfatjb0tu+GWZMnwsmxF9tR5JqKnC7ySF5zyZXIyEiYm5vD3NwcXl5eCAkJgUQiqdE+jhw5gsGDB2PhwoX44Ycf6iipfHmXnQMA0NJownISxfcyJQO9V+xGf1EIZof9ir9SM9mOxAmFRUV48CgRDnZdGeU97Loh5l48S6m4obCoCA8SEuFg141R3qN7N8TEUdsSIk+oMyiDoKAgeHl5AQCcnZ2RnZ2N8+fPy7x9dnY2PD09MXfuXMyfP7+uYsoViUSCVdt2o7N1B7Rra8R2HIVm01oI0Qgn7JroimXDvkDKuxyM2noAGTl5bEdTeOkZGSgpKYGeni6jXF9PB29T01hKxQ1Vtq2uLt6mprKUihBSGeoMfkBiYiJu3bqFESPKfmNVVVUVw4cPR3BwsMz7aNiwIfr3749du3YhISFBpm0KCgqQlZXFWAoKCj/qPbBhxcYdSPzjBdYt9mM7isLrZWEEJ2sztDPUh71Za2wfPxQAcPR32c4l8mG8cncdSiQA3YdYO3go37aSCu1NCGEXdQY/ICgoCMXFxWjRogVUVVWhqqqK7du3IyoqCunp6TLtg8/n4+jRo+jcuTMcHR3x8OHDD24jEomgpaXFWESbd3zq26kXKzbuwIVrN7F3w0oIDfTZjsM5jdTV0M5QD3+mZLAdReHpaGuDz+cjJYU5UpWalg79ciNapGakbVtuFDA1PR36utS2pO6p8ORzkUfUGaxGcXEx9u7di3Xr1iE2Nla6xMXFoU2bNggLC5N5XwKBAFFRUejatSscHR1x//79auv7+/sjMzOTsfj7TP7Ut1SnJBIJlm/YjrNXohG6PhAtDYVsR+KkwuJiPHuTjqYajdmOovDU1dRgaWGOazdvM8qjb96CrY01S6m4QV1NDZbtzXHt5i1GefSNW7DtSG1LiDyhh05X48SJE0hPT8eECROgpaXFeG3YsGEICgqCi4sLgLLp5PI6dOjAWFdXV8fhw4fx1VdfoW/fvjh//jysrSv/UBQIBBUeOyPJUf+Ut1Pnlm/YjhPnLmFr4CI0btgIb1PLRk41mjRCAw49Qqe+rTlxBY7tjWGoo4HU7Dz8dP4WsvMLMfTz9mxH44RxXiPht3gZrDpYwNbGGpFRRyFOfo0Rw9zYjqbwxo3+p23bt4etjRUio46Vta0Hte2nyMnNxctXf0vXXyUlI+HxE2hpaqK5sBmLyYiios5gNYKCgtCvX78KHUEA8PDwwMqVK5GWVnaR+ftrCv/r+fPnFcrU1NRw4MABjBw5UtohtLGxqf3wLIg4dhIAMHamP6N85byZcB/Yj41InPA6Mxtzwk8hPTcPuo0bomNrISKmf4UWOppsR+OEQU79kJ6RiW27gvEmJRXtTNpi56Z1aGFoyHY0hTdoQH+kZ2Zi266g/7Ttj2jRnNr2U9xPSMTYabOk66KNWwEAboMGYFWAf1WbKR2a+pQdT1LTZ6QQ1kjET9iOwEmlt86yHYGz+F9U/JJEagndhFF3CulB+XVGp/6+CPjyKw7kyIMfS+Tv0WDUcSaEEEIIUWI0TUwIIYQQzlGh0XOZ0cggIYQQQogSo84gIYQQQogSo2liQgghhHAOjXbJjtqKEEIIIUSJUWeQEEIIIUSJ0TSxIikqYDsBNz2r+OsxpJb0LWU7AXfx6OO7zqjQOAkXyOvvAMsjOuMJIYQQQpQYdQYJIYQQQpQYzTMQQgghhHNotEt21FaEEEIIIUqMOoOEEEIIIUqMOoOEEEII4RwejyeXS01dvnwZQ4YMQfPmzcHj8XD06NEKdRISEvDll19CS0sLGhoasLOzw8uXL2U+hsJ3BqOjo8Hn8+Hs7Mwof/HiBaPxtbS0YGdnh19++YVRLzQ0FNra2oz199vw+Xzo6OigW7duWL58OTIzMxnbyvIf6L3w8HDw+XxMnjz5k9+zvPopIgrDpvnhsy9Hw95zHKYtWYVnf/3NdixO6LfvIjrsOFVhWXHlIdvROCHsYBT6fjkM1vaOcPcajzsxsWxH4oywA4fQ18UV1nY94T5qLO7cjWE7ksK7fTcWk33nw2GQG8y79sK5i1fYjkTqUE5ODjp27IgtW7ZU+voff/wBBwcHWFhY4OLFi4iLi8PixYvRoEEDmY+h8J3B4OBg+Pj44OrVq5X2gs+dOwexWIybN2+ia9eu8PDwwP3796vdp6amJsRiMV69eoXo6Gh888032Lt3Lzp16oSkpCRpvQ/9Byqf08/PD/v370dubm7N36gCuH3vAUZ96YzITSIEr1qC4pJSTJy/HLl5+WxHU3gHPOxxaayjdNnt8jkAYEDbZiwnU3wnz5yDaN1GTBk/FkfDQtDZ1gaTZsxBUnIy29EU3snTZyFaux5TJozD0fC96GzbCZN8ZiFJTG37KXLz82FuZoKAuTPZjkLqwcCBA/H999/D3d290tcXLlyIQYMGYc2aNbC1tUXbtm0xePBgGBgYyHwMhe4M5uTk4MCBA5gyZQpcXFwQGhpaoY6enh6EQiEsLCwQGBiIoqIi/Pbbb9Xul8fjQSgUwtDQEO3bt8eECRMQHR2N7Oxs+Pn5Set96D/Qey9evEB0dDTmz58PCwsLHDp06KPer7zbLVoM9wF9YWbUGhYmRhDNmYakNyl48OQPtqMpPN2G6mjaSCBdLv35Fq00G6FLc122oym8kLBIeAx1gafrlzAxNsLC2TMhbGaAiENH2I6m8ELCIuDh+iU83YbCpK0xFs71hbBZM0QcOsx2NIXW294Os6ZMgpNjb7ajyDUVOV0KCgqQlZXFWAoKPu5HJUpLS/G///0P7dq1w4ABA2BgYIBu3bpVO1NZGYXuDEZGRsLc3Bzm5ubw8vJCSEgIJBJJpXWLioqwa9cuAICamlqNj2VgYIDRo0fj+PHjKCkpqdG2wcHBGDx4MLS0tODl5YWgoKAaH18RvcspGwHV0tBgOQm3FJaU4pcnSXC3aPFR15+QfxUWFeHBo0Q42HVllPew64qYe9XPIJDqFRYV4UHCIzjYdWOU9+jeFTFx8SylIoR9IpEIWlpajEUkEn3Uvt68eYPs7GysWrUKzs7OOHPmDNzc3ODu7o5Lly7JvB+Ffs5gUFAQvLy8AADOzs7Izs7G+fPn0a9fP2kde3t7qKioIC8vD6WlpTAyMsJXX331UcezsLDAu3fvkJqaKvPwa2lpKUJDQ7F582YAwIgRI+Dr64unT5/C1NT0o3IoAolEglU7QtHZqj3aGbdmOw6nnH/+Gu8KiuFm3oLtKAovPSMDJSUl0NNljrDq6+rgbUoqS6m4Qdq2euXbVg9vU2+wlIoQ9vn7+8PX15dRJhAIPmpfpaVlP/k5dOhQzJo1CwDQqVMnREdHY8eOHejdW7bRY4UdGUxMTMStW7cwYsQIAICqqiqGDx+O4OBgRr3IyEjExMTg+PHjMDU1xe7du6Fb7oNfVu9HHWsyGnPmzBnk5ORg4MCBAAB9fX04OTlVyFle5cPIhR+Vmw0rNu9G4vM/sW7BLLajcE7Uo1fo2VofBo1lvziYVK/8v2mJpGb/zknVeCjfthJqW1IvVHjyuQgEAmhqajKWj+0M6uvrQ1VVFR06dGCUt2/fvkZ3EyvsyGBQUBCKi4vRosW/oyMSiQRqampIT0+XlrVq1QpmZmYwMzNDkyZN4OHhgYcPH9bowsr3EhISoKmpCT09PZm3CQ4ORlpaGho1aiQtKy0tRUxMDFasWAE+n1/pdiKRCMuWLWOUBcycgqWzptY4d31bsWU3Lty4jX3rVkDYVPa2Ih/297s8XP87FRudbNmOwgk62trg8/lISWWOAqamp0Nf7+O+NJIyVbdtGvQ/8gs5IYRJXV0dXbp0QWJiIqP88ePHaNOmjcz7UciRweLiYuzduxfr1q1DbGysdImLi0ObNm0QFhZW6Xa9e/eGlZUVAgMDa3zMN2/eIDw8HK6urlBRka3ZUlNTcezYMezfv5+RMzY2FtnZ2fj111+r3Nbf3x+ZmZmMxX/qxBrnrk8SiQTLN+/C2as3EbpmKVoa0p2ute3Io1fQbShA7zZN2Y7CCepqarC0MMe1m7cZ5dE3b8PWxoqlVNygrqYGy/YWuHbzFqM8+sYt2Ha0ZikVIYonOztb2ncAgOfPnyM2NlY68jd37lxERkZi165dePr0KbZs2YJffvkFU6fKPnikkCODJ06cQHp6OiZMmAAtLS3Ga8OGDUNQUBBcXFwq3Xb27Nnw9PSEn58fY1TxvyQSCZKTkyGRSJCRkYHr169j5cqV0NLSwqpVq6T1srOz8fTpU+n6+/9Aurq6aN26NX7++Wfo6enB09OzQgfSxcWl2pwCgaDCsLEkQ73qRpEDyzfvwokLV7B12Xw0btQQb9PKRmg1GjdCg48cAif/KpVIcCTxb7i2aw5VGb+QkA8bN3o4/AJWwKq9BWxtrBAZdQzi5NcY4eHGdjSFN270SPgtXvpP21ojMuroP21b/RMYSPVycnPx8tW/z3B9lSRGwuMn0NLURHMhfQl/jyufknfu3IGjo6N0/f31hl9//TVCQ0Ph5uaGHTt2QCQSYcaMGTA3N8fhw4fh4OAg8zEUsjMYFBSEfv36VegIAoCHhwdWrlyJtLS0Srd1cXGBkZERAgMDsW3btkrrZGVlwdDQEDweD5qamjA3N8fXX3+N7777DpqamtJ6H/oPFBwcDDc3t0pHEj08PDB8+HC8fv0azZpx4x9vxC+nAQBj5wQwylfOmQb3AX3ZiMQp11+lQpydD3eLlmxH4ZRBTv2QnpmFbbtD8CYlFe1M2mLnxrVoYShkO5rCGzSgP9IzM7FtVzDepKSUte2m9WjR3JDtaArtfkIixk75Trou2lD2rFu3wc5YtWQBW7FIHenTp0+VT0p5b/z48Rg/fvxHH4Mn+dARiNyQvKRHXdSF0sO72I7AWfyJi9mOwF0qCvldXjEUf9wz34gMtOpv8GOZQKfejlUTSwrSP1ypntGnCSGEEEI4R4XuWpcZV6bUCSGEEELIR6DOICGEEEKIEqNpYkIIIYRwDo12yY7aihBCCCFEiVFnkBBCCCFEidE0MSGEEEI4R4VuJpYZdQYVCT1XrG7o6rOdgBBCCGENTRMTQgghhCgxGmoihBBCCOfQaJfsqK0IIYQQQpQYdQYJIYQQQpQYTRMTQgghhHNUQLcTy4pGBgkhhBBClJhSdAajo6PB5/Ph7OzMKH/x4gV4PJ500dLSgp2dHX755RdGvdDQUGhrazPW32/D5/Oho6ODbt26Yfny5cjMzGRsKxKJ0KVLF2hoaMDAwACurq5ITEyss/fKpp/CD2HY1Nn4zGU47D3GYtrilXj21yu2Y3HC63e58Dt2Fd3XH8BnayLgtvt/eCBOZTsWZ4QdjELfL4fB2t4R7l7jcScmlu1InBF24BD6urjC2q4n3EeNxZ27MWxHUni378Zisu98OAxyg3nXXjh38QrbkYiCU4rOYHBwMHx8fHD16lW8fPmywuvnzp2DWCzGzZs30bVrV3h4eOD+/fvV7lNTUxNisRivXr1CdHQ0vvnmG+zduxedOnVCUlKStN6lS5cwbdo03LhxA2fPnkVxcTGcnJyQk5NT6++Tbbfv3ceoLwchcssPCF6zDMUlJZjotxS5eflsR1NomXkFGL33NFT5KvhpeF/88s0Q+H3xGTQaqLMdjRNOnjkH0bqNmDJ+LI6GhaCzrQ0mzZiDpORktqMpvJOnz0K0dj2mTBiHo+F70dm2Eyb5zEKSmNr2U+Tm58PczAQBc2eyHUWuqfDkc5FHnO8M5uTk4MCBA5gyZQpcXFwQGhpaoY6enh6EQiEsLCwQGBiIoqIi/Pbbb9Xul8fjQSgUwtDQEO3bt8eECRMQHR2N7Oxs+Pn5SeudOnUK3t7esLS0RMeOHRESEoKXL1/i999/r+23yrrdq5bC3fkLmBm1hoWJMUR+M5D05i0ePPmD7WgKLejGQwg1GmGliz1smuujhXYTdDc2RGsdDbajcUJIWCQ8hrrA0/VLmBgbYeHsmRA2M0DEoSNsR1N4IWER8HD9Ep5uQ2HS1hgL5/pC2KwZIg4dZjuaQuttb4dZUybBybE321EIR3C+MxgZGQlzc3OYm5vDy8sLISEhkEgkldYtKirCrl27AABqamo1PpaBgQFGjx6N48ePo6SkpNI676eRdXV1a7x/RfMuJxcAoKXRhOUkiu3C41ewMtTDzKjLcNhwEO5B/8PBmCdsx+KEwqIiPHiUCAe7rozyHnZdEXOv+tkBUr3CoiI8SHgEB7tujPIe3bsiJi6epVSEkMpw/m7ioKAgeHl5AQCcnZ2RnZ2N8+fPo1+/ftI69vb2UFFRQV5eHkpLS2FkZISvvvrqo45nYWGBd+/eITU1FQYGBozXJBIJfH194eDgACsrq49/UwpAIpFg1fYgdLbqgHbGbdiOo9BeZbzD/rvv8HW39vjG3grxSSlYefYO1FX5GGrdlu14Ci09IwMlJSXQK/flTF9XB29T6JrMTyFtW73ybauHt6k3WEpFlAnnR7tqEac7g4mJibh16xaioqIAAKqqqhg+fDiCg4MZncHIyEhYWFjg8ePHmDlzJnbs2PHRI3fvRx15vIoXBkyfPh337t3D1atXP7ifgoICFBQUMMrUCwohECjGdWIrNv2ExGd/InyjiO0oCq9UAlgZ6mJWH1sAQAehLp6mZGL/3cfUGawl5f+9SiSV/xsmNcdD+baVUNsSImc43RkMCgpCcXExWrRoIS2TSCRQU1NDenq6tKxVq1YwMzODmZkZmjRpAg8PDzx8+LDCyJ4sEhISoKmpCT09PUa5j48Pjh8/jsuXL6Nly5Yf3I9IJMKyZcsYZQGzpmGp7/QaZ6pvKzbvxIXrt7BvvQjCpvpsx1F4TZs0hIm+FqPMRE8LZx9VvBmK1IyOtjb4fD5SUpmjgKnp6dDX4/6lHHWp6rZNg74SXCZDiCLh7ChqcXEx9u7di3Xr1iE2Nla6xMXFoU2bNggLC6t0u969e8PKygqBgYE1PuabN28QHh4OV1dXqKiUNa1EIsH06dMRFRWFCxcuwNjYWKZ9+fv7IzMzk7H4T/umxpnqk0QiwfJNP+HslesIXfs9Who2YzsSJ3zWsimep2Yxyl6kZaG5VmOWEnGHupoaLC3Mce3mbUZ59M3bsLXh9qUcdU1dTQ2W7S1w7eYtRnn0jVuw7WjNUiqiTNi+a1iR7ibm7MjgiRMnkJ6ejgkTJkBLizmqMmzYMAQFBcHFxaXSbWfPng1PT0/4+fkxRhX/SyKRIDk5GRKJBBkZGbh+/TpWrlwJLS0trFq1Slpv2rRpCA8Px7Fjx6ChoYHkfx5XoaWlhYYNG1aZXyAQQCAQMI+ZJd9TxMs3/YQT5y9j64oFaNyoId6mlY2+ajRuhAbl3guR3diuFhi99zR+unYfzu3bIF6cgoOxT7B0oB3b0Thh3Ojh8AtYAav2FrC1sUJk1DGIk19jhIcb29EU3rjRI+G3eOk/bWuNyKij/7StO9vRFFpObi5evvpbuv4qSYyEx0+gpamJ5kL6Ek5qjrOdwaCgIPTr169CRxAAPDw8sHLlSqSlpVW6rYuLC4yMjBAYGIht27ZVWicrKwuGhobg8XjQ1NSEubk5vv76a3z33XfQ1NSU1tu+fTsAoE+fPoztQ0JC4O3t/XFvTk5FHP8VADDWdyGjfOXcGXB3/oKNSJxg3Vwfmzx6Y/3FWGy/eg8ttZtgfr/PMcRKtlFmUr1BTv2QnpmFbbtD8CYlFe1M2mLnxrVoYShkO5rCGzSgP9IzM7FtVzDepKSUte2m9WjR3JDtaArtfkIixk75Trou2rAFAOA22BmrlixgKxZRYDxJVc9ZIXJH8uoR2xE4qfT8QbYjcBbffQrbEbhLhbPf5dlXXPDhOuTjaNXfyOXmxvJ5zbpPTgrbESrg7DWDhBBCCCHkw6gzSAghhBCixGiegRBCCCGcI6937sojGhkkhBBCCFFi1BkkhBBCCFFiNE1MCCGEEM6h0S7ZUVsRQgghhCgxGhlUJLlZH65Daqzw1Dm2I3BWwy8nsR2Bu+jTu85I8t6xHYGzePX4nEEiO/o4IYQQQgjn0N3EsqNpYkIIIYQQJUadQUIIIYQQJUbTxIQQQgjhHBXQPLGsaGSQEEIIIUSJUWeQEEIIIUSJUWewGtHR0eDz+XB2dgYAeHt7g8fjVbv8t97kyZMr7HPq1Kng8Xjw9vauz7dSLyJOnsOXPgvQ+atJ6PzVJAyfswyX78SxHYsbGjSE2tjpEGyKRIM9Z6C+bCt4bS3YTsUJt2PiMHn2fDi4uMPcrjfOXbrCdiROCTsYhb5fDoO1vSPcvcbjTkws25EU3u24+5g8fxl6uo+BRe/BOHflOtuR5JIKTz4XeUSdwWoEBwfDx8cHV69excuXL7Fx40aIxWLpAgAhISEVygCgVatW2L9/P/Ly8qRl+fn5iIiIQOvWrev9vdSHZvq6mP31Vzi0fjkOrV8OO5sOmBa4Hk/+fMV2NIWn9o0fVKw/R9G2QBT4jUPpvdsQLFwH6OizHU3h5eblwdzMFAGzZ7IdhXNOnjkH0bqNmDJ+LI6GhaCzrQ0mzZiDpORktqMptLy8fFiYGmPxzIoDDoR8DLqBpAo5OTk4cOAAbt++jeTkZISGhiIgIABaWlqMetra2hAKhRW2/+yzz/Ds2TNERUVh9OjRAICoqCi0atUKbdu2rZf3UN/6dv2MsT5rrCf2/3oecYlPYdamJUupOEBNHfyuvVC4biFKH90DABQfDgX/cweo9h+K4gNBLAdUbL3t7dDb3o7tGJwUEhYJj6Eu8HT9EgCwcPZMXL1+CxGHjmD29Cksp1Ncvew+Ry+7z9mOQTiERgarEBkZCXNzc5ibm8PLywshISGQSCQ12se4ceMQEhIiXQ8ODsb48eNrO6pcKikpxf8uX0dufgE6WZixHUex8fng8VWBwkJmeWEhVMyt2clEyAcUFhXhwaNEONh1ZZT3sOuKmHv3WUpFlAlPThd5RJ3BKgQFBcHLywsA4OzsjOzsbJw/f75G+xgzZgyuXr2KFy9e4M8//8S1a9ek+/yQgoICZGVlMZaC8p0BOZT44i985jkRNu7jsHRbKLYs/A6mrVuwHUux5eeh5PF9qLqPBXT0AJ4K+A79wTNtD562HtvpCKlUekYGSkpKoKeryyjX19XB25RUllIRQipDncFKJCYm4tatWxgxYgQAQFVVFcOHD0dwcHCN9qOvr4/Bgwdjz549CAkJweDBg6GvL9s1XiKRCFpaWoxF9NOeGr+X+mbcwhBHNgZi/9olGDGwL+av34mnL/9mO5bCK9oaCPB4aLgtCg1+PgvVAR4oiT4HlJayHY2Qar2/se49iaRiGSGEXXTNYCWCgoJQXFyMFi3+HdGSSCRQU1NDeno6dHR0ZN7X+PHjMX36dADA1q1bZd7O398fvr6+jDL1l/dk3p4t6mqqaNO87IfIrc3a4v6T59h7/DSWT1eO6fG6InmThMLl3wGCBkDDRkBGGtRmLIHkrfjDGxPCAh1tbfD5fKSkMkcBU9PToa+nW8VWhNQeeb1zVx7RyGA5xcXF2Lt3L9atW4fY2FjpEhcXhzZt2iAsLKxG+3N2dkZhYSEKCwsxYMAAmbcTCATQ1NRkLAJ19Zq+HdZJJBIUFhWxHYM7CvKBjDSgcRPwbbqg5M41thMRUil1NTVYWpjj2s3bjPLom7dha2PFUipCSGVoZLCcEydOID09HRMmTKhw5/CwYcMQFBQkHemTBZ/PR0JCgvT/c9mPew+gV+eOEOrrIicvHycv38Ct+wnYtXQu29EUnopNF4DHgyTpJXjCllAbNRkS8V8ouXSS7WgKLyc3Fy9f/Xspw6skMRIeP4GWpiaaC5uxmEzxjRs9HH4BK2DV3gK2NlaIjDoGcfJrjPBwYzuaQsvJzcPLv5Ok66/EyUh48ge0NDXQvJkBi8mIoqLOYDlBQUHo169fhY4gAHh4eGDlypW4e/cuPvvss0q2rpympmZtRpRbqRmZ8PtxB96mZUCjcUOYG7XGrqVz0cOW7nj9VLxGTaA6YhJ4uk2B7HcouXUJRZG7gZIStqMpvPsJiRg7baZ0XbSx7HIOt0HOWBXgz1Iqbhjk1A/pmVnYtjsEb1JS0c6kLXZuXIsWhhUfx0Vkdz/xCb6e+e+5uWrrbgCAq/MXWOXvW9VmSod+m1h2PElNn5dCWCN5fIvtCJyUv4RGLutKw22RbEfgLlX6Ll9XJDkZbEfgLJ7QtN6OtVdLPkdJx2a+YTtCBXTNICGEEEKIEqOvloQQQgjhHLqbWHY0MkgIIYQQosSoM0gIIYQQosRompgQQgghnEOjXbKjtiKEEEIIUWLUGSSEEEIIUWI0TaxAeAZt2I7ASQ1Em9iOwF3qDdhOQEiN8Roqxw8FcB3dTCw7GhkkhBBCCFFi1BkkhBBCCFFiNE1MCCGEEM5R4dFEsaxoZJAQQgghRIlRZ5AQQgghRInRNDEhhBBCOIcmiWWntCOD3t7e4PF40kVPTw/Ozs64d++etM5/X2/SpAk6duyI0NDQSvcXHh4OPp+PyZMnV3jt4sWLjH29XxYtWlRXb48Vt2NiMXn2fDgMdoN5t144d+kK25E44af9RzDMxx+fuY6F/VcTMW3pGjz7K4ntWJwSduAQ+rq4wtquJ9xHjcWduzFsR+IMatu6E3boCPoO/QrWDl/AfewE3ImJYzsSUVBK2xkEAGdnZ4jFYojFYpw/fx6qqqpwcXFh1AkJCYFYLEZcXByGDx+OcePG4fTp0xX2FRwcDD8/P+zfvx+5ubmVHi8xMVF6PLFYjPnz59fJ+2JLbl4+zM1MEDBnJttROOX2vYcYNWQAIjcEIli0CMUlpZi44Hvk5uezHY0TTp4+C9Ha9ZgyYRyOhu9FZ9tOmOQzC0niZLajKTxq27pz8ux5iH7chCnjxuDoz0Ho3KkjJs2ci6Tk12xHIwpIqTuDAoEAQqEQQqEQnTp1wrx58/DXX3/h7du30jra2toQCoUwMTHBggULoKurizNnzjD28+LFC0RHR2P+/PmwsLDAoUOHKj2egYGB9HhCoRBNmjSp0/dX33rb22HW5ElwcuzNdhRO2b1yIdyd+sDMqBUsTIwgmj0VSW9S8ODJM7ajcUJIWAQ8XL+Ep9tQmLQ1xsK5vhA2a4aIQ4fZjqbwqG3rTkh4JDy+HAxP1yEwMTbCQt8ZEDYzQMThI2xHkxs8OV3kkVJ3Bv8rOzsbYWFhMDU1hZ6eXoXXS0pKcODAAaSlpUFNTY3xWnBwMAYPHgwtLS14eXkhKCiovmITJfQup2zkWUuDW18m2FBYVIQHCY/gYNeNUd6je1fExMWzlIobqG3rTmFRER48egyHbl0Z5T26dUHMvfsspSKKTKlvIDlx4oR0dC4nJweGhoY4ceIEVFT+7SOPHDkSfD4f+fn5KCkpga6uLiZOnCh9vbS0FKGhodi8eTMAYMSIEfD19cXTp09hamrKOF7Lli0Z63/++WelHU8AKCgoQEFBAaNMUFAAgUDw8W+YKDyJRIJVO/egs6UF2hm1ZjuOwkvPyEBJSQn09HQZ5fq6enibeoOlVNxAbVt30jMy/2lbHUa5vq4O3qamsZSKKDKlHhl0dHREbGwsYmNjcfPmTTg5OWHgwIH4888/pXXWr1+P2NhYnD17Fp06dcL69esZnbwzZ84gJycHAwcOBADo6+vDyckJwcHBFY535coV6fFiY2Oho6NToc57IpEIWlpajEW0nn5DV9mt2BqExOcvsc7/O7ajcAqv3OSNRCIBjx5YWyuobetOxbYFte1/sD0drEjTxEo9Mti4cWNGx65z587Q0tLCrl278P333wMAhEIhTE1NYWpqioMHD8LW1haff/45OnToAKBsijgtLQ2NGjWS7qe0tBQxMTFYsWIF+Hy+tNzY2Bja2toyZfP394evry+jTJCX8ZHvlHDBiq3BuHD9d+xbtwzCppWPKJOa0dHWBp/PR0pqKqM8NT0N+rq6VWxFZEFtW3d0tLX+aVvmKGBqejr0daseZCCkKko9Mlgej8eDiooK8vLyKn3d1NQUHh4e8Pf3BwCkpqbi2LFj2L9/P2PELzY2FtnZ2fj1118/OotAIICmpiZjoSli5SSRSLB8SxDOXruJ0DUBaCk0YDsSZ6irqcGyvQWu3bzFKI++cQu2Ha1ZSsUN1LZ1R11NDZYW7XDt1m1GefSt27C1sWIpFVFkSj0yWFBQgOTkskccpKenY8uWLcjOzsaQIUOq3Gb27Nno2LEj7ty5g6tXr0JPTw+enp6M6wwBwMXFBUFBQRUeVcNlObm5ePnqb+n6qyQxEh4/gZamJpoLm7GYTLEt3xKEE79dxdalfmjcsCHepmUAADQaN0IDgTq74Thg3OiR8Fu8FFbtLWBrY43IqKMQJ7/GCA93tqMpPGrbujNu1HD4Lfm+rG2tLRF55DjEyW8wwt2V7Whyg6bMZafUncFTp07B0NAQAKChoQELCwscPHgQffr0qXIba2tr9OvXDwEBAXj16hXc3NwqdAQBwMPDA8OHD8fr18rzzKf7CYkYO/Xfa9lEG7YAANwGO2NVwAK2Yim8iBNljzIaO3cpo3zl7Klwd+pT/4E4ZtCA/kjPzMS2XcF4k5KCdiZtsXPTerRobsh2NIVHbVt3BvX/AumZWdgWFIo3KaloZ2KMnevXoIWhkO1oRAHxJBKJhO0QREYZytOxrE+SDHoAbl3hNW3DdgRCaq64kO0E3KVVf5e5HNaVz46xR5r8/c1R6pFBQgghhHATTRLLjm4gIYQQQghRYtQZJIQQQghRYjRNTAghhBDOodEu2VFbEUIIIYQoMeoMEkIIIYTIqcuXL2PIkCFo3rw5eDwejh49WmXdb7/9FjweDxs2bKjRMagzSAghhBDO4fHkc6mpnJwcdOzYEVu2bKm23tGjR3Hz5k00b968xsegawYViRr92kRd4OnW/B8OkRH9AgBRRGr0059EfgwcOBADBw6sts7ff/+N6dOn4/Tp0xg8eHCNj0GdQUIIIYSQelJQUICCggJGmUAggEDwcV9CSktLMWbMGMydOxeWlpYftQ+aJiaEEEII5/Dk9H8ikQhaWlqMRSQSffT7XL16NVRVVTFjxoyP3geNDBJCCCGE1BN/f3/4+voyyj52VPD333/Hxo0bcffuXfA+4bIcGhkkhBBCCKknAoEAmpqajOVjO4NXrlzBmzdv0Lp1a6iqqkJVVRV//vknZs+eDSMjI5n3QyODhBBCCOEcZbh9bcyYMejXrx+jbMCAARgzZgzGjRsn836oM0gIIYQQIqeys7Px9OlT6frz588RGxsLXV1dtG7dGnp6eoz6ampqEAqFMDc3l/kYCjFNnJycDB8fH7Rt2xYCgQCtWrXCkCFDcP78eQCAkZEReDweeDweGjZsCCMjI3z11Ve4cOECYz8vXrwAj8dDbGystOzdu3fo06cPLCws8NdffwEAeDweGjRogD///JOxvaurK7y9vaXrIpEIXbp0gYaGBgwMDODq6orExMRK30N4eDj4fD4mT55cCy0iv8IOHEJfFzdY2/WC+6ivceduLNuROCPsYBT6DvWEdY++cB8zHndi4tiOxBlhBw6h7+ChsO7mAPdRY3HnbgzbkTiD2rbuUNsqhzt37sDW1ha2trYAAF9fX9ja2iIgIKDWjiH3ncEXL16gc+fOuHDhAtasWYP4+HicOnUKjo6OmDZtmrTe8uXLIRaLkZiYiL1790JbWxv9+vVDYGBglft++/YtHB0dkZ2djatXr6JVq1bS13g83gcb+tKlS5g2bRpu3LiBs2fPori4GE5OTsjJyalQNzg4GH5+fti/fz9yc3M/oiXk38nTZyFauwFTJnjjaPgedLbthEk+s5AkTmY7msI7eeY8RD9uwpRxY3F0XzA6d+qISd/NQVIyte2nOnn6LEQ//IgpE8bhaMTPZeft9Jl03tYCatu6Q237YTw5XWqqT58+kEgkFZbQ0NBK67948QIzZ86s0THkvjM4depU8Hg83Lp1C8OGDUO7du1gaWkJX19f3LhxQ1pPQ0MDQqEQrVu3Rq9evbBz504sXrwYAQEBlY7W/fXXX+jZsyc0NDTw22+/QV9fn/G6j48P9u3bh/j4+CqznTp1Ct7e3rC0tETHjh0REhKCly9f4vfff2fUe/HiBaKjozF//nxYWFjg0KFDn9gq8ikkLAIerkPg6TYUJm2NsXDuLAibGSDiUBTb0RReSPh+eAx1gafrEJgYG2Hh7O/+adujbEdTeCH7wuHh+iU83V3/OW99IRQ2Q8TBw2xHU3jUtnWH2pbUJrnuDKalpeHUqVOYNm0aGjduXOF1bW3tarf/7rvvIJFIcOzYMUZ5YmIievToAQsLC5w6dQoaGhoVtrW3t4eLiwv8/f1lzpuZmQkA0NXVZZQHBwdj8ODB0NLSgpeXF4KCgmTep6IoLCrCg4REONh1Y5T36N4NMXFVd6jJhxUWFeHBo8dw6NaFUd6jWxfE3LvPUipuKDtvH8Ghe7nz1q4bYuLusZSKG6ht6w61Laltct0ZfPr0KSQSCSwsLD5qe11dXRgYGODFixeM8rFjx8LExASHDx+u9nZukUiEU6dO4cqVKx88lkQiga+vLxwcHGBlZSUtLy0tRWhoKLy8vAAAI0aMwPXr1xkXg1amoKAAWVlZjKX8E8vlSXpGBkpKSqCnx+wI6+vq4m1qKkupuCE9I7Osbct9ydDXo7b9VOnp/5y3uswLsKltPx21bd2htpWNCk8+F3kk151BiUQCAJ/0IEWJRFJh+6FDh+Lq1as4fLj64fQOHTpg7NixmDdv3gePM336dNy7dw8RERGM8jNnziAnJ0f6u4L6+vpwcnJCcHBwtfur9Anla9d/MAfbeOWuiKis/cnHKd+O1La1p3wzUtvWHmrbukNtS2qLXD9axszMDDweDwkJCXB1da3x9qmpqXj79i2MjY0Z5QsWLICNjQ1Gjx4NiUSC4cOHV7mPZcuWoV27djh69GiVdXx8fHD8+HFcvnwZLVu2ZLwWHByMtLQ0NGrUSFpWWlqKmJgYrFixAnw+v9J9VvqE8mL5vfFER1sbfD4fKeW+laamp0O/3IgWqRkdba3K2zaN2vZT6ehUcd5S234yatu6Q21Laptcjwzq6upiwIAB2Lp1a6V36GZkZFS7/caNG6GiolJpR3LRokVYsWIFRo8eXWE0779atWqF6dOnY8GCBSgpKWG8JpFIMH36dERFReHChQsVOp2pqak4duwY9u/fj9jYWMaSnZ2NX3/9tcrj1uYTyuuDupoaLNub49rNW4zy6Bu3YNvRmqVU3KCupgZLi3a4dvM2ozz61h3Y2lhVsRWRRdl5a4FrNyo7b21YSsUN1LZ1h9pWNmz/BnFV/5NHcj0yCADbtm2Dvb09unbtiuXLl8PGxgbFxcU4e/Ystm/fjoSEBABlzwtMTk5GUVERnj9/jn379mH37t0QiUQwNTWtdN/z588Hn8/HmDFjUFpaitGjR1daz9/fH7t27cLz588Zo4jTpk1DeHg4jh07Bg0NDST/85gPLS0tNGzYED///DP09PTg6ekJFRVmv9vFxQVBQUFwcXGpjWaSC+NGj4Tf4mWwat8etjZWiIw6BnHya4zwcGM7msIbN2oE/JasgFUHC9haWyHyyPF/2taV7WgKb5zXKPgtWgKrDu1ha2ONyKgjECcnY8Qwd7ajKTxq27pDbUtqk9x3Bo2NjXH37l0EBgZi9uzZEIvFaNq0KTp37ozt27dL6wUEBCAgIADq6uoQCoWws7PD+fPn4ejoWO3+586dCz6fj6+//hqlpaUYM2ZMhTq6urqYN28eFixYwCh/f/w+ffowykNCQuDt7Y3g4GC4ublV6AgCgIeHB4YPH47Xr1+jWbNmsjaHXBs0oD/SMzOxbVcQ3qSkop1JW+zc9CNaNDdkO5rCG+T0RVnb7g79p22NsXPDD2hhKGQ7msKTnrc7g/AmJQXtTE2wc/N6Om9rAbVt3aG2JbWJJ3l/lwaRfznpbCfgppJithNwl6o62wkIIfKkkVa9HeqUfvN6O1ZNOKcksR2hArm+ZpAQQgghhNQt6gwSQgghhCgxub9mkBBCCCGkpuiRi7KjkUFCCCGEECVGnUFCCCGEECVG08SEEEII4RyaJZYdjQwSQgghhCgxGhlUJKWlbCfgJElBHtsROIvHV2M7AnfR1fF1p7Tkw3UI4RDqDBJCCCGEc1RoolhmNE1MCCGEEKLEqDNICCGEEKLEaJqYEEIIIZxDk8Syo5FBQgghhBAlxsnOYHJyMnx8fNC2bVsIBAK0atUKQ4YMwfnz5wEARkZG2LBhQ4Xtli5dik6dOjHWeTxehcXCwkJap0+fPpg5cyZjPxs3boRAIEB4eDgAQCQSoUuXLtDQ0ICBgQFcXV2RmJhY6+9bHoQdPIy+X3rA2r4P3L3G4U5MLNuROOF27D1M9luMnkOHw8KhP85dvsZ2JE4JO3AIfV1cYW3XE+6jxuLO3Ri2I3FG2IFD6Dt4KKy7OVDb1rKwA4fRd4g7rLv3hvtob/q8JR+Nc53BFy9eoHPnzrhw4QLWrFmD+Ph4nDp1Co6Ojpg2bVqN92dpaQmxWMxYrl69WmX9JUuWwN/fH0eOHMGoUaMAAJcuXcK0adNw48YNnD17FsXFxXByckJOTs5Hv095dPLMOYjWbcSU8V/jaFgoOtt2xKQZs5GUnMx2NIWXl5cPC9O2WOw7ne0onHPy9FmI1q7HlAnjcDR8LzrbdsIkn1lIEtN5+6lOnj4L0Q8/lrVtxM9lbTt9JrVtLSj7vN2AKeO9cTR8T9nnrY8vte1/8Hjyucgjzl0zOHXqVPB4PNy6dQuNGzeWlltaWmL8+PE13p+qqiqEQuEH60kkEsyYMQM///wzzpw5AwcHB+lrp06dYtQNCQmBgYEBfv/9d/Tq1avGmeRVSNh+eAwdAk/XLwEAC2fPxNXrNxFx6AhmT5/CcjrF1qt7V/Tq3pXtGJwUEhYBD9cv4ek2FACwcK7vP+ftYcz2qfkXSPKvkH3hZW3r7grgfdveQMTBw5g9g9r2U4Tsiyj7vHX75/N2zqx/ztsozPaZynI6omg4NTKYlpaGU6dOYdq0aYyO4Hva2tp1ctzi4mKMGTMGBw8exKVLlxgdwcpkZmYCAHR1deskDxsKi4rw4FEiHOyYHZYedl0Rcy+epVSEVK+wqAgPEh7Bwa4bo7xH966IiaPz9lNI27Z7uba164aYuHsspeKGqj9vu9HnLfkonBoZfPr0KSQSCeOavqrMmzcPixYtYpQVFhaiQ4cOjLL4+Hg0adKEUTZixAjs3r1bur5r1y4AQFxc3AePLZFI4OvrCwcHB1hZWX0wp6JIz8hASUkJ9Mp1cPV1dfE2JY2lVIRUT3re6pU/b/XwNvUGS6m4IT39/WeCHqNcX08Xb1NTWUrFDVWet3o6eJtKn7fvyemMrFziVGdQIpEAAHgyTMrPnTsX3t7ejLJNmzbh8uXLjDJzc3McP36cUaahocFYd3BwQGxsLBYtWoT9+/dDVbXqZp0+fTru3btX7XWHAFBQUICCggJGmaCwAAKBoNrt2Fa+6SUSidxeI0HIe7xyfzbKzls6cWtD5Z8J1La1oXw7SiTUASIfh1PTxGZmZuDxeEhISPhgXX19fZiamjKWyqZt1dXVK9Rr1qwZo461tTXOnz+Pixcv4quvvkJRUVGlx/Tx8cHx48fx22+/oWXLltXmE4lE0NLSYiyidRs++L7YoqOtDT6fj5Ry30pT09Ohr8ed6XDCLf+et8yRqtT0NOhz6DIONujoVNG2aenUtp9Iet6mVNK29HlLPgKnOoO6uroYMGAAtm7dWumduhkZGXV27E6dOuHChQu4evUqPD09GR1CiUSC6dOnIyoqChcuXICxsfEH9+fv74/MzEzG4j97Zp3l/1TqamqwtDDHtZu3GOXRN2/D1saapVSEVE9dTQ2W7S0qnrc3bsG2I523n0Latjcqa1sbllJxw7+ft7cZ5dE3b9Hn7X/w5PR/8ohT08QAsG3bNtjb26Nr165Yvnw5bGxsUFxcjLNnz2L79u0yjRr+V3FxMZLLPRqFx+NVGB0EABsbG/z222/o27cvhg0bhoMHD0JdXR3Tpk1DeHg4jh07Bg0NDen+tLS00LBhw0qPKxAIKk4Jv6t8xFFejBs9An4By2HVvj1sbawQGXUM4uTXGOHhynY0hZeTm4eXf/8tXX8lTkbCk6fQ0tBEc6EBi8kU37jRI+G3eCms2lvA1sYakVFH/zlv3dmOpvDGeY2C36IlsOrQ/p+2PQJxcjJGDKO2/VTjvEbCb/EyWHUod94Oc2M7GlFAnOsMGhsb4+7duwgMDMTs2bMhFovRtGlTdO7cGdu3b6/x/h48eABDQ0NGmUAgQH5+fqX1LS0t8dtvv+GLL76Ah4cHDh8+LD1unz59GHVDQkIqXLeoyAY59UN6Zia27Q7Gm5RUtDNpi50b16JFufYjNXf/0WN8PWOOdH3V5h0AANeB/bFqoR9bsThh0ID+ZeftrmC8SUkpO283rUeL5nTefipp2+4MKmtbUxPs3ExtWxsGOfVDesb78/afz9tN6+jzlnwUnuT9XRdE/r2jO/DqgiSfWw//lie8RppsR+Auugmj7pSWsJ2Au5rU3zWNl5tVf20+W3q9fsV2hAo4dc0gIYQQQgipGeoMEkIIIYQoMc5dM0gIIYQQQhdSyI5GBgkhhBBClBh1BgkhhBBClBhNExNCCCGEc2iaWHY0MkgIIYQQosRoZFCRFBeynYCbMt+ynYC76DmDdYceEVt36DmDRMlQZ5AQQgghnCOvvwMsj2iamBBCCCFEiVFnkBBCCCFEidE0MSGEEEI4h36+W3Y0MkgIIYQQosSoM0gIIYQQosQUpjOYnJwMHx8ftG3bFgKBAK1atcKQIUNw/vx58Hi8apfQ0FBcvHgRPB4PVlZWKClhPjZAW1sboaGh0nUjIyPweDzcuHGDUW/mzJno06ePdH3p0qWM42hpaaFnz564dOkSY7udO3eiT58+0NTUBI/HQ0ZGRm03j1y4HROHybP94eDiAXO7Pjh36QrbkTjhpwPHMGzmInw2bDzsR03GtBXr8OxVEtuxOCXswCH0dXGFtV1PuI8aizt3Y9iOxBnUtnUn7GAU+g71hHWPvnAfMx53YuLYjiRXVOR0kUfymovhxYsX6Ny5My5cuIA1a9YgPj4ep06dgqOjIyZNmgSxWCxdvvrqKzg7OzPKhg8fLt3XH3/8gb17937wmA0aNMC8efM+WM/S0lJ6nOvXr8PMzAwuLi7IzMyU1snNzYWzszMWLFjwcQ2gIHLz8mFuZoKA2d+xHYVTbscnYNTg/ohctxzB3/ujuKQUExetQm5+PtvROOHk6bMQrV2PKRPG4Wj4XnS27YRJPrOQJE5mO5rCo7atOyfPnIfox02YMm4sju4LRudOHTHpuzlISqa2JTWnEJ3BqVOngsfj4datWxg2bBjatWsHS0tL+Pr64u7duxAKhdKlYcOGEAgEFcre8/HxwZIlS5D/gT+k3377LW7cuIGTJ09WW09VVVV6nA4dOmDZsmXIzs7G48ePpXVmzpyJ+fPnw87O7tMaQs71tu+GWZMnwsmxF9tROGX3ivlw798bZm1awqJtG4hmfYuktyl48PQ529E4ISQsAh6uX8LTbShM2hpj4VxfCJs1Q8Shw2xHU3jUtnUnJHw/PIa6wNN1CEyMjbBw9ncQNjNAxKGjbEcjCkjuO4NpaWk4deoUpk2bhsaNG1d4XVtbu0b7mzlzJoqLi7Fly5Zq6xkZGWHy5Mnw9/dHaWmpTPsuKChAaGgotLW1YW5uXqNchMjqXU4uAECrSROWkyi+wqIiPEh4BAe7bozyHt27IiYunqVU3EBtW3cKi4rw4NFjOHTrwijv0a0LYu7dZymV/OHJ6SKP5L4z+PTpU0gkElhYWNTK/ho1aoQlS5ZAJBIxpnIrs2jRIjx//hxhYWFV1omPj0eTJk3QpEkTNGzYEGvXrkVERAQ0NelnuEjtk0gkWLVrHzpbmqOdUSu24yi89IwMlJSUQE9Pl1Gur6uHt6mpLKXiBmrbupOekVnWtrrl2lZPl9qWfBS57wxK/vn9TV4tPjBowoQJ0NfXx+rVq6ut17RpU8yZMwcBAQEoLKz8d4HNzc0RGxuL2NhY/P7775gyZQo8PT1x586dT8pYUFCArKwsxlJQUPBJ+ySKb8X2UCS+eIl1ftPZjsIp5X+2SiKR1OpnjjKjtq075duR2pZ8LLnvDJqZmYHH4yEhIaHW9qmqqorvv/8eGzduRFJS9Xdl+vr6Ii8vD9u2bav0dXV1dZiamsLU1BS2trZYtWoVWrRogQ0bNnxSRpFIBC0tLcYiWr/5k/ZJFNuK7aG4cPN37BUtglBfj+04nKCjrQ0+n4+UcqMpqelp0C836kJqhtq27uhoa1Xetmnp1Lb/8aEnjbC1yCO57wzq6upiwIAB2Lp1K3Jyciq8/rGPafH09ISlpSWWLVtWbb0mTZpg8eLFCAwMRFZWlkz75vP5yMvL+6hc7/n7+yMzM5Ox+M/y+aR9EsUkkUiwfHsIzl6/jdCVC9FSaMB2JM5QV1ODZXsLXLt5i1EefeMWbDtas5SKG6ht6466mhosLdrh2s3bjPLoW3dga2PFUiqiyOS+MwgA27ZtQ0lJCbp27YrDhw/jyZMnSEhIwKZNm9C9e/eP3u+qVasQHBxcaSfzv7755htoaWkhIiKiwmvFxcVITk5GcnIynjx5gu+//x4PHz7E0KFDpXWSk5MRGxuLp0+fAii7zjA2NhZpaWlVHlMgEEBTU5OxCASCj3yn9SMnNxcJj58g4fETAMCrpGQkPH6CpOTXLCdTbMu3heCX365h7dzpaNywId6mZeBtWgbyCyq/dIHUzLjRI3HoyDEcOnocfzx7jpVr10Oc/BojPNzZjqbwqG3rzrhRI3Do2AkcOn4Cfzx/gZU/bvqnbV3ZjkYUkEL8NrGxsTHu3r2LwMBAzJ49G2KxGE2bNkXnzp2xffv2j95v37590bdvX5w5c6baempqalixYgVGjRpV4bUHDx7A0NAQQNnNKSYmJti+fTvGjh0rrbNjxw7GCGSvXmWPXgkJCYG3t/dH55c39xMSMXbaLOm6aONWAIDboAFYFeDPViyFF3HyHABg7PwVjPKVM7+Fe//ebETilEED+iM9MxPbdgXjTUoK2pm0xc5N69GiuSHb0RQetW3dGeT0RVnb7g7Fm5RUtDMxxs4NP6CFoZDtaHJDPidk5RNP8v4ODSL/0sVsJ+AkSSr9mkdd4RmasB2BkJorKWI7AXdpNq23Q902bF1vx6qJLuKXbEeoQCGmiQkhhBBCSN1QiGliQgghhJCaoGli2dHIICGEEEKIEqPOICGEEEKIEqNpYkIIIYRwjrw+4Fke0cggIYQQQogSo84gIYQQQogSo2liRVJSzHYCTip9/SfbETiLb9iW7QiE1Bx91nKCCs0Sy4xGBgkhhBBClBh1BgkhhBBClBhNExNCCCGEc3g0TywzGhkkhBBCCFFi1BkkhBBCCFFiNE1MCCGEEM6hZ07LTilHBr29vcHj8Soszs7O0joxMTEYPnw4DA0NIRAI0KZNG7i4uOCXX36BRCIBALx48aLS/Xh5eVX6upaWFuzs7PDLL7+w8r7r2u3Ye5jstwgOXw6HeY9+OHf5GtuROGnn/y6iw3h/iMK5eR6xIezAIfR1cYO1XS+4j/oad+7Gsh2JM6hta9/tmDhMnu0PBxcPmNv1wblLV9iORBScUnYGAcDZ2RlisZixREREAACOHTsGOzs7ZGdnY8+ePXj48CEOHjwIV1dXLFq0CJmZmYx9nTt3jrGfrVu3Vvr6zZs30bVrV3h4eOD+/fv19l7rS25ePsxN2yLAdzrbUTgr/vlfOHjpFsxbCtmOwhknT5+FaO0GTJngjaPhe9DZthMm+cxCkjiZ7WgKj9q2buTm5cPczAQBs79jOwrhCKWdJhYIBBAKK/5BzcnJwYQJEzB48GBERUVJy01MTNC1a1dMnDhROjL4np6eXqX7Kv+6UChEYGAgNm/ejN9++w1WVla194bkQO/uXdG7e1e2Y3BWTn4B/HZGYtnX7vjpxAW243BGSFgEPFyHwNNtKABg4dxZuHr9BiIORWG2z1SW0yk2atu60du+G3rbd2M7htyjaWLZKe3IYFXOnDmD1NRU+Pn5VVnnY3/8uqioCLt27QIAqKmpfdQ+iPL6ft8x9LaxgL2lKdtROKOwqAgPEhLhYMf8w9qjezfExMWzlIobqG0JURxKOzJ44sQJNGnShFE2b948qKurAwDMzc2l5bdv34ajo6N0ff/+/XBxcZGu29vbQ0Xl3371lStXYGtrW+H1vLw8lJaWwsjICF999VWtvyfCXSdvxuHhn0k4EDCN7Sickp6RgZKSEujp6TLK9XV18TY1laVU3EBtS4jiUNrOoKOjI7Zv384o09XVlY7c/ZeNjQ1iY2MBAGZmZiguZv5uZWRkJNq3by9db9WqVYXXLSws8PjxY8ycORM7duyAri7zA7K8goICFBQUMMoEBQUQCAQffG+EW8RpGRBFnMAu3/EQ0IhyneCBOdovkUg+egaAMFHbErbQeSY7pe0MNm7cGKamFafbzMzMAACJiYmws7MDUHZ9YWV132vVqtUHXzczM4OZmRmaNGkCDw8PPHz4EAYGBlVuIxKJsGzZMkbZkrkzsdTPt9r3RbjnwYu/kZqVDc/lW6RlJaWluPP4BcIv3EDszhXgq9AVHx9DR1sbfD4fKeVGqlLT06H/gS9spHrUtoQoDvoLUo6TkxN0dXWxevXqOtl/7969YWVlhcDAwGrr+fv7IzMzk7H4f0dThMqoe3tTHFv+HaKW+kgXK6MWcLHriKilPtQR/ATqamqwbG+OazdvMcqjb9yCbUdrllJxA7UtIYpDaUcGCwoKkJzMfLyBqqoq9PX1sXv3bgwfPhyDBw/GjBkzYGZmhuzsbJw6dQoAwOfzP+nYs2fPhqenJ/z8/NCiRYtK6wgEgopTwoWZldaVFzm5eXj56m/p+qskMRIeP4WWpgaaC5uxmEyxNW4ogFm5R8k0FKhDu3GjCuWk5saNHgm/xctg1b49bG2sEBl1DOLk1xjh4cZ2NIVHbVs3cnJzy33WJiPh8RNoaWrSZ+1/0Cyx7JS2M3jq1CkYGhoyyszNzfHo0SO4ubkhOjoaq1evxtixY5GWlgYtLS18/vnnFW4e+RguLi4wMjJCYGAgtm3b9kn7kif3HyVirM8c6bpo8w4AgNtAJ6xaVPXd2YSwadCA/kjPzMS2XUF4k5KKdiZtsXPTj2jR3PDDG5NqUdvWjfsJiRg7bZZ0XbSx7Nm2boMGYFWAP1uxSB25fPkyfvjhB/z+++8Qi8U4cuQIXF1dAZQ9pWTRokU4efIknj17Bi0tLfTr1w+rVq1C8+bNZT4GT1L+oXlEfqX8xXYCTipJvM12BM7id3L8cCVC5E1hPtsJuEun/r4I3Dc2rrdj1YTV8+c1qv/rr7/i2rVr+Oyzz+Dh4cHoDGZmZmLYsGGYNGkSOnbsiPT0dMycORPFxcW4c+eOzMdQ2pFBQgghhHAXV+4mHjhwIAYOHFjpa1paWjh79iyjbPPmzejatStevnyJ1q1by3QM6gwSQgghhNSTSh8dV9l9Ah8pMzMTPB4P2traMm9DtyESQgghhNQTkUgELS0txiISiWpl3/n5+Zg/fz5GjRoFTU1NmbejkUFCCCGEcI68zhL7+/vD15f5zODaGBUsKirCiBEjUFpaWuObU6kzSAghhBBST2pzSvi9oqIifPXVV3j+/DkuXLhQo1FBgDqDhBBCCCEK631H8MmTJ/jtt9+gp6dX431QZ5AQQgghnKMir/PENZSdnY2nT59K158/f47Y2Fjo6uqiefPmGDZsGO7evYsTJ06gpKRE+oMaurq6UFdXl+kY9JxBBSJ584LtCJxUSs8ZrDP8z/qxHYGQmqPnDNadenzO4CNTk3o7Vk1YPP2jRvUvXrwIR8eKz2z9+uuvsXTpUhhX8TzF3377DX369JHpGDQySAghhBAip/r06YPqxu1qY0yPOoOEEEII4RyOzBLXC3rOICGEEEKIEqPOICGEEEKIEqNpYkIIIYRwDld+m7g+0MggIYQQQogS41xn0NvbGzwer8Li7OwMADAyMpKWNWzYEBYWFvjhhx8Yd+O8ePGCsa26ujpMTU3x/fffM+otXbqUUU9LSws9e/bEpUuXpHXS0tLg4+MDc3NzNGrUCK1bt8aMGTOQmZlZf41ST27HxmPyvAD0dB0Ji54DcO5yNNuROGnnyUvoMHERRPv/x3YUzgg7cAh9XdxgbdcL7qO+xp27sWxH4gxq29p3OyYOk2f7w8HFA+Z2fXDu0hW2IxEFx7nOIAA4OztDLBYzloiICOnry5cvh1gsRkJCAubMmYMFCxZg586dFfZz7tw5iMViPHnyBMuWLUNgYCCCg4MZdSwtLaXHuH79OszMzODi4iLt7CUlJSEpKQlr165FfHw8QkNDcerUKUyYMKFuG4EFefn5sDBti8WzprEdhbPin7/Cwcu3Yd5SyHYUzjh5+ixEazdgygRvHA3fg862nTDJZxaSxMlsR1N41LZ1IzcvH+ZmJgiY/R3bUeQaT0U+F3kkp7E+jUAggFAoZCw6OjrS1zU0NCAUCmFkZISJEyfCxsYGZ86cqbAfPT09CIVCtGnTBqNHj4a9vT3u3r3LqKOqqio9RocOHbBs2TJkZ2fj8ePHAAArKyscPnwYQ4YMgYmJCfr27YvAwED88ssvKC4urtuGqGe97Lpg5iRvOPV2YDsKJ+XkF8Bv90EsG+sKzUYN2I7DGSFhEfBwHQJPt6EwaWuMhXNnQdjMABGHotiOpvCobetGb/tumDV5Ipwce7EdhXAEJzuDspJIJLh48SISEhKgpqZWbd07d+7g7t276NatW5V1CgoKEBoaCm1tbZibm1dZLzMzE5qamlBVpft3iOy+D/sFva3NYd/BlO0onFFYVIQHCYlwsGP+u+7RvRti4uJZSsUN1LaEKA5O9kZOnDiBJk2aMMrmzZuHxYsXS///okWLUFhYiKKiIjRo0AAzZsyosB97e3uoqKhI633zzTcYO3Yso058fLz0WLm5udDQ0EBkZCQ0NTUrzZaamooVK1bg22+/rfY9FBQUoKCggFGmXlAAgUBQ/ZsnnHTy1j08fCnGgUWT2Y7CKekZGSgpKYGeni6jXF9XF29TU1lKxQ3UtoRtdDex7DjZGXR0dMT27dsZZbq6/34gzZ07F97e3nj79i0WLlyIvn37wt7evsJ+IiMj0b59exQVFSE+Ph4zZsyAjo4OVq1aJa1jbm6O48ePAwDevXuHyMhIeHp64rfffsPnn3/O2F9WVhYGDx6MDh06YMmSJdW+B5FIhGXLljHKAuZ8h6VzZ8rUBoQ7xGkZEEX8D7t8vSH4wAg2+Tg8MP9oSCQS+kNSS6htCZF/nOwMNm7cGKamVU+l6evrw9TUFKampjh8+DBMTU1hZ2eHfv36Meq1atVKup/27dvj2bNnWLx4MZYuXYoGDcqu2Xp/p/F7tra2OHr0KDZs2IB9+/ZJy9+9ewdnZ2c0adIER44c+eC0tL+/P3x9fRll6pli2RqAcMqDP5OQ+i4Hniv+/YJTUlqKO0/+RPiFm4jdsRR8FaW+4uOj6Whrg8/nI6XcSFVqejr0dXWr2IrIgtqWEMXByc5gTejo6MDHxwdz5sxBTExMtd9Y+Xw+iouLUVhYKO0MVlUvLy9Pup6VlYUBAwZAIBDg+PHj1W77nkAgqDAlLMlPk+EdEa7p3t4Ex5b5MMoWhkTBWKiPiQN7UUfwE6irqcGyvTmu3byF/n37SMujb9zCF33o4vxPQW1L2EYD0LLjZGewoKAAycnMRxeoqqpCX1+/0vrTpk3D6tWrcfjwYQwbNkxanpqaiuTkZBQXFyM+Ph4bN26Eo6Mj43rA4uJi6bHeTxM/fPgQ8+bNk5Y5OTkhNzcX+/btQ1ZWFrKysgAATZs2BZ/Pr9X3zqac3Dy8/DtJuv5KnIyEJ39AS1MDzZsZsJhMsTVuIIBZi2aMsobqatBu0qhCOam5caNHwm/xMli1bw9bGytERh2DOPk1Rni4sR1N4VHb1o2c3Fy8fPW3dP1VUjISHj+BlqYmmgvpM4HUHCc7g6dOnYKhoSGjzNzcHI8ePaq0ftOmTTFmzBgsXboU7u7u0vL308Z8Ph+GhoYYNGgQAgMDGds+ePBAeqxGjRrBxMQE27dvl95o8vvvv+PmzZsAUGHq+vnz5zAyMvr4Nypn7ic+xtcz/KTrq7b8BABwde6PVQvnsBWLkGoNGtAf6ZmZ2LYrCG9SUtHOpC12bvoRLZobfnhjUi1q27pxPyERY6fNkq6LNm4FALgNGoBVAf5sxSIKjCf5709qELkmefOC7QicVJp4m+0InMX/rN+HKxEibwrz2U7AXTr190XgmWW7ejtWTbR98JjtCBXQxUaEEEIIIUqMOoOEEEIIIUqMk9cMEkIIIUS50d3EsqORQUIIIYQQJUadQUIIIYQQJUbTxIQQQgjhHBWaJ5YZjQwSQgghhCgxGhlUJPQtp26UlLCdgLvoMaZ1hz4P6g6PxkmIcqHOICGEEEI4h74vyY6+/hBCCCGEKDHqDBJCCCGEKDGaJiaEEEII5/BonlhmNDJICCGEEKLEqDNICCGEEKLEONcZ9Pb2Bo/Hq7A4OzsDAIyMjKRlDRs2hIWFBX744QdI/vMIjBcvXjC2VVdXh6mpKb7//ntGvaVLlzLqaWlpoWfPnrh06VKl2SQSCQYOHAgej4ejR4/WaTuw4XbsPUz2W4yeQ0fAwsEJ5y5fYzsSJ+389TI6fLsEoshf2Y7CGWEHDqPvEHdYd+8N99HeuBMTy3Ykzgg7cAh9XdxgbdcL7qO+xp27sWxHUni3Y2IxefZ8OAx2g3m3Xjh36QrbkeQSjyefizziXGcQAJydnSEWixlLRESE9PXly5dDLBYjISEBc+bMwYIFC7Bz584K+zl37hzEYjGePHmCZcuWITAwEMHBwYw6lpaW0mNcv34dZmZmcHFxQWZmZoX9bdiwgdPXMOTl5cPCtC0W+05nOwpnxb/4Gwev/A7zls3YjsIZJ8+cg2jdBkwZ742j4XvQ2bYjJvn4IkmczHY0hXfy9FmI1m7AlAnv27YTJvnMorb9RLl5+TA3M0HAnJlsRyEcwcnOoEAggFAoZCw6OjrS1zU0NCAUCmFkZISJEyfCxsYGZ86cqbAfPT09CIVCtGnTBqNHj4a9vT3u3r3LqKOqqio9RocOHbBs2TJkZ2fj8ePHjHpxcXH48ccfK3QmuaRX966Y+c04OPV2YDsKJ+XkF8Av6DCWjfkSmo0ash2HM0L2RcBj6BB4un0JE2MjLJwzC8JmBog4FMV2NIUXEhYBD9ch8HQbCpO2xlg4l9q2NvS2t8OsyZPg5Nib7SiEIzjZGZSVRCLBxYsXkZCQADU1tWrr3rlzB3fv3kW3bt2qrFNQUIDQ0FBoa2vD3NxcWp6bm4uRI0diy5YtEAqFtZafKJfvI/6H3tZmsG9vwnYUzigsKsKDR4lwsOvKKO9h1w0x9+JZSsUNhUVFeJCQCAc75mdmj+7dEBNHbUvqHtvTwYo0TczJR8ucOHECTZo0YZTNmzcPixcvlv7/RYsWobCwEEVFRWjQoAFmzJhRYT/29vZQUVGR1vvmm28wduxYRp34+HjpsXJzc6GhoYHIyEhoampK68yaNQv29vYYOnSozO+hoKAABQUFjDL1ggIIBAKZ90G44+TteDx8KcaBBd+wHYVT0jMyUFJSAj09XUa5vp4O3qamsZSKG6psW11dvE1NZSkVIaQynOwMOjo6Yvv27YwyXd1/P5Dmzp0Lb29vvH37FgsXLkTfvn1hb29fYT+RkZFo3749ioqKEB8fjxkzZkBHRwerVq2S1jE3N8fx48cBAO/evUNkZCQ8PT3x22+/4fPPP8fx48dx4cIFxMTE1Og9iEQiLFu2jFEWMOc7LPWbVaP9EMUnTsuEKPJX7PpuLAQfGMEmH6f8tbwSCSCnX+AVDg/l21bC6WunCVFEnOwMNm7cGKamplW+rq+vD1NTU5iamuLw4cMwNTWFnZ0d+vXrx6jXqlUr6X7at2+PZ8+eYfHixVi6dCkaNGgAANI7jd+ztbXF0aNHsWHDBuzbtw8XLlzAH3/8AW1tbca+PTw80LNnT1y8eLHSjP7+/vD19WWUqWfRRdfK6MHLJKS+y4Hnyp+kZSWlpbjz5E+EX7yF2K2LwVdR6is+PpqOtjb4fD5SUpgjValp6dAvN6JFakbatuVGAVPT06GvS21L6h5Phb50yIqTncGa0NHRgY+PD+bMmYOYmJhqv7Hy+XwUFxejsLBQ2hmsql5eXh4AYP78+Zg4cSLjdWtra6xfvx5Dhgypch8CgaDClLCkIF2Wt0Q4prtFWxwLmMooW7jnKIyF+pg4wIE6gp9AXU0NlhbmuHbzNvr37SMtj755C1/07sleMA5QV1ODZXtzXLt5i9m2N27hiz692AtGCKmAk53BgoICJCczR9FUVVWhr69faf1p06Zh9erVOHz4MIYNGyYtT01NRXJyMoqLixEfH4+NGzfC0dGRcT1gcXGx9Fjvp4kfPnyIefPmAYD0TuPyWrduDWNj409+r/IkJzcPL/9Okq6/Eicj4ckf0NLQQHOhAYvJFFvjBgKYtWA+SqahQB3ajRtVKCc1N85rJPwWL4NVBwvY2lgjMuooxMmvMWKYG9vRFN640f+0bfv2sLWxQmTUsbK29aC2/RQ5ubl4+epv6fqrJDESHj+BlqYmmgvpM4HUHCc7g6dOnYKhoSGjzNzcHI8ePaq0ftOmTTFmzBgsXboU7u7u0vL308Z8Ph+GhoYYNGgQAgMDGds+ePBAeqxGjRrBxMQE27dvr3CjiTK4/+gxvp4xV7q+anPZtKbrwP5YtXBuVZsRwqpBTv2QnpGJbbuC8SYlFe1M2mLnpnVoUe4zhNTcoAH9kZ6ZiW27gv7Ttj+iRXNq209xPyERY6d+J10XbdgCAHAb7IxVAQvYiiV36NJU2fEk//1JDSLXJG//ZDsCJ5U+uM52BM7if+7EdgTuor90daeokO0E3KVdfyOX4s/b19uxasLwTgLbESqgi40IIYQQQpQYJ6eJCSGEEKLcVGj0XGY0MkgIIYQQosSoM0gIIYQQosRompgQQgghnEOzxLKjkUFCCCGEECVGnUFCCCGEECVG08QKpPTuBbYjcNKREfPZjsBZw57Fsh2Bu9Sr/klM8okkpWwnILWgup+XJUw0MkgIIYQQosSoM0gIIYQQosRompgQQgghnEOzxLKjkUFCCCGEECVGnUFCCCGEECVG08SEEEII4Ry6m1h2NDL4D29vb/B4vArL06dPAQArV64En8/HqlWrpNsYGRlVus37pU+fPgCAnTt3ok+fPtDU1ASPx0NGRgYL77B+7TxzHR1mrILo8Dm2oygkfXs72Ef8jMEP4zAs/TWaDxpYZd3P1v+AYemvYTr5m3pMyB23Y+IwebY/HFw8YG7XB+cuXWE7EqeEHTiEvi5usLbrBfdRX+PO3Vi2Iyk8OmdJbaPO4H84OztDLBYzFmNjYwBASEgI/Pz8EBwcLK1/+/Ztab3Dhw8DABITE6VlUVFRAIDc3Fw4OztjwYIF9f+mWBD/pxgHo2Nh3rwp21EUlmqjRsi8/wAxfv7V1ms+aCB0O3+GvCRxPSXjnty8fJibmSBg9ndsR+Gck6fPQrR2A6ZM8MbR8D3obNsJk3xmIUmczHY0hUbnLKltNE38HwKBAEKhsEL5pUuXkJeXh+XLl2Pv3r24fPkyevXqhaZN/+3s6OrqAgAMDAygra3N2H7mzJkAgIsXL9ZVdLmRU1AIv73HsWzkQPx0+hrbcRRW8rkLSD5X/UPGGxgK0WnNSlwdNgI9IvfVUzLu6W3fDb3tu7Edg5NCwiLg4ToEnm5DAQAL587C1es3EHEoCrN9prKcTnHROSsbmiWWHY0MyiAoKAgjR46EmpoaRo4ciaCgILYjya3vD55Bb0sT2JsbsR2F23g8dN2xFY83b0PWo0S20xBSQWFRER4kJMLBjtlp6dG9G2Li4llKRQipDHUG/+PEiRNo0qSJdPH09ERWVhYOHz4MLy8vAICXlxcOHTqErKysOs1SUFCArKwsxlJQWFSnx/xUJ39/iId/vcasIX3YjsJ55jN9ICkuxtOfdrEdhZBKpWdkoKSkBHp6uoxyfV1dvE1NZSkVIaQy1Bn8D0dHR8TGxkqXTZs2ITw8HG3btkXHjh0BAJ06dULbtm2xf//+Os0iEomgpaXFWFZF/q9Oj/kpxOlZEEWdw+qxLhCo0dUHdUm7ow3Mvp2E29NmsB2FkA/igTlXJ5FI6C5PUi+qu8GTzUUe0V/t/2jcuDFMTU0ZZcHBwXjw4AFUVf9tqtLSUgQFBeGbb+ru7k1/f3/4+voyylQv1W0H9FM8+CsZqe9y4flDqLSspFSCO3/8hfArvyP2x7ngq9B3j9qg390Ogqb6GBR/V1qmoqqKjt8vhdmUSfi1YxcW0xFSRkdbG3w+HynlRgFT09Ohr6tbxVaEEDZQZ7Aa8fHxuHPnDi5evCi9QQQAMjIy0KtXL9y/fx9WVlZ1cmyBQACBQMAoK1FXq5Nj1Ybu7drg2PwJjLKF4f+DsYEeJvazo45gLXoZeRBvLl1mlPU8tB9/HjiEF2ERLKUihEldTQ2W7c1x7eYt9O/bR1oefeMWvujTi71ghJAKqDNYjaCgIHTt2hW9elX84OrevTuCgoKwfv36D+4nOTkZycnJ0mcWxsfHQ0NDA61bt2Z0MhVZ4wYCmJV7lExDdTVoN25YoZx8GL9xIzT557FGANC4TWtoWVmiMCMDea/+RmF6OqN+aXER8l+/QfbTP+o7qsLLyc3Fy1d/S9dfJSUj4fETaGlqormwGYvJFN+40SPht3gZrNq3h62NFSKjjkGc/BojPNzYjqbQ6JyVDY/GIGRGncEqFBYWYt++fZg3b16lr3t4eEAkEmH16tVQV1evdl87duzAsmXLpOvvO5chISHw9vautcyEO3Q7dULvE0ek6x1XLgcAvAjfjzvT6Nlitel+QiLGTpslXRdt3AoAcBs0AKsCqn/OI6neoAH9kZ6ZiW27gvAmJRXtTNpi56Yf0aK5IdvRFBqds6S28SQSiYTtEEQ2JadD2I7ASUdGzGc7AmcNexbLdgTuUm/AdgLuKsxnOwF36dTfF4GMntb1dqya0L4if49WopFBQgghhHCOvN65K49oRp0QQgghRIlRZ5AQQgghRInRNDEhhBBCuEeFpollRSODhBBCCCFKjDqDhBBCCCFKjKaJCSGEEMI9dDexzKgzqED4doPYjsBJHve6sR2Bu9Qbsp2Au+gPXd2h85YoGZomJoQQQghRYjQySAghhBDOoYdOy45GBgkhhBBClBh1BgkhhBBClBhNExNCCCGEe+ih0zLj7Migt7c3XF1dGWWHDh1CgwYNsGbNGixduhQ8Hq/CYmFhUWFf4eHh4PP5mDx5coXXLl68yNi+adOmGDhwIOLi4qR1oqKiMGDAAOjr64PH4yE2Nra2365cuH03FpN958NhkBvMu/bCuYtX2I7ECT+FH8awqXPx2ZCRsB/2NaYFiPDsr7/ZjsUpYQcOoa+LK6ztesJ91FjcuRvDdiTOCDtwCH0HD4V1Nwdq21pG5y2pLZztDJa3e/dujB49Glu2bIGfnx8AwNLSEmKxmLFcvXq1wrbBwcHw8/PD/v37kZubW+n+ExMTIRaL8b///Q/p6elwdnZGZmYmACAnJwc9evTAqlWr6u4NyoHc/HyYm5kgYO5MtqNwyu17DzBq6EBEbl6N4NVLUVxSgonzliE3L5/taJxw8vRZiNaux5QJ43A0fC8623bCJJ9ZSBInsx1N4Z08fRaiH34sa9uIn8vadvpMattaQOctqU1K0Rlcs2YNpk+fjvDwcEycOFFarqqqCqFQyFj09fUZ27548QLR0dGYP38+LCwscOjQoUqPYWBgAKFQiK5du2LdunVITk7GjRs3AABjxoxBQEAA+vXrV3dvUg70trfDrCmT4OTYm+0onLJ7VQDcB/SFmVFrWJgYQzTXB0lv3uLBkz/YjsYJIWER8HD9Ep5uQ2HS1hgL5/pC2KwZIg4dZjuawgvZF17Wtu6u/7atsBkiDlLbfio6b2XA48nnUkOXL1/GkCFD0Lx5c/B4PBw9epTxukQiwdKlS9G8eXM0bNgQffr0wYMHD2p0DM53BufPn48VK1bgxIkT8PDwqPH2wcHBGDx4MLS0tODl5YWgoKAPbtOwYdkDS4uKimp8PEI+5F1O2ei0lkYTlpMovsKiIjxIeAQHO+aDx3t074qYuHiWUnGDtG27l2tbu26IibvHUipuoPNWueTk5KBjx47YsmVLpa+vWbMGP/74I7Zs2YLbt29DKBSif//+ePfunczH4HRn8Ndff8Xq1atx7NixSkfl4uPj0aRJE8by35HD0tJShIaGwsvLCwAwYsQIXL9+HU+fPq3ymKmpqVi2bBk0NDTQtWvX2n9TRKlJJBKs2hGCzlbt0c64DdtxFF56RgZKSkqgp6fLKNfX1cPb1FSWUnFDevo/baurxyjX19Oltv1EdN4ql4EDB+L777+Hu7t7hdckEgk2bNiAhQsXwt3dHVZWVtizZw9yc3MRHh4u8zE4fTexjY0NUlJSEBAQgC5dukBDQ4Pxurm5OY4fP84o+2+dM2fOICcnBwMHDgQA6Ovrw8nJCcHBwVi5ciVju5YtWwIo68GbmZnh4MGDMDAw+OjsBQUFKCgoYJQJCgogEAg+ep9E8a3YvBOJz14gfMPKD1cmMuOBOXUjkUjogbW1pHwzUtvWHjpvq8eT07uJK/37LhB81N/358+fIzk5GU5OTox99e7dG9HR0fj2229l2g+nRwZbtGiBS5cuQSwWw9nZucKQqbq6OkxNTRlLs2bNpK8HBwcjLS0NjRo1gqqqKlRVVXHy5Ens2bMHJSUljH1duXIFcXFxyMzMxOPHjzFgwIBPyi4SiaClpcVYRD9u+qR9EsW2YvMuXLh+G3vXroCwqf6HNyAfpKOtDT6fj5Ryoymp6WnQ19WtYisiCx2dKto2LZ3a9hPReavYKv37LhJ91L6Sk8tuGPpv3+X9+vvXZMHpziAAtG7dGpcuXcKbN2/g5OSErKwsmbZLTU3FsWPHsH//fsTGxjKW7Oxs/Prrr4z6xsbGMDExgaamZq3k9vf3R2ZmJmPx951RK/smikUikWD55p04e/UGQn9YjpaGzT68EZGJupoaLNtb4NrNW4zy6Bu3YNvRmqVU3CBt2xuVta0NS6m4gc5bxVbp33d//0/aZ/kR4ZqOEnN6mvi9li1b4uLFi3B0dISTkxNOnz4NACguLq7Qc+bxeGjWrBl+/vln6OnpwdPTEyoqzD6zi4sLgoKC4OLiItPx09LS8PLlSyQlJQEoewwNAOkdzJWpdMhYkifT8diSk5uLl6/+ff7dqyQxEh4/gZamJpoLqQPzsZZv2okTFy5j63J/NG7UEG/T0gEAGo0boQFdNvDJxo0eCb/FS2HV3gK2NtaIjDoKcfJrjPCoeH0OqZlxXqPgt2gJrDq0/6dtj0CcnIwRw6htPxWdtzKQ0ynzj50Srsz7PkRycjIMDQ2l5W/evKkwWlgdpegMAv9OGTs6OqJ///6wt7fHgwcPGI0HlP1Hys/PR3BwMNzc3Cp0BAHAw8MDw4cPx+vXr2U69vHjxzFu3Djp+ogRIwAAS5YswdKlSz/+TcmZ+wmJGDvlO+m6aEPZnU9ug52xaskCtmIpvIhfTgEAxs5ezChfOdcH7gP6shGJUwYN6I/0zExs2xWMNykpaGfSFjs3rUeL5oYf3phUS9q2O4PK2tbUBDs3U9vWBjpvCVA2KykUCnH27FnY2toCAAoLC3Hp0iWsXr1a5v3wJBKJpK5CklqWKVvnk9SMJIvuvqsrPN3mbEfgLjkd9eAE+rNYdxpr19uh3jl3qbdj1YTGqds1qp+dnS19iomtrS1+/PFHODo6QldXF61bt8bq1ashEokQEhICMzMzrFy5EhcvXkRiYmKFG2erojQjg4QQQghRHvJ6N3FN3blzB46OjtJ1X19fAMDXX3+N0NBQ+Pn5IS8vD1OnTkV6ejq6deuGM2fOyNwRBGhkULHQyGCdoJHBukMjg3WIRgbrDv1ZrDv1ODKYPUg+n/Xb5OStD1eqZ5y/m5gQQgghhFSNpokJIYQQwj00ei4zGhkkhBBCCFFi1BkkhBBCCFFiNE1MCCGEEO7hyN3E9YE6gwqk5PB2tiNw0rQJP7IdgbN2pCSwHYG71OnXZ+qK5F0a2xE4i1ePdxMT2dE0MSGEEEKIEqORQUIIIYRwDo/uJpYZjQwSQgghhCgx6gwSQgghhCgxmiYmhBBCCPfQ3cQyo5FBQgghhBAlxonOoLe3N1xdXRllhw4dQoMGDbBmzRosXboUPB4Pzs7OFbZds2YNeDwe+vTpAwAwMjICj8ercunTpw/S0tLg4+MDc3NzNGrUCK1bt8aMGTOQmZnJ2Pd/t2vcuDHMzMzg7e2N33//va6agnWv3+XC75dodN94GJ+tOwC3kF/xIJke01BTpj3tMfV4JFb9nYgdkix0HDqY8foOSValS/85M1hKrLhux8Rh8twFcPjSE+b2fXHu0lW2I3FK2MEo9P1yGKztHeHuNR53YmLZjqTwfgo7iGGTZ+GzQV/B3s0L0xZ9j2cvX7EdiygwTnQGy9u9ezdGjx6NLVu2wM/PDwBgaGiI3377Da9eMf/BhISEoHXr1tL127dvQywWQywW4/DhwwCAxMREaVlUVBSSkpKQlJSEtWvXIj4+HqGhoTh16hQmTJhQIUtISAjEYjEePHiArVu3Ijs7G926dcPevXvrsAXYkZlfiNH7zkFVRQU/efbBLxMHwc/RFhoCNbajKRxB48Z4FXcf+6fPqfR1P6EpY9kzbgpKS0sRc/h4PSdVfLn5+TA3NUGArw/bUTjn5JlzEK3biCnjx+JoWAg629pg0ow5SEpOZjuaQrsddx+jXAcjcusPCP5hBYpLSjDRLwC5eflsR5MvPJ58LnKIc9cMrlmzBgEBAQgPD4eHh4e03MDAAJ07d8aePXuwcOFCAEB0dDRSUlLg6emJhw8fAgCaNm0q3UZXV1e6rba2NqP8fUcRAExMTBAYGAgvLy8UFxdDVfXfZtXW1oZQKARQNuro5OSEr7/+GtOnT8eQIUOgo6NT+43AkqAbDyHUbISVg+2kZS20mrCYSHE9OHUWD06drfL1rNdvGOsdhw7G498uI+X5izpOxj29u3dD7+7d2I7BSSFhkfAY6gJP1y8BAAtnz8TV67cQcegIZk+fwnI6xbV7zTLGumjeTNi7eeHB46fo0tGKpVREkXFqZHD+/PlYsWIFTpw4wegIvjd+/HiEhoZK14ODgzF69Gioq6t/8rEzMzOhqanJ6AhWZdasWXj37h3Onq36j70iuvD0b1gJdTHz6FU4bI6Ce8ivOBj7lO1YnKdh0BTWgwfgWtDPbEchRKqwqAgPHiXCwa4ro7yHXVfE3LvPUipuepeTAwDQ0tRgOQlRVJzpDP76669YvXo1jh07hn79+lVax8XFBVlZWbh8+TJycnJw4MABjB8//pOPnZqaihUrVuDbb7+Vqb6FhQUA4MWLF598bHnyKiMb+2OeoI2OBnZ+1QfDbc2w8vxdHLv/nO1onNb961HIf5eNmCiaIibyIz0jAyUlJdD7Z4blPX1dHbxNSWUpFfdIJBKs2haEztYd0M64Ddtx5ApPRT4XecSZaWIbGxukpKQgICAAXbp0gYZGxW9Iampq8PLyQkhICJ49e4Z27drBxsbmk46blZWFwYMHo0OHDliyZIlM20gkEgDVPx29oKAABQUFjDLVomII1OT3P1mpBLAS6mJW744AgA7NdPE0JRP7Y55gqJUxy+m4y378GNwKO4DicucLIfKg/OecREK/DFGbVmzcgcQ/XiB882q2oxAFJqd91Jpr0aIFLl26BLFYDGdnZ7x7967SeuPHj8fBgwexdevWTx4VfPfuHZydndGkSRMcOXIEamqy3SiRkJAAADA2rrqDJBKJoKWlxVhWnZTvuxybNmkAE31NRpmJnibEWbksJeI+U4fuEFq0w9Xde9iOQgiDjrY2+Hw+UlKZo4Cp6enQ19OtYitSEys2/YQL0bewd30ghE312Y5DFBhnOoMA0Lp1a1y6dAlv3ryBk5MTsrKyKtSxtLSEpaUl7t+/j1GjRn30sbKysuDk5AR1dXUcP34cDRo0kHnbDRs2QFNTs8rpbADw9/dHZmYmY5k/yOGj89aHz1o0xfM0Zif8Rdo7NNdszFIi7usxYSz+vHMXf9M1WETOqKupwdLCHNdu3maUR9+8DVsbusnhU0gkEizfuANnr0Qj9MdAtDQUsh1JPrF91zDdTcyeli1b4uLFi3B0dISTkxNOnz5doc6FCxdQVFTEuEO4Jt69ewcnJyfk5uZi3759yMrKknY8mzZtCj6fL62bkZGB5ORkFBQU4PHjx/jpp59w9OhR7N27t9rjCwQCCAQCRlmJHE8RA8DYLuYYve8sfrr+AM4WrREvTsXBuKdYOqDrhzcmDILGjdHUtK10Xd/YCC07WiMnLR3pf5U9HqmBhgY+83TFodkL2YrJCTm5eXj56m/p+iuxGAmPn0JLUwPNhc1YTKb4xo0eDr+AFbBqbwFbGytERh2DOPk1Rni4sR1NoS3fsB0nzl/G1u8XonGjhniblg4A0GjcCA3K/d0gRBby3bv4SO+njB0dHdG/f3/Y29szXm/c+NNGqn7//XfcvHkTAGBqasp47fnz5zAyMpKujxs3DgDQoEEDtGjRAg4ODrh16xY+++yzT8ogj6wN9bDJrSfWX4rD9mv30VKrCeb3/QxDLI3YjqZw2nxuC9+LJ6XrnutFAIDroWHYM67skRyfj/AAj8fD7YhDrGTkivuPEjF2uq90XbRpOwDAbdAArFo0j61YnDDIqR/SM7OwbXcI3qSkop1JW+zcuBYtaCTrk0Qc/xUAMHbWAkb5ynnfwd256hknQqrCk7y/m4HIvZLgpWxH4KRpE35kOwJn7UhJYDsCd6nTCFBdkbyjX02qK7zm7ertWHlf9ay3Y9VEwwNX2I5QAaeuGSSEEEIIITVDnUFCCCGEECXGyWsGCSGEEKLk5PTOXXlEI4OEEEIIIUqMOoOEEEIIIUqMpokJIYQQwj0qNE0sKxoZJIQQQghRYvScQUWSk852Am4qLmQ7AXepyf4zjYTIDfqzWHcaa9fbofJG9q63Y9VEw4hLbEeogKaJCSGEEMI5PLqbWGY0TUwIIYQQosSoM0gIIYQQosRompgQQggh3EN3E8uMRgYJIYQQQpQYdQYJIYQQQpSY0nYGvb294erqyig7dOgQGjRogDVr1mDp0qXg8Xjg8XhQUVFB8+bNMXr0aPz111+V7s/c3Bzq6ur4+++/K7z27NkzjBw5Es2bN0eDBg3QsmVLDB06FI8fP66Lt8aqsAOH0NfFDdZ2veA+6mvcuRvLdiROuH03FpN958NhkBvMu/bCuYtX2I7EKWEHDqHv4KGw7uYA91FjceduDNuROIPatu6Ufd66wtquJ7VtZXg8+VzkkNJ2BsvbvXs3Ro8ejS1btsDPzw8AYGlpCbFYjFevXiEyMhLx8fH46quvKmx79epV5Ofnw9PTE6GhoYzXCgsL0b9/f2RlZSEqKgqJiYmIjIyElZUVMjMz6+Ot1ZuTp89CtHYDpkzwxtHwPehs2wmTfGYhSZzMdjSFl5ufD3MzEwTMncl2FM45efosRD/8iCkTxuFoxM9l5+30mXTe1gJq27pT9nm7vqxtw/fS5y35JNQZBLBmzRpMnz4d4eHhmDhxorRcVVUVQqEQzZs3R8+ePTFp0iTcuHEDWVlZjO2DgoIwatQojBkzBsHBwfjvc7wfPnyIZ8+eYdu2bbCzs0ObNm3Qo0cPBAYGokuXLvX2HutDSFgEPFyHwNNtKEzaGmPh3FkQNjNAxKEotqMpvN72dpg1ZRKcHOXzIaqKLGRfODxcv4Snu+s/560vhMJmiDh4mO1oCo/atu6Ufd5++Z/PW18ImzVDxCFqW1JzSt8ZnD9/PlasWIETJ07Aw8OjynrJycmIiooCn88Hn8+Xlr979w4HDx6El5cX+vfvj5ycHFy8eFH6etOmTaGiooJDhw6hpKSkLt8KqwqLivAgIREOdt0Y5T26d0NMXDxLqQipXtl5+wgO3cudt3bdEBN3j6VU3EBtW3ekbVvh87Yrfd7+x/tLveRtkUdK3Rn89ddfsXr1ahw7dgz9+vWr8Hp8fDyaNGmCRo0awdDQEBcvXsS0adPQuHFjaZ39+/fDzMwMlpaW4PP5GDFiBIKCgqSvt2jRAps2bUJAQAB0dHTQt29frFixAs+ePauX91hf0jMyUFJSAj09XUa5vq4u3qamspSKkOqlp/9z3urqMcr19ei8/VTUtnWn6s9bPWpb8lGUujNoY2MDIyMjBAQE4N27dxVeNzc3R2xsLG7fvo3AwEB06tQJgYGBjDpBQUHw8vKSrnt5eSEqKgoZGRnSsmnTpiE5ORn79u1D9+7dcfDgQVhaWuLs2bNVZisoKEBWVhZjKSgo+PQ3Xcd4YH7rkUgkcvtNiJD3yp+idN7WHmrbukOft6S2KHVnsEWLFrh06RLEYjGcnZ0rdAjV1dVhamoKS0tLLFiwAJ06dcKUKVOkrz98+BA3b96En58fVFVVoaqqCjs7O+Tl5SEiIoKxLw0NDXz55ZcIDAxEXFwcevbsie+//77KbCKRCFpaWoxFtHZ97TZALdLR1gafz0dKuW+lqenp0NfVrWIrQtilo1PFeZtG5+2noratO1V/3qZR2/6XCk8+Fzmk1J1BAGjdujUuXbqEN2/ewMnJqcLNIf+1ePFiRERE4O7duwDKRgV79eqFuLg4xMbGShc/Pz/GVHF5PB4PFhYWyMnJqbKOv78/MjMzGYv/nFkf/0brmLqaGizbm+PazVuM8ugbt2Db0ZqlVIRUr+y8tcC1G5WdtzYspeIGatu6I21b+rwltYR+jg5Ay5YtcfHiRTg6OsLJyQmnT5+utF7btm0xdOhQBAQE4MiRI/j555+xfPlyWFlZMepNnDgRa9asQVxcHCQSCZYsWYIxY8agQ4cOUFdXx6VLlxAcHIx58+ZVmUkgEEAgEDALc+T7BpRxo0fCb/EyWLVvD1sbK0RGHYM4+TVGeLixHU3h5eTm4uWrf59h+SpJjITHT6ClqYnmwmYsJlN847xGwW/RElh1aA9bG2tERh2BODkZI4a5sx1N4VHb1p2yz9ulsGpv8U/bHv3n85baltQcdQb/8X7K2NHREf3794e9vX2l9WbPno0ePXrgxx9/RGpqKtzcKnZ0zMzMYG1tjaCgIAQEBMDIyAjLli3DixcvwOPxpOuzZsnvSN/HGDSgP9IzM7FtVxDepKSinUlb7Nz0I1o0N2Q7msK7n5CIsVO+k66LNmwBALgNdsaqJQvYisUJ0vN2ZxDepKSgnakJdm5eT+dtLaC2rTv/ft4Gl7WtSVvs3ERt+190/aTseJL/PhSPyLecdLYTcFNxIdsJuEutAdsJCKk5+rNYdxpr19uhCsf1r7dj1YR6SNU3j7JF6a8ZJIQQQghRZjRNTAghhBDukdM7d+URjQwSQgghhCixGnUGp06diuzsbOn6zz//zFjPyMjAoEGDai8dIYQQQgipUzW6gYTP50MsFsPAwAAAoKmpidjYWLRt2xYA8Pr1azRv3pzTv8HLKrqBpG7QDSR1h24gIYqIbiCpO/V5A8nEAfV2rJpQ31354+vYVKORwfL9RroRmRBCCCFEsdE1g4QQQgghSozuJlYodGdU3aB2JYQQruHR3cQyq3FnMCAgAI0aNQIAFBYWIjAwEFpaWgCA3Nzc2k1HCCGEEELqVI06g7169UJiYqJ03d7eHs+ePatQhxBCCCGEKIYadQYvXrxYRzEIIYQQQmoR/TaxzOgGEkIIIYQQJVajkcHly5fLVC8gIOCjwhBCCCGEkPpVo4dO29raVr0jHg+JiYnIz8+nh07XlZwMthNwEz10uu6oCdhOQEjN0TN06049PnS6aIp8/iKa2vaTbEeooEbTxDExMZUuISEhMDAwQFFRESZNmlRXWeuct7c3eDweJk+eXOG1qVOngsfjwdvbm1EeHR0NPp8PZ2fnCtucPHkS6urquHv3LqN87dq10NfXR3Jycq3mlwdhBw6hr4srrO16wn3UWNy5G8N2JE64fTcWk33nwWGQK8y79sS5i5fZjsQpYQcOoe/gobDu5kDnbS2jtq079HlLassnXTP4/PlzeHl5oUuXLtDS0sKDBw+wY8eO2srGilatWmH//v3Iy8uTluXn5yMiIgKtW7euUD84OBg+Pj64evUqXr58yXht0KBBGDt2LMaOHYuCggIAQEJCAhYvXoytW7dCKBTW7ZupZydPn4Vo7XpMmTAOR8P3orNtJ0zymYUkMfc6vfUtNz8f5mamCJg7i+0onHPy9FmIfvix7LyN+LnsvJ0+k87bWkBtW3fo85bUpo/qDKakpMDHxwcWFhYQi8WIjo5GZGQkzMzMajtfvfvss8/QunVrREVFScuioqLQqlWrCtPkOTk5OHDgAKZMmQIXFxeEhoZW2N/69euRnZ2NJUuWoLi4GGPHjsWQIUMwfPjwun4r9S4kLAIerl/C020oTNoaY+FcXwibNUPEocNsR1N4ve3tMGvKJDg59mY7CueE7AsvO2/dXf89b4XNEHGQzttPRW1bd+jz9sN4PJ5cLvKoRp3BnJwcLFu2DCYmJoiOjsYvv/yC8+fPo0uXLnWVjxXjxo1DSEiIdD04OBjjx4+vUC8yMhLm5uYwNzeHl5cXQkJCKvxes4aGBoKDg7Fu3TqMHj0af/31F7Zt21bn76G+FRYV4UHCIzjYdWOU9+jeFTFx8SylIqR60vO2e7nz1q4bYuLusZSKG6ht6w593pLaVqO7iU1MTPDu3Tv4+Phg5MiR4PF4uHev4j9qGxubWgvIhjFjxsDf3x8vXrwAj8fDtWvXsH///grPWQwKCoKXlxcAwNnZGdnZ2Th//jz69evHqNe3b18MGzYM+/fvR2RkJPT19T+YoaCgQDq1/J6guAACgXxekJ+ekYGSkhLo6ekyyvV19fA29QZLqQipXnr6P+etrh6jXF9PF29TU1lKxQ3UtnWHPm9JbatRZ/DNmzcAgDVr1uCHH35gjILxeDxIJBLweDyFv5tYX18fgwcPxp49eyCRSDB48OAKHbjExETcunVLOp2sqqqK4cOHIzg4uEJnMCkpCadOnUKjRo1w5coVfPXVVx/MIBKJsGzZMkbZEv95WLpw/ie+u7rFK/c7v+/PCULkWflTlM7b2kNtW3fo8/YD6LeJZVajzuDz58/rKofcGT9+PKZPnw4A2Lp1a4XXg4KCUFxcjBYtWkjLJBIJ1NTUkJ6eDh0dHWn5xIkT0bFjRyxbtgxffPEFhg0bht69q7/2y9/fH76+vowyQXFeFbXZp6OtDT6fj5Ry3/hT09Ogr6tbxVaEsEtHp4rzNi2dzttPRG1bd+jzltS2Gl0z2KZNG5kWLnB2dkZhYSEKCwsxYMAAxmvFxcXYu3cv1q1bh9jYWOkSFxeHNm3aICwsTFp39+7duHLlCkJCQtC7d29Mnz4d48ePR05OTrXHFwgE0NTUZCzyOkUMAOpqarBsb4FrN28xyqNv3IJtR2uWUhFSPel5e6Oy81axL3dhG7Vt3aHPW1LbatQZXLNmDeORK5cvX2Zc1/bu3TtMnTq19tKxiM/nIyEhAQkJCeDz+YzXTpw4gfT0dEyYMAFWVlaMZdiwYQgKCgIAvHz5ErNnz8batWthbGwMAFi5ciVUVFQwf758T/d+jHGjR+LQkWM4dPQ4/nj2HCvXroc4+TVGeLizHU3h5eTmIuHxEyQ8fgIAeJUkRsLjJ0hKfs1yMsU3zmtUufP2R4iTkzFiGJ23n4ratu7Q560MeDz5XORQjX6BhM/nQywWw8DAAACgqamJ2NhYtG3bFgDw+vVrNG/eXGGvGfT29kZGRgaOHj1a6euurq7Q1tZGamoqSktL8b///a9Cnbt376Jz5864c+cO5s2bBz6fj9OnTzPqXL16FX369MH58+c/OF3MoAC/QBJ24BCC9uzDm5QUtDNpC//Zs9Clc9W/XCMXFOAXSG7+HoOxU2ZUKHcb7IxVSxaykEhGCvILJGEHDiEo9Oey89bUBP6zZ6JL58/YjsUJCtm2CvILJAr5eVuPv0BS7DOk3o5VE6qbf2E7QgU16gyqqKggOTlZ2hnU0NBAXFwcZzqDck8BOoMKSQE6gwpLQTqDhDAoSGdQIVFnUC47gzW6gYQQQgghRCHI6ZSsPPqkn6MjhBBCCCGKrcYjg7t370aTJk0AlN1VGxoaKn0G37t372o3HSGEEEIIqVM1umbQyMhIpgdaKtPzCOsVXTNYN+iawbpD1wwSRUTXDNad+rxm8Luh9XasmlDdeIztCBXUaGTwxYsXdRSDEEIIIYSwoUadwfz8fJw7dw4uLi4Ayn4l47/PGVRVVcXy5cvRoEGD2k1JCCGEEELqRI06g3v27MGJEyekncEtW7bA0tISDRs2BAA8evQIQqGwws+oEUIIIYQQ+VSjzmBYWBhmzZrFKAsPD5c+Z3Dfvn3YunUrdQbriCT1b7YjcFJp2Ca2I3AWf9r3bEfgLh49DKKuSPKz2Y7AWbx6vGYQKvRvRFY1aqnHjx+jXbt20vUGDRpA5T+N3bVrVzx8+LD20hFCCCGEKKni4mIsWrQIxsbGaNiwIdq2bYvly5ejtLS0Vo9To5HBzMxMqKr+u8nbt28Zr5eWljKuISSEEEIIIR9n9erV2LFjB/bs2QNLS0vcuXMH48aNg5aWFr777rtaO06NOoMtW7bE/fv3YW5uXunr9+7dQ8uWLWslGCGEEELIR+PAL5Bcv34dQ4cOxeDBgwGUPeIvIiICd+7cqdXj1GiaeNCgQQgICEB+fn6F1/Ly8rBs2TJpYEIIIYQQwlRQUICsrCzGUtWsqoODA86fP4/Hjx8DAOLi4nD16lUMGjSoVjPVqDO4YMECpKWlwdzcHD/88AOOHTuG48ePY82aNTA3N0d6ejoWLFhQqwEJIYQQQrhCJBJBS0uLsYhEokrrzps3DyNHjoSFhQXU1NRga2uLmTNnYuTIkbWaqUbTxM2aNUN0dDSmTJmC+fPn4/2Pl/B4PPTv3x/btm1Ds2bNajUgIYQQQkiNyek0sb+/f4WnrggElf9aU2RkJPbt24fw8HBYWloiNjYWM2fORPPmzfH111/XWqYa/zaxsbExTp06hbS0NDx9+hQAYGpqCl1d3VoLJS+8vb2xZ88eiEQizJ8/X1p+9OhRuLm5QSKR4OLFi3B0dKyw7cKFC/H9998jPz8fkydPxu+//46EhAS4uLjg6NGj9fgu6s9PEYdx9uoNPPvrbzQQqMO2gwVmTxyDtq1asB1N4alMCQRPW79CeenvFyE5E8FCIm4JOxiFoH0ReJuSCrO2Rljg+x0+t+3IdixOCDt4GEE/h//TtsZYMPs7fG7bie1YCu127D0EhR/Eg8QneJuahi0rl6Bfrx5sxyIyEggEVXb+yps7dy7mz5+PESNGAACsra3x559/QiQSsdsZfE9XVxddu3attSDyqkGDBli9ejW+/fZb6OjoVFkvMTERmpqa0vUmTZoAAEpKStCwYUPMmDEDhw8frvO8bLp97wFGfTkQ1uamKCkpwfqQcEycvwwndm9Co4b0qzSfojRUxHxmVtPm4I+cBcmj39kLxREnz5yH6MdNWDJvNj7raI39Uccw6bs5+N+Bn9FcKGQ7nkI7eeYcROs2Ysn8Ofisow32Rx3FpBmz8b+DYdS2nyAvLx8Wpm3hPngAZixcznYcUodyc3MZj/ADAD6fX+uPlqEnMn5Av379IBQKq5zPf8/AwABCoVC6vO8MNm7cGNu3b8ekSZMg5PiH325RANwH9IWZUWtYmBhDNGc6kt6k4MGTP9iOpvjysoGcLOnCM7WBJP0N8PIx28kUXkj4fngMdYGn6xCYGBth4ezvIGxmgIhDR9mOpvBCwvbDY+gQeLp++U/bzvynbY+wHU2h9ereFTO/GQen3g5sR5FvPJ58LjUwZMgQBAYG4n//+x9evHiBI0eO4Mcff4Sbm1utNhV1Bj+Az+dj5cqV2Lx5M169esV2HIXyLicXAKCl0YTlJByjwgfPshskcdFsJ1F4hUVFePDoMRy6dWGU9+jWBTH37rOUihvK2jYRDnbMGaQedl0Rcy+epVSEKJbNmzdj2LBhmDp1Ktq3b485c+bg22+/xYoVK2r1OB89TaxM3Nzc0KlTJyxZsgRBQUGV1in/fMU///wTenp6H33MgoKCCreaqxcUQiBQ/+h91ieJRIJVO0LQ2ao92hm3YTsOp/DadQIaNIQknjqDnyo9IxMlJSXQK3fNs76eLt6mprKUihvSMzIqb1tdXbxNSWMpFSGKRUNDAxs2bMCGDRvq9Dg0Miij1atXY8+ePVX+3N6VK1cQGxsrXaq7vlAWld56vm3XJ+2zPq3YvAuJz//EugWzPlyZ1AivYw/gjwdAdibbUTiDV27qRiKRVCgjH6d8M5a1LTtZiJJRUZHPRQ7RyKCMevXqhQEDBmDBggXw9vau8LqxsTG0tbVr7XiV3Xqu/loxrr1bsWUXLty4jX3rvoewacU7YMkn0NQFjNqjNGoH20k4QUdbC3w+HynlRgFT09Khz8EnJNQnHW3tf9qWOQqYmp4OfT1qW0LkiXx2UeXUqlWr8MsvvyA6uu6n5wQCATQ1NRmLvE8RSyQSLN+8C2ev3kTommVoaUjPnKxtPBt7IPcd8JSuuaoN6mpqsLRoh2s3bzPKo2/dga2NFUupuKGsbc1x7eYtRnn0zduwtbFmKRUhpDI0MlgD1tbWGD16NDZv3lyj7R4+fIjCwkKkpaXh3bt3iI2NBQB06tSp9kOyaPnmnThx4Qq2LvNH40YN8TYtHQCg0bgRGsj4TCVSHR54NvaQxF8HJLX7WAFlNm7UCPgtWQGrDhawtbZC5JHjECe/xggPV7ajKbxxo0fAL2A5rNq3h62NFSKjjlHb1oKc3Dy8/DtJuv5KnIyEJ39AS0MDzYUGLCaTM3Q9gsyoM1hDK1aswIEDB2q0zaBBg/Dnn39K121tbQFA+gsuXBHxy2kAwNg5ixnlK+dMh/uAvmxE4hZjC/C09FB67xrbSThlkNMXSM/MxLbdoXiTkop2JsbYueEHtDDk9qOg6sMgp37/tG3wP23bFjs3rkULQ0O2oym0+48e4+sZc6Xrqzb/BABwHdgfqxbOrWozQqrEk3CtR8JhkpcP2I7ASaVhm9iOwFn8ad+zHYG7eHSVT12R5GezHYGzeE3r7+kSxfOG19uxakJ1dSTbESqgkUFCCCGEcA9NE8uMvloSQgghhCgx6gwSQgghhCgxmiYmhBBCCPfQNLHMaGSQEEIIIUSJUWeQEEIIIUSJ0TQxIYQQQrhHTn8HWB5RZ1CB8LSash2Bk1TGzGQ7Anfx1dhOwF10PVSd4TXSYjsCIfWKus2EEEIIIUqMRgYJIYQQwj00ei4zGhkkhBBCCFFi1BkkhBBCCFFiNE1MCCGEEO6haWKZ0cggIYQQQogSU9jOoLe3N3g8HlatWsUoP3r0KHj/fBu4ePEieDxehWXRokXVvs7j8ZCcnAwA2LVrF3r27AkdHR3o6OigX79+uHXrVoUsrq6uleZMS0uDj48PzM3N0ahRI7Ru3RozZsxAZmZmLbcI+27fjcVk33lwGOQK8649ce7iZbYjccJP4YcwbOocfOYyAvYeX2Pa4pV49tffbMfilLADh9DXxRXWdj3hPmos7tyNYTsSZ4QdOIS+g4fCupsDtW0to/OW1BaF7QwCQIMGDbB69Wqkp6dXWy8xMRFisVi6zJ8/v9rXxWIxDAwMAJR1GEeOHInffvsN169fR+vWreHk5IS//5btj3FSUhKSkpKwdu1axMfHIzQ0FKdOncKECRM+7k3Lsdz8fJibmSJg7iy2o3DK7XsPMOrLgYjcsgbBa5aiuKQUE/2WIjcvn+1onHDy9FmI1q7HlAnjcDR8LzrbdsIkn1lIEiezHU3hnTx9FqIffixr24ify9p2+kxq21pA560MeDz5XOSQQncG+/XrB6FQCJFIVG09AwMDCIVC6dKkSZNqXxcKhVD558nlYWFhmDp1Kjp16gQLCwvs2rULpaWlOH/+vEwZrayscPjwYQwZMgQmJibo27cvAgMD8csvv6C4uPjj3ric6m1vh1lTJsHJsTfbUThl96olcHf+AmZGrWFhYgyRnw+S3rzFgyd/sB2NE0LCIuDh+iU83YbCpK0xFs71/T979x0VxdWGAfxZlqb0ooKKIiJSBCEWsGPFLmAXC6JGjWLBigVsESQq9o6gEbEiIcbPFkWDqGAE7L2gAlGQonRhvz9INlkBBVm4O8P7O2fOce/O7Dz7nnG4e6dBr149hBw7zjoa5wUeOFhcW2fHf2urVw8hR6m2lUXbLZEmTncGhUIhVq9ejc2bN+P169fVss7s7GwUFBRAW1v7mz8jIyMD6urqkJen63dIxX3IygYAaKipfmVO8jX5BQW4e/8BOtrZSrR3aNcWsfG3GaXiB3Ft231WWztbxMbfYpSKH2i7JdLG6c4gADg5OcHa2hre3t5lztOwYUOoqqqKp9TU1C++37x58zI/a+HChWjQoAF69OjxTXlTU1OxcuVKTJ48+Yvz5eXlITMzU2LKy8v7pnUS/hCJRPDdvhetWpjBpElj1nE4Ly09HYWFhdDRkfxxp6utg3ef7SdIxaSl/V1bbR2Jdl0dbaptJdF2Wz4COTmZnGQRL4am1qxZg27dumHOnDmlvv/HH39ATU1N/FpLS+uL75c1Yufn54eQkBBERERAWVm5wjkzMzPRr18/mJubf7HzCgA+Pj5Yvny5RJv3grlY5jmvwusl/LFy0y48fPYCBzd++dQIUjECSJ7HIxKJxBeikcr5vIxUW+mh7ZZICy86g507d4aDgwMWLVoEV1fXEu83adIEmpqaZS7/tfcBYO3atVi9ejXOnz8PKyurCmf88OEDevfuDVVVVZw4cQIKCgpfnN/T0xMeHh4SbUq5/LsCmZTfys27cOFqNA74r4ZeHV3WcXhBS1MTQqEQKZ+NpqSmvYduJU4FIYCWVhm1fZ9Gta0k2m6JtMnmeOU38PX1xa+//oqoqCipf/ZPP/2ElStX4vTp02jdunWFl8/MzESvXr2gqKiI8PDwco0qKikpQV1dXWJSUlL6lviE40QiEVZs2oVzf1xD0NqVaKhfj3Uk3lBUUICFmSmuXJe8XVTUtWjYtLRklIofxLW9VlptK/6DmvyLtttyYn3VMIeuJubFyCAAWFpawsXFBZs3b67wsm/fvkVuruRtOnR0dKCgoAA/Pz8sXboUBw8ehKGhofj+g/+cX/iPjIwMxMXFSXyGtrY2tLS00KtXL2RnZ+PAgQPi8/8AoE6dOhAKhRXOK6uysrOR8PrfW+68TkzC/UePoaGujvp61IH5Vis27cTJ3y9j68pFUKldC+/eF99KSU2lNpTpB0KljXcZiflLl6GFmSlsrCxxODQMScl/YcRgZ9bROG/86FGYv8QbLczN/q7tCSQlJ2PEEKptZdF2S6SJN51BAFi5ciWOHDlS4eVKu2Dk6tWrsLOzw7Zt25Cfn48hQ4ZIvO/t7Y1ly5aJX0dERMDGxkZinnHjxsHV1RXXr18HABgbG0u8//z5cxgaGlY4r6y6c/8hxk6dIX7ts2ELAMCpX2/4ei9mFYvzQsJPAwDGeiyRaF89zx3OvbuziMQrfR16Ii0jA9t278XblBSYNDXCrk3+aFBfn3U0zhPXdldAcW2Nm2LXZqqtNNB2S6RJIBKJRKxDkHLKeMs6AS+JPtDVd1VFoEV/mKqMjB5u4gX6s1h1VDSrbVWFK9yqbV0VIfTayzpCCbw5Z5AQQgghhFQcdQYJIYQQQmowXp0zSAghhBACgE6lqAAaGSSEEEIIqcGoM0gIIYQQUoPRYWJCCCGE8I+MPgdYFlGlCCGEEEJqMBoZ5BDRu1esI/BS4Q4f1hF4S957B+sI/CWg3/JVJi+bdQL+qsb7DJLyo84gIYQQQviHriYuN/ppSQghhBBSg1FnkBBCCCGkBqPDxIQQQgjhHzpMXG40MkgIIYQQUoNRZ5AQQgghpAaT+c6gq6srBAIBfH19JdrDwsIg+M8QcGFhIfz9/WFlZQVlZWVoamqiT58+uHLlingee3t7CASCMidDQ0MAQHJyMtzd3WFkZAQlJSUYGBhgwIAB+P3338WfZWhoiA0bNpTIu2zZMlhbW4tf3717F4MHD4ahoSEEAkGpy/DFziO/YMisJfhuiBvaj5qCaSvX4dnrRNaxuE9ODnK9R0C4aCuEPsEQem6FoOcQOgQiRcFHQ9Ft4BBYtu8K59FuuBEbxzoSbwQfOY5uA5xh2a4LnF1cqbZSEBN7C1PmLUbHgcPQvH13nL8UyTqSbBIIZHOSQTLfGQQAZWVlrFmzBmlpaaW+LxKJMGLECKxYsQIzZszA/fv3cenSJRgYGMDe3h5hYWEAgNDQUCQlJSEpKQnR0dEAgPPnz4vbYmJi8OLFC7Rq1QoXLlyAn58fbt++jdOnT6Nr166YNm1ahbNnZ2fDyMgIvr6+0NPT++YacEHM7fsY1a8nDq9bgb2rPPGpsAgTl/giOzeXdTROE3R1hKBdLxSdCECh3ywU/fYz5LoMgqBDH9bReOHU2fPwWbcRU93GIiw4EK1srDBpxlwkJiezjsZ5xbXdgKlurgg7uA+tbFpikrsHEpOotpWRnZuD5sZN4eXhzjoK4QlOXEDSo0cPPHnyBD4+PvDz8yvx/pEjR3Ds2DGEh4djwIAB4vZdu3YhNTUVEydORM+ePaGtrS1+L/fvDoqOjo5EJ23cuHEQCASIjo6GioqKuN3CwgJubm4Vzt6mTRu0adMGALBw4cIKL88le1ZKfj+f2ZPRftQU3H3yHG1amDFKxX2Cxs0huhMD0f2bAABR2juIrDtCYNAUIsbZ+CAw+DAGD+qPoY4DAQCL58xC5NVohBw7gTnTpzJOx22BB0IweNAADHX6u7ZzZyPy6nWEHAvFHPcfGKfjri7tbNGlnS3rGIRHODEyKBQKsXr1amzevBmvX78u8f7BgwdhYmIi0RH8x5w5c5Camopz5859dT3v37/H6dOnMW3aNImO4D80NTW/KX9N9SGr+C7+GqqqjJNwm+j5fQiaWQK6+sUN+o0haGIq7hySb5dfUIC7Dx6io11bifYOdm0Re+sOo1T8UHZtbRF76zajVKRGkZOTzUkGcWJkEACcnJxgbW0Nb29vBAQESLz36NEjmJmVPvL0T/ujR4++uo4nT55AJBLB1NS0XJkWLFiAJUuWSLTl5+fD3Ny8XMt/SV5eHvLy8iTaFPPyoaSkWOnPrg4ikQi+uw+glUVzmBgasI7DaaKLYRAp14Zw/kZAVAQI5FB0OgSiuCtfX5h8UVp6OgoLC6Hzn6MGAKCrrYV3KamMUvGDuLY6n9VWRwvvUt8zSkUIKY1sdlHLsGbNGuzbtw/37t2r8LKCcpy0KRKJyj0vAMybNw9xcXES05QpUyqcrTQ+Pj7Q0NCQmHx2Bkrls6vDyu1BePgiAevmT2cdhfME1h0gaNUZRQc3otB/PooObYFcl4EQtO7COhpvfP5/XiQq/36AfFmptWWUhRBSOs6MDAJA586d4eDggEWLFsHV1VXcbmJiUmYH8f79+wCAZs2affXzmzVrBoFAgPv378PR0fGr8+vq6sLY2FiiTfuzEYZv5enpCQ8PD4k2xVd3pfLZVW3l9iBcuP4nDqzxgp6uDus4nCfXfwyKLoSJRwJFyQko0qoDuW7OKLxxiXE6btPS1IRQKERKquQoYGpaGnR1pPN/uaYS1/azEdbU91RbUk3oB125cWpkEAB8fX3x66+/IioqStw2YsQIPH78GL/++muJ+detWwcdHR307Nnzq5+tra0NBwcHbN26FVlZWSXeT09Pr1T2ilBSUoK6urrEJOuHiEUiEVZsD8S5qzEIWr0YDfXqso7EDwpKxYeH/0tURDs6KVBUUICFaXNcuR4j0R51PQY2Vi0YpeKHsmsbDRsrS0apCCGl4dTIIABYWlrCxcUFmzdvFreNGDECR48exbhx4/DTTz+he/fuyMzMxNatWxEeHo6jR4+WekFIabZt24b27dujbdu2WLFiBaysrPDp0yecO3cO27dvF480lld+fr541DI/Px9v3rxBXFwcVFVVS4wqct2KbYE4eSkKW5fOgUqtWnj3Ph0AoKZSG8oy3pGVZaJ7NyDXfTCK0lMgSn4FQYMmkOvcH6KYi6yj8cJ4l+GY77USLcxMYWPVAodDf0FS8l8YMdiJdTTOGz96JOYvXY4W5qawsbLE4dCw4toOodpWRlZ2DhJevxG/fp2UjPuPnkBDXQ319eoxTEa4inOdQQBYuXIljhw5In4tEAhw5MgRbNy4Ef7+/pg2bRqUlJTQrl07XLx4ER07diz3Zzdp0gQ3b97Ejz/+iDlz5iApKQl16tRBq1atsH379gpnTUxMhI2Njfj12rVrsXbtWnTp0gUREREV/jxZFnLqPABg7MKVEu2rZ02Gc086v+1bFYUFQM5hBOScJwGq6kBGGkTXzqHo3DHW0Xihb68eSMvIxLY9gXibkgqTpkbYtXEtGujz+76g1aFvrx5IS8/Att17/63tpnVooK/POhqn3XnwEGOnzxG/9tlU/LfJqW8v+C5ZwCqW7KGjJ+UmEP1z1QSReaInf7KOwEuFO3xYR+Atee8drCPwl4BzZ/lwR1426wT8pdOw2lZVuFY2L2AUzt3COkIJtDchhBBCCKnBOHmYmBBCCCHki2T0Bs+yiCpFCCGEEFKDUWeQEEIIIaQGo8PEhBBCCOEfupq43GhkkBBCCCGkBqORQQ4R6NRnHYGX5Nxms47AX3JC1gl4jEY9qoyiMusEhFQr6gwSQgghhH/oMHG50WFiQgghhJAajDqDhBBCCCE1GB0mJoQQQgj/0GHicqORQUIIIYSQGow6g4QQQgghNRgvOoOurq4QCAQQCARQUFCAkZER5s6di6ysLLx48QICgQBxcXEAUOI1AHz48AH29vYwNTXFq1evAAADBw5Eo0aNoKysDH19fYwZMwaJiYkl1n38+HHY29tDQ0MDqqqqsLKywooVK/D+/XuJ+XJycqClpQVtbW3k5ORUWS1YiomNx5Q5nujYfzCa29nj/KU/WEfihZDTFzFolhdaj/oBrUf9gBELfsTlP2+xjsUrwUeOoVt/J1jadYbzqHG4cTOOdSTeKK6tIyztOsF51FjcuBnLOhJvBB8NRbeBQ2DZviucR7vhRmwc60iyRU5ONicZJJupvkHv3r2RlJSEZ8+eYdWqVdi2bRvmzp371eXevXuHrl274uPHj4iMjISBgQEAoGvXrjhy5AgePnyI48eP4+nTpxgyZIjEsosXL8bw4cPRpk0b/O9//8OdO3ewbt06xMfH4+eff5aY9/jx42jRogXMzc0RGhoqvS8uQ7JzctG8WVN4zZnJOgqv6OlowWPMEBz9yQtHf/KCnaUppvtuxuOEN6yj8cKpM+fgs3YDpk5wRdjBfWhlY41J7rORmJTMOhrnFdfWH1MnjEfYwf1UWyk6dfY8fNZtxFS3sQgLDkQrGytMmjEXiclUW1JxApFIJGIdorJcXV2Rnp6OsLAwcdukSZNw8uRJXL16FU2aNEFsbCysra3x4sUL8WsdHR307NkT+vr6CA8Ph5qaWpnrCA8Ph6OjI/Ly8qCgoIDo6GjY2tpiw4YNmDmzZOcnPT0dmpqa4tddu3bFiBEjIBKJcOTIEVy4cKHiXzQtqeLLMNLczh5b16xEjy6dWEf5qqKkZ6wjVJjdGHfMHTcUQ3p0Zh3li+Qam7OO8FVDx7rB3LQ5li9aIG7r4zwcPbp2wRz3Hxgm+xrZPzm+7Np2xhz3aQyTfUXRJ9YJvmrouEkwNzXBcs954rY+Q0ahh30nzJk+lWGyr1DTrbZVFW6eU23rqgih+zrWEUrgzcjg52rVqoWCgoIy33/48CE6dOgAU1NTnD59+osdwffv3yM4OBjt27eHgoICACA4OBiqqqr44YfS/1j8tyP49OlTXL16FcOGDcOwYcMQFRWFZ8+41wEh7BUWFuG3P64jOzcP1s2bso7DefkFBbh7/yE62tlKtHdoZ4vY+NuMUvFDcW0flFLbtlTbSsovKMDdBw/R0a6tRHsHu7aIvXWHUSoZJBDI5iSDeNkZjI6OxsGDB9G9e/cy5xk7diyaNm2K48ePQ0lJqdR5FixYABUVFejo6CAhIQG//PKL+L3Hjx/DyMhI3Dn8kr1796JPnz7icwZ79+6NvXv3VvyLkRrr0cvXaDVyKloO+x7Ld+zH5oXTYWzQgHUszktLT0dhYSF0dLQl2nW1tfEuNZVRKn4ou7Y6VNtKEtdW+/PaauFdCtWWVBxvOoMnT56EqqoqlJWV0a5dO3Tu3BmbN28uc/5BgwYhMjISx48fL3OeefPmITY2FmfPnoVQKMTYsWPxz1F1kUgEQTl6+IWFhdi3bx9Gjx4tbhs9ejT27duHwsLCMpfLy8tDZmamxJSXl/fV9RF+Mqyvh9D1y3BozWKM6N0Vnpv24MkrOmdQWgSfHXIt7/9v8nVU26rzeR1FopJthJQHb2463bVrV2zfvh0KCgqoX7++eMTuxYsXpc6/aNEiWFlZwcXFBSKRCMOHDy8xj66uLnR1dWFiYgIzMzMYGBjg2rVraNeuHUxMTBAZGYmCgoIvjg6eOXMGb968KfH5hYWFOHv2LPr06VPqcj4+Pli+fLlEm/d8Dyxb+PWLYgj/KCrIo7F+PQBAC+MmuP3kOX4+eR7Lp45jnIzbtDQ1IRQKkfLZSFVqWhp0Pxt1IRVTdm3fU20r6YvbrQ7VVow6xuXGm5FBFRUVGBsbo3HjxuU6dAsAS5YswcqVK+Hi4oKQkJAvzvvPiOA/o3OjRo3Cx48fsW3btlLnT09PBwAEBARgxIgRiIuLk5hcXFwQEBBQ5vo8PT2RkZEhMXnOdi/X9yI1gAjIL5D9k9xlnaKCAizMmuPK9WiJ9qhr0bBpackoFT8U19aUalsFFBUUYGHaHFeux0i0R12PgY1VC0apCJfxZmTwWy1cuBBCoRBjxoxBUVERXFxcEB0djejoaHTs2BFaWlp49uwZvLy80LRpU7Rr1w4AYGtri/nz52POnDl48+YNnJycUL9+fTx58gQ7duxAx44dMWrUKPz6668IDw9HixaS/0HHjRuHfv364d27d6hTp06JXEpKSiXPZSzMqrI6SENWdjYSXv976PJ1YjLuP3oMDXV11NerxzAZt/kfOI5O31lCX1cbWTm5OPXHdUTffYBdSz1YR+OF8S4jMX/pcrQwM4ONVQscDv0FScl/YcRgJ9bROK+4tsvQwswUNlaWOBwa9ndtnVlH47zxLsMx32vl37Wl7ZZUTo3vDALF5wYKhUKMGzcORUVFsLa2RmhoKLy9vZGVlQV9fX307t0bhw4dkuigrVmzBq1atcLWrVuxY8cOFBUVoWnTphgyZAjGjRuHgIAAqKiolHohS9euXaGmpoaff/4ZHh78+KN+5/5DjJ02W/zaZ+NWAIBTXwf4enmyisV5KekZWLBhN96lZUCtdi2YGDbErqUe6GBtwToaL/R16Im0jAxs2x2AtympMGlqhF2b1qNBfX3W0Tjv39ruxduUlL9r60+1lYK+vXogLSMT2/YE/rvdblyLBvp6rKPJDgFvDn5WOV7cZ7DG4NB9BrmEi/cZ5Aou3GeQu+h8qCrDgfsMclZ13mdw24Kvz8SA8Ic1rCOUQN1mQgghhJAajDqDhBBCCOEfOYFsThX05s0bjB49Gjo6Oqhduzasra3x559/SrVUdM4gIYQQQogMSktLQ4cOHdC1a1f873//Q926dfH06VOJp5xJA3UGCSGEEEJk0Jo1a2BgYIDAwEBxm6GhodTXQ4eJCSGEEMI/AjmZnCryhLHw8HC0bt0aQ4cORd26dWFjY4Pdu3dLvVTUGSSEEEIIqSY+Pj7Q0NCQmHx8fEqd99mzZ9i+fTuaNWuGM2fOYMqUKZgxYwb2798v1Ux0axkuoVvLVAm6tUzVoVvLVCW6tUyVoVvLVJ3qvLXMzkXVtq6K+OTqXWIksNQHTQBQVFRE69atERUVJW6bMWMGYmJicPXqVallonMGuURRmXUCXpLTN2Idgb/opq+Ei+ToTyMvyOizicvq+JVGX18f5uaSP6rNzMxw/PhxqWaiPTUhhBBCiAzq0KEDHj58KNH26NEjNG7cWKrroc4gIYQQQogMmj17Nq5du4bVq1fjyZMnOHjwIHbt2oVp06ZJdT00Fk4IIYQQ/pHj/nhXmzZtcOLECXh6emLFihVo0qQJNmzYABcXF6muhzqDhBBCCCEyqn///ujfv3+VroP73WZCCCGEEPLNaGSQEEIIIfwjo1cTy6IaOzLo6uoKgUAAgUAABQUFGBkZYe7cuZg3b564vazpxYsXWLZsWanvmZqaitdhb28PgUCAQ4cOSax7w4YNVfI4GVkQfOQYuvV3gqVdZziPGocbN+NYR+KFmNh4TJnjiY79B6O5nT3OX/qDdSReCT5yDN36DYKlbUc4jxqLGzdjWUfiDapt1Sne3zrC0q4T1ZZUSo3tDAJA7969kZSUhGfPnmHVqlXYtm0bUlJSkJSUJJ4aNmyIFStWSLQZGBgAACwsLCTak5KSEBkZKbEOZWVlLFmyBAUFBSy+YrU6deYcfNZuwNQJrgg7uA+tbKwxyX02EpOSWUfjvOycXDRv1hRec2ayjsI7p86cg89P6zF1wniEhfxcvN1On0XbrRRQbatO8f7Wv7i2B/fT/pZUSo3uDCopKUFPTw8GBgYYNWoUXFxccPr0aejp6YknoVAINTW1Em0AIC8vL9Gup6cHXV3Ju6uPHDkSGRkZVfIsQVkTGByCwY4DMNRpEJoaNcHiebOhV68uQo6Fso7GeV3a22L2lIno1bUz6yi8E3jgIAY7DsRQZ8e/t1sP6OnVQ8hR6d7UtSai2lad4v3twP/sbz2gV68eQo5RbcVk4DnEpU4ySDZTMVKrVi2pj+Cpq6tj0aJFWLFiBbKysqT62bIkv6AAd+8/REc7W4n2Du1sERt/m1EqQr6seLt9gI7tPttu7WwRG3+LUSp+oNpWHXFtS+xv29L+lnwT6gz+LTo6GgcPHkT37t3Lvczt27ehqqoqMU2cOLHEfD/88AOUlZWxfv16aUaWKWnp6SgsLISOjrZEu662Nt6lpjJKRciXpaX9vd1q60i06+rQdltZVNuqU/b+VodqS75Jjb6a+OTJk1BVVcWnT59QUFCAQYMGYfPmzeVevnnz5ggPD5doU1NTKzGfkpISVqxYgenTp2Pq1Knl+uy8vLySD7L+lFfu5xmyIoDk1VsikQgCuqKLyLjPN1HabqWHalt1aH/7FVSLcqvRI4Ndu3ZFXFwcHj58iNzcXISGhqJu3brlXl5RURHGxsYSU7169Uqdd/To0TA0NMSqVavK9dk+Pj7Q0NCQmHzW+pc7W3XT0tSEUChEyme/SlPT0qCrrV3GUoSwpaVVxnb7nrbbyqLaVp2y97fvqbbkm9TozqCKigqMjY3RuHFjKCgoVOm65OTk4OPjg+3bt+PFixdfnd/T0xMZGRkSk+fc2VWasTIUFRRgYdYcV65HS7RHXYuGTUtLRqkI+bLi7dYUV66Vtt1aMUrFD1TbqiOuLe1viZTU6MPElfXp0yckJ0texi8QCMocHezXrx9sbW2xc+fOMuf5h5KSUslDwlmFlcpb1ca7jMT8pcvRwswMNlYtcDj0FyQl/4URg51YR+O8rOxsJLx+I379OjEZ9x89hoa6OurrfXlbIl82fvQozF/ijRbmZrCxssTh0BNISk7GiCHOrKNxHtW26hTvb5ehhZnp37UN+3t/S7UV48GziasLdQYr4e7du9DX15doU1JSQm5ubpnLrFmzBu3bt6/qaEz0deiJtIwMbNsdgLcpqTBpaoRdm9ajQX39ry9MvujO/YcYO+3fkWGfjVsBAE59HeDr5ckqFi+It9tdAXibkgIT46bYtdmftlspoNpWnX/3t3uLa9vUCLs2UW3JtxGIRCIR6xCknLLSWCfgp/yyO++kkpRqs05ASMXRn8Wqo6JZbasq3Ley2tZVEcJxS1lHKIFGBgkhhBDCP3Q1cbnRAXVCCCGEkBqMOoOEEEIIITUYHSYmhBBCCP/I6HOAZRFVihBCCCGkBqPOICGEEEJIDUaHiQkhhBDCP3J0NXF5UWeQS+h+eFWiKPEJ6wi8JWfYgnUEQiruUz7rBIRUKzpMTAghhBBSg9HIICGEEEL4h64mLjeqFCGEEEJIDUadQUIIIYSQGowOExNCCCGEf+jZxOVGI4OEEEIIITUYdQbL4OrqCoFAAIFAAHl5eTRq1AhTp05FWlqaeB5DQ0PxPEKhEPXr18eECRMk5snNzYWrqyssLS0hLy8PR0dHBt+mesTExmPKHE907D8Yze3scf7SH6wj8ULI6YsYNNsbrV2mo7XLdIxYuBqXb95mHYtXgo8cQ7f+TrC06wznUeNw42Yc60i8QbWVvpibcZjisRAd+zqhedvOOB9B+1pSOdQZ/ILevXsjKSkJL168wJ49e/Drr7/ihx9+kJhnxYoVSEpKQkJCAoKDg3H58mXMmDFD/H5hYSFq1aqFGTNmoEePHtX9FapVdk4umjdrCq85M1lH4RU9HS14jB6Moz8twdGflsDO0hTTfbfgccIb1tF44dSZc/BZuwFTJ7gi7OA+tLKxxiT32UhMSmYdjfOotlUjO/fvfe28WayjyDaBnGxOMojOGfwCJSUl6OnpAQAaNmyI4cOHIygoSGIeNTU18TwNGjTA2LFjcejQIfH7Kioq2L59OwDgypUrSE9Pr5bsLHRpb4su7W1Zx+Cdrm2sJV7PcnHGoTMRiH/0DM0aNWATikcCg0Mw2HEAhjoNAgAsnjcbkVevIeRYKOa4//CVpcmXUG2rRpf2dujS3o51DMIjstlFlUHPnj3D6dOnoaCgUOY8b968wcmTJ2FrSx0iUjUKC4vwW2Q0snPzYd28Kes4nJdfUIC79x+io53k/9kO7WwRG0+H4iuDaksId9DI4BecPHkSqqqqKCwsRG5u8aPg1q9fLzHPggULsGTJEvE8tra2JeYhpLIevXyNkZ4+yMsvQG1lJWxe8AOMDeqzjsV5aenpKCwshI6OtkS7rrY23qWmMkrFD1Rbwhw9m7jcaGTwC7p27Yq4uDhcv34d7u7ucHBwgLu7u8Q88+bNQ1xcHG7duoXff/8dANCvXz8UFhZWat15eXnIzMyUmPLy8ir1mYS7DOvrIXSdFw75LsKI3vbw3LwXT14lso7FGwJI/tEQiUQQ0G0ppIJqS4jso87gF6ioqMDY2BhWVlbYtGkT8vLysHz5col5dHV1YWxsjGbNmqFbt27YsGEDoqKicPHixUqt28fHBxoaGhKTj//mSn0m4S5FBXk01q+HFsaG8Bg9GM0NDfDzyfOsY3GelqYmhEIhUj4bqUpNS4OutnYZS5HyoNoSwh3UGawAb29vrF27FomJZY/ICIVCAEBOTk6l1uXp6YmMjAyJyXO2+9cXJDWDSIT8T59Yp+A8RQUFWJg1x5Xr0RLtUdeiYdPSklEqfqDaEuYEAtmcZBCdM1gB9vb2sLCwwOrVq7FlyxYAwIcPH5CcnAyRSIRXr15h/vz50NXVRfv27cXL3bt3D/n5+Xj//j0+fPiAuLg4AIC1tXWZ61JSUoKSkpJkY2GWtL+SVGVlZyPh9b+3O3mdmIz7jx5DQ10d9fXqMUzGbf4HQtHpuxbQ19VGVk4uTkVGI/ruQ+xaMot1NF4Y7zIS85cuRwszM9hYtcDh0F+QlPwXRgx2Yh2N86i2VaPkvjaJ9rWkUgQikUjEOoQscnV1RXp6OsLCwiTaDx48iPHjx+PJkyfo1KkTXr58KX6vTp06aNOmDX788UeJjp6hoaHEfP+ocOnTkio2fzW7/mcsxk6bXaLdqa8DfL08GSQqn6LEJ6wjfNHirUG4dus+3qVlQK12LZgYNsREx97oYG3BOtpXyRm2YB2hXIKPHEPAvgN4m5IKk6ZG8JwzC21a2bCOxQucrO2nfNYJvuj6n7EYO7Xk/Vyd+vWGr/ciBokqQKP6OquFR2XzYk7hUA/WEUqgziCXyHhnkKtkvTPIZVzpDBIiQcY7g5xWnZ3BYxuqbV0VIRwyi3WEEuicQUIIIYSQGow6g4QQQgghNRhdQEIIIYQQ/qGbTpcbjQwSQgghhNRg1BkkhBBCCKnB6DAxIYQQQvhHQONd5UWVIoQQQgipwWhkkENEnwpYR+CnnI+sExDyDejk+KpDtSU1C3UGCSGEEMI/MvocYFlEh4kJIYQQQmow6gwSQgghhNRgdJiYEEIIIfxDVxOXG1WKEEIIIaQGo84gIYQQQkgNRp3BL3B1dYVAIIBAIIC8vDwaNWqEqVOnIi0tTTyPoaGheJ7/Tr6+vli2bFmp7/13evHiBbsvKGUxcbcwZf5SdBo0HKYde+L85SusI/FCyPlIDFq4Bq0nLEDrCQswwtsfl+PusY7FK8FHjqFbfydY2nWG86hxuHEzjnUk3iiurSMs7TrBedRY3LgZyzoS58XcjMMUjwXo2NcRzdt2wvmIy6wjySY5gWxOMog6g1/Ru3dvJCUl4cWLF9izZw9+/fVX/PDDDxLzrFixAklJSRKTu7s75s6dK9HWsGHDEvMaGBgw+mbSl5OTC1NjIyz1mM46Cq/oaWvCY8QAHF01B0dXzYGdhQmmrw/A49dJrKPxwqkz5+CzdgOmTnBF2MF9aGVjjUnus5GYlMw6GucV19YfUyeMR9jB/VRbKcnOzUXzZsbwmjebdRTCE3QByVcoKSlBT08PANCwYUMMHz4cQUFBEvOoqamJ5/mcqqqq+N9CofCL83Jd53Zt0bldW9YxeKfrdy0kXs8a1g+Hzl9B/JOXaNZQn1Eq/ggMDsFgxwEY6jQIALB43mxEXr2GkGOhmOP+w1eWJl9SXNuB/6mtByKvXkfIseOY4z6NcTru6tLeDl3a27GOQXiERgYr4NmzZzh9+jQUFBRYRyE1VGFREX67ehPZeXmwNjZkHYfz8gsKcPf+Q3S0s5Vo79DOFrHxtxml4ofi2j4opbZtqbakegjkZHOSQTQy+BUnT56EqqoqCgsLkZubCwBYv369xDwLFizAkiVLSixnb2//zevNy8tDXl6eRJtiXh6UlJS++TMJdz1KSMTIZRuQV/AJtZUVsXn2BBg35OcIc3VKS09HYWEhdHS0Jdp1tbXxLjWVUSp+KLu2OniXeo1RKkJIaWSziypDunbtiri4OFy/fh3u7u5wcHCAu7u7xDzz5s1DXFycxGRra1vGJ5aPj48PNDQ0JCafjdsq9ZmEuwzr10Xo6nk4tHwWRnTvAM8dwXjyms67khbBZ8+iFYlEENCjrKSCakuI7KORwa9QUVGBsbExAGDTpk3o2rUrli9fjpUrV4rn0dXVFc8jLZ6envDw8JBoU8z8S6rrINyhKC+Pxnp1AAAtjBrh9rNX+PnMJSyfMJxxMm7T0tSEUChEymejgKlpadDV1i5jKVIeZdf2PdWWVA/60VFuNDJYQd7e3li7di0SExOrdD1KSkpQV1eXmOgQMfmXCPkFn1iH4DxFBQVYmDXHlevREu1R16Jh09KSUSp+KK6tKdWWEA6gkcEKsre3h4WFBVavXo0tW7YAAD58+IDkZMlDdrVr14a6ujqLiMxkZecg4c0b8evXScm4//gJNNTUUV+vLsNk3OZ/+CQ6tTSDvo4msnLycOpaLKLvPcGuBVNYR+OF8S4jMX/pcrQwM4ONVQscDv0FScl/YcRgJ9bROK+4tsvQwswUNlaWOBwa9ndtnVlH47Ss7GwkvP7PvjYxCfcfPYaGujrq69VjmIxwFXUGv4GHhwfGjx+PBQsWAAC8vLzg5eUlMc/kyZOxY8cOFvGYufPgEcbNmCt+7bu5+Ps79ukJ38XzWcXivJSMD1iw/QDepWdCrXYtmBjUx64FU9DBsjnraLzQ16En0jIysG13AN6mpMKkqRF2bVqPBvXptj2V9W9t9+JtSsrftfWn2lbSnfsPMXbqDPFrnw3FAxNO/XrD13sxq1iyR44OfpaXQCQSiViHIOUjepfAOgIviV7eZR2Bt+TM6F5oVYfOh6oyn/JZJ+Avjeo7SlT4vz3Vtq6KEPaZyDpCCdRtJoQQQgipwegwMSGEEEL4h64mLjcaGSSEEEIIqcGoM0gIIYQQUoPRYWJCCCGE8I+MPgdYFlGlCCGEEEJqMOoMEkIIIYTUYHSYmEtyMlkn4CXRqaOsI/BXs1asE/CXnJB1At4S5eewjsBb1Xp9L11NXG40MkgIIYQQUoNRZ5AQQgghpAajw8SEEEII4R96NnG5UaUIIYQQQmow6gwSQgghhNRgdJiYEEIIIfxDVxOXW40bGXR1dYVAIIBAIIC8vDwaNWqEqVOnIi0tTWK+qKgo9O3bF1paWlBWVoalpSXWrVuHwsJCifkuXryIrl27QltbG7Vr10azZs0wbtw4fPr0CQCQm5sLV1dXWFpaQl5eHo6OjtX1VavdzpBQDJk2H98NdEH7oeMxzdsXz169YR2LF+Rm+EHotbfEJOgzmnU0Xgg+dgLdBg2DZcfucB47ATdi41lH4o3go8fRbeBgWLa3h/Po8bgRG8c6EufFxN3ClPlL0WnQCJh27IXzl6+wjkQ4rsZ1BgGgd+/eSEpKwosXL7Bnzx78+uuv+OGHH8TvnzhxAl26dEHDhg1x8eJFPHjwADNnzsSPP/6IESNGQCQSAQDu3r2LPn36oE2bNrh8+TJu376NzZs3Q0FBAUVFRQCAwsJC1KpVCzNmzECPHj2YfN/qEnPrLkYN7I3Dm3yw19cbnwqLMHHhCmTn5LKOxnlFe1aicN2sf6ef1wIARPdiGCfjvlPnfofP+k2YOn4Mwn4OQCvrlpg0ax4Sk/9iHY3zTp09D591GzHVbRzCgoPQyqYlJs2Yg8TkZNbROC0nJxemxkZY6jGddRTCEzXyMLGSkhL09PQAAA0bNsTw4cMRFBQEAMjKysKkSZMwcOBA7Nq1S7zMxIkTUa9ePQwcOBBHjhzB8OHDce7cOejr68PPz088X9OmTdG7d2/xaxUVFWzfvh0AcOXKFaSnp1f9F2Rkj89Sidc+c6eh/VA33H38FG2sLBil4onsDxIvBc1aQvT+L+DlQ0aB+CPw4GEMHtgPQx0HAAAWe8xA5LVohBw/gTnTpjBOx22BwYcweNAADHUcCABYPGcWIq9eR8ixE5gzfSrjdNzVuV1bdG7XlnUM2UfPJi63Gl+pZ8+e4fTp01BQUAAAnD17FqmpqZg7d26JeQcMGAATExOEhIQAAPT09JCUlITLly9Xa2au+JCVDQDQUFNjnIRn5IQQWNlBFBfJOgnn5RcU4O6DR+hoK/mHtYNtG8TeusMoFT8U1/YhOtp9Vlu7toi9dZtRKkJIaWpkZ/DkyZNQVVVFrVq10LRpU9y7dw8LFiwAADx69AgAYGZmVuqypqam4nmGDh2KkSNHokuXLtDX14eTkxO2bNmCzMzKPzYuLy8PmZmZElNeXn6lP7e6iEQi+O4IQqsWZjBp0oh1HF4RmH4HKNeGKI7OE6qstPQMFBYWQkdHS6JdV1sL71LfM0rFD2np6cW11daWaNfV1sa7FKotId/Cx8cHAoEAs2bNkurn1sjOYNeuXREXF4fr16/D3d0dDg4OcHd3l5jnn/MCPycSiSD4+woloVCIwMBAvH79Gn5+fqhfvz5+/PFHWFhYICkpqVIZfXx8oKGhITH5bNtTqc+sTis378HD5y+xbtFs1lF4R2DTCXhyG/iYzjoKbwg+e2KqSATx/3NSOZ+XsXgfyiYLqWEEAtmcvlFMTAx27doFKysrKRapWI3sDKqoqMDY2BhWVlbYtGkT8vLysHz5cgCAiYkJAOD+/fulLvvgwQM0a9ZMoq1BgwYYM2YMtm7dinv37iE3Nxc7duyoVEZPT09kZGRITJ4/TKzUZ1aXlVv24MK1GOz/aTn06uiwjsMvGjpAE3MU3aRTE6RBS1MDQqEQKZ+NAqampUFXW6uMpUh5aGlqll1bHe0yliKElObjx49wcXHB7t27oaUl/X1TjewMfs7b2xtr165FYmIievXqBW1tbaxbt67EfOHh4Xj8+DFGjhxZ5mdpaWlBX18fWVlZlcqkpKQEdXV1iUlJSbFSn1nVRCIRVmzejXOR1xHktwwN9euxjsQ7AuuOQFYm8PgW6yi8oKigAAtTE1yJlrwqOyo6BjZWLRil4ofi2jbHlevREu1R12NgY2XJKBUh7JV+GljeF5eZNm0a+vXrV2V3JaHOIAB7e3tYWFhg9erVUFFRwc6dO/HLL7/g+++/x61bt/DixQsEBATA1dUVQ4YMwbBhwwAAO3fuxNSpU3H27Fk8ffoUd+/exYIFC3D37l0MGDBA/Pn37t1DXFwc3r9/j4yMDMTFxSEuLo7Rt606Kzbvxq+/X8Zaz1lQqV0L796n4d37NOR+ZSMn5SWAoGUHiG5FAaIi1mF4Y/yo4Tj2y0kcC/8NT5+/wOr1m5CU/BYjnB1ZR+O88S4jcCzsVxz75WRxbddtRFLyXxgx2JF1NE7Lys7B/cdPcf/xUwDA66Rk3H/8FInJbxknkzECOZmcSj0NzMenzK9x6NAh3Lx584vzVFaNvLVMaTw8PDB+/HgsWLAAQ4YMwcWLF7F69Wp07twZOTk5MDY2xuLFizFr1izxuURt27ZFZGQkpkyZgsTERKiqqsLCwgJhYWHo0qWL+LP79u2Lly9fil/b2NgAKPu8RK4K+fUMAGDsXC+J9tVzp8HZoRuLSPxiZA6Bpi6KYv9gnYRX+vbsjrSMTGwLCMLblFSYNG2CXf5+aKCvxzoa5/Xt1QNpGRnYtmfv37U1wq6Na9FAX591NE678+ARxs2YJ37tu3knAMCxT0/4Lp5X1mJERnh6esLDw0OiTUlJqdR5X716hZkzZ+Ls2bNQVlauskwCEd96JDwmSqBbXVSFoqD1rCPwlnCmL+sI/CUnZJ2At0S5H1lH4C1BncbVtq7CiEPVtq6KENqPKPe8YWFhcHJyglD47//3wsJCCAQCyMnJIS8vT+K9b0Ujg4QQQgjhHznuX7bevXt33L4teV/O8ePHw9TUFAsWLJBKRxCgziAhhBBCiExSU1NDixaSF7OpqKhAR0enRHtl0AUkhBBCCCE1GI0MEkIIIYR/ePps4oiICKl/Jj8rRQghhBBCyoU6g4QQQgghNRgdJiaEEEII/9BDsMuNOoNcolh1N5ys0ay+Y52Av4S0i6kydJ/BKiOorcE6AiHVig4TE0IIIYTUYPSznRBCCCH8w9OriasCVYoQQgghpAajziAhhBBCSA1Gh4kJIYQQwjsCupq43GhkkBBCCCGkBuNsZ/Dt27eYPHkyGjVqBCUlJejp6cHBwQFXr14FABgaGkIgEJSYfH19sWzZslLf++/04sWLEvNpaGigU6dOuHTpkkSW/66rdu3aaNGiBXbu3Cl+PykpCaNGjULz5s0hJyeHWbNmVWepqk1M/B1MWbgcnZzHwLRLP5z/4yrrSLyw5dx1mC/YLDF1WhnAOhavBB85jm4DnGHZrgucXVxxIzaOdSTeCD5yDN36DYKlbUc4jxqLGzdjWUfijeAjx9CtvyMs7TpRbUmlcLYzOHjwYMTHx2Pfvn149OgRwsPDYW9vj/fv34vnWbFiBZKSkiQmd3d3zJ07V6KtYcOGJeY1MDAAAFhYWIjbrl69imbNmqF///7IyMiQyPPP8rdu3YKjoyOmTJmCw4cPAwDy8vJQp04dLF68GC1btqy+IlWznJxcmBo3wdJZU1hH4R3jetq4tMRNPP0yexTrSLxx6ux5+KzbgKlurgg7uA+tbFpikrsHEpOSWUfjvFNnzsHnp/WYOmE8wkJ+Risba0yaPotqKwWnzpyDz1r/4toe3F9cW/fZVNv/EsjJ5iSDOHnOYHp6OiIjIxEREYEuXboAABo3boy2bdtKzKempgY9Pb1SP0NVVVX8b6FQWOa88vLy4nY9PT0sX74cgYGBePToEdq0aVPqulatWoUjR44gLCwMw4cPh6GhITZu3AgA2Lt3byW+uWzrbNcane1as47BS0I5OdRRU2Edg5cCD4Rg8KABGOo0EACweO5sRF69jpBjoZjj/gPjdNwWeOAgBjsOxFBnRwDA4nkeiLx6DSFHj2POjGlsw3FcYHBIcW2dBgH4p7bXEXLsOOa4U21JxchmF/UrVFVVoaqqirCwMOTl5VXbevPy8hAUFARNTU00b978i/MqKyujoKCgmpIRvktISUeXVXvR03cf5gSfxqvUjK8vRL4qv6AAdx88REc7yR+SHexsEXvrNqNU/JBfUIC79x+gYztbifYOdraIjb/FKBU/iGtr91lt27VFbDxtt6TiONkZlJeXR1BQEPbt2wdNTU106NABixYtwq1bkjuYBQsWiDuO/0wREREVWtft27fFy9aqVQtr165FSEgI1NXVS53/06dPCAoKwu3bt9G9e/dv/YrIy8tDZmamxFSdHV8iO6wM6sFneE/snjAQywd3RcrHbIzadgzpWTmso3FeWno6CgsLoaOjLdGuq6OFd6nvy1iKlEda2t+11daRaNfV0ca71FRGqfihzO1WW4dq+1+sDwdz6DCxbKYqh8GDByMxMRHh4eFwcHBAREQEvvvuOwQFBYnnmTdvHuLi4iQmW1vbsj+0FM2bNxcv++eff2Lq1KkYOnQobty4ITHfPx3PWrVqYdq0aZg3bx4mT578zd/Px8cHGhoaEpPP5p1fX5DwTmdTQ/SyNIaJvi7aN2uE7eMHAADC/nzAOBl/fH4LCpEIoJtSSMfnd/cQiUR0yw8pEeDz7ZZqS74NJ88Z/IeysjJ69uyJnj17wsvLCxMnToS3tzdcXV0BALq6ujA2Nq7UOhQVFSU+w8bGBmFhYdiwYQMOHDggbp83bx5cXV1Ru3Zt6OvrV/o/pKenJzw8PCSzpL2q1GcSfqitqAATPR28TE1nHYXztDQ1IRQKkZIiOZqS+j4Nup+NupCK0dL6u7appdRWm2pbGeLt9vPapr2n2pJvwtmRwdKYm5sjKyurytcjFAqRkyN5iO6fjmf9+vWl8stMSUkJ6urqEpOSklKlP5dwX/6nQjx7+54uKJECRQUFWJg2x5XrMRLtUdejYWNlySgVPygqKMDCzBRXrkVLtEddi4ZNSytGqfhBXNvrpdWWtlsxOYFsTjKIkyODqampGDp0KNzc3GBlZQU1NTXcuHEDfn5+GDRokHi+Dx8+IDlZ8jL72rVrl3m+X2k+ffok/owPHz7g8OHDuHfvHhYsWFChzHFxcQCAjx8/4t27d4iLi4OioiLMzc0r9DmyLCs7BwlvEsWvXycl4/7jp9BQV0P9enUZJuM2v5OR6GreBPqaqkj9mIOdF2LwMS8fg1qZso7GC+NHj8T8pcvRwtwUNlaWOBwahqTkvzBiiBPraJw3fvQozF/ijRbmZn/X9gSSkpMxYogz62icN95lJOYvXYYWZp9tt4OptqTiONkZVFVVha2tLfz9/fH06VMUFBTAwMAAkyZNwqJFi8TzeXl5wcvLS2LZyZMnY8eOHeVe1927d6Gvrw+guCPZtGlTbN++HWPHjq1QZhsbG/G///zzTxw8eBCNGzfGixcvKvQ5suzOw8cYN8tT/Np36x4AgGPv7vD19ChrMfIVf2V8xNyDZ5CWnQNtlVpo2UgPIdOGoYFW+X/UkLL17dUDaekZ2LZ7L96mpMKkqRF2bVqHBn//vyffrq9DT6RlZGDbrgC8TUmBiXFT7Nrsjwb1qbaVJa7t7r3FtW1qhF2bqLbk2whEIpGIdQhSPqLkJ6wj8FLRtdOsI/CWsAfdHLvKyAlZJ+Av+rNYdVQ0q21VRTGnqm1dFSHXpi/rCCXw6pxBQgghhBBSMdQZJIQQQgipwTh5ziAhhBBCyBfRPRfLjUYGCSGEEEJqMOoMEkIIIYTUYHSYmBBCCCH8I6PPAZZF1BnkkpyPrBPwkuhaJOsI/NV1KOsE/EV/6KrOp3zWCQipVrQ3IYQQQgipwWhkkBBCCCH8Q1cTlxuNDBJCCCGE1GDUGSSEEEIIqcHoMDEhhBBC+Icusio3qhQhhBBCSA3Gy87g27dvMXnyZDRq1AhKSkrQ09ODg4MDrl69CgAwNDSEQCAoMfn6+mLZsmWlvvff6cWLFxLzycvLQ1dXF507d8aGDRuQl5cnzlJQUIAFCxbA0tISKioqqF+/PsaOHYvExERW5akyOw+dwBB3T3znNA7th0/CtOU/4dkr/n3PaicnB7lewyCcvwnClfshnLcRgu7OdHK0lMTcjMMUj4Xo2NcJzdt2xvmIP1hH4pXgI8fQrb8jLO06wXnUWNy4Gcs6EucVb7ML0LGvI5q37YTzEZdZRyIcx8vO4ODBgxEfH499+/bh0aNHCA8Ph729Pd6/fy+eZ8WKFUhKSpKY3N3dMXfuXIm2hg0blpjXwMAAAGBhYYGkpCQkJCTg4sWLGDp0KHx8fNC+fXt8+PABAJCdnY2bN29i6dKluHnzJkJDQ/Ho0SMMHDiQSW2qUszt+xg1wAGH/Vdhr89ifCoswsTFPyI7N5d1NE4TdBkIgW0PFP0SiML1c1D0v4OQ6zwAgva9WUfjhezcXDRv1hRe82axjsI7p86cg89af0ydMB5hB/ejlY01JrnPRmJSMutonFa8zRrDa95s1lFkm5xANicZxLtzBtPT0xEZGYmIiAh06dIFANC4cWO0bdtWYj41NTXo6emV+hmqqqrifwuFwjLnlZeXF7fXr18flpaW6NmzJ1q2bIk1a9Zg1apV0NDQwLlz5ySW27x5M9q2bYuEhAQ0atSoUt9Xluz5cZHEax+PqWg/YhLuPn6GNpbmjFJxn6CRCUT3/oToYfGIiijtHUTW7SFoYAQR42x80KW9Hbq0t2Mdg5cCg0Mw2HEghjoNAgAsnueByKvXEXLsOOa4T2OcjrtomyXSxruRQVVVVaiqqiIsLEzicG11MTU1RZ8+fRAaGlrmPBkZGRAIBNDU1Ky+YAx8yM4GAGioqX5lTvIlohcPIDBuAejqFzfoN4KgcXNx55AQWZRfUIC79x+go52tRHuHdm0RG3+bUSpCSGl41xmUl5dHUFAQ9u3bB01NTXTo0AGLFi3CrVu3JOZbsGCBuOP4zxQRESGVDKampnjx4kWp7+Xm5mLhwoUYNWoU1NXVpbI+WSQSieC7cz9aWZjCxJA/o58siC6FQxR3BUKPdRD+eABCd18UXfkfRPFRrKMRUqa09HQUFhZCR0dbol1XWwfvUlMZpSI1ikBONicZxLvDxEDxOYP9+vXDH3/8gatXr+L06dPw8/PDnj174OrqCgCYN2+e+N//aNCggVTWLxKJICjl5P6CggKMGDECRUVF2LZt2xc/Iy8vr8TIpmJePpSUFKWSsaqt3LoXD58n4OC65ayjcJ7Aqh0ENp1QdGgzRH+9hqC+IeT6j0VRZhpEN+nEcSLbBJDcF5a1fySEsCObXVQpUFZWRs+ePeHl5YWoqCi4urrC29tb/L6uri6MjY0lplq1akll3ffv30eTJk0k2goKCjBs2DA8f/4c586d++qooI+PDzQ0NCQmn+17pZKvqq3cthcXrv2J/X5e0KujwzoO58n1HY2iiF8gunUV+OsVRLF/oOjKKcjZD2IdjZAyaWlqQigUIuWzUcDUtPfQ1dYuYylCCAu87Qx+ztzcHFlZWVW+ngcPHuD06dMYPHiwuO2fjuDjx49x/vx56Oh8vYPk6emJjIwMiclzqltVRq80kUiEFVv34tyVaAStWYqGenVZR+IHBUVA9NmlIkVFMnu4gRAAUFRQgIWZKa5cj5Zoj7oWDZuWloxSkRpFIJDNSQbx7jBxamoqhg4dCjc3N1hZWUFNTQ03btyAn58fBg36dyTlw4cPSE6WvL1B7dq1K3Qe36dPn5CcnIyioiKkpqYiIiICq1atgrW1NebNmyeeZ8iQIbh58yZOnjyJwsJC8Xq1tbWhqFj6YV8lJSUoKSlJtIlSZfsQ8YqtATh58Qq2es+DSq1aePc+HQCgplIbyhw5vC2LRA9uQq6bI4rSUyB6+/dh4o79ILoRwToaL2RlZyPh9Rvx69eJSbj/6DE01NVRX68ew2TcN95lJOYvXYYWZqawsbLE4dAwJCX/hRGDnVlH4zTaZom08a4zqKqqCltbW/j7++Pp06coKCiAgYEBJk2ahEWL/r31iZeXF7y8vCSWnTx5Mnbs2FHudd29exf6+voQCoXQ0NCAubk5PD09MXXqVHFH7vXr1wgPDwcAWFtbSyx/8eJF2Nvbf9sXlUEhJ4tvoTN2vuR5gqs9psK5lz2DRPxQ9Esg5HoNg5yjG6CqAWSmQRR9HkW/H2cdjRfu3H+IsVNnil/7bNgCAHDq1xu+3ovKWoyUQ1+HnkjLyMC23XvxNiUFJk2NsGuTPxrU12cdjdOKt9kZ4teS2+xiVrEIhwlEos+PPxFZJXoexzoCLxXu9GUdgbfkPTeyjsBf8kpfn4d8m0/5rBPwl0b1nT5UdEc2L7CTa9GZdYQS6KQjQgghhJAajDqDhBBCCCE1GO/OGSSEEEIIkdUrd2URjQwSQgghhNRg1BkkhBBCCKnB6DAxIYQQQviHbsxfblQpQgghhJAajEYGOUSgXZ91BF4Szlj+9ZnIt6F74VUdOjm+6gjpTyOpWWiLJ4QQQgj/yNHBz/KiShFCCCGE1GDUGSSEEEIIqcHoMDEhhBBCeEdA59WWG40MEkIIIYTUYNQZJIQQQgipwXjRGXz79i0mT56MRo0aQUlJCXp6enBwcMDVq1cBAIaGhhAIBDh06FCJZS0sLCAQCBAUFFTivdWrV0MoFMLX17fEe0FBQRAIBBAIBBAKhdDS0oKtrS1WrFiBjIyMMrP6+PhAIBBg1qxZ3/x9ZVXMzThM8ViAjn0d0bxtJ5yPuMw6Ei/sDD6KIVM88F3f4WjvNAbTlvyIZwmvWcfileAjx9CtvyMs7TrBedRY3LgZyzoSbwQfOYZu/QbB0rYj1VbKgo8cR7cBzrBs1wXOLq64ERvHOpJsEcjJ5iSDZDNVBQ0ePBjx8fHYt28fHj16hPDwcNjb2+P9+/fieQwMDBAYGCix3LVr15CcnAwVFZVSPzcwMBDz58/H3r17S31fXV0dSUlJeP36NaKiovD9999j//79sLa2RmJiYon5Y2JisGvXLlhZWVXi28qu7NxcNG9mDK95s1lH4ZWY+DsY5dgPh7f+hL0/rcCnwkJMnO+N7Jxc1tF44dSZc/BZ64+pE8Yj7OB+tLKxxiT32UhMSmYdjfNOnTkHn5/WF9c25Ofi2k6fRbWVglNnz8Nn3QZMdXNF2MF9aGXTEpPcPai25JtwvjOYnp6OyMhIrFmzBl27dkXjxo3Rtm1beHp6ol+/fuL5XFxccOnSJbx69UrctnfvXri4uEBevuR1NJcuXUJOTg5WrFiBrKwsXL5ccpRLIBBAT08P+vr6MDMzw4QJExAVFYWPHz9i/vz5EvN+/PgRLi4u2L17N7S0tKRYAdnRpb0dZk+dhF5du7COwit7/JbDuXd3NGvSCKbGTeCzYCYS/3qHu4+esI7GC4HBIRjsOBBDnQahqVETLJ7nAb169RBy7DjraJwXeOBgcW2dHf+trV49hByl2lZW4IEQDB40AEOdBqJpE0MsnjsbevXqIuRYKOtohIM43xlUVVWFqqoqwsLCkJeXV+Z89erVg4ODA/bt2wcAyM7OxuHDh+Hm5lbq/AEBARg5ciQUFBQwcuRIBAQElCtP3bp14eLigvDwcBQWForbp02bhn79+qFHjx4V+HaElPQhKwsAoKGuxjgJ9+UXFODu/QfoaGcr0d6hXVvExt9mlIofxLVt91lt7WwRG3+LUSp+yC8owN0HD9HRrq1Eewc7W8Teou1WTCCQzUkGcb4zKC8vj6CgIOzbtw+ampro0KEDFi1ahFu3Su5s3NzcEBQUBJFIhGPHjqFp06awtrYuMV9mZiaOHz+O0aNHAwBGjx6NY8eOITMzs1yZTE1N8eHDB6SmpgIADh06hJs3b8LHx+fbvyghAEQiEXy37UUrS3OYNGnMOg7npaWno7CwEDo62hLtuto6ePf3/1/ybdLS/q6tto5Eu66ONtW2ksrcbnW08C71fRlLEVI2zncGgeJzBhMTExEeHg4HBwdERETgu+++K3FRSL9+/fDx40dcvnwZe/fuLXNU8ODBgzAyMkLLli0BANbW1jAyMir1ApTSiEQiAMWHkV+9eoWZM2fiwIEDUFZWLvd3ysvLQ2ZmpsT0pZFPUjOs3LgTD5++wLqlc1lH4RUBJH+ti0QiukeZlHxeRqqt9HxeR5EIoMqSb8GLziAAKCsro2fPnvDy8kJUVBRcXV3h7e0tMY+8vDzGjBkDb29vXL9+HS4uLqV+1t69e3H37l3Iy8uLp7t375b7UPH9+/ehrq4OHR0d/Pnnn3j79i1atWol/qxLly5h06ZNkJeXlziU/F8+Pj7Q0NCQmHzWb6pYUQivrNy0ExeiorHffxX06uiyjsMLWpqaEAqFSPlspCo17T10tbXLWIqUh5ZWGbV9n0a1rSTxdptSSm11qLZirK8apquJ2TM3N0fW3+dW/ZebmxsuXbqEQYMGlXohx+3bt3Hjxg1EREQgLi5OPF2+fBkxMTG4c+fOF9f79u1bHDx4EI6OjpCTk0P37t1x+/Ztic9q3bo1XFxcEBcXB6FQWOrneHp6IiMjQ2Ly9JjxbcUgnCYSibBi4w6c++MqgtavQkN9PdaReENRQQEWZqa4cj1aoj3qWjRsWloySsUP4tpeK622/LyjQnVRVFCAhWlzXLkeI9EedT0aNla03ZKK4/zj6FJTUzF06FC4ubnBysoKampquHHjBvz8/DBo0KAS85uZmSElJQW1a9cu9fMCAgLQtm1bdO7cucR77dq1Q0BAAPz9/QEU/5FOTk6GSCRCeno6rl69itWrV0NDQ0N8b0I1NTW0aNFC4nNUVFSgo6NTov2/lJSUoKSkJNkoku1biWRlZyPh9Rvx69eJSbj/6DE01NVRX68ew2TctmLDDpz8/TK2rloMldq18O59GgBATaU2lD/fRkiFjXcZiflLl6GFmSlsrCxxODQMScl/YcRgZ9bROG/86FGYv8QbLczN/q7tCSQlJ2PEEKptZY0fPRLzly5HC/PPttshTqyjEQ7ifGdQVVUVtra28Pf3x9OnT1FQUAADAwNMmjQJixYtKnUZHR2dUtvz8/Nx4MABLFiwoNT3Bw8eDB8fH6xZswZA8YUm+vr6EAgEUFdXR/PmzTFu3DjMnDkT6urq0vmCHHLn/kOMnfrv6KXPhi0AAKd+veHrvZhVLM4LCf8fAGDsbMntefWCmXDu3Z1FJF7p69ATaRkZ2LZ7L96mpMCkqRF2bfJHg/r6rKNxnri2uwKKa2vcFLs2U22loW+vHkhL/2e7Tf17u12HBvpUWzE6N7XcBKJ/rnYgsi/jLesEvCTKSmMdgbcEGjQiXGXoD13VKSr9XG4iBarVd06j6OnNaltXRQiafsc6Qgm8PWeQEEIIIYR8HecPExNCCCGElCBH413lRZUihBBCCKnBqDNICCGEEFKD0WFiQgghhPAPXWRVbjQySAghhBBSg1FnkBBCCCGkBqPDxBwiys5gHYGXiq78xjoCbwn7uLKOwF90h9iqk5fDOgF/qVbjumT0OcCyiCpFCCGEEFKDUWeQEEIIIaQGo8PEhBBCCOEfupq43GhkkBBCCCFEBvn4+KBNmzZQU1ND3bp14ejoiIcPH0p9PdQZJIQQQgiRQZcuXcK0adNw7do1nDt3Dp8+fUKvXr2QlZUl1fXQYWJCCCGE8BD3DxOfPn1a4nVgYCDq1q2LP//8E507d5baepiODL59+xaTJ09Go0aNoKSkBD09PTg4OODq1asAAENDQwgEAhw6dKjEshYWFhAIBAgKCirx3urVqyEUCuHr61vivaCgIAgEAvFUr149DBgwAHfv3pWYLz8/H35+fmjZsiVq164NXV1ddOjQAYGBgSgoKJCYNyoqCkKhEL179/7i901NTUXDhg0hEAiQnp7+lepwz87gIxgyeTa+6zMU7R1dMG3xKjxLeM06Fudt+f0GzJfslJg6+e5nHYtXgo8cR7cBzrBs1wXOLq64ERvHOhJvUG2lLyY2HlPmLULHgUPRvH03nL8UyToSqYC8vDxkZmZKTHl5eeVaNiOj+BZz2traUs3EtDM4ePBgxMfHY9++fXj06BHCw8Nhb2+P9+/fi+cxMDBAYGCgxHLXrl1DcnIyVFRUSv3cwMBAzJ8/H3v37i31fXV1dSQlJSExMRG//fYbsrKy0K9fP+Tn5wMo7gg6ODjA19cX33//PaKiohAdHY1p06Zh8+bNJTqOe/fuhbu7OyIjI5GQkFDm950wYQKsrKzKVRsuiom7g1GO/XB421rsXbsSnwoLMXHeUmTn5LKOxnnGdbVwacEY8fSL+1DWkXjj1Nnz8Fm3AVPdXBF2cB9a2bTEJHcPJCYls47GeVTbqpGdm4vmxk3h5eHOOgr5Bj4+PtDQ0JCYfHx8vrqcSCSCh4cHOnbsiBYtWkg1E7PDxOnp6YiMjERERAS6dOkCAGjcuDHatm0rMZ+Liwv8/f3x6tUrGBgYACjufLm4uGD//pKjI5cuXUJOTg5WrFiB/fv34/LlyyWGUgUCAfT09AAA+vr6mD17NgYOHIiHDx/C0tISGzZswOXLl3Hjxg3Y2NiIlzMyMsLQoUPFnUYAyMrKwpEjRxATE4Pk5GQEBQXBy8urRK7t27cjPT0dXl5e+N///veNVZNte35aIfHaZ+EstHd0wd1HT9CmpXQ33JpGKCeHOmq1WcfgpcADIRg8aACGOg0EACyeOxuRV68j5Fgo5rj/wDgdt1Ftq0aXdrbo0s6WdQzZJ6NXE3t6esLDw0OiTUlJ6avLTZ8+Hbdu3UJkpPRHgpmNDKqqqkJVVRVhYWFfHB6tV68eHBwcsG/fPgBAdnY2Dh8+DDc3t1LnDwgIwMiRI6GgoICRI0ciICDgiznS09Nx8OBBAICCggIAIDg4GD169JDoCP5DQUFBYkTy8OHDaN68OZo3b47Ro0cjMDAQIpHkowHu3bsn7pzKydWca3Y+fCw+wVVDrTpvOc9PCakZ6LLmZ/RcexBzDp/Hq/eZrCPxQn5BAe4+eIiOdpI/QjvY2SL21m1GqfiBaktI6ZSUlKCuri4xfa0z6O7ujvDwcFy8eBENGzaUeiZmPRN5eXkEBQVh37590NTURIcOHbBo0SLcunWrxLxubm4ICgqCSCTCsWPH0LRpU1hbW5eYLzMzE8ePH8fo0aMBAKNHj8axY8eQmSn5hzMjIwOqqqpQUVGBlpYWDh06hIEDB8LU1BQA8PjxY/G/vyYgIEC8vt69e+Pjx4/4/fffxe/n5eVh5MiR+Omnn9CoUaNyfSYfiEQi+G7bg1aW5jAxMmQdh9OsDOrCZ0hX7B7XF8sdOyPlQzZG7QpDejYdfq+stPR0FBYWQkdH8vwbXR0tvEt9X8ZSpDyotoRUnkgkwvTp0xEaGooLFy6gSZMmVbIe5ucMJiYmIjw8HA4ODoiIiMB3331X4qKQfv364ePHj7h8+TL27t1b5qjgwYMHYWRkhJYtWwIArK2tYWRkVOICFDU1NcTFxeHPP//Ejh070LRpU+zYsUP8vkgkgqAcw8sPHz5EdHQ0RowYAaC4gzt8+HCJcxU9PT1hZmYm7jCWV+knmOZ/fUEZsXLjDjx8+gLrls5nHYXzOps0Qi8LI5jo6aC9cUNsH9sHABAW+4hxMv74/P+7SMSH6xBlA9WWMCMQyOZUAdOmTcOBAwdw8OBBqKmpITk5GcnJycjJke7zs5kfs1RWVkbPnj3h5eWFqKgouLq6wtvbW2IeeXl5jBkzBt7e3rh+/TpcXFxK/ay9e/fi7t27kJeXF093794tcahYTk4OxsbGMDU1xeTJkzFmzBgMHz5c/L6JiQnu37//1ewBAQH49OkTGjRoIF7f9u3bERoairS0NADAhQsXcPToUfH73bt3BwDo6uqW+J7/VeoJppt3lDm/LFm5cQcuXLmO/RtWQ6+uLus4vFNbUQEm9bTxMjWDdRTO09LUhFAoREpKqkR76vs06OpI92q9moZqS0jlbd++HRkZGbC3t4e+vr54Onz4sFTXw7wz+Dlzc/NSb6bo5uaGS5cuYdCgQdDS0irx/u3bt3Hjxg1EREQgLi5OPF2+fBkxMTG4c+dOmeucPXs24uPjceLECQDAqFGjcP78ecTGxpaY99OnT8jKysKnT5+wf/9+rFu3TmJ98fHxaNy4MYKDgwEAx48fR3x8vPj9PXv2AAD++OMPTJs2rcxMnp6eyMjIkJg83ad8uXiMiUQirNiwHef+iEKQ/49oqK/HOhIv5X8qxLN36aijSheUVJaiggIsTJvjyvUYifao69GwsbJklIofqLaEVJ5IJCp1cnV1lep6mF1NnJqaiqFDh8LNzQ1WVlZQU1PDjRs34Ofnh0GDBpWY38zMDCkpKahdu/Q/gAEBAWjbtm2pN2Fs164dAgIC4O/vX+qy6urqmDhxIry9veHo6IhZs2bht99+Q/fu3bFy5Up07NhRnG/NmjUICAjAixcvkJaWhgkTJkBDQ0Pi84YMGYKAgABMnz4dTZs2lXgvJSVF/H00NTXLrI+SklKJE0pFWYplzi8LVmzYjpPnL2Hrj0ugUqs23qUWj46qqdaGcjmulCKl8/vfVXQ1bQx9DVWkZuVgZ8RNfMzLxyAbE9bReGH86JGYv3Q5WpibwsbKEodDw5CU/BdGDHFiHY3zqLZVIys7Bwmv34hfv05Kwv1HT6Chrob6evUYJpM1dEJCeTHrDKqqqsLW1hb+/v54+vQpCgoKYGBggEmTJmHRokWlLqOjo1Nqe35+Pg4cOIAFCxaU+v7gwYPh4+ODNWvWlJln5syZ2LRpE44ePYphw4bh3Llz8Pf3x86dOzF37lzUrl0bZmZmmDFjBlq0aIGlS5eiR48eJTqC/6xv9erVuHnzJr777rtyVIMfQn45BQAYO8tTon31gllw7tODRSRe+CszC3OP/I607Fxo11ZGS4N6CJnshAZaaqyj8ULfXj2Qlp6Bbbv34m1KKkyaGmHXpnVooK/POhrnUW2rxp0HDzF2+r+3JvHZtB0A4NTXAb5LSv87SMiXCESf3weFyCxR0mPWEXipKPJX1hF4S9jHlXUEQiouT7on55P/0GlQbasSvX5QbeuqCEHD8t2tpDrRs4kJIYQQwj8yetNpWSRzF5AQQgghhJDqQ51BQgghhJAajA4TE0IIIYR/6ChxudHIICGEEEJIDUadQUIIIYSQGowOExNCCCGEh+g4cXlRZ5BDBKolH8NHKk+uizPrCPwlJ2SdgJCKU6rFOgEh1YoOExNCCCGE1GA0MkgIIYQQ/qGbTpcbjQwSQgghhNRg1BkkhBBCCKnB6DAxIYQQQviHDhOXG40MEkIIIYTUYLzuDL59+xaTJ09Go0aNoKSkBD09PTg4OMDHxwcCgeCLU1BQECIiIiAQCJCeni7+zMTERLRo0QIdO3YUt8+cOROtWrWCkpISrK2tS+TIzc2Fq6srLC0tIS8vD0dHx2r5/iwEHw1Ft4FDYNm+K5xHu+FGbBzrSLwQE3cbUxZ4oZPjSJh2csD5y1GsI/FK8JFj6NZvECxtO8J51FjcuBnLOhJvUG2rTvCR4+g2wBmW7brA2cWV9rfkm/G6Mzh48GDEx8dj3759ePToEcLDw2Fvbw9zc3MkJSWJp2HDhqF3794SbcOHDy/xeU+fPkXHjh3RqFEjnD17FpqamgAAkUgENze3UpcBgMLCQtSqVQszZsxAjx49qvIrM3Xq7Hn4rNuIqW5jERYciFY2Vpg0Yy4Sk5NZR+O8nNxcmBobYensaayj8M6pM+fg89N6TJ0wHmEhP6OVjTUmTZ+FxCTabiuLalt1ive3GzDVzRVhB/ehlU1LTHL3oNpKEMjoJHsEIpFIxDpEVUhPT4eWlhYiIiLQpUuXL87r6uqK9PR0hIWFSbRHRESga9euSEtLQ0JCAhwcHGBvb4/9+/dDQUGhxOcsW7YMYWFhiIuLq/C6yuVDSsWXqUZDx02CuakJlnvOE7f1GTIKPew7Yc70qQyTfZko5yPrCBVi2skBW370Ro/O7VlH+Sou3Ch96JjxMDdtjuWLF4rb+jgPQw/7LpgzgzrflcHZ2hYVsk7wVUPHTiiu7aL54rY+g0egh31nzHH/gWGyr1DVrrZViZKfVtu6KkKg15R1hBJ4OzKoqqoKVVVVhIWFIS8vr1KfFRUVhS5dusDZ2RnBwcGldgRruvyCAtx98BAd7dpKtHewa4vYW3cYpSLky/ILCnD3/gN0bGcr0d7Bzhax8bcYpeIHqm3VKXt/a4vYW7cZpSJcxtvOoLy8PIKCgrBv3z5oamqiQ4cOWLRoEW7dqvhOyMnJCQMGDMDWrVshJ1c9JcvLy0NmZqbEVNlObVVKS09HYWEhdLQlf/XpamvhXUoqo1SEfFla2j/brY5Eu66ONt6l0nZbGVTbqiPe3+p8tr/V0cK71PeMUskggUA2JxnE284gUHzOYGJiIsLDw+Hg4ICIiAh89913CAoKqtDnDBo0CCdOnMAff/xRNUFL4ePjAw0NDYnJZ93Galv/txJ8tqGLRCXbCJE1n2+iIpGItlspodpWnVL3t4yyEG7jdWcQAJSVldGzZ094eXkhKioKrq6u8Pb2rtBn7Ny5EyNHjkSfPn1w6dKlKkoqydPTExkZGRKT55yZ1bLub6GlqQmhUIiUz37xp6alQVen+s4RIaQitLTK2G7fp0FXm7bbyqDaVh3x/jallNrS/pZ8A953Bj9nbm6OrKysCi0jEAiwc+dOjBkzBn379kVERETVhPsPJSUlqKurS0xKSkpVvt5vpaigAAvT5rhyPUaiPep6DGysWjBKRciXKSoowMLMFFeuRUu0R12Lhk1LK0ap+IFqW3XK3t9Gw8bKklEqWcT6qmHuXE3M2yeQpKamYujQoXBzc4OVlRXU1NRw48YN+Pn5YdCgQRX+PIFAgG3btkEoFKJfv3749ddf0a1bNwDAkydP8PHjRyQnJyMnJ0d8NbG5uTkUFRUBAPfu3UN+fj7ev3+PDx8+iOcp7b6EXDXeZTjme61ECzNT2Fi1wOHQX5CU/BdGDHZiHY3zsrJzkPAmUfz6dVIy7j9+Cg11NdSvV5dhMu4bP3oU5i/xRgtzM9hYWeJw6AkkJSdjxBBn1tE4j2pbdcaPHon5S5ejhbnp37UNK97fDqH9Lak43nYGVVVVYWtrC39/fzx9+hQFBQUwMDDApEmTsGjRom/6TIFAgC1btkAoFKJ///4IDw9Hjx49MHHiRInDxzY2NgCA58+fw9DQEADQt29fvHz5ssQ8fLqzT99ePZCWkYltewLxNiUVJk2NsGvjWjTQ12MdjfPuPHyEcTP+vYWE75adAADH3j3hu3guq1i80NehJ9IyMrBtVwDepqTAxLgpdm32R4P6+qyjcR7Vtur07dUDaekZ2LZ777/7203r0ECfaksqjrf3GeQlGb/PIFdx7T6DXMKF+wwSUgIH7jPIWdV5n8G3L6ptXRUhqGvIOkIJNe6cQUIIIYQQ8i/qDBJCCCGE1GC8PWeQEEIIITUY3c+y3GhkkBBCCCGkBqPOICGEEEJIDUaHiQkhhBDCQ3SYuLxoZJAQQgghpAajkUEOEX14zzoCLxVdCmMdgbeE/SewjsBfcvRbvsoU5LNOQEi1os4gIYQQQnhHQFcTlxv9tCSEEEIIqcGoM0gIIYQQUoPRYWJCCCGE8A8dJi43GhkkhBBCCKnBqDNICCGEEFKD8bYzKBAIvji5urqK5wsLCyux/Pfffw+hUIhDhw6VeG/ZsmUQCATo3bt3iff8/PwgEAhgb28vbtu9ezc6deoELS0taGlpoUePHoiOjpbWV5UZO4OPYsiU2fiu7zC0dxqNaUtW4VnCa9axOG/LxZswXxYgMXX66SDrWLwSfPQ4ug0cDMv29nAePR43YuNYR+KN4CPH0K2/EyztOsN51DjcuBnHOhLnxcTGY8qchejY3xnN7brg/KU/WEeSUQIZnWQPbzuDSUlJ4mnDhg1QV1eXaNu4cWOZy2ZnZ+Pw4cOYN28eAgICSp1HX18fFy9exOvXkp2dwMBANGrUSKItIiICI0eOxMWLF3H16lU0atQIvXr1wps3byr/RWVITPwdjHLsh8Nbf8Len1biU2EhJs73QnZOLutonGdcRxOX5owUT7/84MQ6Em+cOnsePus2YqrbOIQFB6GVTUtMmjEHicnJrKNx3qkz5+CzdgOmTnBF2MF9aGVjjUnus5GYRLWtjOycHDRvZgyvObNYRyE8wdvOoJ6ennjS0NCAQCAo0VaWo0ePwtzcHJ6enrhy5QpevHhRYp66deuiV69e2Ldvn7gtKioKKSkp6Nevn8S8wcHB+OGHH2BtbQ1TU1Ps3r0bRUVF+P3336X2fWXBHr/lcO7dA82aNIapcRP4LJiFxL/e4e6jJ6yjcZ5QTg511GqLJ22VWqwj8UZg8CEMHjQAQx0HomkTQyyeMwt69eoi5NgJ1tE4LzA4BIMdB2Co0yA0NWqCxfNm/13bUNbROK1LezvMnjIRvbp2Zh2F8ARvO4OVERAQgNGjR0NDQwN9+/ZFYGBgqfO5ubkhKChI/Hrv3r1wcXGBoqLiFz8/OzsbBQUF0NbWlmZsmfMhKwsAoKGuxjgJ9yW8z0SXtSHoueEw5hy9gFfvM1lH4oX8ggLcffAQHe3aSrR3sGuL2Fu3GaXih/yCAty9/xAd7Wwl2ju0s0VsPNWWVAOBQDYnGUSdwc88fvwY165dw/DhwwEAo0ePRmBgIIqKikrM279/f2RmZuLy5cvIysrCkSNH4Obm9tV1LFy4EA0aNECPHj3KnCcvLw+ZmZkSU14edx6RJBKJ4LstAK0szWHSpDHrOJxm1bAOfJw6Y/cYBywf0BEpH3MwKuAk0rPp8HtlpaWno7CwEDqf/TDT1dbGuxR6/GNliGurU0ptU1MZpSKElIY6g58JCAiAg4MDdHV1AQB9+/ZFVlYWzp8/X2JeBQUFcWfx6NGjMDExgZWV1Rc/38/PDyEhIQgNDYWysnKZ8/n4+EBDQ0Ni8tmys3Jfrhqt3LgDD5++wLql81hH4bzOzQzQy7wJTOppo33TBtju0gsAEBb3mHEy/vj8x7pIJJLVH/CcI/jshPni2lJxCZEldNPp/ygsLMT+/fuRnJwMeXl5ifaAgAD06tWrxDJubm6wtbXFnTt3vjoquHbtWqxevRrnz5//aqfR09MTHh4eEm2KqQkV+DbsrNy0ExeionFgow/06uiyjsM7tRUVYFJPCy/pUHGlaWlqQigUIiVVchQwNS0Nujr8Po2jqv1bW8lRwNS0NOjy/BQZIiPoR0e5UWfwP06dOoUPHz4gNjYWQqFQ3P7gwQO4uLggNTUVOjo6EstYWFjAwsICt27dwqhRo8r87J9++gmrVq3CmTNn0Lp1669mUVJSgpKSkkSb6OOXz0VkTSQSYeWmnTgfeRX7/X3QUF+PdSReyv9UiGfv0tGqEdW3shQVFGBh2hxXrkejZ9cu4vao6zHo3qUTw2Tcp6igAAuzv2vbzV7cHnUtGt3t6cIHQmQJdQb/IyAgAP369UPLli0l2i0sLDBr1iwcOHAAM2fOLLHchQsXUFBQAE1NzVI/18/PD0uXLsXBgwdhaGiI5L9vWaGqqgpVVVWpfw9WVmzYjpO/X8bWVYuhUrsW3r1PAwCoqdSG8mcdW1J+fmeuo2vzRtDXUEVqVg52Xo7Dx7wCDLI2Zh2NF8a7jMB8rxVoYWYGG6sWOBz6C5KS/8KIwY6so3HeeJeRmL90eSm1pVsjVUZWdjYSXv97a7LXiUm4/+gxNNTVUV+vHsNkhKuoM/i3v/76C7/99hsOHix5M1+BQABnZ2cEBASU2hlUUVH54mdv27YN+fn5GDJkiES7t7c3li1bVqncsiQk/H8AgLGzF0m0r14wE869y75YhnzZX5lZmHssAmnZudBWUUbLhnURMnEAGmjSVdrS0LdXD6RlZGDbnr14m5IKk6ZG2LVxLRro67OOxnl9HXoW13Z3wL+13bQeDepTbSvjzv2HGDttlvi1z8atAACnvr3h6+XJKJUsosPE5SUQiUQi1iFI+YgSH7GOwEtFl8JYR+AtYf8JrCPwlxxd/1dl8vNYJ+AvrWo8vSVNRm9uXp01KCfamxBCCCGE1GB0mJgQQggh/ENXE5cbjQwSQgghhNRg1BkkhBBCCKnB6DAxIYQQQviHjhKXG40MEkIIIYTUYNQZJIQQQgipweg+g1zyIYV1Al4S5WaxjsBbAhVN1hEIqbiiQtYJ+Eu1Gp9LnfG2+tZVERp1WScogUYGCSGEEEJqMOoMEkIIIYTUYHQ1MSGEEEL4h246XW40MkgIIYQQUoNRZ5AQQgghpAajw8SEEEII4R86TFxuNXZkUCAQfHFydXUtMZ+amhpat26N0NBQ8ecEBQWVunxubq54nu3bt8PKygrq6upQV1dHu3bt8L///a+6v3K1CD4aim4Dh8CyfVc4j3bDjdg41pF4ISbuFqbMX4pOg0bAtGMvnL98hXUkXgk+cgzd+g2CpW1HOI8aixs3Y1lH4g2qbdUJPnIc3QY4w7JdFzi7uNL+lnyzGtsZTEpKEk8bNmyAurq6RNvGjRvF8wYGBiIpKQkxMTFo2bIlhg4diqtXr4rf/3zZpKQkKCsri99v2LAhfH19cePGDdy4cQPdunXDoEGDcPfu3Wr9zlXt1Nnz8Fm3EVPdxiIsOBCtbKwwacZcJCYns47GeTk5uTA1NsJSj+mso/DOqTPn4PPTekydMB5hIT+jlY01Jk2fhcQk2m4ri2pbdYr3txsw1c0VYQf3oZVNS0xy96Dakm9SYzuDenp64klDQwMCgaBE2z80NTWhp6cHU1NT7NixA8rKyggPDxe///myenp6EusaMGAA+vbtCxMTE5iYmODHH3+Eqqoqrl27Vm3ftzoEBh/G4EH9MdRxIJo2McTiObOgV68uQo6dYB2N8zq3a4tZ349Hry4dWUfhncADBzHYcSCGOjuiqVETLJ7nAT29egg5epx1NM6j2ladwAMhGDxoAIY6/b2/nTv77/1t6NcXrjEEMjrJnhrbGfxWCgoKkJeXR0FBgbjt48ePaNy4MRo2bIj+/fsjNrbswyCFhYU4dOgQsrKy0K5du+qIXC3yCwpw98FDdLRrK9Hewa4tYm/dYZSKkC/LLyjA3fsP0LGdrUR7BztbxMbfYpSKH6i2Vafs/a0tYm/dZpSKcBldQFIBeXl5+Omnn5CZmYnu3bsDAExNTREUFARLS0tkZmZi48aN6NChA+Lj49GsWTPxsrdv30a7du2Qm5sLVVVVnDhxAubm5l9cV15enkSbUn4elJSUqubLVVJaejoKCwuhoy35qCFdbS28S0lllIqQL0tL+2e71ZFo19XRxrtU2m4rg2pbdcT7W53P9rc6WniX+p5RKsJlNDJYDiNHjoSqqipq166N9evXY+3atejTpw8AwM7ODqNHj0bLli3RqVMnHDlyBCYmJti8ebPEZzRv3hxxcXG4du0apk6dinHjxuHevXtlrtPHxwcaGhoSk8+6jWXOLysEn129JRKVbCNE1ny+iYpEItpupYRqW3VK3d8yyiKTBALZnGQQjQyWg7+/P3r06AF1dXXUrfvlB0zLycmhTZs2ePz4sUS7oqIijI2NAQCtW7dGTEwMNm7ciJ07d5b6OZ6envDw8JBoU8r/UIlvUbW0NDUhFAqR8tkv/tS0NOjqVOODyQmpAC2tMrbb92nQ1abttjKotlVHvL9NKaW2tL8l34BGBstBT08PxsbGX+0IAsW/euPi4qCvr//V+T4/DPxfSkpK4lvR/DPJ6iFiAFBUUICFaXNcuR4j0R51PQY2Vi0YpSLkyxQVFGBhZoor16Il2qOuRcOmpRWjVPxAta06Ze9vo2FjZckoFeEyGhmspOXLl8POzg7NmjVDZmYmNm3ahLi4OGzdulU8z6JFi9CnTx8YGBjgw4cPOHToECIiInD69GmGyaVvvMtwzPdaiRZmprCxaoHDob8gKfkvjBjsxDoa52Vl5yDhTaL49eukZNx//BQaamqor/f1HymkbONHj8L8Jd5oYW4GGytLHA49gaTkZIwY4sw6GudRbavO+NEjMX/pcrQwN/27tmHF+9shtL8Vk9FDsrKIOoOVlJ6eju+//x7JycnQ0NCAjY0NLl++jLZt/73K66+//sKYMWOQlJQEDQ0NWFlZ4fTp0+jZsyfD5NLXt1cPpGVkYtueQLxNSYVJUyPs2rgWDfT1vr4w+aI7Dx5h3Ix54te+m4tPL3Ds0xO+i+eVtRgph74OPZGWkYFtuwLwNiUFJsZNsWuzPxrU//LoPvk6qm3V6durB9LSM7Bt995/97eb1qHBV45KEVIagUgkErEOQcrpQwrrBLwkys1iHYG3BCqarCMQUnFFhawT8JdqNZ7TmJVWfeuqCBUt1glKoJFBQgghhPAQHSYuL7qAhBBCCCGkBqPOICGEEEJIDUaHiQkhhBDCP3Q1cbnRyCAhhBBCSA1GnUFCCCGEkBqMbi1DpC4vLw8+Pj7w9PSU6aemcBHVtupQbasG1bXqUG2JtFBnkEhdZmYmNDQ0kJGRAXV1ddZxeIVqW3WotlWD6lp1qLZEWugwMSGEEEJIDUadQUIIIYSQGow6g4QQQgghNRh1BonUKSkpwdvbm05orgJU26pDta0aVNeqQ7Ul0kIXkBBCCCGE1GA0MkgIIYQQUoNRZ5AQQgghpAajziAhhBBCSA1GnUFCCCGEkBqMOoOEEEIIITUYdQYJIYRUmdzcXNYRCCFfQZ1BIhWnT59GZGSk+PXWrVthbW2NUaNGIS0tjWEyQkh1KyoqwsqVK9GgQQOoqqri2bNnAIClS5ciICCAcTpuMzIyQmpqaon29PR0GBkZMUhE+IA6g0Qq5s2bh8zMTADA7du3MWfOHPTt2xfPnj2Dh4cH43Tc98cff2D06NFo164d3rx5AwD4+eefJTrg5Ns9efIEZ86cQU5ODgCAbr9aOatWrUJQUBD8/PygqKgobre0tMSePXsYJuO+Fy9eoLCwsER7Xl6eeN9ASEXJsw5A+OH58+cwNzcHABw/fhz9+/fH6tWrcfPmTfTt25dxOm47fvw4xowZAxcXF8TGxiIvLw8A8OHDB6xevRqnTp1inJC7UlNTMXz4cFy4cAECgQCPHz+GkZERJk6cCE1NTaxbt451RE7av38/du3ahe7du2PKlCnidisrKzx48IBhMu4KDw8X//vMmTPQ0NAQvy4sLMTvv/8OQ0NDBskIH1BnkEiFoqIisrOzAQDnz5/H2LFjAQDa2triEUPybVatWoUdO3Zg7NixOHTokLi9ffv2WLFiBcNk3Dd79mzIy8sjISEBZmZm4vbhw4dj9uzZ1Bn8Rm/evIGxsXGJ9qKiIhQUFDBIxH2Ojo4AAIFAgHHjxkm8p6CgAENDQ9peyTejziCRio4dO8LDwwMdOnRAdHQ0Dh8+DAB49OgRGjZsyDgdtz18+BCdO3cu0a6uro709PTqD8QjZ8+exZkzZ0pso82aNcPLly8ZpeI+CwsL/PHHH2jcuLFE+9GjR2FjY8MoFbcVFRUBAJo0aYKYmBjo6uoyTkT4hDqDRCq2bNmCH374AceOHcP27dvRoEEDAMD//vc/9O7dm3E6btPX18eTJ09KHAKKjIykE8YrKSsrC7Vr1y7RnpKSAiUlJQaJ+MHb2xtjxozBmzdvUFRUhNDQUDx8+BD79+/HyZMnWcfjtOfPn7OOQHhIIKIzpQmRaX5+fti3bx/27t2Lnj174tSpU3j58iVmz54NLy8vTJ8+nXVEzurXrx++++47rFy5Empqarh16xYaN26MESNGoKioCMeOHWMdkbPOnDmD1atX488//0RRURG+++47eHl5oVevXqyjcd7vv/+O33//HW/fvhWPGP5j7969jFIRLqPOIJGKhISEL77fqFGjakrCT4sXL4a/v7/4nm1KSkqYO3cuVq5cyTgZt927dw/29vZo1aoVLly4gIEDB+Lu3bt4//49rly5gqZNm7KOSIiE5cuXY8WKFWjdujX09fUhEAgk3j9x4gSjZITLqDNIpEJOTq7ETum/SrsVAqmY7Oxs3Lt3D0VFRTA3N4eqqirrSLyQnJyM7du3S4xgTZs2Dfr6+qyjcVZMTAyKiopga2sr0X79+nUIhUK0bt2aUTLu09fXh5+fH8aMGcM6CuER6gwSqYiPj5d4XVBQgNjYWKxfvx4//vgjnJ2dGSXjn8zMTFy4cAHNmzeXuAKWEFnRtm1bzJ8/H0OGDJFoDw0NxZo1a3D9+nVGybhPR0cH0dHRNGpNpIo6g6RK/fbbb/jpp58QERHBOgpnDRs2DJ07d8b06dORk5MDa2trPH/+HCKRCIcOHcLgwYNZR+Ss06dPQ1VVFR07dgRQ/OSc3bt3w9zcHFu3boWWlhbjhNykqqqKW7dulbjA6fnz57CyssKHDx8YJeO+BQsWQFVVFUuXLmUdhfAIPYGEVCkTExPExMSwjsFply9fRqdOnQAUnw9UVFSE9PR0bNq0CatWrWKcjts+f3KOh4cHPTlHCpSUlPDXX3+VaE9KSoK8PN3EojJyc3Oxfv16dOnSBe7u7vDw8JCYCPkWNDJIpOLzG0uLRCIkJSVh2bJlePDgAeLi4tgE44FatWrh0aNHMDAwwNixY1G/fn34+voiISEB5ubm+PjxI+uInKWqqoo7d+7A0NAQy5Ytw507d3Ds2DHxk3OSk5NZR+SkESNGIDk5Gb/88ov4SRnp6elwdHRE3bp1ceTIEcYJuatr165lvicQCHDhwoVqTEP4gn6iEanQ1NQscQGJSCSCgYGBxFMzSMUZGBjg6tWr0NbWxunTp8X1TEtLg7KyMuN03EZPzqka69atQ+fOndG4cWPxTabj4uJQr149/Pzzz4zTcdvFixdZRyA8RJ1BIhWf76Dk5ORQp04dGBsb02GhSpo1axZcXFygqqqKxo0bw97eHkDx4WNLS0u24TiOnpxTNRo0aIBbt24hODgY8fHxqFWrFsaPH4+RI0dCQUGBdTxCyGfoMDEhHHDjxg28evUKPXv2FN9S5rfffoOmpiY6dOjAOB13JSQk4IcffsCrV68wY8YMTJgwAUDxM4sLCwuxadMmxgkJKSkmJgZHjx5FQkIC8vPzJd4LDQ1llIpwGXUGidQ8ffoUGzZswP379yEQCGBmZoaZM2fSLRAIqQHCw8PRp08fKCgoIDw8/IvzDhw4sJpS8c+hQ4cwduxY9OrVC+fOnUOvXr3w+PFjJCcnw8nJCYGBgawjEg6iziCRijNnzmDgwIGwtrZGhw4dIBKJEBUVhfj4ePz666/o2bMn64icVVhYiKCgoDIfP0UnjFdOUVERnjx5UmptO3fuzCgV98jJySE5ORl169aFnFzZN6oQCAR0E/pKsLKywuTJkzFt2jSoqakhPj4eTZo0weTJk6Gvr4/ly5ezjkg4iDqDRCpsbGzg4OAAX19fifaFCxfi7NmzuHnzJqNk3Dd9+nQEBQWhX79+pT5+yt/fn1Ey7rt27RpGjRqFly9f4vNdIXVaiCxSUVHB3bt3YWhoCF1dXVy8eBGWlpa4f/8+unXrhqSkJNYRCQfRmf1EKu7fv1/q7SLc3NywYcOG6g/EI4cOHcKRI0fQt29f1lF4Z8qUKWjdujV+++23UjvahMgabW1t8U27GzRogDt37sDS0hLp6eniK+MJqSjqDBKpqFOnDuLi4tCsWTOJ9ri4ONStW5dRKn5QVFSEsbEx6xi89PjxYxw7dozqKwUVudhmxowZVZiE3zp16oRz587B0tISw4YNw8yZM3HhwgWcO3cO3bt3Zx2PcBR1BolUTJo0Cd9//z2ePXuG9u3bQyAQIDIyEmvWrMGcOXNYx+O0OXPmYOPGjdiyZQuNXEmZra0tnjx5Qp1BKSjv6QoCgYA6g5WwZcsW5ObmAgA8PT2hoKCAyMhIODs70yPqyDejcwaJVIhEImzYsAHr1q1DYmIiAKB+/fqYN28eZsyYQSvoxdwAACNqSURBVJ2YSnBycsLFixehra0NCwuLEvdpo1tJfLsTJ05gyZIlmDdvHiwtLUvU1srKilEyQgipPtQZJFL3z/ksampqjJPww/jx47/4Pt1K4tuVdtWrQCCASCSiC0ik5J8/MfSD8NtlZmZCXV1d/O8v+Wc+QiqCOoOEkBrr5cuXX3y/cePG1ZSEfwICAuDv74/Hjx8DAJo1a4ZZs2Zh4sSJjJNxj1AoRFJSkvi2PaV1rOkHDKkMOmeQSEVqaiq8vLxw8eLFUu/X9v79e0bJ+OHTp0+IiIjA06dPMWrUKKipqSExMRHq6uriJ5KQiqPOXtVYunQp/P394e7ujnbt2gEArl69itmzZ+PFixdYtWoV44TccuHCBWhrawOgZxOTqkEjg0Qq+vTpg6dPn2LChAmoV69eiV+u48aNY5SM+16+fInevXsjISEBeXl5ePToEYyMjDBr1izk5uZix44drCNy2s8//4wdO3bg+fPnuHr1Kho3bowNGzagSZMmGDRoEOt4nKSrq4vNmzdj5MiREu0hISFwd3dHSkoKo2Tc9unTJ/z4449wc3ODgYEB6ziER2hkkEhFZGQkIiMj0bJlS9ZReGfmzJlo3bo14uPjoaOjI253cnKiQ26VtH37dnh5eWHWrFn48ccfxYfYNDU1sWHDBuoMfqPCwkK0bt26RHurVq3w6dMnBon4QV5eHmvXrqUf10Tqyn5mECEVYGpqipycHNYxeCkyMhJLliyBoqKiRHvjxo3x5s0bRqn4YfPmzdi9ezcWL14MoVAobm/dujVu377NMBm3jR49Gtu3by/RvmvXLri4uDBIxB/du3dHREQE6xiEZ2hkkEjFtm3bsHDhQnh5eaFFixYlbtFBV7h9u6KiolJPCn/9+jVdsV1Jz58/h42NTYl2JSUlZGVlMUjEHwEBATh79izs7OwAFD/679WrVxg7diw8PDzE861fv55VRE7q06cPPD09cefOHbRq1QoqKioS7w8cOJBRMsJl1BkkUqGpqYmMjAx069ZNop2ucKu8nj17YsOGDdi1axeA4lt0fPz4Ed7e3vSIukpq0qQJ4uLiSlxI8r///Q/m5uaMUnHfnTt38N133wEAnj59CqD4KUV16tTBnTt3xPPR7WYqburUqQBK70TTvpZ8K+oMEqlwcXGBoqIiDh48WOoFJOTb+fv7o2vXrjA3N0dubi5GjRqFx48fQ1dXFyEhIazjcdq8efMwbdo05ObmQiQSITo6GiEhIfDx8cGePXtYx+MsuuK16nx+pwZCpIGuJiZSUbt2bcTGxqJ58+aso/BSTk4ODh06hD///BNFRUX47rvv4OLiglq1arGOxnm7d+/GqlWr8OrVKwBAgwYNsGzZMkyYMIFxMv7IzMzEhQsXYGpqClNTU9ZxCCGfoc4gkYrOnTvDy8sLPXr0YB2FkG+SkpKCoqIi1K1bl3UUzhs2bBg6d+6M6dOnIycnBy1btsSLFy8gEolw6NAhDB48mHVETsvKysKlS5eQkJCA/Px8iffouc/kW1BnkEjF0aNHsWzZMnrGaxXYt28fdHV10a9fPwDA/PnzsWvXLpibmyMkJIRunFwJOTk5EIlEqF27NoDiezqeOHEC5ubm6NWrF+N03KWnp4czZ86gZcuWOHjwILy9vREfH499+/Zh165diI2NZR2Rs2JjY9G3b19kZ2cjKysL2traSElJQe3atVG3bl08e/aMdUTCQdQZJFJBz3itOs2bN8f27dvRrVs3XL16Fd27d8eGDRtw8uRJyMvLIzQ0lHVEzurVqxecnZ0xZcoUpKeno3nz5lBUVERKSgrWr18vPlmfVEytWrXw6NEjGBgYYOzYsahfvz58fX2RkJAAc3NzfPz4kXVEzrK3t4eJiQm2b98OTU1NxMfHQ0FBAaNHj8bMmTPh7OzMOiLhILqAhEjF8+fPWUfgrVevXsHY2BgAEBYWhiFDhuD7779Hhw4dYG9vzzYcx928eRP+/v4AgGPHjkFPTw+xsbE4fvw4vLy8qDP4jQwMDHD16lVoa2vj9OnTOHToEAAgLS0NysrKjNNxW1xcHHbu3AmhUAihUIi8vDwYGRnBz88P48aNo84g+SbUGSRSoaurW+J+V0Q6VFVVkZqaikaNGuHs2bOYPXs2AEBZWZlu9F1J2dnZ4ns1nj17Fs7OzpCTk4OdnR1evnzJOB13zZo1Cy4uLlBVVUXjxo3FP1ouX74MS0tLtuE4TkFBQXy3hnr16iEhIQFmZmbQ0NBAQkIC43SEq+gJJEQq6tWrBzc3N0RGRrKOwjs9e/bExIkTMXHiRDx69Eh87uDdu3dhaGjINhzHGRsbIywsDK9evcKZM2fE5wm+ffuWbpReCT/88AOuXbuGvXv3IjIyUnwaiZGREVatWsU4HbfZ2Njgxo0bAICuXbvCy8sLwcHBmDVrFnW0yTejziCRipCQEGRkZKB79+4wMTGBr68vEhMTWcfiha1bt6Jdu3Z49+4djh8/Ln4+8Z9//omRI0cyTsdtXl5emDt3LgwNDWFra4t27doBKB4lLO3JJKT8WrVqBScnJ6iqqorb+vXrhw4dOohfq6ur0wUPFbR69Wro6+sDAFauXAkdHR1MnToVb9++Fd+YnpCKogtIiFSlpqZi//79CAoKwr179+Dg4AA3NzcMHDgQ8vJ0VgKRPcnJyUhKSkLLli3FI1jR0dFQV1ene+JVMTU1NcTHx8PIyIh1FEJqNBoZJFKlo6OD2bNnIz4+HuvXr8f58+cxZMgQ1K9fH15eXsjOzmYdkXNOnz4tcfh969atsLa2xqhRo5CWlsYwGT/o6enBxsZG4or4tm3bUkeQyKTly5eLH/FHiLRQZ5BIVXJyMvz8/GBmZoaFCxdiyJAh+P333+Hv748TJ07A0dGRdUTOmTdvHjIzMwEAt2/fxpw5c9C3b188e/YMHh4ejNNxW1ZWFpYuXYr27dvD2NgYRkZGEhMhsub48eMwMTGBnZ0dtmzZgnfv3rGORHiAjtsRqQgNDUVgYCDOnDkDc3NzTJs2DaNHj4ampqZ4HmtrazoP6xs8f/4c5ubmAIr/EPTv3x+rV6/GzZs30bdvX8bpuG3ixIm4dOkSxowZA319fXqmNpF5t27dwt27dxEcHIz169fDw8MDPXr0wOjRo+Ho6Ci+gTohFUHnDBKp0NDQwIgRIzBx4kS0adOm1HlycnLg5+cHb2/vak7Hbdra2oiMjIS5uTk6duyIsWPH4vvvv8eLFy9gbm5Oh94rQVNTE7/99pvERQ2k+qirqyMuLo5GYSvhypUrOHjwII4ePYrc3FzxUQRCKoJGBolUJCUlffUXaa1atagj+A06duwIDw8PdOjQAdHR0Th8+DAA4NGjR2jYsCHjdNympaUFbW1t1jFqLBqLqDwVFRXUqlULioqK+PDhA+s4hKNoZJBITVFREZ48eYK3b9+iqKhI4r3OnTszSsV9CQkJ+OGHH/Dq1SvMmDEDEyZMAADMnj0bhYWF2LRpE+OE3HXgwAH88ssv2LdvHx1eYyAyMhJt2rSBkpIS6yic8vz5cxw8eBDBwcF49OgROnfujFGjRmHo0KHQ0NBgHY9wEHUGiVRcu3YNo0aNwsuXL0v82qdnExNZZWNjg6dPn0IkEsHQ0BAKCgoS79+8eZNRMu6pyMVM69evr8Ik/NauXTtER0fD0tISLi4uGDVqFBo0aMA6FuE4OkxMpGLKlClo3bo1fvvtNzoRXwoyMzPFT8D42jlA9KSMb0dXt0tPbGxsueajfUPldO3aFXv27IGFhQXrKIRHaGSQSIWKigri4+NhbGzMOgovCIVCJCUloW7dupCTkyv1D6hIJKJRV0JIqejiHFIRNDJIpMLW1hZPnjyhzqCUXLhwQXxhw8WLFxmn4b8///wT9+/fh0AggLm5Od0CiXAejfOQiqDOIJEKd3d3zJkzB8nJybC0tCxx7pWVlRWjZNzUpUsXiX/n5ubi1q1bpV6cQ77d27dvMWLECEREREBTUxMikQgZGRno2rUrDh06hDp16rCOyFkxMTE4evQoEhISkJ+fL/FeaGgoo1SEkNLQYWIiFf99lNc/BAIBHcqUgtOnT2Ps2LFISUkp8R7VtnKGDx+Op0+f4ueff4aZmRkA4N69exg3bhyMjY0REhLCOCE3HTp0CGPHjkWvXr1w7tw59OrVC48fP0ZycjKcnJwQGBjIOiLv0XOfSUVQZ5BIxcuXL7/4fuPGjaspCf8YGxvDwcEBXl5eqFevHus4vKKhoYHz58+XuFF6dHQ0evXqhfT0dDbBOM7KygqTJ0/GtGnTxJ2SJk2aYPLkydDX18fy5ctZR+Q96gySiqDDxEQqqLNXdd6+fQsPDw/qCFaBoqKiEqc0AICCggIdjq+Ep0+fol+/fgAAJSUlZGVlQSAQYPbs2ejWrRt1BqsBXbVNKqLksT1CvtHTp0/h7u6OHj16oGfPnpgxYwaePn3KOhbnDRkyBBEREaxj8FK3bt0wc+ZMJCYmitvevHmD2bNno3v37gyTcZu2trb4aRgNGjTAnTt3AADp6en0+MRqQgf9SEXQyCCRijNnzmDgwIGwtrZGhw4dIBKJEBUVBQsLC/z666/o2bMn64ictWXLFgwdOhR//PFHqRfnzJgxg1Ey7tuyZQsGDRoEQ0NDGBgYQCAQ4OXLl7CyssLPP//MOh5nderUCefOnYOlpSWGDRuGmTNn4sKFCzh37hx1sqvJ//73P7oZNSk3OmeQSIWNjQ0cHBzg6+sr0b5w4UKcPXuWnuRQCXv27MGUKVNQq1Yt6OjoSBz+EQgEePbsGcN0/HD+/Hncv38fIpEI5ubm6NGjB+tInPb+/Xvk5uaifv36KCoqwtq1axEZGQljY2MsXboUWlparCNyVmFhIYKCgvD777+XeneBCxcuMEpGuIw6g0QqlJWVcfv2bTRr1kyi/dGjR7CyskJubi6jZNynp6eHGTNmYOHChaVetU0q5/fffy/zD+vevXsZpSKkdNOnT0dQUBD69etX6tOe/P39GSUjXEaHiYlU1KlTB3FxcSU6g3Fxcahbty6jVPyQn5+P4cOHU0ewCixfvhwrVqxA69at6TGKlUSPUKwehw4dwpEjR9C3b1/WUQiPUGeQSMWkSZPw/fff49mzZ2jfvj0EAgEiIyOxZs0azJkzh3U8Ths3bhwOHz6MRYsWsY7COzt27EBQUBDGjBnDOgrnaWlpiR+hqKmpSY9QrCKKior0pCciddQZJFKxdOlSqKmpYd26dfD09AQA1K9fH8uWLaMLHCqpsLAQfn5+OHPmDKysrEpcQLJ+/XpGybgvPz8f7du3Zx2DF+gRitVjzpw52LhxI7Zs2UIj2URq6JxBInX/3FJCTU2NcRJ+6Nq1a5nvCQQCOmG8EhYsWABVVVUsXbqUdRReSUhIEF+d/V8ikQivXr1Co0aNGCXjPicnJ1y8eBHa2tqwsLAo8eOQHvVHvgWNDBKpo06gdNEoS9XJzc3Frl27cP78eRp1laImTZqIDxn/1/v379GkSRM6TFwJmpqacHJyYh2D8Ax1BolU/PXXX5g7d674qszPB5xp509k0a1bt2BtbQ0A4hsj/4MOwX27f84N/NzHjx+hrKzMIBF/0HOdSVWgziCRCldXVyQkJGDp0qV0VSbhDBp1lS4PDw8AxR3ppUuXonbt2uL3CgsLcf36dXHnm1TOu3fv8PDhQwgEApiYmKBOnTqsIxEOo84gkYrIyEj88ccftKMnpAaLjY0FUDwyePv2bSgqKorfU1RURMuWLTF37lxW8XghKysL7u7u2L9/v/i+mEKhEGPHjsXmzZslOuCElBd1BolUGBgY0LMwCanh/hlpHT9+PDZu3Ej3E6wCHh4euHTpEn799Vd06NABQPGP8RkzZmDOnDnYvn0744SEi+hqYiIVZ8+exbp167Bz504YGhqyjkMIIbykq6uLY8eOwd7eXqL94sWLGDZsGN69e8cmGOE0GhkkUjF8+HBkZ2ejadOmqF27domrMt+/f88oGSGkumVlZcHX17fMx/zR87S/XXZ2NurVq1eivW7dusjOzmaQiPABdQaJVPj7+9NFI4QQAMDEiRNx6dIljBkzhi4ok7J27drB29sb+/fvF1+ZnZOTg+XLl6Ndu3aM0xGuosPEhBBCpEpTUxO//fab+Jw2Ij137txB7969kZubi5YtW0IgECAuLg7Kyso4c+YMLCwsWEckHESdQSIV9vb2cHNzw9ChQ1GrVi3WcQghDDVp0gSnTp2CmZkZ6yi8lJOTgwMHDuDBgwcQiUQwNzeHi4sL7XvJN6POIJGKOXPmIDg4GDk5ORg2bBgmTJgAOzs71rEIIQwcOHAAv/zyC/bt20e3OiGEA6gzSKSmsLAQJ0+eRGBgIE6dOgVjY2O4ublhzJgxpZ7wTAjhJxsbGzx9+hQikQiGhoYlLii7efMmo2TcFB4ejj59+kBBQQHh4eFfnHfgwIHVlIrwCXUGSZV49+4ddu7ciR9//BGFhYXo27cvZsyYgW7durGORgipYsuXL//i+97e3tWUhB/k5OSQnJyMunXrQk5Orsz5BAIBPfqTfBPqDBKpi46ORmBgIEJCQqChoQFXV1ckJSUhODgYU6dOxdq1a1lHJIQQQsjfqDNIpOLt27f4+eefERgYiMePH2PAgAGYOHEiHBwcxLeVOH/+PBwdHfHx40fGaQkhVS09PR3Hjh3D06dPMW/ePGhra+PmzZuoV68eGjRowDoeIeQ/6D6DRCoaNmyIpk2bws3NDa6urqU+NL1t27Zo06YNg3SEkOp069Yt9OjRAxoaGnjx4gUmTZoEbW1tnDhxAi9fvsT+/ftZR+SsTZs2ldouEAigrKwMY2NjdO7cGUKhsJqTES6jkUEiFZcvX0arVq2goqICAHj58iVOnDgBMzMzODg4ME5HCKlOPXr0wHfffQc/Pz+oqakhPj4eRkZGiIqKwqhRo/DixQvWETmrSZMmePfuHbKzs6GlpQWRSIT09HTUrl0bqqqqePv2LYyMjHDx4kUYGBiwjks4ouwzUQmpgFWrVuHnn38GUHx4qG3btli3bh0cHR3pwemE1DAxMTGYPHlyifYGDRogOTmZQSL+WL16Ndq0aYPHjx8jNTUV79+/x6NHj2Bra4uNGzciISEBenp6mD17NuuohEOoM0ik4ubNm+jUqRMA4NixY9DT0xMfDirrsAYhhJ+UlZWRmZlZov3hw4elnkJCym/JkiXw9/dH06ZNxW3GxsZYu3YtPD090bBhQ/j5+eHKlSsMUxKuoc4gkYrs7GyoqakBAM6ePQtnZ2fIycnBzs4OL1++ZJyOEFKdBg0ahBUrVqCgoABA8flsCQkJWLhwIQYPHsw4HbclJSXh06dPJdo/ffp/e/ceU3X9+HH8dVLk6oGWiZoXbiZiC5VZwpQyDFA2r5uUl81LahrTlmi6strUtJZldkHnCMUushTPHwoBmZdUyhhIDTVEvOCCMjecAQoC3z9c5/fjS32zOvrxzXk+trN53ufzkdfhD/fy/Xl/3p8bzlnXXr166erVq3c6GgxGGYRLhIWFyeFwqKqqSnl5eYqPj5d08y5ju91ucToAd9Jbb72lS5cuqXv37mpoaNBjjz2msLAwde3aVWvWrLE6ntFGjRql+fPnq6SkxDlWUlKiBQsWOPdx/eGHHxQcHGxVRBiIG0jgEjt37tTUqVPV3NysuLg45efnS5LWrl2rQ4cOKTc31+KEAO60r776SsXFxWppadHQoUM1evRoqyMZr6amRjNmzNC+ffucT3a5ceOG4uLitH37dgUGBmr//v1qampy/qcc+CuUQbhMTU2NqqurFRkZ6dwl/9ixY7Lb7QoPD7c4HQAr1dbWKiAgwOoYHcapU6dUXl6u1tZWhYeHa8CAAVZHgsEogwAAl3rjjTcUFBSk5ORkSdKUKVO0a9cu9ejRQzk5OYqMjLQ4ofkaGxt19uxZhYaGqnNntgzGv8OaQQCAS23evNm5x11BQYEKCgqUm5urMWPGaOnSpRanM1t9fb3mzJkjHx8fDRo0SBcuXJAkLVq0SOvWrbM4HUxFGQQAuFR1dbWzDO7Zs0dTpkxRfHy8li1bpu+++87idGZbsWKFSktLdeDAAXl5eTnHR48eraysLAuTwWSUQQCAS917772qqqqSJH3xxRfOG0daW1vV3NxsZTTjORwOvf/++xoxYoTzue+SFBERoTNnzliYDCZjoQEAwKUmTZqkqVOnqn///rp8+bLGjBkjSTp+/LjCwsIsTme237fs+W91dXVtyiHwdzAzCABwqXfeeUcpKSmKiIhQQUGB/Pz8JN28fLxw4UKL05lt2LBh2rt3r/P97wVwy5Ytio6OtioWDMfdxAAAGOLo0aNKTEzUtGnTtHXrVs2fP19lZWUqLCzUwYMHFRUVZXVEGIiZQQCAS23btq3N7NWyZcsUEBCgmJgYHk/5L8XExOjo0aOqr69XaGio8vPzFRgYqMLCQoog/jFmBgEALjVgwAClpaXpiSeeUGFhoeLi4rRhwwbt2bNHnTt3VnZ2ttURjdTU1KR58+Zp5cqVCgkJsToOOhDKIADApXx8fHTq1Cn17dtXL774oqqrq5WZmamysjI9/vjjunTpktURjRUQEKDi4mLKIFyKy8QAAJfy8/PT5cuXJUn5+fnOrWW8vLzU0NBgZTTjTZw4UQ6Hw+oY6GDYWgYA4FJPPvmknnnmGQ0ZMkTl5eVKSkqSJJWVlalfv34WpzNbWFiYVq1apaNHjyoqKkq+vr5tPl+0aJFFyWAyLhMDAFyqtrZWK1euVFVVlRYsWKCEhARJ0quvvqouXbropZdesjihuYKDg//0M5vNpsrKyjuYBh0FZRAA4HKHDh3S5s2bVVlZqZ07d+qBBx5QZmamQkJCNGLECKvjAfh/WDMIAHCpXbt2KTExUT4+PiopKdH169clSb/99ptef/11i9O5B7vdziwhbhllEADgUqtXr9amTZu0ZcsWeXh4OMdjYmJUXFxsYTL3wUU//B2UQQCAS/3444+KjY1tN26321VbW3vnAwH4nyiDAACX6tmzpyoqKtqNHz58mP3xgLsQZRAA4FLz58/X4sWL9e2338pms+mnn37SJ598otTUVC1cuNDqeAD+C/sMAgBcatmyZbpy5YpGjRqla9euKTY2Vp6enkpNTVVKSorV8dyCzWazOgIMwtYyAIDbor6+XidOnFBLS4siIiLk5+dndSS30bVrV5WWlnJZHreEMggAgGEaGxt19uxZhYaGqnPn9hf5Dh8+rGHDhsnT09OCdDANawYBADBEfX295syZIx8fHw0aNEgXLlyQdPMxdOvWrXMeN2LECIogbhllEAAAQ6xYsUKlpaU6cOCAvLy8nOOjR49WVlaWhclgMm4gAQDAEA6HQ1lZWRo+fHibm0QiIiJ05swZC5PBZMwMAgBgiEuXLql79+7txuvq6riDGP8YZRAAAEMMGzZMe/fudb7/vQBu2bJF0dHRVsWC4bhMDACAIdauXavExESdOHFCN27c0LvvvquysjIVFhbq4MGDVseDoZgZBADAEDExMTpy5Ijq6+sVGhqq/Px8BQYGqrCwUFFRUVbHg6HYZxAAAMCNMTMIAIAhcnJylJeX1248Ly9Pubm5FiRCR0AZBADAEMuXL1dzc3O78dbWVi1fvtyCROgIKIMAABji9OnTioiIaDceHh6uiooKCxKhI6AMAgBgCH9/f1VWVrYbr6iokK+vrwWJ0BFQBgEAMMS4ceP0/PPPt3naSEVFhZYsWaJx48ZZmAwm425iAAAMceXKFSUmJqqoqEi9e/eWJF28eFEjR45Udna2AgICrA0II1EGAQAwSGtrqwoKClRaWipvb289/PDDio2NtToWDEYZBAAAcGM8jg4AgLvYxo0bNW/ePHl5eWnjxo3/89hFixbdoVToSJgZBADgLhYcHKyioiLdd999Cg4O/tPjbDbbH95pDPwVyiAAAIAbY2sZAAAAN8aaQQAA7mIvvPDCLR/79ttv38Yk6KgogwAA3MVKSkpu6TibzXabk6CjYs0gAACAG2PNIAAABqqqqtLFixetjoEOgDIIAIAhbty4oZUrV8rf319BQUHq16+f/P399fLLL6upqcnqeDAUawYBADBESkqKdu/erTfffFPR0dGSpMLCQr322mv69ddftWnTJosTwkSsGQQAwBD+/v7asWOHxowZ02Y8NzdXTz31lK5cuWJRMpiMy8QAABjCy8tLQUFB7caDgoLUpUuXOx8IHQJlEAAAQzz33HNatWqVrl+/7hy7fv261qxZo5SUFAuTwWRcJgYAwBATJ07Uvn375OnpqcjISElSaWmpGhsbFRcX1+bY7OxsKyLCQNxAAgCAIQICAjR58uQ2Y3369LEoDToKZgYBADBEQ0ODWlpa5OvrK0k6d+6cHA6HBg4cqISEBIvTwVSsGQQAwBDjx4/X9u3bJUm1tbUaPny41q9frwkTJigtLc3idDAVZRAAAEMUFxdr5MiRkqSdO3cqMDBQ58+fV2ZmpjZu3GhxOpiKMggAgCHq6+vVtWtXSVJ+fr4mTZqke+65R8OHD9f58+ctTgdTUQYBADBEWFiYHA6HqqqqlJeXp/j4eEnSL7/8IrvdbnE6mIoyCACAIV555RWlpqYqKChIjz76qPORdPn5+RoyZIjF6WAq7iYGAMAgNTU1qq6uVmRkpO655+aczrFjx2S32xUeHm5xOpiIMggAAODGuEwMAADgxiiDAAAAbowyCAAA4MYogwAAAG6MMggAAODGKIMAjDBz5kxNmDDB+WebzaZnn3223XELFy6UzWbTzJkz25xrs9lks9nk4eGhkJAQpaamqq6urs258+bNU6dOnbRjx44/zFBRUaFZs2apd+/e8vT0VHBwsJ5++mkVFRVp69atzp/xZ68DBw646tcBAC5DGQRgpD59+mjHjh1qaGhwjl27dk2fffaZ+vbt2+74xMREVVdXq7KyUqtXr9aHH36o1NRU5+f19fXKysrS0qVLlZ6e3u78oqIiRUVFqby8XJs3b9aJEye0e/duhYeHa8mSJUpOTlZ1dbXzFR0drblz57YZi4mJuT2/DAD4FzpbHQAA/omhQ4eqsrJS2dnZmjZtmiQpOztbffr0UUhISLvjPT091aNHD0nS1KlTtX//fjkcDqWlpUmSPv/8c0VERGjFihXq2bOnzp07p6CgIElSa2urZs6cqf79++vrr792bvQrSYMHD9bixYvl7e0tb29v53iXLl3k4+Pj/JkAcLdiZhCAsWbNmqWMjAzn+48++kizZ8++pXO9vb3V1NTkfJ+enq7p06fL399fY8eObfP3Hj9+XGVlZVqyZEmbIvi7gICAf/4lAMBilEEAxpoxY4YOHz6sc+fO6fz58zpy5IimT5/+l+cdO3ZMn376qeLi4iRJp0+f1jfffKPk5GRJ0vTp05WRkaGWlhbn55J41BeADokyCMBY3bp1U1JSkrZt26aMjAwlJSWpW7duf3jsnj175OfnJy8vL0VHRys2NlbvvfeepJuzggkJCc5zx44dq7q6On355ZeSbl4mliSbzXYHvhUA3FmsGQRgtNmzZyslJUWS9MEHH/zpcaNGjVJaWpo8PDzUq1cveXh4SJKam5uVmZmpmpoade78f/8kNjc3Kz09XfHx8XrwwQclSSdPntTgwYNv35cBAAtQBgEYLTExUY2NjZKkhISEPz3O19dXYWFh7cZzcnJ09epVlZSUqFOnTs7xU6dOadq0abp8+bIGDx6siIgIrV+/XsnJye3WDdbW1rJuEICxuEwMwGidOnXSyZMndfLkyTZl7lalp6crKSlJkZGReuihh5yvyZMn6/7779fHH38sm82mjIwMlZeXKzY2Vjk5OaqsrNT333+vNWvWaPz48bfhmwHAnUEZBGA8u90uu93+t8/7+eeftXfvXk2ePLndZzabTZMmTXLuOfjII4+oqKhIoaGhmjt3rgYOHKhx48aprKxMGzZs+LdfAQAsY2v9fWU0AAAA3A4zgwAAAG6MMggAAODGKIMAAABujDIIAADgxiiDAAAAbowyCAAA4MYogwAAAG6MMggAAODGKIMAAABujDIIAADgxiiDAAAAbuw/fWIBcfpBRNAAAAAASUVORK5CYII=", + "text/plain": [ + "
" + ] + }, + "metadata": {}, + "output_type": "display_data" + } + ], + "source": [ + "#plot mutations_per_gene_consequence as a heatmap\n", + "data_to_plot = mutations_per_gene_consequence.unstack(level=0).loc[consequences_to_plot]\n", + "data_to_plot.columns = data_to_plot.columns.droplevel(0)\n", + "data_to_plot = data_to_plot[[x for x in data_to_plot.columns if not data_to_plot[x].isna().all()]].T\n", + "plt.figure(figsize=(7, 10))\n", + "sns.heatmap(data_to_plot, annot=True, fmt=\".0f\", cmap=\"Reds\")\n", + "plt.show()" + ] + } + ], + "metadata": { + "kernelspec": { + "display_name": "notebooks_env", + "language": "python", + "name": "python3" + }, + "language_info": { + "codemirror_mode": { + "name": "ipython", + "version": 3 + }, + "file_extension": ".py", + "mimetype": "text/x-python", + "name": "python", + "nbconvert_exporter": "python", + "pygments_lexer": "ipython3", + "version": "3.10.16" + } + }, + "nbformat": 4, + "nbformat_minor": 5 +} diff --git a/assets/build_datasets/dndscv/biomartQuery.MANE.txt b/assets/build_datasets/dndscv/biomartQuery.MANE.txt new file mode 100644 index 00000000..7f8897e0 --- /dev/null +++ b/assets/build_datasets/dndscv/biomartQuery.MANE.txt @@ -0,0 +1,22 @@ + + + + + + + + + + + + + + + + + + + + + + \ No newline at end of file diff --git a/assets/build_datasets/dndscv/biomartQuery.canonical.txt b/assets/build_datasets/dndscv/biomartQuery.canonical.txt new file mode 100644 index 00000000..7cb945b6 --- /dev/null +++ b/assets/build_datasets/dndscv/biomartQuery.canonical.txt @@ -0,0 +1,22 @@ + + + + + + + + + + + + + + + + + + + + + + \ No newline at end of file diff --git a/assets/build_datasets/dndscv/instructions.txt b/assets/build_datasets/dndscv/instructions.txt new file mode 100644 index 00000000..1886537e --- /dev/null +++ b/assets/build_datasets/dndscv/instructions.txt @@ -0,0 +1,26 @@ + +# Biomart Query +biomart_cds_query_file="biomartQuery.MANE.txt" +BIOMART_CDS="homo_sapiens.v111.MANE.biomart.tsv" + +biomart_cds_query_file="biomartQuery.canonical.txt" +BIOMART_CDS="homo_sapiens.v111.canonical.biomart.tsv" + + +biomart_url="http://jan2024.archive.ensembl.org/biomart/martservice" + +# Read XML query from file +biomart_cds_query=$(cat "${biomart_cds_query_file}") + +# URL encode it (remove newlines first) +biomart_cds_query_encoded=$(python3 -c " +from urllib.parse import quote_plus +query = '''${biomart_cds_query}''' +print(quote_plus(query.replace('\n', ''))) +") + +echo "Downloading biomart..." + +curl -L -s "${biomart_url}?query=${biomart_cds_query_encoded}" | \ + tail -n +2 | \ + awk -F'\t' '($5 != "") { print $0 }' > "${BIOMART_CDS}" \ No newline at end of file diff --git a/assets/omega_consequences_groupings.json b/assets/omega_consequences_groupings.json new file mode 100644 index 00000000..ebfc54eb --- /dev/null +++ b/assets/omega_consequences_groupings.json @@ -0,0 +1,7 @@ +{ + "missense": ["missense"], + "nonsense": ["nonsense"], + "essential_splice": ["essential_splice"], + "truncating": ["nonsense", "essential_splice"], + "nonsynonymous_splice": ["missense", "nonsense", "essential_splice"] +} \ No newline at end of file diff --git a/assets/placeholder_no_file.tsv b/assets/placeholder_no_file.tsv new file mode 100644 index 00000000..4f2e970e --- /dev/null +++ b/assets/placeholder_no_file.tsv @@ -0,0 +1 @@ +PLACEHOLDER diff --git a/bin/__plot_selectionfeatures.py b/bin/__plot_selectionfeatures.py deleted file mode 100755 index 51afae8d..00000000 --- a/bin/__plot_selectionfeatures.py +++ /dev/null @@ -1,899 +0,0 @@ -#!/usr/bin/env python - - -import pandas as pd -import json -import os, sys -import matplotlib.pyplot as plt -import seaborn as sns -import numpy as np -from matplotlib.ticker import MaxNLocator -from mpl_toolkits.axes_grid1.inset_locator import inset_axes -from read_utils import custom_na_values - -pd.set_option('display.max_columns', None) - -# Suppress warnings -import warnings -pd.options.mode.chained_assignment = None -warnings.filterwarnings("ignore", message="FixedFormatter should only be used together with FixedLocator") - - -import pandas as pd -import os, sys -import matplotlib.pyplot as plt -import seaborn as sns -import numpy as np -from scipy.stats import linregress,norm -import tabix -import matplotlib.cm as cm -import matplotlib.colors as mcolors -from matplotlib.patches import Patch -from matplotlib.lines import Line2D - -import sys, os -import pandas as pd -import numpy as np -import matplotlib.pyplot as plt - - -def plot_single_needle(gene, snvs_maf_obs, - seq_info_df, - sample = '' - ): - - snvs_maf_obs_gene = snvs_maf_obs[(snvs_maf_obs["canonical_SYMBOL"] == gene) - & (snvs_maf_obs["TYPE"] == "SNV")] - - # Subset gene - seq_info_df_gene = seq_info_df[seq_info_df["Gene"] == gene] - gene_len = len(seq_info_df_gene.Seq.values[0]) - pos_gene = pd.DataFrame({"Pos": range(1, gene_len + 1)}) - - # Get per-position SNV mutations count - obs_snv_count_gene = snvs_maf_obs_gene.groupby("canonical_Protein_position").size().reset_index(name='Count').astype(int) - obs_snv_count_gene.columns = ["Pos", "Count"] - obs_snv_count_gene = pos_gene.merge(obs_snv_count_gene, how="left", on="Pos") - - - # Determine the maximum y-value for this gene - max_y_gene = obs_snv_count_gene["Count"].max() - - - # Create a single figure with two subplots - fig, axs = plt.subplots(1, 1, figsize=(10, 3), sharex=True) - - - # Plot the data for observed and randomized mutations in separate subplots - plotting_needle_from_counts(obs_snv_count_gene, axs, max_y=max_y_gene) - - # Set titles for each subplot - axs.set_title(f'{gene} {sample} - {int(obs_snv_count_gene["Count"].sum()):,} observed SNVs') - - # Adjust layout to prevent overlap - plt.tight_layout() - - return fig - - -def plotting_needle_from_counts(data_gene, - ax = None, - max_y=None, - col_pos_track='#003366', - label_pos_track='observed', - col_hv_lines='grey'): - - # Determine max_y if not provided - if max_y is None: - max_y = data_gene["Count"].max() - - # Calculate the precise margin to add to the y-axis - marker_size = 60 - marker_radius = np.sqrt(marker_size / np.pi) - marker_margin = marker_radius * 0.5 # Convert marker radius to a suitable margin - - # Add the margin to max_y - max_y += marker_margin - - - # If ax is not provided, create a new figure and axis - if ax is None: - fig, ax = plt.subplots(figsize=(10, 3)) - else: - fig = None # No need to create a figure if ax is provided - - - ax.vlines(data_gene["Pos"], ymin=0, ymax=data_gene["Count"], color=col_hv_lines, lw=1, zorder=1, alpha=0.5) - ax.scatter(data_gene["Pos"], data_gene["Count"], color='white', zorder=3, lw=1, ec="white") # To cover the overlapping needle top part - ax.scatter(data_gene["Pos"].values, data_gene["Count"].values, color=col_pos_track, zorder=4, - alpha=0.7, lw=0.1, ec="black", s=30, label=label_pos_track) - - # Remove the right and top spines - ax.spines['right'].set_visible(False) - ax.spines['top'].set_visible(False) - - # Set the y-axis limit - ax.set_ylim(0, max_y) - - # Add labels - ax.set_xlabel('Position') - ax.set_ylabel('Count') - - # If a new figure was created, show it - if fig is not None: - plt.show() - - return fig - - -def manager(mutations_file, o3d_seq_file, sample_name, sample_name_out): - # Load your MAF DataFrame (raw_annotated_maf) - maf = pd.read_csv(mutations_file, sep = "\t", header = 0, na_values = custom_na_values) - o3d_seq_df = pd.read_csv(o3d_seq_file, sep = "\t", header = 0) - - maf_f = maf[maf["canonical_Protein_position"] != '-'].reset_index(drop = True) - del maf - - gene_order = sorted(pd.unique(maf_f["canonical_SYMBOL"])) - - - os.makedirs(f"{sample_name_out}") - - # Loop over each gene to plot - for geneeee in gene_order: - print(geneeee) - fig_needles = plot_single_needle(geneeee, maf_f, - o3d_seq_df, - sample=sample_name - ) - - fig_needles.savefig(f"{sample_name_out}/{geneeee}.needles.pdf", bbox_inches='tight') - plt.close() - - - - -# @click.command() -# @click.option('--sample_name', type=str, help='Name of the sample being processed.') -# @click.option('--mut_file', type=click.Path(exists=True), help='Input mutation file') -# @click.option('--out_maf', type=click.Path(), help='Output MAF file') -# @click.option('--json_filters', type=click.Path(exists=True), help='Input mutation filtering criteria file') -# @click.option('--req_plots', type=click.Path(exists=True), help='Column names to output') -# # @click.option('--plot', is_flag=True, help='Generate plot and save as PDF') - -# def main(sample_name, mut_file, out_maf, json_filters, req_plots): # , plot): -# click.echo(f"Subsetting MAF file...") -# subset_mutation_dataframe(sample_name, mut_file, out_maf, json_filters, req_plots) - -# if __name__ == '__main__': -# main() - - -sample_name_ = sys.argv[1] -mut_file = sys.argv[2] -o3d_seq_file_ = sys.argv[3] -sample_name_out_ = sys.argv[4] -# out_maf = sys.argv[3] -# json_filters = sys.argv[4] -# req_plots = sys.argv[5] - - - -if __name__ == '__main__': - # maf = subset_mutation_dataframe(mut_file, json_filters) - # plot_manager(sample_name, maf, req_plots) - manager(mut_file, o3d_seq_file_, sample_name_, sample_name_out_) - - -# Init -run_name = "all_samples" - -deepcsa_run_dir = "/workspace/nobackup/bladder_ts/results/2024-06-20_deepCSA" -maf_file = os.path.join(deepcsa_run_dir, f"writemaf/{run_name}.filtered.tsv.gz") -o3d_datasets = "/workspace/nobackup/scratch/oncodrive3d/datasets_240506" -o3d_annotations = "/workspace/nobackup/scratch/oncodrive3d/annotations_240506" -fig3_data = "/workspace/projects/bladder_ts/notebooks/manuscript_figures/Fig3/data" - -gene_order = ["KMT2D","EP300","ARID1A","CREBBP","NOTCH2","KMT2C","STAG2","RB1", - "RBM10","KDM6A","TP53","FGFR3","CDKN1A","FOXQ1", - # "PIK3CA","TERT" - ] - - - -# Oncodrive3D - -o3d_seq_df = pd.read_table(f"{o3d_datasets}/seq_for_mut_prob.tsv") -o3d_annot_df = pd.read_csv(f"{o3d_annotations}/uniprot_feat.tsv", sep="\t") - -o3d_prob = f"{deepcsa_run_dir}/oncodrive3d/run/{run_name}/{run_name}.miss_prob.processed.json" -o3d_prob = json.load(open(o3d_prob, encoding="utf-8")) - -o3d_score = f"{deepcsa_run_dir}/oncodrive3d/run/{run_name}/{run_name}.3d_clustering_pos.csv" -o3d_score = pd.read_csv(o3d_score)[["Gene", "Pos", "Score", "Score_obs_sim", "pval", "C", "C_ext"]] - - -# MAF - -maf_df = pd.read_csv(maf_file, sep = "\t", na_values = custom_na_values) -maf_df_f, snv_df, trunc_df, synon_df, miss_df = preprocess_maf(maf_df) - - - - - -for gene in gene_order: - - plot_pars = init_plot_pars() - gene_len = len(o3d_seq_df[o3d_seq_df["Gene"] == gene].Seq.values[0]) - if gene_len > 400: - # if gene_len > 2500: - # gene_len = 2500 - ref_ratio = 500 / gene_len - fsize_x = round(gene_len / 26) - else: - ref_ratio = 1 - fsize_x = 15 - - plot_pars["fsize"] = fsize_x, 10 - plot_pars["ofml_cbar_coord"] = (0.04, 0.36, 1.1*ref_ratio, 0.85) - - - generate_plot(gene, - o3d_seq_df, - o3d_annot_df, - o3d_prob, - o3d_score, - maf_df_f, - snv_df, - trunc_df, - synon_df, - miss_df, - plot_pars, - output_path="Fig3_plots/Fig3b") - - - - - - - -# Plot -# ==== - -def init_plot_pars(): - - plot_pars = {"fsize" : (20,10), - "hspace" : 0.1, # General space between all tracks - "ofml_cbar_coord" : (0.01, 0.35, 0.9, 1), # Box to anchor coordinate for OncodriveFML color bar - "track_title_x_coord" : 0.83, # x-coordinate (respect to protein len) for track txt title - "score_txt_x_coord" : 1.13, # as track title but for track score txt - "track_title_fontsize" : 14, - "ylabel_fontsize" : 13.5, - "xlabel_fontsize" : 13.5, - "ylabel_pad" : 38, - "ticksize" : 10.5, - "legend_fontsize" : 12, - "legend_frameon" : True, - "dpi" : 300 - } - plot_pars["colors"] = {"ofml" : "viridis_r", - "omega_trunc" : "#FA5E32", - "omega_synon" : "#89E4A2", - "omega_miss" : "#FABE4A", - "o3d_score" : "#6DBDCC", - "o3d_cluster" : "skyblue", - "o3d_prob" : "darkgray", - "frameshift" : "#E4ACF4", - "inframe" : "C5", - "hv_lines" : "lightgray" # General horizontal and vertical lines (e.g., needle plot vline) - } - plot_pars["h_ratios"] = {"omega_trunc" : 0.3, - "omega_synon" : 0.3, - "omega_miss" : 0.3, - "space1" : 0.015, - "o3d" : 0.5, - "space2" : 0.015, - "ofml" : 0.5, - "space3" : 0.015, - "indels" : 0.5, - "space4" : 0.015, - "domain" : 0.07 - } - - return plot_pars - - -def plot_count_track(pos_track, - axes, - ax=0, - neg_track=None, - gene_len=None, - col_pos_track="gray", - col_neg_track="gray", - label_pos_track=None, - label_neg_track=None, - ymargin=None, - hv_lines="lightgray"): - - axes[ax].vlines(pos_track["Pos"], ymin=0, ymax=pos_track["Count"], color=hv_lines, lw=1, zorder=1, alpha=0.5) - axes[ax].scatter(pos_track["Pos"], pos_track["Count"], color='white', zorder=3, lw=1, ec="white") # To cover the overlapping needle top part - axes[ax].scatter(pos_track["Pos"].values, pos_track["Count"].values, color=col_pos_track, zorder=4, - alpha=0.7, lw=0.1, ec="black", s=60, label=label_pos_track) - - if isinstance(neg_track, pd.DataFrame): - if isinstance(gene_len, int): - axes[ax].hlines(0, xmin=0, xmax=gene_len, color="gray", lw=0.6, zorder=1) - axes[ax].vlines(neg_track["Pos"], ymin=-neg_track["Count"], ymax=0, color=hv_lines, lw=1, zorder=1, alpha=0.5) - axes[ax].scatter(neg_track["Pos"], -neg_track["Count"], color='white', zorder=3, lw=1, ec="white") # To cover the overlapping needle top part - axes[ax].scatter(pos_track["Pos"].values, -neg_track["Count"].values, color=col_neg_track, zorder=4, - alpha=0.7, lw=0.1, ec="black", s=60, label=label_neg_track) - axes[ax].set_yticklabels(abs(axes[ax].get_yticks())) - - if ymargin is not None: - axes[ax].set_ylim(-np.max(neg_track["Count"])-ymargin, np.max(pos_track["Count"])+ymargin) - - -def add_ax_text(score, pvalue, y_text, x_text, ax, y_shift=0.2, y_adjust=0, equal_less=True): - - ax.text(x_text, y_text+(y_text*y_shift)+y_adjust, - fr'$\mathit{{Score}}$ = {np.round(score, 2)}', ha='center', va='center', fontsize=13.5, color="black") - - equal = "≤" if equal_less else "=" - ax.text(x_text, y_text-(y_text*y_shift)+y_adjust, - fr'$\mathit{{p}}$-value {equal} {pvalue}', ha='center', va='center', fontsize=13.5, color="black") - - -def get_transcript_ids(gene, maf_df_f, o3d_seq_df): - - canonical_tr = maf_df_f[maf_df_f["SYMBOL"] == gene].canonical_Feature.unique()[0] - o3d_tr = o3d_seq_df[o3d_seq_df["Gene"] == gene].Ens_Transcr_ID.values[0] - - return canonical_tr, o3d_tr - - -def plot_gene_selection_signals(maf_trunc_count_gene, - maf_miss_count_gene, - maf_synon_count_gene, - frameshift_indels_count_gene, - inframe_indels_count_gene, - ofml_muts_score_gene, - o3d_score_gene, - o3d_prob_gene, - plot_pars, - domain_df=None, - add_track_title_text=True, - add_score_text=False, - rm_spines=False, - light_spines=False, - output_path=None): - - - gene_len = len(o3d_prob_gene) - fig, axes = plt.subplots(len(plot_pars["h_ratios"]), 1, - figsize=plot_pars["fsize"], - sharex=True, - gridspec_kw={'hspace': plot_pars["hspace"], - 'height_ratios': plot_pars["h_ratios"].values()}) - colors = plot_pars["colors"] - - - # Omega - # ===== - - # Omega trunc - ax=0 - plot_count_track(maf_trunc_count_gene, axes, ax=ax, gene_len=gene_len, col_pos_track=colors["omega_trunc"]) - if add_score_text: - add_ax_text(score=np.round(omega_truncating_score, 2), - pvalue=omega_truncating_pvalue, - y_text=np.max(maf_synon_count_gene["Count"])/2, - x_text=gene_len*plot_pars["score_txt_x_coord"], - ax=axes[ax], - y_shift=0.9, - y_adjust=2.4) - if add_track_title_text: - axes[ax].text(gene_len*plot_pars["track_title_x_coord"], 4.7, - fr'$\mathbf{{Omega}}$ $\mathbf{{Truncating}}$', - ha='center', va='center', fontsize=plot_pars["track_title_fontsize"], color="black") - axes[ax].set_ylabel('Truncating\ncount', fontsize=plot_pars["ylabel_fontsize"], rotation=45, labelpad=plot_pars["ylabel_pad"], va='center') - - n_max = np.max(maf_trunc_count_gene["Count"]) - j = 12 - i = 24 - axes[ax].set_ylim(-n_max/i, n_max + n_max/j) - - axes[ax].yaxis.set_major_locator(MaxNLocator(integer=True, nbins=3)) - axes[ax].tick_params(axis='y', labelsize=plot_pars["ticksize"]) - if rm_spines: - axes[ax].spines['top'].set_visible(False) - axes[ax].spines['right'].set_visible(False) - elif light_spines: - axes[ax].spines['top'].set_color(colors["hv_lines"]) - axes[ax].spines['right'].set_color(colors["hv_lines"]) - - # Omega synonym - ax=1 - plot_count_track(maf_synon_count_gene, axes, ax=ax, gene_len=gene_len, col_pos_track=colors["omega_synon"], hv_lines=colors["hv_lines"]) - axes[ax].set_ylabel('Synonymous\ncount', fontsize=plot_pars["ylabel_fontsize"], rotation=45, labelpad=plot_pars["ylabel_pad"], va='center') - - n_max = np.max(maf_synon_count_gene["Count"]) - j = 10 - i = 85.714 - axes[ax].set_ylim(-n_max/i, n_max + n_max/j) - - axes[ax].tick_params(axis='y', labelsize=plot_pars["ticksize"]) - axes[ax].yaxis.set_major_locator(MaxNLocator(integer=True, nbins=3)) - axes[ax].spines['top'].set_visible(False) - if rm_spines: - axes[ax].spines['right'].set_visible(False) - elif light_spines: - axes[ax].spines['right'].set_color(colors["hv_lines"]) - - # Omega miss - ax=2 - plot_count_track(maf_miss_count_gene, axes, ax=ax, gene_len=gene_len, col_pos_track=colors["omega_miss"], hv_lines=colors["hv_lines"]) - y_text=np.max(maf_miss_count_gene["Count"])/2 - if add_score_text: - add_ax_text(score=np.round(omega_misss_score, 2), - pvalue=omega_misss_pvalue, - y_text=y_text, - x_text=gene_len*plot_pars["score_txt_x_coord"], - y_shift=0.3, - ax=axes[ax], - y_adjust=4) - if add_track_title_text: - axes[ax].text(gene_len*plot_pars["track_title_x_coord"], y_text+y_text*0.6935, - fr'$\mathbf{{Omega}}$ $\mathbf{{Missense}}$', - ha='center', va='center', fontsize=plot_pars["track_title_fontsize"], color="black") - axes[ax].set_ylabel('Missense\ncount', fontsize=plot_pars["ylabel_fontsize"], rotation=45, labelpad=plot_pars["ylabel_pad"], va='center') - - n_max = np.max(maf_miss_count_gene["Count"]) - j = 10.74 - i = 14.5 - axes[ax].set_ylim(-n_max/i, n_max + n_max/j) - - axes[ax].spines['top'].set_visible(False) - axes[ax].yaxis.set_major_locator(MaxNLocator(integer=True, nbins=3)) - axes[ax].tick_params(axis='y', labelsize=plot_pars["ticksize"]) - if rm_spines: - axes[ax].spines['right'].set_visible(False) - elif light_spines: - axes[ax].spines['right'].set_color(colors["hv_lines"]) - - - # Oncodrive3D - # =========== - - ax=4 - axes[ax].plot(range(1, gene_len+1), o3d_score_gene["O3D_score_norm"], zorder=2, color=colors["o3d_score"], lw=1, label="Clustering score") - axes[ax].plot(range(1, gene_len+1), -o3d_prob_gene, zorder=3, color=colors["o3d_prob"], lw=1, label="Missense mut probability") - axes[ax].fill_between(o3d_score_gene['Pos'], 0, o3d_score_gene["O3D_score_norm"], where=(o3d_score_gene['Cluster'] == 1), - color=colors["o3d_cluster"], alpha=0.3, label='Cluster', zorder=0, lw=2) - if add_score_text: - add_ax_text(score=np.round(o3d_score, 2), - pvalue=o3d_pvalue, - y_text=(np.max(o3d_score_gene["O3D_score_norm"]) + np.max(o3d_prob_gene))/2, - x_text=gene_len*plot_pars["score_txt_x_coord"], - ax=axes[ax], - y_shift=0.15, - y_adjust=-0.0075, - equal_less=True) - if add_track_title_text: - axes[ax].text(gene_len*plot_pars["track_title_x_coord"], 0.015, - fr'$\mathbf{{3D}}$-$\mathbf{{Clustering}}$', - ha='center', va='center', fontsize=plot_pars["track_title_fontsize"], color="black") - - axes[ax].set_ylabel('Score &\nprobability', fontsize=plot_pars["ylabel_fontsize"], rotation=45, labelpad=plot_pars["ylabel_pad"], va='center') - axes[ax].yaxis.set_major_locator(MaxNLocator(integer=False, nbins=4)) - tick_labels = [np.round(abs(tick), 3) for tick in axes[ax].get_yticks()] - axes[ax].set_yticklabels(tick_labels) - axes[ax].tick_params(axis='y', labelsize=plot_pars["ticksize"]) - - axes[ax].legend(loc="upper left", fontsize=plot_pars["legend_fontsize"], frameon=plot_pars["legend_frameon"]) - #axes[ax].legend(bbox_to_anchor=[0, 0.35, 1, 0]) - if rm_spines: - axes[ax].spines['top'].set_visible(False) - axes[ax].spines['right'].set_visible(False) - elif light_spines: - axes[ax].spines['top'].set_color(colors["hv_lines"]) - axes[ax].spines['right'].set_color(colors["hv_lines"]) - - - # OncodriveFML - # ============ - - ax=6 - axes[ax].vlines(ofml_muts_score_gene["Pos"].values, ymin=0, ymax=ofml_muts_score_gene["Count"], color=colors["hv_lines"], lw=1, zorder=1, alpha=0.5) - fml_scatter = axes[ax].scatter(ofml_muts_score_gene["Pos"].values, ofml_muts_score_gene["Count"].values, - c=ofml_muts_score_gene["CADD_score"].values, cmap=colors["ofml"], zorder=4, - alpha=0.7, lw=0.1, ec="black", s=60, label="SNV") - - y_text=np.round(np.max(ofml_muts_score_gene["Count"])/2) - if add_score_text: - add_ax_text(score=np.round(ofml_score, 2), - pvalue=ofml_pvalue, - y_text=y_text, - x_text=gene_len*plot_pars["score_txt_x_coord"], - ax=axes[ax], - y_adjust=3, - equal_less=True) - - if add_track_title_text: - axes[ax].text(gene_len*plot_pars["track_title_x_coord"], y_text+y_text*0.8, - fr'$\mathbf{{Functional}}$ $\mathbf{{Impact}}$', - ha='center', va='center', fontsize=plot_pars["track_title_fontsize"], color="black") - - inset_ax = inset_axes(axes[ax], width="10%", height="10%", loc='center left', - bbox_to_anchor=plot_pars["ofml_cbar_coord"], bbox_transform=axes[ax].transAxes, borderpad=0) - cbar = plt.colorbar(fml_scatter, cax=inset_ax, orientation='horizontal') - cbar.set_label('Impact score', fontsize=plot_pars["legend_fontsize"]) - cbar.ax.tick_params(labelsize=plot_pars["ticksize"]) - - n_max = np.max(ofml_muts_score_gene["Count"]) - j = 14.5 - i = 19.33 - axes[ax].set_ylim(-n_max/i, n_max + n_max/j) - - axes[ax].set_ylabel('SNV count', fontsize=plot_pars["ylabel_fontsize"], rotation=45, labelpad=plot_pars["ylabel_pad"], va='center') - axes[ax].tick_params(axis='y', labelsize=plot_pars["ticksize"]) - - if rm_spines: - axes[ax].spines['top'].set_visible(False) - axes[ax].spines['right'].set_visible(False) - elif light_spines: - axes[ax].spines['top'].set_color(colors["hv_lines"]) - axes[ax].spines['right'].set_color(colors["hv_lines"]) - - - # Frameshift enrichment - # ===================== - - ax=8 - plot_count_track(frameshift_indels_count_gene, axes, ax=ax, neg_track=inframe_indels_count_gene, gene_len=gene_len, - col_pos_track=colors["frameshift"], col_neg_track=colors["inframe"], label_pos_track="Frameshift", label_neg_track="Inframe", - ymargin=1, hv_lines=colors["hv_lines"]) - if add_score_text: - add_ax_text(score=np.round(indels_score, 2), - pvalue=np.round(indels_pvalue, 2), - y_text=(np.max(frameshift_indels_count_gene["Count"]) + np.max(inframe_indels_count_gene["Count"]))/2, - x_text=gene_len*plot_pars["score_txt_x_coord"], - ax=axes[ax], - y_shift=0.24, - y_adjust=0.4, - equal_less=False) - if add_track_title_text: - axes[ax].text(gene_len*plot_pars["track_title_x_coord"], 3.9, - fr'$\mathbf{{Frameshift}}$ $\mathbf{{enrichment}}$', - ha='center', va='center', fontsize=plot_pars["track_title_fontsize"], color="black") - axes[ax].legend(loc="upper left", fontsize=plot_pars["legend_fontsize"], frameon=plot_pars["legend_frameon"]) - axes[ax].set_ylabel('Indels count', fontsize=plot_pars["ylabel_fontsize"], rotation=45, labelpad=plot_pars["ylabel_pad"], va='center') - - axes[ax].yaxis.set_major_locator(MaxNLocator(integer=True, nbins=5)) - tick_labels = [abs(int(tick)) for tick in axes[ax].get_yticks()] - axes[ax].set_yticklabels(tick_labels) - - if rm_spines: - axes[ax].spines['top'].set_visible(False) - axes[ax].spines['right'].set_visible(False) - elif light_spines: - axes[ax].spines['top'].set_color(colors["hv_lines"]) - axes[ax].spines['right'].set_color(colors["hv_lines"]) - - - # Domain - # ====== - - if isinstance(domain_df, pd.DataFrame): - ax=10 - domain_color_dict = {} - - for n, name in enumerate(domain_df["Description"].unique()): - domain_color_dict[name] = f"C{n}" - - n = 0 - added_domain = [] - for i, row in domain_df.iterrows(): - if pd.Series([row["Description"], row["Begin"], row["End"]]).isnull().any(): - continue - - name = row["Description"] - start = int(row["Begin"]) - end = int(row["End"]) - axes[ax].fill_between(range(start, end+1), -0.45, 0.45, alpha=0.5, color=domain_color_dict[name]) - if name not in added_domain: - y = -0.04 - axes[ax].text(((start + end) / 2)+0.5, y, name, ha='center', va='center', fontsize=10, color="black") - added_domain.append(name) - axes[ax].set_yticks([]) - else: - axes[10].remove() - - - # Spaces - # ===== - axes[3].remove() - axes[5].remove() - axes[7].remove() - axes[9].remove() - - - # X axes - # ====== - axes[ax].tick_params(axis='y', labelsize=plot_pars["ticksize"]) - axes[ax].tick_params(axis='x', labelsize=plot_pars["ticksize"]) - axes[ax].set_xlabel('Protein position', fontsize=plot_pars["xlabel_fontsize"]) - - if output_path is not None: - print(f"Saving {output_path}") - plt.savefig(output_path, dpi=plot_pars["dpi"], bbox_inches='tight') - plt.show() - - -def generate_plot(gene, - o3d_seq_df, - o3d_annot_df, - o3d_prob, - o3d_score, - maf_df_f, - snv_df, - trunc_df, - synon_df, - miss_df, - plot_pars, - domain_df=None, - add_track_title_text=False, - add_score_text=False, - rm_spines=True, - light_spines=False, - output_path=None): - - # Prepare data - # ============ - - # O3D gene - o3d_score_df_gene, o3d_score_gene, o3d_prob_gene, o3d_top_pos_score_gene, uni_id, gene_pos = get_o3d_gene_data(gene, - o3d_seq_df, - o3d_prob, - o3d_score) - - # MAF gene - snv_count_gene = get_maf_pos_count(gene, snv_df, gene_pos) - trunc_count_gene = get_maf_pos_count(gene, trunc_df, gene_pos) - synon_count_gene = get_maf_pos_count(gene, synon_df, gene_pos) - miss_count_gene = get_maf_pos_count(gene, miss_df, gene_pos) - - # OFML gene - ofml_score_gene = get_ofml_score_gene(gene, fig3_data) - - # Indels - frameshift_indels_df, inframe_indels_df = get_frameshift_indels_maf(maf_df_f) - frameshift_indels_count_gene, inframe_indels_count_gene = get_frameshift_indels_gene(gene, - frameshift_indels_df, - inframe_indels_df, - gene_pos) - - # Domain - domain_gene = o3d_annot_df[(o3d_annot_df["Gene"] == gene)].reset_index(drop=True) - - # Transcripts and Uniprot ID - canonical_tr, o3d_tr = get_transcript_ids(gene, maf_df_f, o3d_seq_df) - print(f"> {gene} - {canonical_tr} - {o3d_tr} - {uni_id}") - - - # Plot - # ==== - - if output_path is not None: - output_path = f"{output_path}/{gene}.png" - plot_gene_selection_signals(trunc_count_gene, - miss_count_gene, - synon_count_gene, - frameshift_indels_count_gene, - inframe_indels_count_gene, - ofml_score_gene, - o3d_score_df_gene, - o3d_prob_gene, - plot_pars, - domain_gene, - add_track_title_text, - add_score_text, - rm_spines, - light_spines, - output_path) - - - - - -# Prepare data -# ============ - -def get_o3d_gene_data(gene, - seq_df, - prob_dict, - score_df): - - # Subset gene - seq_df_gene = seq_df[seq_df["Gene"] == gene] - gene_len = len(seq_df_gene.Seq.values[0]) - gene_pos = pd.DataFrame({"Pos" : range(1, gene_len+1)}) - uni_id, af_f = seq_df_gene[["Uniprot_ID", "F"]].values[0] - prob_gene = np.array(prob_dict[f"{uni_id}-F{af_f}"]) - score_gene_df = score_df[score_df["Gene"] == gene].reset_index(drop=True) - score_gene, top_pos_score_gene = score_df[["Score", "Score_obs_sim"]].values[0] - - ## O3D score vector - score_gene_df = gene_pos.merge(score_gene_df[["Pos", "Score_obs_sim", "C", "C_ext"]], how="left", on="Pos") - - # Don't include Extended clusters - score_gene_df["C"] = (score_gene_df["C"] == 1) & (score_gene_df["C_ext"] == 0) - score_gene_df["C"] = score_gene_df["C"].astype(int) - score_gene_df = score_gene_df.drop(columns=["C_ext"]) - - score_gene_df.columns = "Pos", "O3D_score", "Cluster" - score_gene_df["O3D_score"] = score_gene_df["O3D_score"].fillna(0) - score_gene_df["Cluster"] = score_gene_df["Cluster"].fillna(0) - score_gene_df["O3D_score_norm"] = score_gene_df["O3D_score"] / sum(score_gene_df["O3D_score"]) - - return score_gene_df, score_gene, prob_gene, top_pos_score_gene, uni_id, gene_pos - - -def preprocess_maf(maf_df): - - maf_df["CLEAN_SAMPLE_ID"] = maf_df["SAMPLE_ID"].apply(lambda x: "_".join(x.split("_")[1:3])) - - maf_df_f = maf_df.loc[(maf_df["VAF"] <= 0.35) & - (~maf_df["FILTER.not_in_exons"]) & - (~maf_df["FILTER.not_covered"]) & - (~maf_df["FILTER.no_pileup_support"]) & # avoid variants w/o VAF recomputed - (~maf_df["FILTER.n_rich"]) & - (~maf_df["FILTER.low_mappability"]) - ].reset_index(drop = True) - - # SNV - snvs_maf = maf_df_f[(maf_df_f["TYPE"] == 'SNV') & - (maf_df_f["canonical_SYMBOL"].isin(gene_order)) & - (maf_df_f["canonical_Protein_position"] != '-' ) - ].reset_index(drop = True) - snvs_maf["canonical_Protein_position"] = snvs_maf["canonical_Protein_position"].astype(int) - - # Omega - maf_df_trunc = snvs_maf[snvs_maf["canonical_Consequence_broader"].isin(['nonsense', "essential_splice"])] - maf_df_synon = snvs_maf[snvs_maf["canonical_Consequence_broader"] == 'synonymous'] - maf_df_miss = snvs_maf[snvs_maf["canonical_Consequence_broader"] == 'missense'] - - return maf_df_f, snvs_maf, maf_df_trunc, maf_df_synon, maf_df_miss - - -def get_maf_pos_count(gene, maf_df, gene_pos): - - # Subset gene and cols - cols = ["SYMBOL", - "TYPE", - "Feature", - "canonical_Feature", - "canonical_Consequence", - "canonical_Consequence_broader", - "canonical_Protein_position"] - - df_gene = maf_df.loc[maf_df["SYMBOL"] == gene, cols] - - # Get per-position mutations count - df_gene_count = df_gene.groupby("canonical_Protein_position").apply(lambda x: len(x)).reset_index().astype(int) - df_gene_count.columns = "Pos", "Count" - df_gene_count = gene_pos.merge(df_gene_count, how="left", on="Pos") - - return df_gene_count - - -def get_ofml_score_gene(gene, fig3_data_path): - - mutations_in_gene = pd.read_table(f"{fig3_data_path}/mutations_scored.{gene}.tsv", na_values = custom_na_values) - ofml_muts_score_gene = mutations_in_gene.groupby(by = "canonical_Protein_position").agg( { "MUT_ID" : 'count', - "CADDscore" : 'mean'}).reset_index() - ofml_muts_score_gene.columns = "Pos", "Count", "CADD_score" - - return ofml_muts_score_gene - - -def get_frameshift_indels_maf(maf_df_f): - - # Filter and somatic only - indel_maf_df = maf_df_f.loc[ - (maf_df_f["TYPE"].isin(["INSERTION", "DELETION"])) - ].reset_index(drop = True) - indel_maf_df["INDEL_LENGTH"] = (indel_maf_df["REF"].str.len() - indel_maf_df["ALT"].str.len()).abs() - indel_maf_df["INDEL_INFRAME"] = [ x % 3 == 0 for x in indel_maf_df["INDEL_LENGTH"] ] - indel_maf_df.loc[indel_maf_df["INDEL_LENGTH"] >= 15 , "INDEL_INFRAME"] = False - indel_maf_df.loc[indel_maf_df["canonical_Consequence_broader"] == 'nonsense' , "INDEL_INFRAME"] = False - indel_maf_df.loc[indel_maf_df["canonical_Consequence_broader"] == 'essential_splice' , "INDEL_INFRAME"] = False - - # Observed frameshift indels + inframes of length >= 5 AA - frameshift_indels = indel_maf_df[(~indel_maf_df["INDEL_INFRAME"]) & - (indel_maf_df["canonical_Protein_position"] != '-' )].reset_index(drop = True) - - # Inframe indels - inframe_indels = indel_maf_df[(indel_maf_df["INDEL_INFRAME"]) & - (indel_maf_df["canonical_Protein_position"] != '-' )].reset_index(drop = True) - - return frameshift_indels, inframe_indels - - -def get_frameshift_indels_gene(gene, frameshift_indels_df, inframe_indels_df, gene_pos): - - ## Frameshift indels - frameshift_indels_gene = frameshift_indels_df[frameshift_indels_df["canonical_SYMBOL"] == gene] - - # Count by pos considering first pos as protein pos - frameshift_indels_gene.canonical_Protein_position = frameshift_indels_gene.canonical_Protein_position.apply(lambda x: x.split("-")[0]) - frameshift_indels_gene.canonical_Protein_position = frameshift_indels_gene.canonical_Protein_position[ - frameshift_indels_gene.canonical_Protein_position.apply(lambda x: x.isdigit())] - frameshift_indels_count_gene = frameshift_indels_gene.groupby("canonical_Protein_position").apply(lambda x: len(x)).reset_index().astype(int) - frameshift_indels_count_gene.columns = "Pos", "Count" - frameshift_indels_count_gene = gene_pos.merge(frameshift_indels_count_gene, how="left", on="Pos") - - ## Inframe indels - inframe_indels_gene = inframe_indels_df[inframe_indels_df["canonical_SYMBOL"] == gene] - - # Count - inframe_indels_gene.canonical_Protein_position = inframe_indels_gene.canonical_Protein_position.apply(lambda x: x.split("-")[0]) - inframe_indels_gene.canonical_Protein_position = inframe_indels_gene.canonical_Protein_position[ - inframe_indels_gene.canonical_Protein_position.apply(lambda x: x.isdigit())] - inframe_indels_count_gene = inframe_indels_gene.groupby("canonical_Protein_position").apply(lambda x: len(x)).reset_index().astype(int) - inframe_indels_count_gene.columns = "Pos", "Count" - inframe_indels_count_gene = gene_pos.merge(inframe_indels_count_gene, how="left", on="Pos") - - return frameshift_indels_count_gene, inframe_indels_count_gene - -seq_regions_df = pd.read_table("/workspace/nobackup/scratch/oncodrive3d/datasets_240506/seq_for_mut_prob.tsv") -tb = tabix.open("/workspace/datasets/CADD/v1.6/hg38/whole_genome_SNVs.tsv.gz") - - -def get_scores_gene(gene_info): - chromosome = gene_info["Chr"].values[0] - reverse_strand = gene_info["Reverse_strand"].values[0] == 1 - exon_coords = eval(gene_info["Exons_coord"].values[0]) - if reverse_strand: - exon_coords = [(y, x) for x, y in exon_coords] - - - position_mutation = dict() - for start, stop in exon_coords: - for e in tb.query(str(chromosome), start, stop): - if int(e[1]) not in position_mutation: - position_mutation[int(e[1])] = dict() - position_mutation[int(e[1])][e[3]] = float(e[5]) - - start2end_coords = { x:y for x, y in enumerate(sorted(position_mutation.keys(), reverse = not reverse_strand))} - - return position_mutation, start2end_coords - - - - -def score_mutations_from_gene(snv_data): - scores_muts = [] - for ind, row in snv_data.iterrows(): - try: - scores_muts.append(scores_all_mutations_in_gene[row["POS"]].get(row["ALT"], 0)) - print("found") - except: - scores_muts.append(0) - print(row["Consequence"], "not found") - snv_data["CADDscore"] = scores_muts - return snv_data - - - -# for gene in ["TP53"]: -for gene in gene_order: - gene_info = seq_regions_df[seq_regions_df["Gene"] == gene] - scores_all_mutations_in_gene, scores_of_protein_position = get_scores_gene(gene_info) - - mutations_in_gene = snvs_maf[snvs_maf["canonical_SYMBOL"] == gene].copy() - scored_mutations_in_gene = score_mutations_from_gene(mutations_in_gene) - - scored_mutations_in_gene.to_csv(f"{data_dir}/mutations_scored.{gene}.tsv", sep = '\t', header= True, index = False) - - diff --git a/bin/annotate_omega_failing.py b/bin/annotate_omega_failing.py index af6f6ae6..7cc1ce1e 100755 --- a/bin/annotate_omega_failing.py +++ b/bin/annotate_omega_failing.py @@ -179,6 +179,7 @@ def annotate(omegas: pd.DataFrame, gene_flagged: pd.DataFrame, sample_flagged: p return annotated_omegas[['gene', 'sample', 'impact', 'mutations', 'dnds', 'pvalue', 'lower', 'upper', + 'pvalue_adj', 'flagged', 'flag_reason']] @@ -271,26 +272,8 @@ def main(omegas_file: str, compiled_flagged_files: str, output: str) -> None: lines = [ln.strip() for ln in fh if ln.strip()] flagged_paths = [Path(l) for l in lines] - # Read omegas with resilience to missing header lines - # Some aggregation steps may drop the header; if so, re-read with explicit names - def _read_omegas(path: Path) -> pd.DataFrame: - try: - df = pd.read_csv(path, sep="\t", header=0, dtype=str, skip_blank_lines=True) - except pd.errors.EmptyDataError: - return pd.DataFrame(columns=["gene","sample","impact","mutations","dnds","pvalue","lower","upper"]) # empty - # If expected columns are missing (e.g., header was dropped), re-read with names - expected = {"gene","sample","impact","mutations","dnds","pvalue","lower","upper"} - if not expected.issubset(set(map(str, df.columns))): - df = pd.read_csv(path, - sep="\t", - header=None, - names=["gene","sample","impact","mutations","dnds","pvalue","lower","upper"], - dtype=str, - skip_blank_lines=True) - return df.fillna("") - # Read omegas - omegas = _read_omegas(omegas_path) + omegas = pd.read_csv(omegas_path, sep="\t", header=0) syn_flagged_sample, syn_flagged_gene, npa_flagged_sample, npa_flagged_gene = load_flagged_tables(flagged_paths) diff --git a/bin/check_contamination.py b/bin/check_contamination.py index a004360a..a882a136 100755 --- a/bin/check_contamination.py +++ b/bin/check_contamination.py @@ -6,435 +6,661 @@ import click -import pandas as pd import matplotlib.pyplot as plt +import pandas as pd +import logging import seaborn as sns +import operator from read_utils import custom_na_values - +from utils_filter import germline_mask, somatic_mask + +# Logging +logging.basicConfig( + format="%(asctime)s | %(levelname)s | %(name)s - %(message)s", + level=logging.INFO, + datefmt="%m/%d/%Y %I:%M:%S %p" +) + +LOG = logging.getLogger("check_contamination") + +# Constants +GERMLINE_LABEL = "Germline Samples" +SOMATIC_LABEL = "Somatic Samples" +CONTAMINATION_PROPORTION_THRESHOLD = 0.5 +# Deliberately more restrictive than the --somatic-vaf-boundary used for the between-samples +# analysis: this is the numerator of a QC metric, so only confidently low-VAF calls should count. +SNP_CONTAMINATION_VAF_THRESHOLD = 0.05 +VAF_COLUMNS = ["VAF", "vd_VAF", "VAF_AM"] +VARIANT_DETAIL_COLUMNS = [ + "SAMPLE_ID", + "MUT_ID", + "canonical_SYMBOL", + "ALT_DEPTH", + "DEPTH", + "VAF", + "canonical_Consequence_broader", + "FILTER", +] +OPS = { + ">": operator.gt, + "<": operator.lt, + ">=": operator.ge, + "<=": operator.le +} # Assuming somatic_variants and germline_variants are loaded as pandas DataFrames def compute_shared_variants(somatic_variants, germline_variants): - """ - # Example usage: - # shared_variants_matrix = compute_shared_variants(somatic_variants, germline_variants) + """Count mutations shared between each somatic sample and each germline sample. + + Parameters + ---------- + somatic_variants : pd.DataFrame + Variant table with at least ``SAMPLE_ID`` and ``MUT_ID`` columns, + providing the somatic side of the comparison. + germline_variants : pd.DataFrame + Variant table with at least ``SAMPLE_ID`` and ``MUT_ID`` columns, + providing the germline side of the comparison. + + Returns + ------- + pd.DataFrame + Integer matrix indexed by somatic ``SAMPLE_ID`` (rows) and germline + ``SAMPLE_ID`` (columns); each cell is the number of ``MUT_ID`` values + shared between that pair of samples. """ - unique_somatic_samples = sorted(somatic_variants['SAMPLE_ID'].unique()) - unique_germline_samples = sorted(germline_variants['SAMPLE_ID'].unique()) + unique_somatic_samples = sorted(somatic_variants["SAMPLE_ID"].unique()) + unique_germline_samples = sorted(germline_variants["SAMPLE_ID"].unique()) # Create a DataFrame to store counts (avoid .fillna downcasting warning) shared_counts = pd.DataFrame(0, index=unique_somatic_samples, columns=unique_germline_samples, dtype=int) # Iterate through each somatic sample for somatic_sample in unique_somatic_samples: - somatic_mutations = set(somatic_variants[somatic_variants['SAMPLE_ID'] == somatic_sample]['MUT_ID']) + somatic_mutations = set(somatic_variants[somatic_variants["SAMPLE_ID"] == somatic_sample]["MUT_ID"]) # Compare with germline mutations of all other samples for germline_sample in unique_germline_samples: - germline_mutations = set(germline_variants[germline_variants['SAMPLE_ID'] == germline_sample]['MUT_ID']) + germline_mutations = set(germline_variants[germline_variants["SAMPLE_ID"] == germline_sample]["MUT_ID"]) # Count shared mutations shared_counts.loc[somatic_sample, germline_sample] = len(somatic_mutations & germline_mutations) return shared_counts - - - -def contamination_detection_between_samples(maf_df, somatic_maf_df): - - # this is if we were to consider both unique and no-unique variants - vaf_threshold = 0.2 - germline_vars_all_samples = maf_df.loc[(maf_df["VAF"] > vaf_threshold) & (maf_df["vd_VAF"] > vaf_threshold) & (maf_df["VAF_AM"] > vaf_threshold), - ["SAMPLE_ID", "MUT_ID"]].drop_duplicates() - - print(germline_vars_all_samples["MUT_ID"].shape) - print(len(germline_vars_all_samples["MUT_ID"].unique())) - - - somatic_variants = somatic_maf_df[["SAMPLE_ID", "MUT_ID"]] - print(somatic_variants.shape) - - - all_variants = maf_df[["SAMPLE_ID", "MUT_ID"]] - print(all_variants.shape) - - - ## Somatic vs Germline - - shared_variants_somatic2germline_matrix = compute_shared_variants(somatic_variants, germline_vars_all_samples) - - plt.figure(figsize=(18, 15)) - - # Compute total number of germline mutations per sample - germline_counts = germline_vars_all_samples["SAMPLE_ID"].value_counts().reindex(shared_variants_somatic2germline_matrix.columns) - - # Create custom column labels with germline mutation counts - col_labels = [f"(n={germline_counts[col]}) {col}" for col in shared_variants_somatic2germline_matrix.columns] - - # Build annotation DataFrame without using deprecated DataFrame.applymap - mask = shared_variants_somatic2germline_matrix > 30 - annot = shared_variants_somatic2germline_matrix.where(mask) - # convert selected values to nullable int then to string, replace missing with empty string - annot = annot.round(0).astype('Int64').astype(str).replace('', '').fillna('') - +def create_heatmap(variants_matrix: pd.DataFrame, + annot: pd.DataFrame, + col_labels: list | None, + xlabel: str, + ylabel: str, + title: str, + output_file: str, + annot_kws_color: str = "black", + size: tuple[int, int] = (18, 15)): + """Create heatmap for specified set of mutations.""" + plt.figure(figsize=size) sns.heatmap( - shared_variants_somatic2germline_matrix, + variants_matrix, annot=annot, fmt="", cmap="Blues", - cbar_kws={'label': 'Shared Mutations'}, - xticklabels=col_labels, - yticklabels=shared_variants_somatic2germline_matrix.index, + cbar_kws={"label": "Shared Mutations"}, + xticklabels=col_labels if col_labels is not None else "auto", + yticklabels=variants_matrix.index, linewidths=0.5, - annot_kws={"color": "black", "fontsize": 10} + annot_kws={"color": annot_kws_color, "fontsize": 10}, ) - plt.xlabel("Germline Samples", fontsize=14) - plt.ylabel("Somatic Samples", fontsize=14) - plt.title("Somatic mutations that are germline in other samples", fontsize=16) - plt.savefig("somatic_vs_germline.pdf", bbox_inches = 'tight', dpi = 100) - plt.show() - - + plt.xlabel(xlabel, fontsize=14) + plt.ylabel(ylabel, fontsize=14) + plt.title(title, fontsize=16) + plt.savefig(output_file, bbox_inches="tight", dpi=100) + plt.close() +def prepare_datasets(maf_df, somatic_maf_df, somatic_vaf_boundary): + """Prepare datasets for contamination analysis. + + Parameters + ---------- + maf_df : pd.DataFrame + Full mutation table for all samples (used to derive germline variants and the + all-variants set), with at least ``SAMPLE_ID``, ``MUT_ID``, ``VAF``, ``vd_VAF`` and + ``VAF_AM`` columns. + somatic_maf_df : pd.DataFrame + Filtered somatic mutation table, with at least ``SAMPLE_ID`` and ``MUT_ID`` columns. + somatic_vaf_boundary : float + VAF threshold passed to ``germline_mask`` to identify germline variants (a variant is + germline when all of ``VAF``/``vd_VAF``/``VAF_AM`` exceed it). + + Returns + ------- + tuple + Tuple containing: + - germline_vars_all_samples: DataFrame of germline variants across all samples. + - somatic_variants: DataFrame of somatic variants. + - all_variants: DataFrame of all variants. + """ + # Consider both unique and non-unique variants when collecting germline variants + germline_vars_all_samples = maf_df.loc[ + germline_mask(maf_df, somatic_vaf_boundary), ["SAMPLE_ID", "MUT_ID"] + ].drop_duplicates() + LOG.info(f"Total variants: {germline_vars_all_samples.shape}") + LOG.info(f"Unique germline variants: {len(germline_vars_all_samples['MUT_ID'].unique())}") + + somatic_variants = somatic_maf_df[["SAMPLE_ID", "MUT_ID"]] + LOG.info(f"Somatic variants: {somatic_variants.shape}") + all_variants = maf_df[["SAMPLE_ID", "MUT_ID"]] + LOG.info(f"All variants: {all_variants.shape}") + + return germline_vars_all_samples, somatic_variants, all_variants + + +def two_way_comparison(df_a: pd.DataFrame, df_b: pd.DataFrame, annotation_threshold: float, operator_str: str, normalize: bool) -> tuple[pd.DataFrame, pd.DataFrame, list[str]]: + """Detect cross-sample contamination, compare one set of mutations with another set of mutations. + + Parameters + ---------- + df_a : pd.DataFrame + First DataFrame of variants (e.g. somatic); its samples become the rows of the matrix. + df_b : pd.DataFrame + Second DataFrame of variants (e.g. germline); its samples become the columns of the + matrix and provide the per-sample counts used for the labels and the normalization. + annotation_threshold : float + Threshold for annotating shared variants. Ignored when ``normalize`` is True, where + cells are annotated when the proportion falls between 0.8 and 1. + operator_str : str + String representing the comparison operator (e.g., ">", "<", ">=", "<="). + normalize : bool + Whether to divide the shared counts by the number of variants of each df_b sample. + """ + shared_df = compute_shared_variants(df_a, df_b) - ## All vs Germline + # Compute total number of counts of mutations per sample -> the columns of shared_df are the samples of df_b + counts = ( + df_b["SAMPLE_ID"].value_counts().reindex(shared_df.columns) + ) - shared_all_vs_germline_variants_matrix = compute_shared_variants(all_variants, germline_vars_all_samples) + if normalize: + shared_df = shared_df.divide( + counts, axis=1 + ) - # Compute total number of germline mutations per sample - germline_counts = germline_vars_all_samples["SAMPLE_ID"].value_counts().reindex(shared_all_vs_germline_variants_matrix.columns) + # Create custom column labels with mutation counts + col_labels = [f"(n={counts[col]}) {col}" for col in shared_df.columns] + # Build annotation DataFrame without using deprecated DataFrame.applymap + if normalize: + mask = (shared_df > 0.8) & ( + shared_df < 1) + else: + mask = OPS[operator_str](shared_df, annotation_threshold) + + annot = shared_df.where(mask) + if normalize: + annot = annot.round(2).astype("string").fillna("") + else: + # convert selected values to nullable int then to string, replace missing with empty string + annot = annot.round(0).astype("Int64").astype("string").replace("", "").fillna("") + + return shared_df, annot, col_labels + +def find_contaminated_pairs(proportion_matrix: pd.DataFrame, counts_matrix: pd.DataFrame, threshold: float) -> pd.DataFrame: + """Identify receiver samples carrying the germline variants of another sample. + + A sample is flagged as a receiver when its highest proportion of non-shared germline + variants coming from any single other sample exceeds ``threshold``. All sources tied at + that maximum are reported. + + Parameters + ---------- + proportion_matrix : pd.DataFrame + Matrix of the proportion of each source sample's non-shared germline variants that are + present as non-germline variants in each receiver sample (receivers as rows, sources as + columns). + counts_matrix : pd.DataFrame + Matrix of the corresponding raw shared variant counts, with the same shape and labels. + threshold : float + Minimum proportion above which a receiver sample is considered contaminated. + + Returns + ------- + pd.DataFrame + One row per contaminated receiver, with columns ``SAMPLE_ID``, + ``MAX_PROPORTION_GERMLINE_FROM_SOURCE`` and ``SOURCE_SAMPLEID_COUNTS``, the latter + holding the list of ``(shared_variant_count, source_sample_id)`` pairs tied at the + maximum proportion. + """ + max_prop_per_sample = proportion_matrix.max(axis="columns") - normalized_shared_all_vs_germline_variants_matrix = shared_all_vs_germline_variants_matrix.divide(germline_counts, axis=1) + receiver_source_pairs = [] + for sample, max_val in max_prop_per_sample[max_prop_per_sample > threshold].items(): + sample_vals = proportion_matrix.loc[sample, :] + sample_vals_count = counts_matrix.loc[sample, :] + source_sampleids = sample_vals[sample_vals == max_val].index.values + receiver_source_pairs.append( + ( + sample, + round(max_val, 3), + list(zip([sample_vals_count[x].item() for x in source_sampleids], source_sampleids)), + ) + ) + LOG.info( + f"{sample} has {max_val:.2f} proportion of the germline variants of {source_sampleids[0]} " + f"with a VAF not corresponding to germline variants " + f"(shared variants count: {sample_vals_count[source_sampleids[0]]})." + ) - # Count shared mutations between somatic and germline samples + return pd.DataFrame( + receiver_source_pairs, columns=["SAMPLE_ID", "MAX_PROPORTION_GERMLINE_FROM_SOURCE", "SOURCE_SAMPLEID_COUNTS"] + ) - plt.figure(figsize=(18, 15)) +def contaminated_pairs_to_long(contaminated_pairs: pd.DataFrame) -> pd.DataFrame: + """Expand the tied-source lists into one row per receiver/source pair. - # Create custom column labels with germline mutation counts - col_labels = [f"(n={germline_counts[col]}) {col}" for col in normalized_shared_all_vs_germline_variants_matrix.columns] + Parameters + ---------- + contaminated_pairs : pd.DataFrame + Output of ``find_contaminated_pairs``; must be non-empty. + Returns + ------- + pd.DataFrame + Long-format table with columns ``SAMPLE_ID``, ``MAX_PROPORTION_GERMLINE_FROM_SOURCE``, + ``SHARED_VARIANT_COUNT`` and ``SOURCE_SAMPLEID``. + """ + contaminated_pairs_long = contaminated_pairs.explode("SOURCE_SAMPLEID_COUNTS") + contaminated_pairs_long[["SHARED_VARIANT_COUNT", "SOURCE_SAMPLEID"]] = pd.DataFrame( + contaminated_pairs_long["SOURCE_SAMPLEID_COUNTS"].tolist(), index=contaminated_pairs_long.index + ) - # Annotation: show rounded values only when 0.8 < x < 1 - cond = (normalized_shared_all_vs_germline_variants_matrix > 0.8) & (normalized_shared_all_vs_germline_variants_matrix < 1) - annot = normalized_shared_all_vs_germline_variants_matrix.where(cond) - annot = annot.round(2).astype('string').fillna('') + return contaminated_pairs_long.drop(columns="SOURCE_SAMPLEID_COUNTS") + + +def export_germline_variants_in_receiver(maf_df: pd.DataFrame, + germline_vars_all_samples: pd.DataFrame, + receiver_sample: str, + source_sample: str): + """Write the source sample's germline variants alongside their VAF in the receiver sample. + + Parameters + ---------- + maf_df : pd.DataFrame + Full mutation table for all samples, with at least the ``VARIANT_DETAIL_COLUMNS``. + germline_vars_all_samples : pd.DataFrame + Germline variants across all samples, with at least ``SAMPLE_ID`` and ``MUT_ID`` columns. + receiver_sample : str + Sample suspected of being contaminated. + source_sample : str + Sample suspected of being the contamination source. + """ + variant_details = maf_df[VARIANT_DETAIL_COLUMNS] - sns.heatmap(normalized_shared_all_vs_germline_variants_matrix, - annot=annot, - fmt="", - cmap="Blues", - cbar_kws={'label': 'Shared Mutations'}, - xticklabels=col_labels, yticklabels=normalized_shared_all_vs_germline_variants_matrix.index, - annot_kws={"color": "white", "fontsize": 10}, - linewidths=0.5) + receiver_variants = variant_details[variant_details["SAMPLE_ID"] == receiver_sample].drop("SAMPLE_ID", axis=1) - plt.xlabel("Germline Samples", fontsize = 14) - plt.ylabel("All mutations samples", fontsize = 14) - plt.title("All mutations that are germline in other samples", fontsize = 16) - plt.savefig("allmutations_vs_germline.pdf", bbox_inches = 'tight', dpi = 100) - plt.show() + source_germline = germline_vars_all_samples[germline_vars_all_samples["SAMPLE_ID"] == source_sample] + source_variants = variant_details[ + (variant_details["SAMPLE_ID"] == source_sample) & (variant_details["MUT_ID"].isin(source_germline["MUT_ID"].values)) + ].drop("SAMPLE_ID", axis=1) + merged_samples = receiver_variants.merge( + source_variants, + on=["MUT_ID", "canonical_SYMBOL", "canonical_Consequence_broader"], + suffixes=("_dest", "_source"), + how="right", + ) + merged_samples.sort_values(by=["VAF_dest"], ascending=False).to_csv( + f"{source_sample}.germline_variants_in.{receiver_sample}.tsv", header=True, sep="\t", index=False + ) +def contamination_detection_between_samples(maf_df, somatic_maf_df, somatic_vaf_boundary): + """Detect cross-sample contamination by comparing somatic and germline mutations. + + Parameters + ---------- + maf_df : pd.DataFrame + Full mutation table for all samples (used to derive germline variants and the + all-variants set), with at least ``SAMPLE_ID``, ``MUT_ID``, ``VAF``, ``vd_VAF`` and + ``VAF_AM`` columns. + somatic_maf_df : pd.DataFrame + Filtered somatic mutation table, with at least ``SAMPLE_ID`` and ``MUT_ID`` columns. + somatic_vaf_boundary : float + VAF threshold passed to ``germline_mask`` to identify germline variants (a variant is + germline when all of ``VAF``/``vd_VAF``/``VAF_AM`` exceed it). + """ + # Prepare datasets + germline_vars_all_samples, somatic_variants, all_variants = prepare_datasets(maf_df, somatic_maf_df, somatic_vaf_boundary) + ## Somatic vs Germline + shared_variants_somatic2germline_matrix, somatic_vs_germline_annot, somatic_vs_germline_col_labels = two_way_comparison(somatic_variants, germline_vars_all_samples, annotation_threshold=30, operator_str=">", normalize=False) + create_heatmap( + variants_matrix=shared_variants_somatic2germline_matrix, + annot=somatic_vs_germline_annot, + col_labels=somatic_vs_germline_col_labels, + xlabel=GERMLINE_LABEL, + ylabel=SOMATIC_LABEL, + title="Somatic mutations that are germline in other samples", + output_file="somatic_vs_germline.pdf") ## Germline vs Germline + shared_germline_variants_matrix, germline_vs_germline_annot, germline_vs_germline_col_labels = two_way_comparison(germline_vars_all_samples, germline_vars_all_samples, annotation_threshold=0, operator_str="<", normalize=False) - shared_germline_variants_matrix = compute_shared_variants(germline_vars_all_samples, germline_vars_all_samples) - - plt.figure(figsize=(18, 15)) - - # Compute total number of germline mutations per sample - germline_counts = germline_vars_all_samples["SAMPLE_ID"].value_counts().reindex(shared_germline_variants_matrix.columns) - - # Create custom column labels with germline mutation counts - col_labels = [f"(n={germline_counts[col]}) {col}" for col in shared_germline_variants_matrix.columns] - - - # Annotation: follow original logic (keep values where < 0, else blank) - mask = shared_germline_variants_matrix < 0 - annot = shared_germline_variants_matrix.where(mask) - annot = annot.astype('string').fillna('') - - sns.heatmap(shared_germline_variants_matrix, - annot=annot, - fmt="", - cmap="Blues", - cbar_kws={'label': 'Shared Mutations'}, - xticklabels=col_labels, yticklabels=shared_germline_variants_matrix.index, - linewidths=0.5, - annot_kws={"fontsize": 8} - ) - - plt.xlabel("Germline Samples", fontsize = 14) - plt.ylabel("Germline Samples", fontsize = 14) - plt.title("Germline mutations that are germline in other samples", fontsize = 16) - plt.savefig("germline_vs_germline.pdf", bbox_inches = 'tight', dpi = 100) - plt.show() - - - - - - # Compute total number of germline mutations per sample - germline_counts = germline_vars_all_samples["SAMPLE_ID"].value_counts().reindex(shared_germline_variants_matrix.columns) - - normalized_share_germline_vs_germline_variants_matrix = shared_germline_variants_matrix.divide(germline_counts, axis=1) - - - plt.figure(figsize=(18, 15)) - - - # Create custom column labels with germline mutation counts - col_labels = [f"(n={germline_counts[col]}) {col}" for col in normalized_share_germline_vs_germline_variants_matrix.columns] - - - cond = (normalized_share_germline_vs_germline_variants_matrix > 0.8) & (normalized_share_germline_vs_germline_variants_matrix < 1) - annot = normalized_share_germline_vs_germline_variants_matrix.where(cond) - annot = annot.round(2).astype('string').fillna('') - - sns.heatmap(normalized_share_germline_vs_germline_variants_matrix, - annot=annot, - fmt="", - cmap="Blues", - cbar_kws={'label': 'Shared Mutations'}, - xticklabels=col_labels, yticklabels=normalized_share_germline_vs_germline_variants_matrix.index, - annot_kws={"color": "white", "fontsize": 10}, - linewidths=0.5) - - plt.xlabel("Germline Samples", fontsize = 14) - plt.ylabel("Germline samples", fontsize = 14) - plt.savefig("normalized.germline_vs_germline.pdf", bbox_inches = 'tight', dpi = 100) - plt.show() + create_heatmap( + variants_matrix=shared_germline_variants_matrix, + annot=germline_vs_germline_annot, + col_labels=germline_vs_germline_col_labels, + xlabel=GERMLINE_LABEL, + ylabel=GERMLINE_LABEL, + title="Germline mutations that are germline in other samples", + output_file="germline_vs_germline.pdf") + ## All vs Germline + normalized_shared_all_vs_germline_variants_matrix, all_vs_germline_annot, all_vs_germline_col_labels = two_way_comparison(all_variants, germline_vars_all_samples, annotation_threshold=0.3, operator_str=">", normalize=True) + + create_heatmap( + variants_matrix=normalized_shared_all_vs_germline_variants_matrix, + annot=all_vs_germline_annot, + col_labels=all_vs_germline_col_labels, + xlabel=GERMLINE_LABEL, + ylabel="All mutations samples", + title="All mutations that are germline in other samples", + output_file="allmutations_vs_germline.pdf", + annot_kws_color="white") + + # Germline mutations + normalized_germline_vs_germline_variants_matrix, norm_germline_vs_germline_annot, norm_germline_vs_germline_col_labels = two_way_comparison(germline_vars_all_samples, germline_vars_all_samples, annotation_threshold=0, operator_str="<", normalize=True) + create_heatmap( + variants_matrix=normalized_germline_vs_germline_variants_matrix, + annot=norm_germline_vs_germline_annot, + col_labels=norm_germline_vs_germline_col_labels, + xlabel=GERMLINE_LABEL, + ylabel=GERMLINE_LABEL, + title="Normalized germline mutations that are germline in other samples", + output_file="normalized.germline_vs_germline.pdf", + annot_kws_color="white") ## Somatic vs Remaining Germline + shared_all_vs_germline_variants_matrix = compute_shared_variants(all_variants, germline_vars_all_samples) shared_somatic_to_non_shared_germline = shared_all_vs_germline_variants_matrix - shared_germline_variants_matrix # Those cases where the number of mutations is smaller than 5 are set to 0 shared_somatic_to_non_shared_germline[shared_somatic_to_non_shared_germline < 5] = 0 # Compute total number of germline mutations per sample - germline_counts = germline_vars_all_samples["SAMPLE_ID"].value_counts().reindex(shared_somatic_to_non_shared_germline.columns) - - - total_germline_available_per_sample = (germline_counts - shared_germline_variants_matrix) - - shared_somatic_to_non_shared_germline_proportion = (shared_somatic_to_non_shared_germline / total_germline_available_per_sample).fillna(0) - + germline_counts = ( + germline_vars_all_samples["SAMPLE_ID"].value_counts().reindex(shared_somatic_to_non_shared_germline.columns) + ) + total_germline_available_per_sample = germline_counts - shared_germline_variants_matrix - plt.figure(figsize=(22, 18)) + shared_somatic_to_non_shared_germline_proportion = ( + shared_somatic_to_non_shared_germline / total_germline_available_per_sample + ).fillna(0) cond = shared_somatic_to_non_shared_germline_proportion > 0.45 annot = shared_somatic_to_non_shared_germline_proportion.where(cond) - annot = annot.round(2).astype('string').fillna('') - - sns.heatmap(shared_somatic_to_non_shared_germline_proportion, - annot=annot, - fmt="", - cmap="Blues", - cbar_kws={'label': 'Shared Mutations'}, - # xticklabels=col_labels, - yticklabels=shared_somatic_to_non_shared_germline_proportion.index, - annot_kws={"color": "black", "fontsize": 10}, - linewidths=0.5) - - plt.xlabel("Non-shared germline", fontsize = 14) - plt.ylabel("Somatic", fontsize = 14) - plt.title("Somatic mutations that are germline in other samples", fontsize = 16) - plt.savefig("contamination.somatic_vs_remaininggermline.pdf", bbox_inches = 'tight', dpi = 100) - plt.show() - plt.close() - - - plt.figure(figsize=(22, 18)) + annot = annot.round(2).astype("string").fillna("") + create_heatmap( + variants_matrix=shared_somatic_to_non_shared_germline_proportion, + annot=annot, + col_labels=None, + xlabel="Non-shared germline", + ylabel="Somatic", + title="Somatic mutations that are germline in other samples", + output_file="contamination.somatic_vs_remaininggermline.pdf", + annot_kws_color="white", + size=(22, 18) + ) + cond = shared_somatic_to_non_shared_germline > 0 annot = shared_somatic_to_non_shared_germline.where(cond) # convert to nullable int then string, replace missing with empty string - annot = annot.round(0).astype('Int64').astype('string').replace('', '').fillna('') - - sns.heatmap(shared_somatic_to_non_shared_germline, - annot=annot, - fmt="", - cmap="Blues", - cbar_kws={'label': 'Shared Mutations'}, - # xticklabels=col_labels, - yticklabels=shared_somatic_to_non_shared_germline.index, - annot_kws={"color": "black", "fontsize": 10}, - linewidths=0.5) - - plt.xlabel("Non-shared germline", fontsize = 14) - plt.ylabel("Somatic", fontsize = 14) - plt.title("Somatic mutations that are germline in other samples (count)", fontsize = 16) - plt.savefig("contamination.somatic_vs_remaininggermline.numbers.pdf", bbox_inches = 'tight', dpi = 100) - plt.show() - plt.close() - + annot = annot.round(0).astype("Int64").astype("string").replace("", "").fillna("") - max_prop_per_sample = shared_somatic_to_non_shared_germline_proportion.max(axis = 'columns') + create_heatmap( + variants_matrix=shared_somatic_to_non_shared_germline, + annot=annot, + col_labels=None, + xlabel="Non-shared germline", + ylabel="Somatic", + title="Somatic mutations that are germline in other samples (count)", + output_file="contamination.somatic_vs_remaininggermline.numbers.pdf", + size=(22, 18) + ) ## Exploration of contaminated samples - receiver_source_pairs = [] - for sample, max_val in max_prop_per_sample[max_prop_per_sample>0.5].reset_index().values: - sample_vals = shared_somatic_to_non_shared_germline_proportion.loc[sample,:] - sample_vals_count = shared_somatic_to_non_shared_germline.loc[sample,:] + contaminated_pairs = find_contaminated_pairs( + shared_somatic_to_non_shared_germline_proportion, + shared_somatic_to_non_shared_germline, + CONTAMINATION_PROPORTION_THRESHOLD, + ) + contaminated_pairs.to_csv("contaminated_samples.detailed.tsv", header=True, sep="\t", index=False) - source_sampleids = sample_vals[sample_vals == max_val].index.values - source_sampleid = source_sampleids[0] - receiver_source_pairs.append((sample, round(max_val,3), - list(zip([sample_vals_count[x].item() for x in source_sampleids], source_sampleids)))) - - print(f'{sample} has {max_val:.2f} proportion of the germline variants of {source_sampleid} as with a VAF not corresponding to germline variants.') - print(f'Shared variants count: {sample_vals_count[source_sampleid]}') - print() - - - subseeeet = maf_df[["SAMPLE_ID", "MUT_ID", 'canonical_SYMBOL', "ALT_DEPTH", "DEPTH", "VAF", 'canonical_Consequence_broader', 'FILTER']] - p_dest = subseeeet[subseeeet["SAMPLE_ID"] == sample].drop("SAMPLE_ID", axis = 1) - - p_source_germ = germline_vars_all_samples[germline_vars_all_samples["SAMPLE_ID"] == source_sampleid] - p_source = subseeeet[(subseeeet["SAMPLE_ID"] == source_sampleid) - & (subseeeet["MUT_ID"].isin(p_source_germ["MUT_ID"].values)) - ].drop("SAMPLE_ID", axis = 1) - - merged_samples = p_dest.merge(p_source, - on = ["MUT_ID", 'canonical_SYMBOL', 'canonical_Consequence_broader'], - suffixes = ("_dest", "_source"), - how = 'right' - ) - - merged_samples.sort_values(by =["VAF_dest"], ascending=False - ).to_csv(f"{source_sampleid}.germline_variants_in.{sample}.tsv", - header = True, - sep = '\t', - index = False) - - # plt.figure(figsize=(8, 6)) - # plt.scatter(x = merged_samples["VAF_dest"].fillna(0), - # y = merged_samples["VAF_source"].fillna(0), - # # color = ['blue' if x == 0 else 'red' for x in merged_samples["VAF_dest"].fillna(0)] - # ) - - # plt.xscale('log') - # # plt.yscale('log') - # plt.xlabel("VAF_dest " + sample) - # plt.ylabel("VAF_source " + source_sampleid) - # plt.savefig(f"{source_sampleid}_germline_in_{sample}_VAF_scatter.pdf", bbox_inches = 'tight', dpi = 100) - # plt.show() - - # Store contamination results - contamination_detailed_df = pd.DataFrame(receiver_source_pairs, - columns=["SAMPLE_ID", "MAX_PROPORTION_GERMLINE_FROM_SOURCE", "SOURCE_SAMPLEID_COUNTS"]) - contamination_detailed_df.to_csv(f"contaminated_samples.detailed.tsv", - header = True, - sep = '\t', - index = False) - - if contamination_detailed_df.empty: - print("No contaminated samples detected.") + if contaminated_pairs.empty: + LOG.info("No contaminated samples detected.") return - contamination_detailed_df_long = contamination_detailed_df.explode("SOURCE_SAMPLEID_COUNTS") - expanded_df = pd.DataFrame(contamination_detailed_df_long["SOURCE_SAMPLEID_COUNTS"].tolist()) - expanded_df.columns = ["SHARED_VARIANT_COUNT", "SOURCE_SAMPLEID"] - contamination_detailed_df_long["SHARED_VARIANT_COUNT"] = expanded_df["SHARED_VARIANT_COUNT"].values - contamination_detailed_df_long["SOURCE_SAMPLEID"] = expanded_df["SOURCE_SAMPLEID"].values - contamination_detailed_df_long = contamination_detailed_df_long.drop("SOURCE_SAMPLEID_COUNTS", axis = 1) - contamination_detailed_df_long.to_csv(f"contaminated_samples.detailed.long.tsv", - header = True, - sep = '\t', - index = False) + for receiver in contaminated_pairs.itertuples(): + # Only the first of the sources tied at the maximum proportion is reported in detail + source_sampleid = receiver.SOURCE_SAMPLEID_COUNTS[0][1] + export_germline_variants_in_receiver(maf_df, germline_vars_all_samples, receiver.SAMPLE_ID, source_sampleid) + + contaminated_pairs_to_long(contaminated_pairs).to_csv( + "contaminated_samples.detailed.long.tsv", header=True, sep="\t", index=False + ) def data_loading(maf_path, somatic_maf_path): + """Load the full and somatic MAF tables, keeping only covered SNVs. + + Parameters + ---------- + maf_path : str + Path to the full MAF file; rows flagged ``FILTER.not_covered`` are dropped and only + ``TYPE == "SNV"`` rows are kept. + somatic_maf_path : str + Path to the filtered somatic MAF file; only ``TYPE == "SNV"`` rows are kept. + + Returns + ------- + tuple + ``(maf_df, somatic_maf_df)`` — the filtered full mutation table and the filtered + somatic mutation table, both as ``pd.DataFrame``. + """ maf_df = pd.read_table(maf_path, na_values=custom_na_values) print(maf_df.shape) - maf_df = maf_df[~(maf_df["FILTER.not_covered"]) - & (maf_df["TYPE"] == 'SNV') - ].reset_index() + maf_df = maf_df[~(maf_df["FILTER.not_covered"]) & (maf_df["TYPE"] == "SNV")].reset_index() print(maf_df.shape) somatic_maf_df = pd.read_table(somatic_maf_path, na_values=custom_na_values) print(somatic_maf_df.shape) - somatic_maf_df = somatic_maf_df[(somatic_maf_df["TYPE"] == 'SNV')] + somatic_maf_df = somatic_maf_df[(somatic_maf_df["TYPE"] == "SNV")] print(somatic_maf_df.shape) return maf_df, somatic_maf_df -def contamination_detection_in_snps(maf): - - snp_positions_maf = maf[maf["FILTER.gnomAD_SNP"]][ - ["SAMPLE_ID", "MUT_ID", "VAF"] - ].reset_index(drop = True) - - # being very restrictive in the VAF to count the occurrences of potentially contaminated mutations - somatic_snp_positions_maf = snp_positions_maf[snp_positions_maf["VAF"] < 0.05].reset_index(drop = True) - germline_snp_positions_maf = snp_positions_maf[snp_positions_maf["VAF"] >= 0.05].reset_index(drop = True) - - unique_SNP_positions = snp_positions_maf["MUT_ID"].unique() - number_unique_SNP_positions = len(unique_SNP_positions) - - sample_SNP_mutation_freq = [] - for sample in snp_positions_maf["SAMPLE_ID"].unique(): - germline_count = len(germline_snp_positions_maf[germline_snp_positions_maf["SAMPLE_ID"] == sample]) - somatic_count = len(somatic_snp_positions_maf[somatic_snp_positions_maf["SAMPLE_ID"] == sample]) - remaining_germline = number_unique_SNP_positions-germline_count - sample_SNP_mutation_freq.append([sample, - germline_count, - remaining_germline, - somatic_count, - somatic_count / remaining_germline if remaining_germline > 0 else 1 - ]) - sample_SNP_mutation_freq_df = pd.DataFrame(sample_SNP_mutation_freq) - sample_SNP_mutation_freq_df.columns = ["SAMPLE_ID", "germline_count", "remaining_germline", "somatic_count", "prop_somatic_SNPs"] - - # identify outliers in the "prop_somatic_SNPs" column - sample_SNP_mutation_freq_df = sample_SNP_mutation_freq_df.sort_values(by = "prop_somatic_SNPs", ascending = False) - sample_SNP_mutation_freq_df.to_csv("sample_SNP_mutation_freq.tsv", header = True, sep = '\t', index = False) - +def prepare_snp_datasets(maf: pd.DataFrame, vaf_threshold: float) -> tuple[pd.DataFrame, pd.DataFrame, pd.DataFrame]: + """Split the known SNP positions into somatic-looking and germline-looking variants. + + Germline is defined as the complement of somatic rather than through ``germline_mask``, so + that the two sets partition the SNP rows. A variant whose VAF estimates disagree (some at or + below the threshold, some above) is therefore counted as germline: it is not a confident + somatic call, so it must not be left out of both sets and silently inflate the denominator of + ``prop_somatic_SNPs``. + + Rows with an undefined VAF cannot be classified either way and are dropped, since + ``somatic_mask`` is False for them and the germline complement would otherwise absorb them. + + Parameters + ---------- + maf : pd.DataFrame + Full mutation table with at least ``SAMPLE_ID``, ``MUT_ID``, the ``VAF_COLUMNS`` and the + boolean ``FILTER.gnomAD_SNP`` column. + vaf_threshold : float + Upper bound (inclusive) on all VAF estimates for a variant to be called somatic. + + Returns + ------- + tuple + Tuple containing: + - snp_positions_maf: DataFrame of all classifiable variants at known SNP positions. + - somatic_snp_positions_maf: DataFrame of the somatic ones. + - germline_snp_positions_maf: DataFrame of the remaining ones. + """ + snp_positions_maf = maf.loc[maf["FILTER.gnomAD_SNP"], ["SAMPLE_ID", "MUT_ID", *VAF_COLUMNS]] + + undefined_vaf = snp_positions_maf[VAF_COLUMNS].isna().any(axis="columns") + if undefined_vaf.any(): + LOG.warning(f"Discarding {undefined_vaf.sum()} SNP variants with an undefined {VAF_COLUMNS} VAF.") + snp_positions_maf = snp_positions_maf.loc[~undefined_vaf] + LOG.info(f"SNP variants: {snp_positions_maf.shape}") + + is_somatic = somatic_mask(snp_positions_maf, vaf_threshold) + somatic_snp_positions_maf = snp_positions_maf.loc[is_somatic] + germline_snp_positions_maf = snp_positions_maf.loc[~is_somatic] + LOG.info(f"Somatic SNP variants: {somatic_snp_positions_maf.shape}") + LOG.info(f"Germline SNP variants: {germline_snp_positions_maf.shape}") + + return snp_positions_maf, somatic_snp_positions_maf, germline_snp_positions_maf + + +def compute_snp_somatic_proportion(snp_positions_maf: pd.DataFrame, + somatic_snp_positions_maf: pd.DataFrame, + germline_snp_positions_maf: pd.DataFrame) -> pd.DataFrame: + """Compute, per sample, the proportion of non-germline SNP positions that look somatic. + + Parameters + ---------- + snp_positions_maf : pd.DataFrame + All classifiable variants at known SNP positions, as returned by ``prepare_snp_datasets``. + somatic_snp_positions_maf : pd.DataFrame + The somatic subset of ``snp_positions_maf``. + germline_snp_positions_maf : pd.DataFrame + The germline subset of ``snp_positions_maf``. + + Returns + ------- + pd.DataFrame + One row per sample with columns ``SAMPLE_ID``, ``germline_count``, ``remaining_germline``, + ``somatic_count`` and ``prop_somatic_SNPs``, sorted by the latter in descending order. + """ + samples = snp_positions_maf["SAMPLE_ID"].unique() + number_unique_snp_positions = snp_positions_maf["MUT_ID"].nunique() + + germline_count = germline_snp_positions_maf["SAMPLE_ID"].value_counts().reindex(samples, fill_value=0) + somatic_count = somatic_snp_positions_maf["SAMPLE_ID"].value_counts().reindex(samples, fill_value=0) + + # NOTE: the denominator counts every SNP position of the cohort that is not germline in this + # sample, including positions where the sample has no variant at all. Those can never end up in + # the numerator, so `prop_somatic_SNPs` is driven as much by how many SNP positions were called + # in a sample as by how many of them look somatic. It is comparable across samples of a cohort + # sequenced with the same panel, not an absolute contamination rate. + remaining_germline = number_unique_snp_positions - germline_count + + sample_snp_mutation_freq_df = pd.DataFrame( + { + "germline_count": germline_count, + "remaining_germline": remaining_germline, + "somatic_count": somatic_count, + "prop_somatic_SNPs": (somatic_count / remaining_germline).where(remaining_germline > 0, 1), + } + ).rename_axis("SAMPLE_ID").reset_index() + + return sample_snp_mutation_freq_df.sort_values(by="prop_somatic_SNPs", ascending=False) + + +def create_snp_proportion_plot(sample_snp_mutation_freq_df: pd.DataFrame, output_file: str): + """Plot the distribution of the per-sample proportion of somatic SNPs across the cohort. + + Parameters + ---------- + sample_snp_mutation_freq_df : pd.DataFrame + Per-sample table with a ``prop_somatic_SNPs`` column, as returned by + ``compute_snp_somatic_proportion``. + output_file : str + Path of the plot to write. + """ plt.figure(figsize=(6, 3)) - sns.violinplot(data=sample_SNP_mutation_freq_df, x="prop_somatic_SNPs", - fill= False, color="lightgray", inner=None) - sns.swarmplot(data=sample_SNP_mutation_freq_df, x="prop_somatic_SNPs", color="black", size=3) + sns.violinplot(data=sample_snp_mutation_freq_df, x="prop_somatic_SNPs", fill=False, color="lightgray", inner=None) + sns.swarmplot(data=sample_snp_mutation_freq_df, x="prop_somatic_SNPs", color="black", size=3) plt.title("Proportion of all SNPs across samples\ndetected as somatic") plt.xlabel("Proportion of somatic SNPs per sample") plt.ylabel("Density") - plt.savefig("sample_SNP_mutation_freq.pdf", dpi=300, bbox_inches="tight") + plt.savefig(output_file, dpi=300, bbox_inches="tight") plt.close() -@click.command() -@click.option('--maf_path', type=click.Path(exists=True), required=True, help='Path to the MAF file.') -@click.option('--somatic_maf', type=click.Path(exists=True), required=True, help='Path to the filtered somatic mutations file.') -def main(maf_path, somatic_maf): +def contamination_detection_in_snps(maf): + """Estimate per-sample contamination from the VAF distribution at known SNP positions. + + Restricts to gnomAD SNP positions, splits them into somatic-looking and germline-looking + sets by VAF, computes the per-sample proportion of SNP positions that look somatic, writes + the resulting table, and plots its distribution across samples. + + Parameters + ---------- + maf : pd.DataFrame + Full mutation table with at least ``SAMPLE_ID``, ``MUT_ID``, the ``VAF_COLUMNS`` and the + boolean ``FILTER.gnomAD_SNP`` column. """ - CLI entry point for assessing contamination between samples using germline and somatic mutations. + snp_positions_maf, somatic_snp_positions_maf, germline_snp_positions_maf = prepare_snp_datasets( + maf, SNP_CONTAMINATION_VAF_THRESHOLD + ) + + sample_snp_mutation_freq_df = compute_snp_somatic_proportion( + snp_positions_maf, somatic_snp_positions_maf, germline_snp_positions_maf + ) + sample_snp_mutation_freq_df.to_csv("sample_SNP_mutation_freq.tsv", header=True, sep="\t", index=False) + + create_snp_proportion_plot(sample_snp_mutation_freq_df, "sample_SNP_mutation_freq.pdf") + + +@click.command() +@click.option("--maf_path", type=click.Path(exists=True), required=True, help="Path to the MAF file.") +@click.option( + "--somatic_maf", type=click.Path(exists=True), required=True, help="Path to the filtered somatic mutations file." +) +@click.option( + "--somatic-vaf-boundary", + type=float, + default=0.3, + show_default=True, + help="VAF boundary for somatic variants; a variant with all of VAF/vd_VAF/VAF_AM above it is germline.", +) +def main(maf_path, somatic_maf, somatic_vaf_boundary): + """Assess cross-sample contamination using germline and somatic mutations. + + Loads the input tables and runs both the between-samples and the SNP-based contamination + analyses, writing their tables and plots to the current working directory. + + Parameters + ---------- + maf_path : str + Path to the full MAF file. + somatic_maf : str + Path to the filtered somatic mutations file. + somatic_vaf_boundary : float, optional + VAF boundary separating somatic from germline variants. Default is 0.3. """ - + maf_df, somatic_maf_df = data_loading(maf_path, somatic_maf) print("Running contamination analysis between samples") - contamination_detection_between_samples(maf_df, somatic_maf_df) + contamination_detection_between_samples(maf_df, somatic_maf_df, somatic_vaf_boundary) print("Running general contamination analysis") contamination_detection_in_snps(maf_df) - -if __name__ == '__main__': - +if __name__ == "__main__": main() diff --git a/bin/compare_trinucleotide_proportions.py b/bin/compare_trinucleotide_proportions.py index b88a306a..d07dc7e2 100755 --- a/bin/compare_trinucleotide_proportions.py +++ b/bin/compare_trinucleotide_proportions.py @@ -22,14 +22,19 @@ def plot_trinucleotide_proportions(wgs_counts_file): counts_all = wgs_counts.copy() for cnsq in ["all", "non_protein_affecting", "introns_intergenic", "exons_splice_sites"]: - wgs_counts_cnsq = pd.read_table(f"consensus.{cnsq}.tsv", - header = 0, - sep = '\t', - usecols = ["CHROM", "POS", "CONTEXT_MUT", "CONTEXT"] - ) - wgs_counts_cnsq = wgs_counts_cnsq.drop_duplicates() - counts_panel_cnsq = wgs_counts_cnsq["CONTEXT"].value_counts().to_frame(name = f'COUNT_{cnsq}').reset_index() - counts_all = counts_all.merge(counts_panel_cnsq, on = 'CONTEXT') + try: + wgs_counts_cnsq = pd.read_table(f"consensus.{cnsq}.tsv", + header = 0, + sep = '\t', + usecols = ["CHROM", "POS", "CONTEXT_MUT", "CONTEXT"] + ) + wgs_counts_cnsq = wgs_counts_cnsq.drop_duplicates() + counts_panel_cnsq = wgs_counts_cnsq["CONTEXT"].value_counts().to_frame(name = f'COUNT_{cnsq}').reset_index() + counts_all = counts_all.merge(counts_panel_cnsq, on = 'CONTEXT') + + except FileNotFoundError: + print(f"File not found for consensus panel: {cnsq}") + counts_all = counts_all.set_index("CONTEXT") proportions_all = counts_all / counts_all.sum() @@ -42,8 +47,12 @@ def plot_trinucleotide_proportions(wgs_counts_file): for i, cnsq in enumerate(["all", "non_protein_affecting", "introns_intergenic", "exons_splice_sites"]): ax = axs[i] - - rmse = np.sqrt(((proportions_all_plot["COUNT_WGS"] - proportions_all_plot[f"COUNT_{cnsq}"])**2).mean()) + try : + rmse = np.sqrt(((proportions_all_plot["COUNT_WGS"] - proportions_all_plot[f"COUNT_{cnsq}"])**2).mean()) + + except KeyError: + print(f"KeyError: 'COUNT_{cnsq}' not found in proportions_all_plot. Skipping RMSE calculation for this panel.") + continue # Scatter plot sns.scatterplot(data=proportions_all_plot, @@ -87,14 +96,17 @@ def plot_trinucleotide_proportions(wgs_counts_file): ax = axs[i] # rmse = np.sqrt(((counts_all_plot["COUNT_WGS"] - counts_all_plot[f"COUNT_{cnsq}"])**2).mean()) - - sns.scatterplot(data=counts_all_plot, - x="COUNT_WGS", - y=f"COUNT_{cnsq}", - hue="CONTEXT", - legend=False, - ax=ax) - + try: + sns.scatterplot(data=counts_all_plot, + x="COUNT_WGS", + y=f"COUNT_{cnsq}", + hue="CONTEXT", + legend=False, + ax=ax) + except KeyError: + print(f"KeyError: 'COUNT_{cnsq}' not found in counts_all_plot. Skipping scatter plot for this panel.") + continue + # Annotate points with CONTEXT for namee, row in counts_all_plot.iterrows(): ax.text(row["COUNT_WGS"], row[f"COUNT_{cnsq}"], namee, @@ -113,7 +125,10 @@ def plot_trinucleotide_proportions(wgs_counts_file): @click.option('--wgs-trinucleotide', type=click.Path(exists=True), help='Input trinucleotide counts file for WGS') def main(wgs_trinucleotide): click.echo("Comparing the trinucleotide proportions...") - plot_trinucleotide_proportions(wgs_trinucleotide) + try: + plot_trinucleotide_proportions(wgs_trinucleotide) + except Exception as e: + click.echo(f"Error occurred: {e}") if __name__ == '__main__': main() diff --git a/bin/compute_hotspots_selection.py b/bin/compute_hotspots_selection.py new file mode 100755 index 00000000..059ee2ab --- /dev/null +++ b/bin/compute_hotspots_selection.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python + +import click +import numpy as np +import pandas as pd +import scipy.stats as stats +import glob +import os + +from utils import MIN_NONZERO_PVALUE +from omega_comparison_per_site import poisson_pvalue, benjamini_hochberg + +def load_panel_hotspots(panel_file, hotspots_file): + panel = pd.read_csv(panel_file, sep="\t", compression="gzip") + panel["MUTTYPE"] = panel["CONTEXT_MUT"].str[1] + panel["CONTEXT_MUT"].str[-2:] + + hotspots = pd.read_csv(hotspots_file, sep="\t").drop_duplicates() + + # Some hotspots might have MUTTYPE = '-' for non-SNV, but let's assume standard format for now + hotspots_panel = panel.merge(hotspots, on=['CHROM', 'POS', 'MUTTYPE'], how='inner') + + return hotspots_panel + +def process_comparison(comparison_file, hotspots_panel, size_type, output_prefix): + comp_df = pd.read_csv(comparison_file, sep="\t", compression="gzip") + + if size_type == "site": + group_cols = ['CHROM', 'POS', 'REF', 'ALT', 'GENE'] + elif size_type == "aminoacid": + group_cols = ['GENE', 'Feature', 'Protein_position'] + elif size_type == "aminoacid_change": + group_cols = ['GENE', 'Feature', 'Protein_position', 'Amino_acids'] + else: + raise ValueError(f"Unknown size type: {size_type}") + + # Determine which entries are hotspots + # Get the unique hotspots for the current grouping + hotspots_grouped = hotspots_panel[group_cols].drop_duplicates() + hotspots_grouped['Hotspot'] = 'Yes' + + # Merge with the comparison data + comp_df = comp_df.merge(hotspots_grouped, on=group_cols, how='left') + comp_df['Hotspot'] = comp_df['Hotspot'].fillna('No') + + # Group by GENE and Hotspot + grouped_size = comp_df.groupby(['GENE', 'Hotspot']).size().to_frame(name='Count') + grouped_size_mut = comp_df.groupby(['GENE', 'Hotspot']).agg({'OBSERVED_MUTS': lambda x: (x != 0).sum()}).reset_index().rename(columns={'OBSERVED_MUTS': 'CountDiffMutatedSites'}) + grouped_size = grouped_size.merge(grouped_size_mut, on=['GENE', 'Hotspot']) + grouped = comp_df.groupby(['GENE', 'Hotspot'])[['OBSERVED_MUTS', 'EXPECTED_MUTS']].sum().reset_index() + + grouped["OBS/EXP"] = (grouped["OBSERVED_MUTS"] / grouped["EXPECTED_MUTS"]).fillna(0) + grouped["OBS/EXP"] = grouped["OBS/EXP"].replace([np.inf, -np.inf], 0) + grouped["p_value"] = grouped.apply(lambda row: poisson_pvalue(row["OBSERVED_MUTS"], row["EXPECTED_MUTS"]), axis=1) + + grouped["p_value"] = grouped["p_value"].replace(0, MIN_NONZERO_PVALUE) + grouped = grouped.merge(grouped_size, on=['GENE', 'Hotspot']) + grouped = grouped[["GENE", "Hotspot", "Count", "CountDiffMutatedSites", "OBSERVED_MUTS", "EXPECTED_MUTS", "OBS/EXP", "p_value"]] + + # TO BE FIXED: Adjusted p-values are not being calculated correctly. The following code is commented out for now. + # grouped["p_value_adj"] = np.nan + # for gene, gene_df in grouped.groupby("GENE", dropna=False): + # valid_mask = gene_df["p_value"].notna() + # if not valid_mask.any(): + # continue + # ## something is wrong here with the adjusted p-values, they are not being assigned correctly. Let's fix that. + # adjusted = benjamini_hochberg(gene_df.loc[valid_mask, "p_value"].to_numpy()) + # grouped.loc[gene_df.index[valid_mask], "p_value_adj"] = adjusted + + # Write output + output_file = f"{output_prefix}.{size_type}.hotspots_selection.tsv.gz" + grouped.to_csv(output_file, sep="\t", index=False) + click.echo(f"Results for size '{size_type}' written to {output_file}") + + +@click.command() +@click.option('--comparisons', multiple=True, type=click.Path(exists=True), required=True, help="Path to comparison files.") +@click.option('--panel-file', type=click.Path(exists=True), required=True, help="Path to captured panel file (gzip compressed).") +@click.option('--hotspots-file', type=click.Path(exists=True), required=True, help="Path to hotspots file.") +@click.option('--output-prefix', type=str, required=True, help="Output file prefix.") +def main(comparisons, panel_file, hotspots_file, output_prefix): + """Compute selection for known hotspots based on site comparison output.""" + hotspots_panel = load_panel_hotspots(panel_file, hotspots_file) + + for comp_file in comparisons: + filename = os.path.basename(comp_file) + if ".site.comparison." in filename: + size_type = "site" + elif ".aminoacid_change.comparison." in filename: + size_type = "aminoacid_change" + elif ".aminoacid.comparison." in filename: + size_type = "aminoacid" + else: + click.echo(f"Warning: Could not determine size type from filename {filename}. Skipping.") + continue + + click.echo(f"Processing size: {size_type} from {comp_file}") + process_comparison(comp_file, hotspots_panel, size_type, output_prefix) + +if __name__ == "__main__": + main() diff --git a/bin/concat_profiles.py b/bin/concat_profiles.py index 621978f1..476569ad 100755 --- a/bin/concat_profiles.py +++ b/bin/concat_profiles.py @@ -98,8 +98,9 @@ def compile_profiles(mutation_profile_files, groups_json): if len(keys_sizegt1) > 1: plot_similarity_heatmaps(mut_profile_matrix, mode, keys_sizegt1, "groups") - - plot_similarity_heatmaps(mut_profile_matrix, mode, all_keys, "all") + + if len(all_keys) > 1: + plot_similarity_heatmaps(mut_profile_matrix, mode, all_keys, "all") @click.command() diff --git a/bin/cordblood_mutrate_genome_trint_corrected.R b/bin/cordblood_mutrate_genome_trint_corrected.R deleted file mode 100644 index 3b9dccde..00000000 --- a/bin/cordblood_mutrate_genome_trint_corrected.R +++ /dev/null @@ -1,186 +0,0 @@ -library(Hmisc) -library(tidyr) -library(stringr) -library(dplyr, warn = FALSE) -library(ggplot2) -library(jsonlite) -library(Biostrings) -library(data.table) - -genome_content_json <- "../data/genome_counts_tribases.json" - -deepCSA_run <- "2025-12-15" -add_nanoseq_for_comparison = TRUE #if you want to add cord blood sequenced with Nanoseq for comparison - -# root_dir <- "../data/cord_blood_run" -root_dir <- "/data/bbg/nobackup2/prominent/duplex_seq_tests/error_rate/cord_blood/bbg/deepCSA/2026-03-26_deepUMIcaller_and_dupcaller_4_paper" -deepCSA_run_dir <- root_dir # paste0(root_dir,"deepCSA/", deepCSA_run) - -path2sites = paste0( deepCSA_run_dir, "/depths/individual/") -path2mutations = paste0( deepCSA_run_dir, "/mutations/clean_somatic/") -consensus_bed <- paste0( deepCSA_run_dir, "/regions/consensuspanels/consensus.all.bed") - -path2out = "results/cordblood_mutrate" - -get_genome_content <- function(genome_content_json){ - #' Get genome trinucleotide content - #' - #' This function reads trinucleotide counts from a JSON genome composition file - #' and computes the 96 pyrimidine-centered contexts. - - #' As input uses path to the json file with genome trinucleotide content - - genome_json <- read_json(genome_content_json) - df_sites_genome <- data.frame(Context = names(genome_json), sites = unlist(genome_json), stringsAsFactors = FALSE) - df_sites_genome <- df_sites_genome %>% - filter(!grepl("N",Context)) - df_sites_genome$Compl_context <- as.character(reverseComplement(DNAStringSet(df_sites_genome$Context))) - df_sites_genome$CONTEXT <- ifelse(substr(df_sites_genome$Context,2,2) %in% c("T","C"), - df_sites_genome$Context, - df_sites_genome$Compl_context) - df_sites_genome_agg = df_sites_genome %>% - group_by(CONTEXT) %>% - summarise(N_sites_genome= sum(sites)) - df_sites_genome_agg <- as.data.frame(df_sites_genome_agg) - message("Contexts in genome sites = ", nrow(df_sites_genome_agg)) - return(df_sites_genome_agg) -} - - -get_consensus_sites_depth <- function(sample, depth_path, consensus_bed){ - #' Get depth per position for positions in consensus panel - #' - #' This function intersects consensus panel with file with annotated depth per position - - # Load depth data - dt_pos <- fread(paste0(depth_path, sample, ".depths.annotated.tsv.gz")) - colnames(dt_pos) <- c("CHROM", "POS", "CONTEXT", "DEPTH") - dt_pos[, `:=`( - start = POS, - end = POS - )] - setkey(dt_pos, CHROM, start, end) - # Load consensus bed - dt_bed <- fread(consensus_bed, col.names = c("CHROM", "start", "end")) - setkey(dt_bed, CHROM, start, end) - # Overlap - hits <- foverlaps(dt_pos, dt_bed, type = "any", nomatch = 0L) - consensus_depth <- hits[, .(CHROM, POS, CONTEXT, DEPTH)] - return(consensus_depth) -} - -get_mutations_and_sites <- function(path2sites, path2mutations, consensus_bed, genome_sites_df){ - #' Get number of mutations per sample and normalize panel content to whole genome content - #' - #' This function gets number of mutations per sample in each context and normalizes panel content to whole genome content - result <- NULL - sites_files <- list.files(path=path2sites, pattern=glob2rx("*.depths.annotated.tsv.gz")) - sites_files <- sites_files[sites_files != 'all_samples.depths.annotated.tsv.gz'] - for(file in sites_files){ - sample_name = str_split_i(file, ".depths", 1) - print(sample_name) - - df_sites = get_consensus_sites_depth(sample_name, path2sites, consensus_bed) - df_sites$depth = as.numeric(df_sites$DEPTH) - df_sites_panel_agg = df_sites %>% - group_by(CONTEXT) %>% - summarise(N = sum(DEPTH)) - df_sites_panel_agg = as.data.frame(df_sites_panel_agg) - colnames(df_sites_panel_agg) <- c("CONTEXT", "N_sites_panel") - message("Contexts in panel sites = ", nrow(df_sites_panel_agg)) - print(head(df_sites_panel_agg)) - - df_sites = merge(genome_sites_df, df_sites_panel_agg, by="CONTEXT") - print(head(df_sites)) - message("Contexts in panel and genome sites = ", nrow(df_sites)) - df_sites <- df_sites %>% mutate(proportion_genome=N_sites_genome/sum(N_sites_genome)) - df_sites <- df_sites %>% mutate(proportion_panel=N_sites_panel/sum(N_sites_panel)) - df_sites$ratio2genome = df_sites$proportion_panel/df_sites$proportion_genome - print(head(df_sites)) - print(sum(df_sites$N_sites_panel)) - - df_mutations = read.table(paste(path2mutations, sample_name, ".somatic.mutations.tsv", sep=""), header=TRUE, sep="\t") - df_mutations = df_mutations[df_mutations$TYPE=="SNV",] - df_mutations$CONTEXT = str_split_i(df_mutations$CONTEXT_MUT, ">", 1) - df_mutations_agg = df_mutations %>% - group_by(CONTEXT) %>% - summarise(N_mut = n()) - print(head(df_mutations_agg)) - df_mutations = as.data.frame(df_mutations) - result_sample = merge(df_sites, df_mutations_agg, by="CONTEXT", all=TRUE) - # if some contexts are absent in mutataions - keep them but put mutation number to 0 - if (nrow(result_sample[is.na(result_sample$N_mut),]) > 0){ - result_sample[is.na(result_sample$N_mut),]$N_mut <- 0 - } - print(head(result_sample)) - message("Contexts in df with mutations = ", nrow(result_sample)) - result_sample$N_mut_corrected = result_sample$N_mut * result_sample$ratio2genome - sample_out <- c(sample_name, sum(result_sample$N_mut), sum(result_sample$N_mut_corrected), sum(df_sites$N_sites_panel), "sample", sample_name) - result <- rbind(result, sample_out) - - } - return(result) -} - -#Download df with genome trinucleotide contexts -df_sites_genome_agg <- get_genome_content(genome_content_json) - -#Get number of mutations and normalize panel contetnt to genome content -result <- get_mutations_and_sites(path2sites, path2mutations, consensus_bed, df_sites_genome_agg) -print(head(result)) - -# Add nanoseq rate for comparison if needed -if (add_nanoseq_for_comparison == TRUE){ - result <- rbind(result, c("PD48442_cordblood_nanoseqv2", 41, 38.43131656, 2799554062, "Nanoseq_Sanger", "PD48442")) - result <- rbind(result, c("PD47269_cordblood_nanoseqv2", 26, 28.29844724, 2003725667, "Nanoseq_Sanger", "PD47269")) -} -result <- as.data.frame(result) -colnames(result) <- c("sample", "N_mut", "N_mut_corrected", "DEPTH", "protocol", "donor_id") -result$N_mut_corrected = as.numeric(result$N_mut_corrected) -result$DEPTH = as.numeric(result$DEPTH) -result$N_mut = as.numeric(result$N_mut) -result$mutrate_observed = result$N_mut_corrected/result$DEPTH -result$mutrate_observed_per_MB <- result$mutrate_observed * 10**6 - -# why the indices of the two boundaries are different? is this desired or a typo? -result$mutrate_CI_high <- apply(result, 1, function(x) binconf(as.numeric(x["N_mut_corrected"]), as.numeric(x["DEPTH"]), alpha=0.05, method=c("wilson","exact","asymptotic","all"), include.x=FALSE, include.n=FALSE, return.df=FALSE)[3]) -result$mutrate_CI_low <- apply(result, 1, function(x) binconf(as.numeric(x["N_mut_corrected"]), as.numeric(x["DEPTH"]), alpha=0.05, method=c("wilson","exact","asymptotic","all"), include.x=FALSE, include.n=FALSE, return.df=FALSE)[2]) -result$Muts_per_cell <- result$mutrate_observed*2*sum(df_sites_genome_agg$N_sites_genome) -print((result)) - -message("Output will be written to ", path2out) -if (!dir.exists(path2out)) { - dir.create(path2out, recursive = TRUE) -} - -write.csv(result, paste0(path2out,"/mutrates_results.tsv"),row.names = FALSE, quote = FALSE) - -setwd(path2out) - -jpeg(filename=paste("mutrate_trint_corrected.with_nanoseq.jpeg", sep=""), width=30, height=15, res=300, units='cm') - -ggplot(result, aes(x=sample, y=mutrate_observed)) + - geom_bar(stat="identity", position="dodge", fill="grey") + - geom_errorbar(aes(x=sample, ymin=mutrate_CI_low, ymax=mutrate_CI_high), width=0.4, alpha=0.9, linewidth=1, position=position_dodge(.9)) + - theme_bw() + - facet_grid(~factor(protocol, levels=c("tests","IDT","TWS","Nanoseq_Sanger")), scales="free_x", space="free_x") + - theme(axis.text.x = element_text(angle = 90, hjust=1)) + - geom_text(aes(label = N_mut, x = sample, y = mutrate_observed), position = position_dodge(width = 0.9), vjust = -0.5, hjust = -0.1) + - xlab("") + - ylab("Mutation rate") + - theme(legend.position="bottom") -dev.off() - -jpeg(filename=paste("mutrate_trint_corrected.jpeg", sep=""), width=30, height=15, res=300, units='cm') -result<-result[result$protocol != "Nanoseq_Sanger",] -ggplot(result, aes(x=sample, y=mutrate_observed)) + - geom_bar(stat="identity", position="dodge", fill="grey") + - geom_errorbar(aes(x=sample, ymin=mutrate_CI_low, ymax=mutrate_CI_high), width=0.4, alpha=0.9, linewidth=1, position=position_dodge(.9)) + - theme_bw() + - facet_grid(~factor(protocol, levels=c("tests","IDT","TWS","Nanoseq_Sanger")), scales="free_x", space="free_x") + - theme(axis.text.x = element_text(angle = 90, hjust=1)) + - geom_text(aes(label = N_mut, x = sample, y = mutrate_observed), position = position_dodge(width = 0.9), vjust = -0.5, hjust = -0.1) + - xlab("") + - ylab("Mutation rate") + - theme(legend.position="bottom") -dev.off() \ No newline at end of file diff --git a/bin/dNdS_run.R b/bin/dNdS_run.R index 46cf20bc..fa410bc5 100755 --- a/bin/dNdS_run.R +++ b/bin/dNdS_run.R @@ -22,13 +22,13 @@ option_list = list( help="sample name/identifier of the run", metavar="character"), make_option(c("-i", "--inputfile"), type="character", default=NULL, help="mutation dataset file name", metavar="character"), - make_option(c("-o", "--outputfile"), type="character", default=NULL, + make_option(c("-o", "--outputprefix"), type="character", default="dNdScv_output", help="output file name [default= %default]", metavar="character"), make_option(c("-r", "--referencetranscripts"), type="character", - default="/workspace/projects/prominent/analysis/dNdScv/data/reference_files/RefCDS_human_latest_intogen.rda", + default="RefCDS.rda", help="Annotation reference file [default= %default]", metavar="character"), make_option(c("-c", "--covariates"), type="character", - default="/workspace/projects/prominent/analysis/dNdScv/data/reference_files/covariates_hg19_hg38_epigenome_pcawg.rda", + default="covariates_hg19_hg38_epigenome_pcawg.rda", help="Human GRCh38 covariates file [default= %default]", metavar="character"), make_option(c("-g", "--genelist"), type="character", default=NULL, @@ -93,12 +93,25 @@ if (!is.null(opt$genelist)){ } - +# CDKN2A.p16INK4a # Loads the covs object load(opt$covariates) load(opt$referencetranscripts) -reference_genes <- intersect(rownames(covs), unique(gr_genes$names)) +# Remove CDKN2A.p14arf row and +# rename CDKN2A.p16INK4a to CDKN2A +if ("CDKN2A.p14arf" %in% rownames(covs)) { + covs <- covs[rownames(covs) != "CDKN2A.p14arf", ] +} +if ("CDKN2A.p16INK4a" %in% rownames(covs)) { + rownames(covs)[rownames(covs) == "CDKN2A.p16INK4a"] <- "CDKN2A" +} + +reference_genes <- intersect( + unique(rownames(covs)), + unique(gr_genes$names) + ) + # Identify genes that are in 'genes' but not in the row names of 'covs' missing_genes <- setdiff(genes, reference_genes) @@ -172,7 +185,7 @@ if (!is.null(dnds_genes) && nrow(dnds_genes) > 0) { # Write to file if dnds_genes is still valid if (nrow(dnds_genes) > 0) { write.table(dnds_genes, - file = opt$outputfile, + file = paste(opt$outputprefix, '.dNdScv.cv.tsv', sep = ''), sep = "\t", row.names = FALSE, quote = FALSE) @@ -185,7 +198,7 @@ if (!is.null(dnds_genes) && nrow(dnds_genes) > 0) { dnds_genes <- cbind(list("sample" = opt$samplename), dnds_genes) write.table(dnds_genes, - file = paste(opt$outputfile, 'globaldnds', sep = ''), + file = paste(opt$outputprefix, '.dNdScv.globaldnds.tsv', sep = ''), sep = "\t", row.names = FALSE, quote = FALSE) @@ -197,7 +210,7 @@ if (!is.null(dnds_genes) && nrow(dnds_genes) > 0) { dnds_genes <- cbind(list("sample" = opt$samplename), dnds_genes) write.table(dnds_genes, - file = paste(opt$outputfile, 'loc', sep = ''), + file = paste(opt$outputprefix, '.dNdScv.loc.tsv', sep = ''), sep = "\t", row.names = FALSE, quote = FALSE) diff --git a/bin/dNdScv_panel_prep.py b/bin/dNdScv_panel_prep.py new file mode 100755 index 00000000..066eab9c --- /dev/null +++ b/bin/dNdScv_panel_prep.py @@ -0,0 +1,715 @@ +#!/usr/bin/env python3 +""" +Codon-align panel BED regions for dndscv build_refCDS.R compatibility. + +copied from : https://github.com/bbglab/intogen-plus-dsl2/blob/dev/build-refcds/bin/panel_reformat.py + +""" + +from __future__ import annotations + +import bisect +import csv +import gzip +import logging +import sys +from collections import defaultdict +from dataclasses import dataclass, field + +import click + +FORMAT = "%(asctime)s - %(name)s - %(levelname)s - %(message)s" +logging.basicConfig(level=logging.INFO, format=FORMAT) +logger = logging.getLogger(__name__) + +OUTPUT_FIELDS = [ + "GENE_ID", + "SYMBOL", + "PROTEIN_ID", + "CHR", + "START", + "END", + "CDS_START", + "CDS_END", + "CDS_LEN", + "STRAND", + "TRANSCRIPT_ID", + "EXON_CHR_START", + "EXON_CHR_END", +] + + +@dataclass +class Exon: + """ + A single CDS exon with genomic and transcript-relative coordinates. + """ + + start: int # CDS region genomic start (1-based, closed) + end: int # CDS region genomic end (1-based, closed) + cds_start: int # first position within transcript CDS (1-based) + cds_end: int # last position within transcript CDS (1-based) + exon_chr_start: int # full exon genomic start (may extend into UTR) + exon_chr_end: int # full exon genomic end (may extend into UTR) + + +@dataclass +class Gene: + """A gene with metadata and CDS exons in 5'→3' transcript order. + + Exons are sorted by ``cds_start`` ascending, which gives transcript + order for both forward and reverse strand genes because ``cds_start=1`` + always marks the 5'-most exon in the transcript. + """ + + gene_id: str + symbol: str + protein_id: str + chrom: str + strand: int + cds_len: int + transcript_id: str + exons: list[Exon] = field(default_factory=list) + + @classmethod + def from_rows(cls, gene_id: str, rows: list[dict[str, str]]) -> Gene: + """Build a Gene from annotation TSV rows sharing the same symbol. + + Gene-level fields are taken from the first row. One Exon is created + per row and the list is sorted by ``cds_start`` to enforce transcript + order. + """ + if not rows: + raise ValueError(f"Gene {gene_id!r} has no annotation rows") + first = rows[0] + exons: list[Exon] = [] + for r in rows: + genomic_start = int(r["START"]) + genomic_end = int(r["END"]) + chr_start_raw = r.get("EXON_CHR_START") + chr_end_raw = r.get("EXON_CHR_END") + exons.append( + Exon( + start=genomic_start, + end=genomic_end, + cds_start=int(r["CDS_START"]), + cds_end=int(r["CDS_END"]), + exon_chr_start=int(chr_start_raw) if chr_start_raw else genomic_start, + exon_chr_end=int(chr_end_raw) if chr_end_raw else genomic_end, + ) + ) + exons.sort(key=lambda e: e.cds_start) + return cls( + gene_id=gene_id, + symbol=first.get("SYMBOL", gene_id), + protein_id=first.get("PROTEIN_ID", ""), + chrom=first["CHR"], + strand=int(first["STRAND"]), + cds_len=int(first["CDS_LEN"]), + transcript_id=first.get("TRANSCRIPT_ID", ""), + exons=exons, + ) + + +@dataclass +class Fragment: + """An aligned genomic fragment ready for mock-transcript construction.""" + + start: int + end: int + exon: Exon + + @property + def length(self) -> int: + """Fragment length in bases (1-based closed).""" + return self.end - self.start + 1 + + @property + def at_exon_start(self) -> bool: + """True if this fragment reaches the exon's CDS start boundary.""" + return self.start <= self.exon.start + + @property + def at_exon_end(self) -> bool: + """True if this fragment reaches the exon's CDS end boundary.""" + return self.end >= self.exon.end + + +@dataclass +class PanelGene: + """Panel-covered fragments of a gene in 5'→3' transcript order. + + Built by intersecting a Gene's exons with panel regions and aligning + each overlap's 5' boundary to a codon boundary. The ``gene`` reference + provides access to gene-level metadata for downstream output. + """ + + gene: Gene + fragments: list[Fragment] + + @classmethod + def from_gene( + cls, + gene: Gene, + panel_regions: list[tuple[int, int]], + aligner: "CodonAligner", + ) -> "PanelGene": + """Intersect gene exons with panel regions and build a PanelGene. + + Exons are iterated in CDS order. For each overlap the natural + junction state with the previously collected fragment is computed + before calling the aligner, so the 5' decision is made once and + correctly. + + Within-exon hits are sorted in transcript direction before being + appended, ensuring the full list stays in 5'→3' order. + """ + fragments: list[Fragment] = [] + last_fragment: Fragment | None = None + + # Panel regions are sorted by start coordinate. Extract start + # coordinates once for bisect lookups. + panel_starts = [r[0] for r in panel_regions] + + for exon in gene.exons: + exon_fragments: list[Fragment] = [] + + # Use bisect to skip panel regions that end before this exon. + # Any region with start > exon.end cannot overlap, so we + # only scan from the first region whose start could reach + # back to exon.start. + lo = bisect.bisect_right(panel_starts, exon.end) + for idx in range(lo): + p_start, p_end = panel_regions[idx] + if p_end < exon.start: + continue + ov_start = max(p_start, exon.start) + ov_end = min(p_end, exon.end) + + is_junction = cls._check_5prime_junction( + last_fragment, ov_start, ov_end, exon, gene.strand, + ) + + if gene.strand == 1: + result = aligner.align_forward(ov_start, ov_end, exon, is_junction) + else: + result = aligner.align_reverse(ov_start, ov_end, exon, is_junction) + + if result is None: + logger.debug( + "Discarded fragment: overlap (%d, %d) collapsed after alignment", + ov_start, ov_end, + ) + continue + exon_fragments.append(Fragment(start=result[0], end=result[1], exon=exon)) + + # Sort within-exon hits in transcript order before appending. + if gene.strand == 1: + exon_fragments.sort(key=lambda f: f.start) + else: + exon_fragments.sort(key=lambda f: -f.end) + if exon_fragments: + last_fragment = exon_fragments[-1] + fragments.extend(exon_fragments) + + return cls(gene=gene, fragments=fragments) + + @staticmethod + def _check_5prime_junction( + prev: Fragment | None, + ov_start: int, + ov_end: int, + exon: Exon, + strand: int, + ) -> bool: + """Return True if the overlap forms a natural 5' junction with prev. + + All three conditions must hold: + 1. The previous fragment ends at its exon's natural 3' boundary. + 2. The overlap starts at the current exon's natural 5' boundary. + 3. The two exons are consecutive in the transcript CDS. + """ + if prev is None: + return False + if strand == 1: + prev_at_3prime = prev.end == prev.exon.end + curr_at_5prime = ov_start == exon.start + else: + prev_at_3prime = prev.start == prev.exon.start + curr_at_5prime = ov_end == exon.end + cds_adjacent = prev.exon.cds_end + 1 == exon.cds_start + return prev_at_3prime and curr_at_5prime and cds_adjacent + + +class CodonAligner: + """Align the 5' end of panel-exon overlaps to codon boundaries. + + Separate methods for forward and reverse strands. The general strategy + is pad-first (include uncovered CDS bases within the exon span) with + trim as fallback. At natural exon junctions the 5' coordinate is left + untouched so that the split codon spanning the junction is preserved. + + Only the 5' boundary is adjusted here. The 3' boundary is handled + during the sequential junction pass in TranscriptBuilder. + """ + + @staticmethod + def _phase(cds_anchor: int, distance: int) -> int: + """Return the codon phase (0, 1, or 2) of a coordinate.""" + return (cds_anchor + distance - 1) % 3 + + # -- Forward strand (5' = low genomic coordinate) ------------------------- + + @staticmethod + def _pad_or_trim_start_fwd(start: int, phase: int, exon_chr_start: int) -> int: + """Pad start backward into uncovered CDS or trim forward. + + Padding is preferred: extend by 1-2 bases to reach phase 0 within + the full exon span. If padding would exceed the exon boundary, trim + the covered region instead. + """ + if phase == 1: + padded = start - 1 + return padded if padded >= exon_chr_start else start + 2 + if phase == 2: + padded = start - 2 + return padded if padded >= exon_chr_start else start + 1 + return start + + def align_forward( + self, + panel_start: int, + panel_end: int, + exon: Exon, + is_5prime_junction: bool, + ) -> tuple[int, int] | None: + """Align the 5' (start) of a forward-strand overlap. + + Three cases: + 1. Natural junction: keep start (split codon continuation). + 2. At CDS boundary, no junction: trim leading orphan bases. + 3. Interior cut: pad backward within exon, trim as fallback. + + Returns adjusted (start, end) or None if the region collapses. + """ + if is_5prime_junction: + adj_start = panel_start + elif panel_start == exon.start: + phase = (exon.cds_start - 1) % 3 + trim = (3 - phase) % 3 + adj_start = panel_start + trim + else: + phase = self._phase(exon.cds_start, panel_start - exon.start) + adj_start = self._pad_or_trim_start_fwd(panel_start, phase, exon.exon_chr_start) + + if adj_start > panel_end: + return None + return (adj_start, panel_end) + + # -- Reverse strand (5' = high genomic coordinate) ------------------------ + + @staticmethod + def _pad_or_trim_end_rev(end: int, phase: int, exon_chr_end: int) -> int: + """Pad end forward into uncovered CDS or trim backward. + + Mirror of ``_pad_or_trim_start_fwd`` for the reverse strand. + """ + if phase == 1: + padded = end + 1 + return padded if padded <= exon_chr_end else end - 2 + if phase == 2: + padded = end + 2 + return padded if padded <= exon_chr_end else end - 1 + return end + + def align_reverse( + self, + panel_start: int, + panel_end: int, + exon: Exon, + is_5prime_junction: bool, + ) -> tuple[int, int] | None: + """Align the 5' (end) of a reverse-strand overlap. + + Three cases (mirrored from forward): + 1. Natural junction: keep end (split codon continuation). + 2. At CDS boundary, no junction: trim leading orphan bases. + 3. Interior cut: pad forward within exon, trim as fallback. + + Returns adjusted (start, end) or None if the region collapses. + """ + if is_5prime_junction: + adj_end = panel_end + elif panel_end == exon.end: + phase = (exon.cds_start - 1) % 3 + trim = (3 - phase) % 3 + adj_end = panel_end - trim + else: + phase = self._phase(exon.cds_start, exon.end - panel_end) + adj_end = self._pad_or_trim_end_rev(panel_end, phase, exon.exon_chr_end) + + if panel_start > adj_end: + return None + return (panel_start, adj_end) + + +class PanelLoader: + """Load panel regions from a BED file.""" + + @staticmethod + def load(path: str) -> dict[str, list[tuple[int, int]]]: + r"""Read a BED file and return regions indexed by chromosome. + + Coordinates are converted from 0-based half-open (] to 1-based closed []. + For downstream compatibility with dndscv, chromosome names have "chr" prefix removed if present. + + Notes + ----- + e.g. a BED line "chr1\t0\t100" becomes (1, 100) for chromosome "1". + """ + regions: dict[str, list[tuple[int, int]]] = defaultdict(list) + opener = gzip.open if path.endswith(".gz") else open + with opener(path, "rt", encoding="utf-8") as fh: + for line in fh: + line = line.strip() + if not line or line.startswith("#") or line.startswith("track"): + continue + parts = line.split("\t") + if len(parts) < 3: + continue + if not parts[1].lstrip("-").isdigit() or not parts[2].lstrip("-").isdigit(): + logger.debug("Skipping header line: %s", line) + continue + regions[parts[0].replace("chr", "")].append((int(parts[1]), int(parts[2]))) + # Sort regions by start coordinate on each chromosome so that + # downstream code can use bisect to find overlaps efficiently. + for chrom in regions: + regions[chrom].sort() + return dict(regions) + + +class GeneLoader: + """Load gene annotations from a TSV file.""" + + @staticmethod + def load(path: str) -> dict[str, Gene]: + """Read a gene annotation TSV and return Gene objects keyed by GENE_ID. + + Rows are grouped by GENE_ID first, then each group is passed to + ``Gene.from_rows`` which extracts gene-level metadata from the first + row and builds Exon objects sorted in transcript order. + """ + raw_rows: dict[str, list[dict[str, str]]] = defaultdict(list) + with open(path, encoding="utf-8") as fh: + for row in csv.DictReader(fh, delimiter="\t", fieldnames=OUTPUT_FIELDS): + raw_rows[row["GENE_ID"]].append(row) + genes: dict[str, Gene] = {} + for gene_id, rows in raw_rows.items(): + genes[gene_id] = Gene.from_rows(gene_id, rows) + return genes + + +class TranscriptBuilder: + """Align panel fragments to codon boundaries and build mock transcripts.""" + + def __init__(self, aligner: CodonAligner) -> None: + self._aligner = aligner + + @staticmethod + def _is_natural_junction(prev: Fragment, frag: Fragment, strand: int) -> bool: + """Check if two consecutive fragments form a natural CDS continuation. + + A natural junction means: + 1. The previous fragment's 3' end is at the exon's natural boundary. + 2. The current fragment's 5' start is at the exon's natural boundary. + 3. The two exons are adjacent in the original transcript's CDS. + """ + if strand == 1: + prev_3_natural = prev.end == prev.exon.end + curr_5_natural = frag.start == frag.exon.start + else: + prev_3_natural = prev.start == prev.exon.start + curr_5_natural = frag.end == frag.exon.end + cds_adjacent = prev.exon.cds_end + 1 == frag.exon.cds_start + return prev_3_natural and curr_5_natural and cds_adjacent + + @staticmethod + def _ensure_codon_junctions(fragments: list[Fragment], strand: int) -> list[Fragment]: + """Trim 3' orphan bases at artificial junctions and enforce total CDS length % 3 == 0. + + By the time fragments reach this method, ``align()`` (called from + ``PanelGene.from_gene``) has already resolved every 5' boundary: + - Natural junctions: 5' left untouched (split codon preserved). + - Natural CDS boundary without junction: leading orphans trimmed. + - Panel cuts into exon interior: pad/trim to phase 0. + + This method only needs to handle the 3' side: + 1. At an artificial junction (panel cut or non-adjacent exons), trim + trailing orphan bases from the 3' end of the preceding group so + the accumulated length is a multiple of 3 before the next phase-0 + fragment begins. + 2. At a natural junction (CDS-adjacent exons, both boundaries natural), + the split codon is biologically real; no trimming is applied. + 3. After the full sequential pass, trim the 3' end of the last + fragment to make the total CDS length a multiple of 3. + """ + if not fragments: + return fragments + + first = fragments[0] + running_bases = first.length if first.start <= first.end else 0 + + for i in range(1, len(fragments)): + frag = fragments[i] + prev = fragments[i - 1] + + natural = TranscriptBuilder._is_natural_junction(prev, frag, strand) + + if not natural: + # Trim trailing orphan bases from the 3' end of the preceding + # group. The next fragment already starts at phase 0 (handled + # by align()), so no 5' adjustment of frag is needed here. + orphan = running_bases % 3 + if orphan > 0: + if strand == 1: + prev.end -= orphan + else: + prev.start += orphan + running_bases -= orphan + logger.debug("Trimmed %d orphan base(s) at artificial junction", orphan) + + running_bases += frag.length + + # Final: trim total to multiple of 3 + remainder = running_bases % 3 + if remainder > 0 and fragments: + last = fragments[-1] + if strand == 1: + last.end -= remainder + else: + last.start += remainder + logger.debug("Trimmed %d trailing base(s) for total codon alignment", remainder) + + return [f for f in fragments if f.start <= f.end] + + @staticmethod + def _reindex(gene: Gene, fragments: list[Fragment]) -> list[dict[str, object]]: + """Assign continuous CDS pointers to ordered fragments.""" + rows: list[dict[str, object]] = [] + cumulative = 0 + for frag in fragments: + new_cds_start = cumulative + 1 + new_cds_end = cumulative + frag.length + cumulative = new_cds_end + rows.append( + { + "GENE_ID": gene.gene_id, + "SYMBOL": gene.symbol, + "PROTEIN_ID": gene.protein_id, + "CHR": "chr" + gene.chrom, + "START": frag.start, + "END": frag.end, + "CDS_START": new_cds_start, + "CDS_END": new_cds_end, + "CDS_LEN": 0, + "STRAND": gene.strand, + "TRANSCRIPT_ID": gene.transcript_id, + "EXON_CHR_START": frag.start, + "EXON_CHR_END": frag.end, + } + ) + for row in rows: + row["CDS_LEN"] = cumulative + return rows + + def build( + self, + gene: Gene, + panel_by_chr: dict[str, list[tuple[int, int]]], + ) -> tuple[list[dict[str, object]], list[Fragment]]: + """Process a single gene and return mock-transcript rows and fragments. + + The pipeline is three sequential steps, each operating on the + ``PanelGene`` produced in the first step: + + 1. Intersect the gene's exons with the panel and 5'-align each + overlap to a codon boundary → ``PanelGene`` + 2. Fix codon junctions between consecutive fragments in transcript + order, trimming orphan bases at artificial cuts. + 3. Reindex fragments into a continuous mock CDS. + + Returns (output_rows, final_fragments) so callers can inspect which + fragment boundaries are real vs. artificial for splice site reporting. + """ + panel_regions = panel_by_chr.get(gene.chrom, []) + + panel_gene = PanelGene.from_gene(gene, panel_regions, self._aligner) + if not panel_gene.fragments: + return [], [] + + panel_gene.fragments = self._ensure_codon_junctions(panel_gene.fragments, panel_gene.gene.strand) + if not panel_gene.fragments: + return [], [] + + rows = self._reindex(panel_gene.gene, panel_gene.fragments) + logger.debug("%s: %d fragment(s), total CDS length = %d", panel_gene.gene.symbol, len(rows), rows[0]["CDS_LEN"]) + return rows, panel_gene.fragments + + +SPLICE_REPORT_FIELDS = [ + "CHR", + "POSITION", + "GENE_SYMBOL", + "GENE_ID", + "JUNCTION_TYPE", + "SPLICE_TYPE", + "FRAGMENT_BOUNDARY", +] + + +class SpliceClassifier: + """Classify splice sites that dndscv will infer from the mock transcript. + + dndscv computes essential splice sites from consecutive CDS exon rows: + - Donor (5'ss): prev_END +1, +2, +5 (forward) or next_START -1, -2, -5 (reverse) + - Acceptor (3'ss): next_START -1, -2 (forward) or prev_END +1, +2 (reverse) + + Every junction between consecutive output fragments becomes a splice + site source. This class determines whether each junction is a REAL + exon-intron boundary or an ARTIFICIAL panel cut, allowing users to + post-filter dndscv splice mutation calls. + """ + + @staticmethod + def classify( + gene: Gene, + fragments: list[Fragment], + ) -> list[dict[str, object]]: + """Return a list of splice site records for every inter-fragment junction. + + Each record contains the genomic position, gene info, whether the + junction is real or artificial, and the splice type (donor/acceptor). + """ + if len(fragments) < 2: + return [] + + records: list[dict[str, object]] = [] + for i in range(len(fragments) - 1): + prev_frag = fragments[i] + next_frag = fragments[i + 1] + + # Determine if this junction is a real exon-intron boundary. + # Real requires: prev fragment reaches its exon's 3' CDS end, + # next fragment reaches its exon's 5' CDS start, and the two + # exons are CDS-adjacent in the original transcript. + if gene.strand == 1: + prev_at_3prime = prev_frag.at_exon_end + next_at_5prime = next_frag.at_exon_start + else: + prev_at_3prime = prev_frag.at_exon_start + next_at_5prime = next_frag.at_exon_end + cds_adjacent = prev_frag.exon.cds_end + 1 == next_frag.exon.cds_start + is_real = prev_at_3prime and next_at_5prime and cds_adjacent + junction_type = "real" if is_real else "artificial" + + # Compute the positions dndscv will use, mirroring its + # get_splicesites() logic. + if gene.strand == 1: + donor_positions = [prev_frag.end + 1, prev_frag.end + 2, prev_frag.end + 5] + acceptor_positions = [next_frag.start - 1, next_frag.start - 2] + else: + donor_positions = [next_frag.start - 1, next_frag.start - 2, next_frag.start - 5] + acceptor_positions = [prev_frag.end + 1, prev_frag.end + 2] + + boundary_desc_donor = f"{prev_frag.end}|{next_frag.start}" + for pos in donor_positions: + records.append({ + "CHR": "chr" + gene.chrom, + "POSITION": pos, + "GENE_SYMBOL": gene.symbol, + "GENE_ID": gene.gene_id, + "JUNCTION_TYPE": junction_type, + "SPLICE_TYPE": "donor", + "FRAGMENT_BOUNDARY": boundary_desc_donor, + }) + for pos in acceptor_positions: + records.append({ + "CHR": "chr" + gene.chrom, + "POSITION": pos, + "GENE_SYMBOL": gene.symbol, + "GENE_ID": gene.gene_id, + "JUNCTION_TYPE": junction_type, + "SPLICE_TYPE": "acceptor", + "FRAGMENT_BOUNDARY": boundary_desc_donor, + }) + return records + +@click.command(context_settings={"help_option_names": ["-h", "--help"], "show_default": True}) +@click.version_option(version="0.1.0", prog_name="panel_reformat") +@click.option("-b", "--bed", type=click.Path(exists=True), required=True, help="Panel BED file (CHR, START, END).") +@click.option("-g", "--genes", type=click.Path(exists=True), required=True, help="Gene annotation TSV file.") +@click.option("-o", "--output", type=click.Path(), default="-", help="Output TSV path.") +@click.option("-s", "--splice-report", type=click.Path(), default="splice_sites.tsv", + help="Output TSV with splice site classification (real vs artificial).") +@click.option("-v", "--verbose", is_flag=True, default=False, help="Enable debug logging.") +def cli(bed: str, genes: str, output: str, splice_report: str | None, verbose: bool) -> None: + """Align panel BED regions to codon boundaries for dndscv build_refCDS.R.""" + if verbose: + logging.getLogger().setLevel(logging.DEBUG) + + panel_by_chr = PanelLoader.load(bed) + gene_map = GeneLoader.load(genes) + logger.info("Chromosomes in panel: %s", ", ".join(sorted(panel_by_chr))) + logger.info("Reference cds region: %d genes.", len(gene_map)) + + # Only process genes on chromosomes covered by the panel. + panel_chroms = set(panel_by_chr) + candidate_genes = {gid: g for gid, g in gene_map.items() if g.chrom in panel_chroms} + logger.info("Genes on panel chromosomes: %d (skipping %d).", + len(candidate_genes), len(gene_map) - len(candidate_genes)) + + builder = TranscriptBuilder(CodonAligner()) + splice_classifier = SpliceClassifier() + + out_fh = sys.stdout if output == "-" else open(output, "w", encoding="utf-8") + try: + writer = csv.DictWriter(out_fh, fieldnames=OUTPUT_FIELDS, delimiter="\t") + writer.writeheader() + + gene_count = 0 + fragment_count = 0 + all_splice_records: list[dict[str, object]] = [] + for gene_id in sorted(candidate_genes): + gene = candidate_genes[gene_id] + rows, fragments = builder.build(gene, panel_by_chr) + for row in rows: + writer.writerow(row) + if rows: + gene_count += 1 + fragment_count += len(rows) + if splice_report: + all_splice_records.extend( + splice_classifier.classify(gene, fragments) + ) + + logger.info("Wrote %d fragment(s) across %d gene(s).", fragment_count, gene_count) + finally: + if output != "-": + out_fh.close() + + if splice_report and all_splice_records: + real_count = sum(1 for r in all_splice_records if r["JUNCTION_TYPE"] == "real") + artificial_count = len(all_splice_records) - real_count + logger.info( + "Splice report: %d positions (%d real, %d artificial).", + len(all_splice_records), real_count, artificial_count, + ) + with open(splice_report, "w", encoding="utf-8") as sfh: + sw = csv.DictWriter(sfh, fieldnames=SPLICE_REPORT_FIELDS, delimiter="\t") + sw.writeheader() + for rec in all_splice_records: + sw.writerow(rec) + + +if __name__ == "__main__": + cli() # pylint: disable=no-value-for-parameter \ No newline at end of file diff --git a/bin/depth_group_comparison.py b/bin/depth_group_comparison.py index 92a2c4d7..b9e73ecd 100755 --- a/bin/depth_group_comparison.py +++ b/bin/depth_group_comparison.py @@ -123,7 +123,7 @@ def main(table_filename, depth_gene_sample, depth_sample, unique_identifier, sep # Process groups # First clean the string (adds quotes if the shell stripped them) - cleaned = re.sub(r'(? pd.DataFrame: + +def flag_repetitive_variants( + maf_df: pd.DataFrame, repetitive_variant_threshold: int, somatic_vaf_boundary: float +) -> pd.DataFrame: """ Flags filter column for repetitive variants from the MAF dataframe. A variant is considered repetitive if it appears in at least ``repetitive_variant_threshold`` samples. Additionally, variants that consistently appear at the same position in reads @@ -87,13 +87,17 @@ def flag_repetitive_variants(maf_df: pd.DataFrame, # Work with already filtered df + somatic only to explore potential artifacts # take only variant and sample info from the df - maf_df_f_somatic = maf_df.loc[maf_df["VAF"] <= somatic_vaf_boundary][["MUT_ID","SAMPLE_ID", "PMEAN", "PSTD"]].reset_index(drop = True) + maf_df_f_somatic = maf_df.loc[somatic_mask(maf_df, somatic_vaf_boundary)][ + ["MUT_ID", "SAMPLE_ID", "PMEAN", "PSTD"] + ].reset_index(drop=True) # Group by 'MUT_ID' and count occurrences maf_df_f_somatic_pivot = maf_df_f_somatic.groupby("MUT_ID").size().reset_index(name="count") # Store repetitive variants - repetitive_variants = maf_df_f_somatic_pivot[maf_df_f_somatic_pivot["count"] >= repetitive_variant_threshold]["MUT_ID"] + repetitive_variants = maf_df_f_somatic_pivot[maf_df_f_somatic_pivot["count"] >= repetitive_variant_threshold][ + "MUT_ID" + ] LOG.info("%s repetitive_variants", len(repetitive_variants)) # Flag repetitive variants in the original dataframe @@ -104,10 +108,10 @@ def flag_repetitive_variants(maf_df: pd.DataFrame, maf_df = maf_df.drop("repetitive_variant", axis=1) # Use the position in read information to filter repetitive variants with a fixed position (likely artifacts) - maf_df_f_somatic_pos_info = maf_df_f_somatic[~(maf_df_f_somatic["PMEAN"].isna()) & - (maf_df_f_somatic["PMEAN"] != -1) & - (maf_df_f_somatic["PSTD"] == 0)] - + maf_df_f_somatic_pos_info = maf_df_f_somatic[ + ~(maf_df_f_somatic["PMEAN"].isna()) & (maf_df_f_somatic["PMEAN"] != -1) & (maf_df_f_somatic["PSTD"] == 0) + ] + # Check if there are any repetitive variants with a fixed position if maf_df_f_somatic_pos_info.shape[0] == 0: LOG.info("No repetitive variants with fixed position found.") @@ -127,19 +131,20 @@ def flag_repetitive_variants(maf_df: pd.DataFrame, # Flag these variants in the maf dataframe maf_df["repetitive_mapping_variant"] = maf_df["MUT_ID"].isin(variants_with_rep_position) LOG.info("%s muts flagged as repetitive_mapping_variant", maf_df["repetitive_mapping_variant"].sum()) - - maf_df["FILTER"] = maf_df[["FILTER","repetitive_mapping_variant"]].apply(lambda x: add_filter(x["FILTER"], x["repetitive_mapping_variant"], "repetitive_mapping_variant"), - axis = 1 - ) - maf_df = maf_df.drop("repetitive_mapping_variant", axis = 1) + + maf_df["FILTER"] = maf_df[["FILTER", "repetitive_mapping_variant"]].apply( + lambda x: add_filter(x["FILTER"], x["repetitive_mapping_variant"], "repetitive_mapping_variant"), axis=1 + ) + maf_df = maf_df.drop("repetitive_mapping_variant", axis=1) return maf_df -def flag_cohort_n_rich(maf_df: pd.DataFrame, - n_rich_cohort_proportion: float, - somatic_vaf_boundary: float) -> pd.DataFrame: + +def flag_cohort_n_rich( + maf_df: pd.DataFrame, n_rich_cohort_proportion: float, somatic_vaf_boundary: float +) -> pd.DataFrame: """ - Flags FILTER column for cohort_n_rich variants from the MAF dataframe. + Flags FILTER column for cohort_n_rich variants from the MAF dataframe. Parameters ---------- @@ -161,62 +166,59 @@ def flag_cohort_n_rich(maf_df: pd.DataFrame, if max_samples < 2: LOG.warning("Not enough samples to identify cohort_n_rich mutations!") return maf_df - + number_of_samples = max(2, (max_samples * n_rich_cohort_proportion) // 1) LOG.info(f"Flagging mutations that are n_rich in at least: {number_of_samples} samples as cohort_n_rich") # Work with already filtered df to explore potential artifacts # take only variant and sample info from the df. - maf_df_f = maf_df[["MUT_ID", "SAMPLE_ID", "VAF_Ns", "FILTER"]].reset_index(drop = True) + maf_df_f = maf_df[["MUT_ID", "SAMPLE_ID", "VAF_Ns", "FILTER"]].reset_index(drop=True) # Aggregate n_rich variants n_rich_vars_df = ( maf_df_f[maf_df_f["FILTER"].str.contains("n_rich")] .groupby("MUT_ID") - .agg( - N_rich_frequency=('SAMPLE_ID', 'count'), - VAF_Ns_threshold=('VAF_Ns', 'min') - ) - ) - + .agg(N_rich_frequency=("SAMPLE_ID", "count"), VAF_Ns_threshold=("VAF_Ns", "min")) + ) + # Flag variants that are n_rich in at least number_of_samples samples -> cohort_n_rich - n_rich_vars = set(n_rich_vars_df[n_rich_vars_df['N_rich_frequency'] >= number_of_samples].index) + n_rich_vars = set(n_rich_vars_df[n_rich_vars_df["N_rich_frequency"] >= number_of_samples].index) maf_df["cohort_n_rich"] = maf_df["MUT_ID"].isin(n_rich_vars) LOG.info("%s muts flagged as cohort_n_rich", maf_df["cohort_n_rich"].sum()) - maf_df["FILTER"] = maf_df[["FILTER","cohort_n_rich"]].apply(lambda x: add_filter(x["FILTER"], x["cohort_n_rich"], "cohort_n_rich"), - axis = 1 - ) - + maf_df["FILTER"] = maf_df[["FILTER", "cohort_n_rich"]].apply( + lambda x: add_filter(x["FILTER"], x["cohort_n_rich"], "cohort_n_rich"), axis=1 + ) + # Flag variants that are n_rich in at least 1 sample -> cohort_n_rich_uni - n_rich_vars_uni = set(n_rich_vars_df[n_rich_vars_df['N_rich_frequency'] > 0].index) + n_rich_vars_uni = set(n_rich_vars_df[n_rich_vars_df["N_rich_frequency"] > 0].index) maf_df["cohort_n_rich_uni"] = maf_df["MUT_ID"].isin(n_rich_vars_uni) LOG.info("%s muts flagged as cohort_n_rich_uni", maf_df["cohort_n_rich_uni"].sum()) - maf_df["FILTER"] = maf_df[["FILTER","cohort_n_rich_uni"]].apply(lambda x: add_filter(x["FILTER"], x["cohort_n_rich_uni"], "cohort_n_rich_uni"), - axis = 1 - ) - + maf_df["FILTER"] = maf_df[["FILTER", "cohort_n_rich_uni"]].apply( + lambda x: add_filter(x["FILTER"], x["cohort_n_rich_uni"], "cohort_n_rich_uni"), axis=1 + ) + # Flag variants that exceed the VAF_Ns threshold -> cohort_n_rich_threshold - maf_df = maf_df.merge(n_rich_vars_df, on = 'MUT_ID', how = 'left') - maf_df['N_rich_frequency'] = maf_df['N_rich_frequency'].fillna(0) - maf_df['VAF_Ns_threshold'] = maf_df['VAF_Ns_threshold'].fillna(1.1) + maf_df = maf_df.merge(n_rich_vars_df, on="MUT_ID", how="left") + maf_df["N_rich_frequency"] = maf_df["N_rich_frequency"].fillna(0) + maf_df["VAF_Ns_threshold"] = maf_df["VAF_Ns_threshold"].fillna(1.1) - maf_df["cohort_n_rich_threshold"] = maf_df["VAF_Ns"] >= maf_df['VAF_Ns_threshold'] + maf_df["cohort_n_rich_threshold"] = maf_df["VAF_Ns"] >= maf_df["VAF_Ns_threshold"] LOG.info("%s muts flagged as cohort_n_rich_threshold", maf_df["cohort_n_rich_threshold"].sum()) - maf_df["FILTER"] = maf_df[["FILTER","cohort_n_rich_threshold"]].apply(lambda x: add_filter(x["FILTER"], x["cohort_n_rich_threshold"], "cohort_n_rich_threshold"), - axis = 1 - ) + maf_df["FILTER"] = maf_df[["FILTER", "cohort_n_rich_threshold"]].apply( + lambda x: add_filter(x["FILTER"], x["cohort_n_rich_threshold"], "cohort_n_rich_threshold"), axis=1 + ) # Drop temporary columns - maf_df = maf_df.drop(["cohort_n_rich", "cohort_n_rich_uni", "cohort_n_rich_threshold"], axis = 1) - + maf_df = maf_df.drop(["cohort_n_rich", "cohort_n_rich_uni", "cohort_n_rich_threshold"], axis=1) + return maf_df -def flag_other_samples_snp(maf_df, - somatic_vaf_boundary: float) -> pd.DataFrame: + +def flag_other_samples_snp(maf_df, somatic_vaf_boundary: float) -> pd.DataFrame: """ Filters out SNPs from other samples from the MAF dataframe @@ -234,28 +236,27 @@ def flag_other_samples_snp(maf_df, """ LOG.info("Flagging SNPs from other samples...") # Get all germline variants from all samples, consider both unique and non-unique variants - germline_vars_all_samples = maf_df.loc[(maf_df["VAF"] > somatic_vaf_boundary) & - (maf_df["VAF_AM"] > somatic_vaf_boundary) & - (maf_df["vd_VAF"] > somatic_vaf_boundary), - "MUT_ID"].unique() - + germline_vars_all_samples = maf_df.loc[germline_mask(maf_df, somatic_vaf_boundary), "MUT_ID"].unique() + LOG.info(f"Using all germline variants of all samples, total: {len(germline_vars_all_samples)} variants.") # Identify variants that are germline in other samples but somatic in the current sample maf_df["other_sample_SNP"] = False - maf_df.loc[(maf_df["MUT_ID"].isin(germline_vars_all_samples)) & - (maf_df["VAF"] <= somatic_vaf_boundary), "other_sample_SNP"] = True - LOG.info("%s muts flagged as other_sample_SNP", maf_df['other_sample_SNP'].sum()) + maf_df.loc[ + (maf_df["MUT_ID"].isin(germline_vars_all_samples)) & somatic_mask(maf_df, somatic_vaf_boundary), + "other_sample_SNP", + ] = True + LOG.info("%s muts flagged as other_sample_SNP", maf_df["other_sample_SNP"].sum()) # Flag variants that are germline in other samples but somatic in the current sample - maf_df["FILTER"] = maf_df[["FILTER","other_sample_SNP"]].apply( - lambda x: add_filter(x["FILTER"], x["other_sample_SNP"], "other_sample_SNP"), - axis = 1 - ) - maf_df = maf_df.drop("other_sample_SNP", axis = 1) + maf_df["FILTER"] = maf_df[["FILTER", "other_sample_SNP"]].apply( + lambda x: add_filter(x["FILTER"], x["other_sample_SNP"], "other_sample_SNP"), axis=1 + ) + maf_df = maf_df.drop("other_sample_SNP", axis=1) return maf_df + def flag_gnomad_snp(maf_df: pd.DataFrame) -> pd.DataFrame: """ Flags gnomAD SNPs in the MAF dataframe @@ -274,19 +275,20 @@ def flag_gnomad_snp(maf_df: pd.DataFrame) -> pd.DataFrame: # Flag gnomAD SNPs if "gnomAD_SNP" in maf_df.columns: - maf_df["gnomAD_SNP"] = maf_df["gnomAD_SNP"].replace({"True": True, "False": False, '-' : False}).fillna(False).astype(bool) + maf_df["gnomAD_SNP"] = ( + maf_df["gnomAD_SNP"].replace({"True": True, "False": False, "-": False}).fillna(False).astype(bool) + ) LOG.info("Out of %d positions, %d are gnomAD SNPs", maf_df["gnomAD_SNP"].shape[0], maf_df["gnomAD_SNP"].sum()) - - maf_df["FILTER"] = maf_df[["FILTER","gnomAD_SNP"]].apply( - lambda x: add_filter(x["FILTER"], x["gnomAD_SNP"], "gnomAD_SNP"), - axis = 1 - ) - maf_df = maf_df.drop("gnomAD_SNP", axis = 1) + + maf_df["FILTER"] = maf_df[["FILTER", "gnomAD_SNP"]].apply( + lambda x: add_filter(x["FILTER"], x["gnomAD_SNP"], "gnomAD_SNP"), axis=1 + ) + maf_df = maf_df.drop("gnomAD_SNP", axis=1) return maf_df -def flag_vaf_ns_threshold(maf_df: pd.DataFrame, vaf_ns_threshold: float) -> pd.DataFrame: +def flag_vaf_ns_threshold(maf_df: pd.DataFrame, vaf_ns_threshold: float) -> pd.DataFrame: """ Flag variants that have a proportion of Ns higher than vaf_ns_threshold @@ -307,14 +309,14 @@ def flag_vaf_ns_threshold(maf_df: pd.DataFrame, vaf_ns_threshold: float) -> pd.D maf_df["high_n_vaf"] = maf_df[["VAF_Ns", "VAF_Ns_AM"]].ge(vaf_ns_threshold).any(axis=1) LOG.info("%s muts flagged as high_n_vaf", maf_df["high_n_vaf"].sum()) - maf_df["FILTER"] = maf_df[["FILTER","high_n_vaf"]].apply( - lambda x: add_filter(x["FILTER"], x["high_n_vaf"], "high_n_vaf"), - axis = 1 - ) - maf_df = maf_df.drop("high_n_vaf", axis = 1) + maf_df["FILTER"] = maf_df[["FILTER", "high_n_vaf"]].apply( + lambda x: add_filter(x["FILTER"], x["high_n_vaf"], "high_n_vaf"), axis=1 + ) + maf_df = maf_df.drop("high_n_vaf", axis=1) return maf_df + def flag_distorted_expanded(maf_df: pd.DataFrame) -> pd.DataFrame: """ If there is a column named VAF_distorted_expanded_sq, add a filter flag for variants with distorted VAF distribution. @@ -331,19 +333,22 @@ def flag_distorted_expanded(maf_df: pd.DataFrame) -> pd.DataFrame: """ LOG.info("Flagging variants with distorted VAF distribution...") - if 'VAF_distorted_expanded_sq' in maf_df.columns: - maf_df["FILTER"] = maf_df[["FILTER","VAF_distorted_expanded_sq"]].apply( - lambda x: add_filter(x["FILTER"], x["VAF_distorted_expanded_sq"], "VAF_distorted_expanded_sq"), - axis = 1 - ) + if "VAF_distorted_expanded_sq" in maf_df.columns: + maf_df["FILTER"] = maf_df[["FILTER", "VAF_distorted_expanded_sq"]].apply( + lambda x: add_filter(x["FILTER"], x["VAF_distorted_expanded_sq"], "VAF_distorted_expanded_sq"), axis=1 + ) return maf_df -def flag_maf(maf_df: pd.DataFrame, sample_name: str, - repetitive_variant_threshold: int, - somatic_vaf_boundary: float, - n_rich_cohort_proportion: float, - vaf_ns_threshold: float) -> None: + +def flag_maf( + maf_df: pd.DataFrame, + sample_name: str, + repetitive_variant_threshold: int, + somatic_vaf_boundary: float, + n_rich_cohort_proportion: float, + vaf_ns_threshold: float, +) -> None: """ Script to process a MAF (Mutation Annotation Format) file. It filters out repetitive variants, cohort_n_rich variants, and SNPs from other samples. @@ -386,36 +391,44 @@ def flag_maf(maf_df: pd.DataFrame, sample_name: str, maf_df = expand_filter_column(maf_df) ## Save final filtered MAF - maf_df.to_csv(f"{sample_name}.cohort.filtered.tsv.gz", - sep = "\t", - header = True, - index = False) - + maf_df.to_csv(f"{sample_name}.cohort.filtered.tsv.gz", sep="\t", header=True, index=False) + LOG.info("Cohort flagging complete!") + @click.command() -@click.option('--maf-df-file', required=True, type=click.Path(exists=True), help='Input gzipped MAF file (TSV)') -@click.option('--sample-name', required=True, type=str, help='Sample name for output file') -@click.option('--repetitive-variant-threshold', required=True, type=int, help='Threshold for repetitive variants') -@click.option('--somatic-vaf-boundary', required=True, type=float, help='VAF boundary for somatic variants') -@click.option('--n-rich-cohort-proportion', required=True, type=float, help='Proportion for n-rich cohort filtering') -@click.option('--vaf-ns-threshold', required=False, type=float, default=0.1, help='VAF of Ns threshold for filtering variants') -def main(maf_df_file: str, sample_name: str, repetitive_variant_threshold: int, - somatic_vaf_boundary: float, n_rich_cohort_proportion: float, vaf_ns_threshold: float): +@click.option("--maf-df-file", required=True, type=click.Path(exists=True), help="Input gzipped MAF file (TSV)") +@click.option("--sample-name", required=True, type=str, help="Sample name for output file") +@click.option("--repetitive-variant-threshold", required=True, type=int, help="Threshold for repetitive variants") +@click.option("--somatic-vaf-boundary", required=True, type=float, help="VAF boundary for somatic variants") +@click.option("--n-rich-cohort-proportion", required=True, type=float, help="Proportion for n-rich cohort filtering") +@click.option( + "--vaf-ns-threshold", required=False, type=float, default=0.1, help="VAF of Ns threshold for filtering variants" +) +def main( + maf_df_file: str, + sample_name: str, + repetitive_variant_threshold: int, + somatic_vaf_boundary: float, + n_rich_cohort_proportion: float, + vaf_ns_threshold: float, +): """ CLI wrapper for flag_maf function. """ # Load MAF dataframe - maf_df = pd.read_csv(maf_df_file, compression='gzip', header=0, sep='\t', na_values=custom_na_values) + maf_df = pd.read_csv(maf_df_file, compression="gzip", header=0, sep="\t", na_values=custom_na_values) LOG.debug(f"{maf_df_file}") # Flag MAF file - flag_maf(maf_df, + flag_maf( + maf_df, sample_name, repetitive_variant_threshold, - somatic_vaf_boundary, + somatic_vaf_boundary, n_rich_cohort_proportion, - vaf_ns_threshold) - + vaf_ns_threshold, + ) + -if __name__ == '__main__': - main() \ No newline at end of file +if __name__ == "__main__": + main() diff --git a/bin/metrics_vs_depth_qc.py b/bin/metrics_vs_depth_qc.py new file mode 100755 index 00000000..fd5d6631 --- /dev/null +++ b/bin/metrics_vs_depth_qc.py @@ -0,0 +1,304 @@ +#!/usr/bin/env python + +import json +from pathlib import Path + +import click +import matplotlib.pyplot as plt +import numpy as np +import pandas as pd +import seaborn as sns +from matplotlib.backends.backend_pdf import PdfPages +from scipy.stats import linregress + + +def load_group_samples(group_definition_file, group_name): + with open(group_definition_file, "r") as f: + groups = json.load(f) + if group_name not in groups: + raise ValueError(f"Group '{group_name}' not present in {group_definition_file}") + return [str(x) for x in groups[group_name]] + + +def _find_column(columns, candidates): + lower = {c.lower(): c for c in columns} + for cand in candidates: + if cand.lower() in lower: + return lower[cand.lower()] + return None + + +def load_depth_table(depth_gene_sample_file, samples): + depth_df = pd.read_csv(depth_gene_sample_file, sep="\t", header=0) + sample_col = _find_column(depth_df.columns, ["SAMPLE_ID", "sample"]) + gene_col = _find_column(depth_df.columns, ["GENE", "gene"]) + depth_col = _find_column(depth_df.columns, ["MEAN_GENE_DEPTH", "mean_gene_depth"]) + + if sample_col is None or gene_col is None or depth_col is None: + raise ValueError("Depth table must contain SAMPLE_ID, GENE and MEAN_GENE_DEPTH columns.") + + depth_df = depth_df.rename(columns={sample_col: "SAMPLE_ID", gene_col: "GENE", depth_col: "MEAN_GENE_DEPTH"}) + depth_df["SAMPLE_ID"] = depth_df["SAMPLE_ID"].astype(str) + depth_df["GENE"] = depth_df["GENE"].astype(str) + depth_df = depth_df[depth_df["SAMPLE_ID"].isin(samples)].copy().reset_index(drop=True) + depth_df["MEAN_GENE_DEPTH"] = pd.to_numeric(depth_df["MEAN_GENE_DEPTH"], errors="coerce") + return depth_df + + +def load_mutdensity(mutdensity_file, samples, metric_name): + mut = pd.read_csv(mutdensity_file, sep="\t", header=0) + required = {"SAMPLE_ID", "GENE", "MUTDENSITY_MB"} + if not required.issubset(set(mut.columns)): + raise ValueError(f"Mutdensity file {mutdensity_file} must contain {sorted(required)}") + + mut = mut[(mut["SAMPLE_ID"].astype(str).isin(samples)) + & (mut["GENE"] != "ALL_GENES") + & (mut["MUTTYPES"] == "SNV") + ].copy() + mut["SAMPLE_ID"] = mut["SAMPLE_ID"].astype(str) + mut["GENE"] = mut["GENE"].astype(str) + mut["metric_value"] = pd.to_numeric(mut["MUTDENSITY_MB"], errors="coerce") + mut["metric_name"] = metric_name + if "REGIONS" in mut.columns: + mut["metric_name"] = mut["metric_name"] + "." + mut["REGIONS"].fillna("unknown").astype(str) + return mut[["SAMPLE_ID", "GENE", "metric_name", "metric_value"]] + +def load_adjmutdensity(mutdensity_file, samples, metric_name): + mut = pd.read_csv(mutdensity_file, sep="\t", header=0) + required = {"SAMPLE", "GENE", "synonymous", "missense", "nonsense", "essential_splice", "truncating", "nonsynonymous_splice", "all_impacts"} + if not required.issubset(set(mut.columns)): + raise ValueError(f"Mutdensity file {mutdensity_file} must contain {sorted(required)}") + + mut = mut[(mut["SAMPLE"].astype(str).isin(samples))].copy() + mut["SAMPLE_ID"] = mut["SAMPLE"].astype(str) + mut["GENE"] = mut["GENE"].astype(str) + + mut_dfs = [] + for impact in ["synonymous", "missense", "nonsense", "essential_splice", "truncating", "nonsynonymous_splice", "all_impacts"]: + subset_mut = mut[["SAMPLE_ID", "GENE", impact]].copy() + subset_mut["metric_value"] = subset_mut[impact] + subset_mut["metric_name"] = f"{impact}_density" + mut_dfs.append(subset_mut[["SAMPLE_ID", "GENE", "metric_name", "metric_value"]]) + + return mut_dfs + + +def load_omegas(omegas_file, samples): + omega = pd.read_csv(omegas_file, sep="\t", header=0) + sample_col = "sample" + gene_col = "gene" + dnds_col = "dnds" + if sample_col is None or gene_col is None or dnds_col is None: + raise ValueError(f"Omega file {omegas_file} must contain sample, gene and dnds/omega columns") + + omegas_dfs = [] + omega = omega.rename(columns={sample_col: "SAMPLE_ID", gene_col: "GENE", dnds_col: "metric_value"}) + for impact in ["missense", "truncating"]: + subset_omega = omega[(omega["impact"] == impact) + & ~(omega["GENE"].astype(str).str.contains("--")) + ].copy() + subset_omega["SAMPLE_ID"] = subset_omega["SAMPLE_ID"].astype(str) + subset_omega = subset_omega[subset_omega["SAMPLE_ID"].isin(samples)].copy() + subset_omega["metric_value"] = pd.to_numeric(subset_omega["metric_value"], errors="coerce") + subset_omega["metric_name"] = f"omega_gloc_{impact}" + omegas_dfs.append(subset_omega[["SAMPLE_ID", "GENE", "metric_name", "metric_value"]]) + + return omegas_dfs + + +def summarize_effect_by_gene(df): + rows = [] + work = df.dropna(subset=["MEAN_GENE_DEPTH", "metric_value"]).copy() + all_groups = [("all_samples", work)] + [(s, g) for s, g in work.groupby("GENE")] + for group_name, gdf in all_groups: + n = len(gdf) + if n < 3: + rows.append( + { + "sample_scope": group_name, + "n_points": n, + "pearson_r": np.nan, + "spearman_r": np.nan, + "slope_metric_per_depth": np.nan, + "linreg_pvalue": np.nan, + } + ) + continue + lr = linregress(gdf["MEAN_GENE_DEPTH"], gdf["metric_value"]) + rows.append( + { + "sample_scope": group_name, + "n_points": n, + "pearson_r": gdf["MEAN_GENE_DEPTH"].corr(gdf["metric_value"], method="pearson"), + "spearman_r": gdf["MEAN_GENE_DEPTH"].corr(gdf["metric_value"], method="spearman"), + "slope_metric_per_depth": lr.slope, + "linreg_pvalue": lr.pvalue, + } + ) + return pd.DataFrame(rows) + + +def summarize_missingness(depth_df, merged_df): + m = depth_df.merge( + merged_df[["SAMPLE_ID", "GENE", "metric_name", "metric_value"]], + on=["SAMPLE_ID", "GENE"], + how="left", + ) + m["is_missing_metric"] = m["metric_value"].isna() + by_gene = ( + m.groupby(by = ["GENE", "metric_name"]) + .apply( + lambda g: pd.Series( + { + "n_samples_total": int(len(g)), + "n_missing_metric": int(g["is_missing_metric"].sum()), + "missing_fraction": float(g["is_missing_metric"].mean()), + "median_depth_missing": g.loc[g["is_missing_metric"], "MEAN_GENE_DEPTH"].median(), + "median_depth_nonmissing": g.loc[~g["is_missing_metric"], "MEAN_GENE_DEPTH"].median(), + "min_depth_missing": g.loc[g["is_missing_metric"], "MEAN_GENE_DEPTH"].min(), + "max_depth_missing": g.loc[g["is_missing_metric"], "MEAN_GENE_DEPTH"].max(), + } + ) + ) + .reset_index() + ) + return by_gene.sort_values(["missing_fraction", "n_missing_metric"], ascending=[False, False]).reset_index(drop=True) + + +def plot_scatter_per_gene(df, group_name, metric_name, output_pdf): + genes = sorted(df["GENE"].dropna().unique().tolist()) + if not genes: + return + + # keep only the top 200 genes + genes = genes[:200] + + sns.set_style("whitegrid") + per_page = 12 + ncols = 3 + nrows = 4 + with PdfPages(output_pdf) as pdf: + for start in range(0, len(genes), per_page): + page_genes = genes[start : start + per_page] + fig, axes = plt.subplots(nrows, ncols, figsize=(12, 16), sharex=False, sharey=False) + axes = axes.flatten() + for i, gene in enumerate(page_genes): + ax = axes[i] + sdf = df[df["GENE"] == gene].dropna(subset=["MEAN_GENE_DEPTH", "metric_value"]) + if sdf.empty: + ax.set_title(f"{gene} (no data)") + continue + sns.scatterplot( + data=sdf, + x="MEAN_GENE_DEPTH", + y="metric_value", + alpha=0.5, + s=16, + linewidth=0, + ax=ax, + ) + title = f"{gene} (n={len(sdf)})" + if len(sdf) >= 3: + lr = linregress(sdf["MEAN_GENE_DEPTH"], sdf["metric_value"]) + title += f" | slope={lr.slope:.2e}, p={lr.pvalue:.2e}" + line_color = "darkred" if lr.pvalue < 0.05 else "darkgrey" + sns.regplot( + data=sdf, + x="MEAN_GENE_DEPTH", + y="metric_value", + scatter=False, + line_kws={"color": line_color, "linewidth": 1.2}, + ax=ax, + ) + + y_min, y_max = sdf["metric_value"].min(), sdf["metric_value"].max() + val_range = y_max - y_min if y_max > y_min else (abs(y_max) * 0.1 if y_max != 0 else 1.0) + ax.set_ylim(bottom=-0.05 * val_range if y_min >= 0 else y_min - 0.05 * val_range) + + ax.set_title(title, fontsize=9) + ax.set_xlabel("MEAN_GENE_DEPTH") + ax.set_ylabel(metric_name) + + for j in range(len(page_genes), len(axes)): + axes[j].axis("off") + + fig.suptitle(f"{group_name} | {metric_name} vs depth (per gene)", fontsize=14) + fig.tight_layout(rect=[0, 0, 1, 0.97]) + pdf.savefig(fig) + plt.close(fig) + + +@click.command() +@click.option("--mutdensity-file", required=True, type=click.Path(exists=True)) +@click.option("--depth-gene-sample-file", required=True, type=click.Path(exists=True)) +@click.option("--group-definition", required=True, type=click.Path(exists=True)) +@click.option("--group-name", required=True, type=str) +@click.option("--output-dir", required=True, type=click.Path()) +@click.option("--adjusted-mutdensity-file", required=False, type=click.Path(exists=True)) +@click.option("--omegas-file", required=False, type=click.Path(exists=True)) +def main( + mutdensity_file, + depth_gene_sample_file, + group_definition, + group_name, + output_dir, + adjusted_mutdensity_file, + omegas_file, +): + outdir = Path(output_dir) + outdir.mkdir(parents=True, exist_ok=True) + + samples = load_group_samples(group_definition, group_name) + depth_df = load_depth_table(depth_gene_sample_file, samples) + + metric_frames = [load_mutdensity(mutdensity_file, samples, "mutdensity")] + if adjusted_mutdensity_file: + try: + metric_frames.extend(load_adjmutdensity(adjusted_mutdensity_file, samples, "adjusted_mutdensity")) + except Exception as e: + print(f"Warning: skipping adjusted mutdensity file {adjusted_mutdensity_file}: {e}") + if omegas_file: + try: + metric_frames.extend(load_omegas(omegas_file, samples)) + except Exception as e: + print(f"Warning: skipping omegas file {omegas_file}: {e}") + + metrics_df = pd.concat(metric_frames, ignore_index=True) + + status_rows = [] + for metric_name in metrics_df["metric_name"].unique(): + print(f"Processing metric '{metric_name}'...") + + metric_specific_df = metrics_df[metrics_df["metric_name"] == metric_name].copy() + merged = depth_df.merge(metric_specific_df, on=["SAMPLE_ID", "GENE"], how="left") + merged["metric_name"] = metric_name + + summary_effect = summarize_effect_by_gene(merged) + summary_effect.insert(0, "metric_name", metric_name) + summary_effect.to_csv(outdir / f"{group_name}.{metric_name}.depth_effect_summary.tsv", sep="\t", index=False) + + missingness = summarize_missingness(depth_df, merged) + missingness.to_csv(outdir / f"{group_name}.{metric_name}.depth_missingness_by_gene.tsv", sep="\t", index=False) + + plot_scatter_per_gene( + merged, + group_name=group_name, + metric_name=metric_name, + output_pdf=outdir / f"{group_name}.{metric_name}.depth_scatter_per_gene.pdf", + ) + + status_rows.append( + { + "group_name": group_name, + "metric_name": metric_name, + "n_depth_rows": int(len(depth_df)), + "n_metric_rows": int(len(metric_specific_df)), + "n_merged_nonmissing": int(merged["metric_value"].notna().sum()), + } + ) + + pd.DataFrame(status_rows).to_csv(outdir / f"{group_name}.metrics_vs_depth_qc.status.tsv", sep="\t", index=False) + + +if __name__ == "__main__": + main() diff --git a/bin/mut_density.py b/bin/mut_density_adjusted.py similarity index 82% rename from bin/mut_density.py rename to bin/mut_density_adjusted.py index 308b2953..ebf684f9 100755 --- a/bin/mut_density.py +++ b/bin/mut_density_adjusted.py @@ -31,6 +31,8 @@ def get_correction_factor(sample_name, trinucleotide_counts_df, mutability_df, f triplet_counts = np.array(l) # genome length in Mb + # accounting for the fact that each position contributes: + # 3*depth because of the 3 mutations available at each position genome_length = sum(triplet_counts) / (3 * 1e6) # vector of relative mutabilities in 96-channel canonical sorting @@ -61,13 +63,17 @@ def mutation_density(sample_name, depths_file, somatic_mutations_file, mutabilit for csqn, csqn_set in broadimpact_grouping_dict_with_synonymous.items(): - for gene in panel_df['GENE'].unique(): + for gene in list(panel_df['GENE'].unique()) + ["ALL_GENES"]: # compute vector of sum of depths per trinucleotide context # tailored to the specific gene-impact target - region_df = panel_df[(panel_df['IMPACT'].isin(csqn_set)) & (panel_df['GENE'] == gene)].copy() + if gene == 'ALL_GENES': + region_df = panel_df[(panel_df['IMPACT'].isin(csqn_set))][['CHROM', 'POS', 'REF', 'ALT']].drop_duplicates() + else: + region_df = panel_df[(panel_df['IMPACT'].isin(csqn_set)) & (panel_df['GENE'] == gene)][['CHROM', 'POS', 'REF', 'ALT']].drop_duplicates() - # counting every position once + # counting every position as many times as the number of possible + # mutations of the selected consequences at that position (1,2 or 3) dh = pd.merge(region_df[['CHROM', 'POS']], depths_df[['CHROM', 'POS', 'CONTEXT', sample_name]], on=['CHROM', 'POS'], how='left') @@ -88,10 +94,13 @@ def mutation_density(sample_name, depths_file, somatic_mutations_file, mutabilit except AssertionError: res.loc[gene, csqn] = None continue - - # observed somatic mutations - n = somatic_mutations_df[(somatic_mutations_df['IMPACT'].isin(csqn_set)) & (somatic_mutations_df['GENE'] == gene)].shape[0] + + # observed somatic mutations + if gene == 'ALL_GENES': + n = somatic_mutations_df[(somatic_mutations_df['IMPACT'].isin(csqn_set))].shape[0] + else: + n = somatic_mutations_df[(somatic_mutations_df['IMPACT'].isin(csqn_set)) & (somatic_mutations_df['GENE'] == gene)].shape[0] res.loc[gene, csqn] = n / effective_length @@ -149,12 +158,12 @@ def main(sample_name, depths_file, somatic_mutations_file, mutability_file, pane logfoldchange_plot(sample_name, res, res_flat) # save results - res["SAMPLE"] = sample_name - res_flat["SAMPLE"] = sample_name + res["SAMPLE_ID"] = sample_name + res_flat["SAMPLE_ID"] = sample_name res.index.name = 'GENE' res_flat.index.name = 'GENE' - res[['SAMPLE'] + [col for col in res.columns if col != 'SAMPLE']].to_csv(f'{sample_name}.mutdensities.tsv', sep='\t') - res_flat[['SAMPLE'] + [col for col in res_flat.columns if col != 'SAMPLE']].to_csv(f'{sample_name}.mutdensities_flat.tsv', sep='\t') + res[['SAMPLE_ID'] + [col for col in res.columns if col != 'SAMPLE_ID']].to_csv(f'{sample_name}.mutdensities.tsv', sep='\t') + res_flat[['SAMPLE_ID'] + [col for col in res_flat.columns if col != 'SAMPLE_ID']].to_csv(f'{sample_name}.mutdensities_flat.tsv', sep='\t') diff --git a/bin/mut_density_adjusted_dnds.py b/bin/mut_density_adjusted_dnds.py new file mode 100755 index 00000000..a39836ac --- /dev/null +++ b/bin/mut_density_adjusted_dnds.py @@ -0,0 +1,64 @@ +#!/usr/bin/env python + + +import click +import pandas as pd +import numpy as np +from read_utils import custom_na_values + + +def compute_dnds_proxy(mutdensity_file, cohort_syn_mutdensities_file, output_file, mode): + """ + TODO: explain what this function does + TODO 2: store a log file that is also outputted and can be used to check some basic statistics + + right now the use of mode is not implemented, + since we only compute one type of synonymous mutation densities. + """ + + mutdensity_df_init = pd.read_csv(mutdensity_file, sep = "\t", header = 0, na_values = custom_na_values) + mutdensity_df_init["synonymous"] = mutdensity_df_init["synonymous"].replace(0, np.nan) + all_possible_genes = list(mutdensity_df_init["GENE"].unique()) + + cohort_syn_mutdensity_df = pd.read_csv(cohort_syn_mutdensities_file, sep = "\t", header = 0, na_values = custom_na_values) + cohort_syn_mutdensity_df.columns = ['GENE', 'cohort_synonymous'] + cohort_syn_mutdensity_df = cohort_syn_mutdensity_df.set_index("GENE") + + init_cohort_syn_df = pd.DataFrame(index = all_possible_genes) + cohort_syn_df = pd.concat((init_cohort_syn_df, cohort_syn_mutdensity_df), axis = 1) + + # filling the null mutation densities with the value of the 1st decile + cohort_syn_df = cohort_syn_df.fillna(cohort_syn_df[~(cohort_syn_df.isna())].quantile(.1)).reset_index() + cohort_syn_df.columns = ['GENE', 'cohort_synonymous'] + + mutdensity_df = mutdensity_df_init.merge(cohort_syn_df, on = "GENE") + for impact in ["missense", "truncating", "nonsynonymous_splice"]: + mutdensity_df[f"d_{impact}/d_synonymous"] = mutdensity_df[impact] / mutdensity_df["synonymous"] + mutdensity_df[f"d_{impact}/d_cohort_synonymous"] = mutdensity_df[impact] / mutdensity_df["cohort_synonymous"] + mutdensity_df.iloc[:,2:] = mutdensity_df.iloc[:,2:].round(5) + + # summary at all_samples level + subset_mutdensities = mutdensity_df[(mutdensity_df["SAMPLE_ID"] == 'all_samples')] + for impact in ["missense", "truncating"]: + print(subset_mutdensities.sort_values(by=f"d_{impact}/d_synonymous", ascending=False)[ + ["GENE", "SAMPLE_ID", impact, "synonymous", f"d_{impact}/d_synonymous"] + ].head(10)) + + mutdensity_df.to_csv(f"{output_file}", + header=True, + index=False, + sep="\t") + + +@click.command() +@click.option('--mutdensities', type=click.Path(exists=True), help='Input mutation density file') +@click.option('--cohort-syn-mutdensities', type=click.Path(exists=True), help='Input cohort synonymous mutation densities') +@click.option('--output', type=click.Path(), help='Output file') +@click.option('--mode', type=click.Choice(['mutations', 'mutated_reads']), default='mutations') +def main(mutdensities, cohort_syn_mutdensities, output, mode): + click.echo("Selecting the gene synonymous mutation densities...") + compute_dnds_proxy(mutdensities, cohort_syn_mutdensities, output, mode) + +if __name__ == '__main__': + main() + diff --git a/bin/compute_mutdensity.py b/bin/mut_density_simple.py similarity index 99% rename from bin/compute_mutdensity.py rename to bin/mut_density_simple.py index 880138e7..f0e97b4a 100755 --- a/bin/compute_mutdensity.py +++ b/bin/mut_density_simple.py @@ -208,4 +208,4 @@ def main(maf_path, depths_path, annot_panel_path, sample_name, panel_version): if __name__ == '__main__': - main() + main() \ No newline at end of file diff --git a/bin/mutation_densities_qc.py b/bin/mutation_densities_qc.py index 84b1e1b6..7a3266b3 100755 --- a/bin/mutation_densities_qc.py +++ b/bin/mutation_densities_qc.py @@ -178,6 +178,8 @@ def main(input_file, output_dir, panel, group_definition, group_name): for reg in regions : for mod in mode_list: filt_mutden = filter_mutdensities(mutden_df_panel, reg, group_name, mod, sample_names) + if filt_mutden.shape[0] < 2: + continue zero_cases_flag, mutden_zscore = z_score_log10(filt_mutden, mod) ## Generate csv with zscore result and store as csv @@ -223,6 +225,10 @@ def main(input_file, output_dir, panel, group_definition, group_name): # Close to free memory before next loop plt.close(fig) + if compile_all_flagged.empty: + print("No flagged cases found across all regions and modes.") + compile_all_flagged = pd.DataFrame(columns=['ID', 'reason_exclusion', 'cohort', 'regions', 'criteria']) + # Compile all flagged cases into a single csv compile_all_flagged.to_csv(f"{output_dir}/compiled_all_flagged_cases.{group_name}.tsv", index = False, sep='\t') diff --git a/bin/mutgenomes_summary_tables.py b/bin/mutgenomes_summary_tables.py index db070392..61b30069 100755 --- a/bin/mutgenomes_summary_tables.py +++ b/bin/mutgenomes_summary_tables.py @@ -13,6 +13,7 @@ @click.command() @click.option('--metadata-file') def gather(metadata_file): + remove_sex_chromosome = False # parse all samples @@ -35,7 +36,13 @@ def gather(metadata_file): except Exception: continue else: - raise ValueError("Input file must contain 'SAMPLE_ID' and 'SEX' columns") + print("Metadata input file does not contain 'SAMPLE_ID' and 'SEX' columns") + print("Without proper metadata no values can be computed in the X chromosome") + clinical_df = df_all.copy() + clinical_df["SEX"] = "UNKNOWN" + clinical_df = clinical_df[['SAMPLE', 'SEX']].copy().drop_duplicates() + clinical_df.rename(columns={'SAMPLE': 'SAMPLE_ID'}, inplace=True) + remove_sex_chromosome = True df_all = df_all[~(df_all['SAMPLE'] == 'all_samples')] @@ -51,9 +58,9 @@ def gather(metadata_file): for label in ['GENOMES_SNV_AM', 'GENOMES_SNV_ND']: for k in ['LOWER', 'MEAN', 'UPPER', 'TOTAL']: - df_all_annotated[f'CELLS_DOUBLE_HIT_{'_'.join(label.split('_')[1:])}_{k}'] = df_all_annotated[f'{label}_{k}'].values + df_all_annotated[f"CELLS_DOUBLE_HIT_{'_'.join(label.split('_')[1:])}_{k}"] = df_all_annotated[f'{label}_{k}'].values - df_all_annotated[f'CELLS_SINGLE_HIT_{'_'.join(label.split('_')[1:])}_{k}'] = df_all_annotated.apply( + df_all_annotated[f"CELLS_SINGLE_HIT_{'_'.join(label.split('_')[1:])}_{k}"] = df_all_annotated.apply( lambda r: r[f'{label}_{k}'] if ((r['chr'] == 'chrX') and (r['SEX'] == 'M')) else 2 * r[f'{label}_{k}'], axis=1) @@ -62,9 +69,9 @@ def gather(metadata_file): label = 'GENOMES_INDEL_AM' k = 'TOTAL' - df_all_annotated[f'CELLS_DOUBLE_HIT_{'_'.join(label.split('_')[1:])}_{k}'] = df_all_annotated[f'{label}_{k}'].values + df_all_annotated[f"CELLS_DOUBLE_HIT_{'_'.join(label.split('_')[1:])}_{k}"] = df_all_annotated[f'{label}_{k}'].values - df_all_annotated[f'CELLS_SINGLE_HIT_{'_'.join(label.split('_')[1:])}_{k}'] = df_all_annotated.apply( + df_all_annotated[f"CELLS_SINGLE_HIT_{'_'.join(label.split('_')[1:])}_{k}"] = df_all_annotated.apply( lambda r: r[f'{label}_{k}'] if ((r['chr'] == 'chrX') and (r['SEX'] == 'M')) else 2 * r[f'{label}_{k}'], axis=1) @@ -77,41 +84,43 @@ def gather(metadata_file): df_gene_annotated = pd.merge(df_gene, clinical_df, left_on='SAMPLE', right_on='SAMPLE_ID', how='left') # snvs - for label in ['GENOMES_SNV_AM', 'GENOMES_SNV_ND']: for k in ['LOWER', 'MEAN', 'UPPER', 'TOTAL']: - df_gene_annotated[f'CELLS_DOUBLE_HIT_{'_'.join(label.split('_')[1:])}_{k}'] = df_gene_annotated[f'{label}_{k}'].values + df_gene_annotated[f"CELLS_DOUBLE_HIT_{'_'.join(label.split('_')[1:])}_{k}"] = df_gene_annotated[f'{label}_{k}'].values - df_gene_annotated[f'CELLS_SINGLE_HIT_{'_'.join(label.split('_')[1:])}_{k}'] = df_gene_annotated.apply( + df_gene_annotated[f"CELLS_SINGLE_HIT_{'_'.join(label.split('_')[1:])}_{k}"] = df_gene_annotated.apply( lambda r: r[f'{label}_{k}'] if ((r['chr'] == 'chrX') and (r['SEX'] == 'M')) else 2 * r[f'{label}_{k}'], axis=1) # indels - label = 'GENOMES_INDEL_AM' k = 'TOTAL' - df_gene_annotated[f'CELLS_DOUBLE_HIT_{'_'.join(label.split('_')[1:])}_{k}'] = df_gene_annotated[f'{label}_{k}'].values + df_gene_annotated[f"CELLS_DOUBLE_HIT_{'_'.join(label.split('_')[1:])}_{k}"] = df_gene_annotated[f'{label}_{k}'].values - df_gene_annotated[f'CELLS_SINGLE_HIT_{'_'.join(label.split('_')[1:])}_{k}'] = df_gene_annotated.apply( + df_gene_annotated[f"CELLS_SINGLE_HIT_{'_'.join(label.split('_')[1:])}_{k}"] = df_gene_annotated.apply( lambda r: r[f'{label}_{k}'] if ((r['chr'] == 'chrX') and (r['SEX'] == 'M')) else 2 * r[f'{label}_{k}'], axis=1) - ## collapse genes + # whenever the information of sex is not available, set chromosomes cannot be trusted and we remove them from the tables + if remove_sex_chromosome: + df_all_annotated = df_all_annotated[~(df_all_annotated['chr'] == 'chrX')] + df_gene_annotated = df_gene_annotated[~(df_gene_annotated['chr'] == 'chrX')] + # df_all_annotated = df_all_annotated[~(df_all_annotated['chr'].isin(['chrX', 'chrY']))] + # df_gene_annotated = df_gene_annotated[~(df_gene_annotated['chr'].isin(['chrX', 'chrY']))] + ## collapse genes df_sample_annotated = df_gene_annotated.groupby(by=['SAMPLE']).agg({ k: lambda x: inclusion_exclusion(list(x)) for k in df_gene_annotated.columns if (k.startswith('GENOMES') or k.startswith('CELLS')) }).reset_index(drop=False) # sort - df_all_annotated.sort_values(by=['SAMPLE', 'gene', 'impact'], inplace=True) df_gene_annotated.sort_values(by=['SAMPLE', 'gene'], inplace=True) df_sample_annotated.sort_values(by=['SAMPLE'], inplace=True) # dump - df_all_annotated.to_csv('./covered_genomes_cells.tsv', sep='\t', index=False) df_gene_annotated.to_csv('./covered_genomes_cells.genewise.grouped.tsv', sep='\t', index=False) df_sample_annotated.to_csv('./covered_genomes_cells.samplewise.tsv', sep='\t', index=False) diff --git a/bin/mutrate_genome_trinuc_corrected.R b/bin/mutrate_genome_trinuc_corrected.R new file mode 100755 index 00000000..a5859496 --- /dev/null +++ b/bin/mutrate_genome_trinuc_corrected.R @@ -0,0 +1,197 @@ +#!/opt/conda/bin/Rscript --vanilla + + +library(Hmisc) +library(tidyr) +library(stringr) +library(dplyr, warn = FALSE) +library(ggplot2) +library(jsonlite) +library(Biostrings) +library(optparse) +library(R.utils) +library(data.table) + +## Read it from a TSV and format it as json + +## command line arguments +option_list = list( + make_option(c("-n", "--samplename"), type="character", default=NULL, + help="sample name/identifier of the run", metavar="character"), + make_option(c("-m", "--mutations"), type="character", default=NULL, + help="mutation dataset file name", metavar="character"), + make_option(c("-o", "--outputprefix"), type="character", default="dNdScv_output", + help="output file name [default= %default]", metavar="character"), + make_option(c("-d", "--depths"), type="character", default=NULL, + help="depths dataset file name", metavar="character"), + make_option(c("-c", "--consensus_bed"), type="character", default=NULL, + help="consensus bed file name", metavar="character"), + make_option(c("-w", "--wgs_counts"), type="character", default=NULL, + help="WGS trinucleotide counts file name", metavar="character"), + make_option(c("-p", "--panel_version"), type="character", default=NULL, + help="panel version", metavar="character") +); + +opt_parser = OptionParser(option_list=option_list); +opt = parse_args(opt_parser); + + +sample_name <- opt$samplename +mutations <- opt$mutations +depths <- opt$depths +consensus_bed <- opt$consensus_bed +wgs_counts <- opt$wgs_counts +output_name <- opt$outputprefix +panel_version <- opt$panel_version + +print(paste("Sample name: ", sample_name)) +print(paste("Mutations file: ", mutations)) +print(paste("Depths file: ", depths)) +print(paste("Consensus bed file: ", consensus_bed)) +print(paste("WGS counts file: ", wgs_counts)) +print(paste("Output name: ", output_name)) +print(paste("Panel version: ", panel_version)) + +path2out = "wgs_mutrate" + +get_genome_content <- function(genome_content){ + #' Get genome trinucleotide content + #' + #' This function reads trinucleotide counts from a JSON genome composition file + #' and computes the 96 pyrimidine-centered contexts. + + #' As input uses path to the json file with genome trinucleotide content + + counts_df <- read.table(genome_content, header=TRUE, sep="\t") + colnames(counts_df) <- c("CONTEXT", "N_sites_genome") + message("Contexts in genome sites = ", nrow(counts_df)) + return(counts_df) +} + + +get_consensus_sites_depth <- function(sample, depths_file, consensus_bed){ + #' Get depth per position for positions in consensus panel + #' + #' This function intersects consensus panel with file with annotated depth per position + + # Load depth data + dt_pos <- fread(depths_file) + colnames(dt_pos) <- c("CHROM", "POS", "CONTEXT", "DEPTH") + dt_pos[, `:=`( + start = POS, + end = POS + )] + setkey(dt_pos, CHROM, start, end) + # Load consensus bed + dt_bed <- fread(consensus_bed, col.names = c("CHROM", "start", "end")) + setkey(dt_bed, CHROM, start, end) + # Overlap + hits <- foverlaps(dt_pos, dt_bed, type = "any", nomatch = 0L) + consensus_depth <- hits[, .(CHROM, POS, CONTEXT, DEPTH)] + return(consensus_depth) +} + +get_mutations_and_sites <- function(sample_name, sites_file, mutations_file, consensus_bed, genome_sites_df){ + #' Get number of mutations per sample and normalize panel content to whole genome content + #' + #' This function gets number of mutations per sample in each context and normalizes panel content to whole genome content + result <- NULL + print(sample_name) + + df_sites = get_consensus_sites_depth(sample_name, sites_file, consensus_bed) + print("1") + df_sites$depth = as.numeric(df_sites$DEPTH) + print("2") + df_sites_panel_agg = df_sites %>% + group_by(CONTEXT) %>% + summarise(N = sum(DEPTH)) + print("3") + df_sites_panel_agg = as.data.frame(df_sites_panel_agg) + colnames(df_sites_panel_agg) <- c("CONTEXT", "N_sites_panel") + message("Contexts in panel sites = ", nrow(df_sites_panel_agg)) + print(head(df_sites_panel_agg)) + + df_sites = merge(genome_sites_df, df_sites_panel_agg, by="CONTEXT") + print(head(df_sites)) + message("Contexts in panel and genome sites = ", nrow(df_sites)) + df_sites <- df_sites %>% mutate(proportion_genome=N_sites_genome/sum(N_sites_genome)) + df_sites <- df_sites %>% mutate(proportion_panel=N_sites_panel/sum(N_sites_panel)) + df_sites$ratio2genome = df_sites$proportion_panel/df_sites$proportion_genome + print(head(df_sites)) + print(sum(df_sites$N_sites_panel)) + + df_mutations = read.table(mutations_file, header=TRUE, sep="\t") + df_mutations = df_mutations[df_mutations$TYPE=="SNV",] + df_mutations$CONTEXT = str_split_i(df_mutations$CONTEXT_MUT, ">", 1) + df_mutations_agg = df_mutations %>% + group_by(CONTEXT) %>% + summarise(N_mut = n()) + print(head(df_mutations_agg)) + df_mutations = as.data.frame(df_mutations) + result_sample = merge(df_sites, df_mutations_agg, by="CONTEXT", all=TRUE) + # if some contexts are absent in mutataions - keep them but put mutation number to 0 + if (nrow(result_sample[is.na(result_sample$N_mut),]) > 0){ + result_sample[is.na(result_sample$N_mut),]$N_mut <- 0 + } + print(head(result_sample)) + message("Contexts in df with mutations = ", nrow(result_sample)) + result_sample$N_mut_corrected = result_sample$N_mut * result_sample$ratio2genome + sample_out <- c(sample_name, sum(result_sample$N_mut), sum(result_sample$N_mut_corrected), sum(df_sites$N_sites_panel)) + result <- rbind(result, sample_out) + + return(result) +} + +#Download df with genome trinucleotide contexts +df_sites_genome_agg <- get_genome_content(wgs_counts) + +#Get number of mutations and normalize panel contetnt to genome content +result <- get_mutations_and_sites(sample_name, depths, mutations, consensus_bed, df_sites_genome_agg) +print(head(result)) + + +result <- as.data.frame(result) +colnames(result) <- c("sample", "N_mut", "N_mut_corrected", "DEPTH") +result$N_mut_corrected = as.numeric(result$N_mut_corrected) +result$DEPTH = as.numeric(result$DEPTH) +result$N_mut = as.numeric(result$N_mut) +result$mutrate_observed = result$N_mut_corrected/result$DEPTH +result$mutrate_observed_per_MB <- result$mutrate_observed * 10**6 + +# why the indices of the two boundaries are different? is this desired or a typo? +result$mutrate_CI_high <- apply(result, 1, function(x) binconf(as.numeric(x["N_mut_corrected"]), as.numeric(x["DEPTH"]), alpha=0.05, method=c("wilson","exact","asymptotic","all"), include.x=FALSE, include.n=FALSE, return.df=FALSE)[3]) +result$mutrate_CI_low <- apply(result, 1, function(x) binconf(as.numeric(x["N_mut_corrected"]), as.numeric(x["DEPTH"]), alpha=0.05, method=c("wilson","exact","asymptotic","all"), include.x=FALSE, include.n=FALSE, return.df=FALSE)[2]) +result$Muts_per_cell <- result$mutrate_observed*2*sum(df_sites_genome_agg$N_sites_genome) +result$panel_version <- panel_version +print((result)) + +message("Output will be written to ", paste0(output_name, "_mutrates_results.tsv")) +write.table(result, paste0(output_name, "_mutrates_results.tsv"), sep="\t", row.names = FALSE, quote = FALSE) + +# setwd(path2out) + +# jpeg(filename=paste("mutrate_trint_corrected.with_nanoseq.jpeg", sep=""), width=30, height=15, res=300, units='cm') + +# ggplot(result, aes(x=sample, y=mutrate_observed)) + +# geom_bar(stat="identity", position="dodge", fill="grey") + +# geom_errorbar(aes(x=sample, ymin=mutrate_CI_low, ymax=mutrate_CI_high), width=0.4, alpha=0.9, linewidth=1, position=position_dodge(.9)) + +# theme_bw() + +# theme(axis.text.x = element_text(angle = 90, hjust=1)) + +# geom_text(aes(label = N_mut, x = sample, y = mutrate_observed), position = position_dodge(width = 0.9), vjust = -0.5, hjust = -0.1) + +# xlab("") + +# ylab("Mutation rate") + +# theme(legend.position="bottom") +# dev.off() + +# jpeg(filename=paste("mutrate_trint_corrected.jpeg", sep=""), width=30, height=15, res=300, units='cm') +# result<-result[result$protocol != "Nanoseq_Sanger",] +# ggplot(result, aes(x=sample, y=mutrate_observed)) + +# geom_bar(stat="identity", position="dodge", fill="grey") + +# geom_errorbar(aes(x=sample, ymin=mutrate_CI_low, ymax=mutrate_CI_high), width=0.4, alpha=0.9, linewidth=1, position=position_dodge(.9)) + +# theme_bw() + +# theme(axis.text.x = element_text(angle = 90, hjust=1)) + +# geom_text(aes(label = N_mut, x = sample, y = mutrate_observed), position = position_dodge(width = 0.9), vjust = -0.5, hjust = -0.1) + +# xlab("") + +# ylab("Mutation rate") + +# theme(legend.position="bottom") +# dev.off() diff --git a/bin/omega_comparison_per_site.py b/bin/omega_comparison_per_site.py index 4b44d73f..488b0661 100755 --- a/bin/omega_comparison_per_site.py +++ b/bin/omega_comparison_per_site.py @@ -6,9 +6,14 @@ import pandas as pd import scipy.stats as stats -def load_data(panel_file, mutations_file, mutabilities_file): +from utils import MIN_NONZERO_PVALUE + +def load_data(panel_file, mutations_file, mutabilities_file, genes_subset=None): # Load captured panel panel = pd.read_csv(panel_file, sep="\t", compression="gzip") + if genes_subset: + genes_list = [gene.strip() for gene in genes_subset.split(',')] + panel = panel[panel['GENE'].isin(genes_list)].reset_index(drop=True) # Load mutations mutations = pd.read_csv(mutations_file, sep="\t") if "EFFECTIVE_MUTS" not in mutations.columns: @@ -32,6 +37,17 @@ def poisson_pvalue(observed, expected): """ return 1 - stats.poisson.cdf(observed - 1, expected) +def benjamini_hochberg(pvals): + if pvals.size == 0: + return pvals + order = np.argsort(pvals) + ranked = np.arange(1, pvals.size + 1) + adjusted = np.empty_like(pvals, dtype=float) + adjusted[order] = pvals[order] * pvals.size / ranked + adjusted_sorted = np.minimum.accumulate(adjusted[order][::-1])[::-1] + adjusted[order] = adjusted_sorted + return np.clip(adjusted, 0.0, 1.0) + def compute_by_size(panel, mutations, mutabilities, size, sample_column): # Merge mutations and panel to associate protein positions mutations = mutations.merge(panel[['CHROM', 'POS', 'REF', 'ALT', 'Protein_position', 'Amino_acids', 'GENE', 'Feature']], @@ -66,6 +82,17 @@ def compute_by_size(panel, mutations, mutabilities, size, sample_column): result["OBS/EXP"] = (result["OBSERVED_MUTS"] / result["EXPECTED_MUTS"]).fillna(0) result["OBS/EXP"] = result["OBS/EXP"].replace([np.inf, -np.inf], 0) result["p_value"] = result[["OBSERVED_MUTS", "EXPECTED_MUTS"]].apply(lambda row: poisson_pvalue(row["OBSERVED_MUTS"], row["EXPECTED_MUTS"]), axis=1) + + # fill pvalues == 0 with the value or resolution limit + result["p_value"] = result["p_value"].replace(0, MIN_NONZERO_PVALUE) + + result["p_value_adj"] = np.nan + for gene, gene_df in result.groupby("GENE", dropna=False): + valid_mask = gene_df["p_value"].notna() + if not valid_mask.any(): + continue + adjusted = benjamini_hochberg(gene_df.loc[valid_mask, "p_value"].to_numpy()) + result.loc[gene_df.index[valid_mask], "p_value_adj"] = adjusted return result @@ -78,9 +105,10 @@ def compute_by_size(panel, mutations, mutabilities, size, sample_column): @click.option('--panel-file', type=click.Path(exists=True), required=True, help="Path to captured panel file (gzip compressed).") @click.option('--size', type=str, default='aminoacid_change', help="Comma-separated list of sizes or 'all' for all sizes. The options are: 'site', 'aminoacid', 'aminoacid_change'") @click.option('--output-prefix', type=str, required=True, help="Output file prefix.") -def main(panel_file, mutations_file, mutabilities_file, size, output_prefix): +@click.option('--genes', type=str, default= '', required=False, help="List of genes to subset (comma-separated). If not provided, all genes will be included.") +def main(panel_file, mutations_file, mutabilities_file, size, output_prefix, genes): """Compute comparison between observed mutations and accumulated mutabilities.""" - panel, mutations, mutabilities = load_data(panel_file, mutations_file, mutabilities_file) + panel, mutations, mutabilities = load_data(panel_file, mutations_file, mutabilities_file, genes if genes != '' else None) # Get sample column dynamically sample_column = get_sample_column(mutabilities) diff --git a/bin/omega_multiple_testing.py b/bin/omega_multiple_testing.py new file mode 100755 index 00000000..08d65077 --- /dev/null +++ b/bin/omega_multiple_testing.py @@ -0,0 +1,87 @@ +#!/usr/bin/env python + +import json +from typing import Iterable + +import click +import numpy as np +import pandas as pd + +MIN_NONZERO_PVALUE = 1.17e-38 + + +def _load_json_keys(path: str) -> set[str]: + with open(path, "r", encoding="utf-8") as handle: + data = json.load(handle) + return set(data.keys()) + + +def _benjamini_hochberg(pvals: np.ndarray) -> np.ndarray: + if pvals.size == 0: + return pvals + order = np.argsort(pvals) + ranked = np.arange(1, pvals.size + 1) + adjusted = np.empty_like(pvals, dtype=float) + adjusted[order] = pvals[order] * pvals.size / ranked + adjusted_sorted = np.minimum.accumulate(adjusted[order][::-1])[::-1] + adjusted[order] = adjusted_sorted + return np.clip(adjusted, 0.0, 1.0) + + +def _sample_category(sample: str, group_names: set[str], sample_names: set[str]) -> str: + if sample == "all_samples": + return "all_samples" + if sample in group_names: + return "group" + if sample in sample_names: + return "sample" + return "unknown" + + +@click.command() +@click.option("--omegas-file", type=click.Path(exists=True), required=True) +@click.option("--samples-json", type=click.Path(exists=True), required=True) +@click.option("--groups-json", type=click.Path(exists=True), required=True) +@click.option("--output", type=click.Path(), required=True) +def main(omegas_file: str, samples_json: str, groups_json: str, output: str) -> None: + df = pd.read_table(omegas_file) + + pvalues = pd.to_numeric(df["pvalue"], errors="coerce") + df["pvalue"] = pvalues.mask(pvalues <= 0, MIN_NONZERO_PVALUE) + df["pvalue_adj"] = np.nan + + sample_names = _load_json_keys(samples_json) + group_names = _load_json_keys(groups_json) - {"all_samples"} + + sample_categories = df["sample"].astype(str).map( + lambda sample: _sample_category(sample, group_names, sample_names) + ) + region_levels = df["gene"].astype(str).str.contains("--", na=False).map( + lambda has_subgenic: "subgenic" if has_subgenic else "gene" + ) + + for sample_group in ("all_samples", "group", "sample"): + for region_level in ("gene", "subgenic"): + click.echo( + "Correcting omega values for: " + f"{sample_group} samples, {region_level} regions" + ) + mask = (sample_categories == sample_group) & (region_levels == region_level) + valid_mask = mask & df["pvalue"].notna() + if not valid_mask.any(): + continue + adjusted = _benjamini_hochberg(df.loc[valid_mask, "pvalue"].to_numpy()) + df.loc[valid_mask, "pvalue_adj"] = adjusted + + unknown_samples = sample_categories[sample_categories == "unknown"].unique() + if len(unknown_samples) > 0: + click.echo( + "Warning: omega results contain samples not found in groups/samples JSON: " + f"{', '.join(sorted(unknown_samples))}" + ) + + df.to_csv(output, sep="\t", index=False) + + +if __name__ == "__main__": + main() diff --git a/bin/omega_select_mutdensity.py b/bin/omega_select_mutdensity.py index 912afbbf..9162d6a3 100755 --- a/bin/omega_select_mutdensity.py +++ b/bin/omega_select_mutdensity.py @@ -8,7 +8,40 @@ def select_syn_mutdensity(mutdensity_file, output_file, mode): """ - INFO + This function selects the synonymous mutation densities for all genes + from the mutation density file of all samples computed in the adjusted mode. + Accounting for the trinucleotide mutation probabilities and sequencing depth. + + right now the use of mode is not implemented, + since we only compute one type of synonymous mutation densities. + """ + + mutdensity_df = pd.read_csv(mutdensity_file, sep = "\t", header = 0, na_values = custom_na_values) + + synonymous_mutdensities_all_samples = mutdensity_df[(mutdensity_df["SAMPLE_ID"] == 'all_samples') & + ~(mutdensity_df["GENE"].str.contains("--")) + ].reset_index(drop = True) + + synonymous_mutdensities_genes = synonymous_mutdensities_all_samples[['GENE', 'synonymous']].copy() + + # TODO implement these different modes if appropriate + # if mode == 'mutations': + # synonymous_mutdensities_genes = synonymous_mutdensities_all_samples[['GENE', 'synonymous']] + # elif mode == 'mutated_reads': + # synonymous_mutdensities_genes = synonymous_mutdensities_all_samples[['GENE', 'synonymous']] + + synonymous_mutdensities_genes.columns = ["GENE", "MUTDENSITY"] + synonymous_mutdensities_genes.to_csv(f"{output_file}", + header=True, + index=False, + sep="\t") + + +def select_syn_mutdensity_old(mutdensity_file, output_file, mode): + """ + From simple adjusted version + This function needs to be removed eventually when the use of the + updated adjusted mutation density is functional in omega """ mutdensity_df = pd.read_csv(mutdensity_file, sep = "\t", header = 0, na_values = custom_na_values) @@ -29,13 +62,17 @@ def select_syn_mutdensity(mutdensity_file, output_file, mode): sep="\t") + @click.command() @click.option('--mutdensities', type=click.Path(exists=True), help='Input mutation density file') @click.option('--output', type=click.Path(), help='Output file') -@click.option('--mode', type=click.Choice(['mutations', 'mutated_reads']), default='mutations') +@click.option('--mode', type=click.Choice(['mutations', 'mutated_reads', 'new']), default='mutations') def main(mutdensities, output, mode): click.echo("Selecting the gene synonymous mutation densities...") - select_syn_mutdensity(mutdensities, output, mode) + if mode == 'new': + select_syn_mutdensity(mutdensities, output, mode) + else: + select_syn_mutdensity_old(mutdensities, output, mode) if __name__ == '__main__': main() diff --git a/bin/omega_vs_global_vs_dndscv_qc.py b/bin/omega_vs_global_vs_dndscv_qc.py new file mode 100755 index 00000000..03cd6c07 --- /dev/null +++ b/bin/omega_vs_global_vs_dndscv_qc.py @@ -0,0 +1,235 @@ +#!/usr/bin/env python + +# Needed basic packages +import click +import os +import pandas as pd +import matplotlib.pyplot as plt +from matplotlib.backends.backend_pdf import PdfPages +import numpy as np +import seaborn as sns +import json +from scipy import stats +from scipy.stats import pearsonr + + +# Functions +def filter_omega_tables(df): + # Apply basic filter for omega/omegaglobal tables + filtered_omega = df[~(df['gene'].str.contains("--")) & # discards exons if they are included in the analysis + (df['impact'].isin(['missense', 'truncating']))].copy() + print('Filtered omega/omegaglobal table:') + print(filtered_omega.shape) + return filtered_omega + +def process_dndscv_table(dndscv_df): + + # Separate missense and truncating variants for dNdScv table in two separated tables + missense_df = dndscv_df[['sample', 'gene_name', 'n_mis', 'wmis_cv', 'qmis_cv']].copy() + missense_df = missense_df.rename(columns={'n_mis': 'mutations', 'wmis_cv': 'dnds', 'qmis_cv': 'pvalue', 'gene_name': 'gene'}) + missense_df['impact'] = 'missense' + + truncating_df = dndscv_df[['sample', 'gene_name', 'n_non', 'n_spl', 'wnon_cv', 'qtrunc_cv']].copy() + truncating_df['n_trunc'] = truncating_df['n_non'] + truncating_df['n_spl'] + truncating_df = truncating_df.rename(columns={'n_trunc': 'mutations', 'wnon_cv': 'dnds', 'qtrunc_cv': 'pvalue', 'gene_name': 'gene'}) + truncating_df['impact'] = 'truncating' + + # Then concat the two tables + dndscv_cv_df = pd.concat([missense_df, truncating_df.drop(columns=['n_non', 'n_spl'])], axis=0) + print('Processed dndscv table: ') + print(dndscv_cv_df.shape) + + return dndscv_cv_df + +def process_and_analyze_comparison( + primary_df, + comparison_df, + flagged_genes_df, + sample_groups, + output_dir, + comp_label +): + """ + Unified helper to merge, filter flagged genes, export TSV, and apply plotting/correlation. + """ + suffix = f'_{comp_label}' + tsv_filename = f"filtered_omega_{comp_label}_table.tsv" + + # Merge comparison tables (note primary_df is always omega_df) + merged_df = primary_df.merge( + comparison_df, + on=['sample', 'gene', 'impact'], + how='outer', + suffixes=('_omega', suffix) + ) + print(f'Merged omega and {comp_label} table shape:', merged_df.shape) + + # Filter flagged genes + filtered_df = filter_flagged_genes_per_group(merged_df, flagged_genes_df) + print(f'Filtered table shape ({comp_label}):', filtered_df.shape) + + # Export filtered table + filtered_output_path = os.path.join(output_dir, tsv_filename) + filtered_df.to_csv(filtered_output_path, sep='\t', index=False) + + # Compute correlation and plot (passing comp_label to handle dynamic columns and filenames) + apply_correlation_and_plotting( + df=filtered_df, + samples_group=sample_groups, + output_dir=output_dir, + comp_label=comp_label + ) + + return + + +def filter_flagged_genes_per_group(df, flagged_cases_df): + # Optimized: Vectorized filtering using a multi-index instead of .apply row-by-row + flagged_index = pd.MultiIndex.from_frame(flagged_cases_df[['cohort', 'gene']].rename(columns={'cohort': 'sample'})) + df_index = pd.MultiIndex.from_frame(df[['sample', 'gene']]) + + filtered_df = df[~df_index.isin(flagged_index)].copy() + return filtered_df + + +def apply_correlation_and_plotting(df, samples_group, output_dir, comp_label): + """ + Plots dNdS comparisons dynamically based on the target dataset (comp_label). + + comp_label: e.g. 'omegaglobal' or 'dndscv' + Column expected: 'dnds_omega' vs f'dnds_{comp_label}' + """ + + os.makedirs(output_dir, exist_ok=True) # Ensure output dir exists + + # Define dynamic column name and display labels + x_col = 'dnds_omega' + y_col = f'dnds_{comp_label}' + target_display_name = "dNdScv" if comp_label == "dndscv" else "Omega Global" + + for group in sorted(samples_group): + subset = df[df['sample'] == group].copy() + + # Fill missing dNdS values with 0 dynamically + if x_col in subset.columns and y_col in subset.columns: + subset[[x_col, y_col]] = subset[[x_col, y_col]].fillna(0) + print(f'Subset shape for {group} ({target_display_name}): {subset.shape[0]}') + + if subset.empty: + print(f'No data available for sample: {group}') + continue + + fig, axes = plt.subplots(1, 2, figsize=(8, 3), sharex=False, sharey=False) + + impacts = ['missense', 'truncating'] + + for i, impact in enumerate(impacts): + ax = axes[i] + impact_subset = subset[subset['impact'] == impact] + impact_subset_reg = impact_subset[(impact_subset[x_col] > 0) & (impact_subset[y_col] > 0)] + + if impact_subset.empty: + ax.set_title(f'No {impact} data for {group}') + continue + + # Scatter plot + sns.scatterplot(data=impact_subset, x=x_col, y=y_col, hue='gene', ax=ax, legend=False) + + # Regression line + sns.regplot(data=impact_subset_reg, x=x_col, y=y_col, scatter=False, color='red', label='Regression Line', ax=ax) + + # X=Y line + all_vals = pd.concat([impact_subset[x_col], impact_subset[y_col]]) + min_val, max_val = all_vals.min(), all_vals.max() + ax.plot([min_val, max_val], [min_val, max_val], color='black', linestyle='--', label='x=y') + + if impact_subset_reg.shape[0] > 2: + # Compute correlation + corr, pval = pearsonr(impact_subset_reg[x_col], impact_subset_reg[y_col]) + + ax.set_title(f'{impact.capitalize()} mutations\nPearson R: {corr:.3f} (p={pval:.3e})') + else: + ax.set_title(f'{impact.capitalize()} mutations\nNot enough data for correlation') + + ax.set_xlabel('dN/dS (Omega)') + ax.set_ylabel(f'dN/dS ({target_display_name})') + + fig.suptitle(f'Comparison for {group}') + fig.tight_layout() + + output_path = os.path.join(output_dir, f"{group}_omega_vs_{comp_label}_qc_summary_plot.pdf") + with PdfPages(output_path) as pdf: + pdf.savefig(fig) + + plt.close(fig) + + return + + + +@click.command() +@click.option("--input-omega-file", required=True, type=click.Path(exists=True), + help="Directory containing selection/omega/all_omegas.tsv") + +@click.option("--input-omegaglobal-file", required=True, type=click.Path(exists=True), + help="Directory containing selection/omegaglobal/all_omegaglobal.tsv") + +@click.option("--input-dndscv-file", required=False, default=None, type=click.Path(exists=True), #set to default and required false/none to avoid processing when is missing + help="Directory containing selection/dndscv/cv/all_dNdScv.cv.tsv") + +@click.option("--output-dir", required=True, type=click.Path(writable=True), + help="Directory where output files will be written") + +@click.option("--flagged-genes-omega", required=True, type=click.Path(exists=True), + help="Directory where flagged_omega genes were stored from /qc/omega_flagged/debug.syn_flagged_gene.tsv") + +@click.option("--defined-groups", required=True, type=click.Path(exists=True), + help="User defined sample groups in the analysis") + + + +# Main function +def main(input_omega_file, input_omegaglobal_file, input_dndscv_file, output_dir, flagged_genes_omega, defined_groups): + + # Read tables; initialize dndscv to avoid errors + omega_df = pd.read_table(input_omega_file, sep='\t') + omegaglobal_df = pd.read_table(input_omegaglobal_file, sep='\t') + flagged_genes_df = pd.read_table(flagged_genes_omega, sep='\t') + + # Read groups from imported json file and load it as list + with open(defined_groups, 'r') as file: + json_data = json.load(file) + sample_groups = list(json_data.keys()) + print(sample_groups) + + # Apply basic filters to omega/omegaglobal tables + omega_filtered_df = filter_omega_tables(omega_df) + omegaglobal_filtered_df = filter_omega_tables(omegaglobal_df) + + # Initialize a dictionary for comparisons (label used : table to compare) + comparisons = { + 'omegaglobal': omegaglobal_filtered_df + } + + # Add dndscv to comparisons dictionary if provided + if input_dndscv_file: + print(f"Processing optional dNdScv file: {input_dndscv_file}") + dndscv_df = pd.read_table(input_dndscv_file, sep='\t') + dndscv_cv_df = process_dndscv_table(dndscv_df) + comparisons['dndscv'] = dndscv_cv_df + + # Execute helper function for each comparison in dictionary: + + for comp_label, comp_df in comparisons.items(): + process_and_analyze_comparison( + primary_df=omega_filtered_df, + comparison_df=comp_df, + flagged_genes_df=flagged_genes_df, + sample_groups=sample_groups, + output_dir=output_dir, + comp_label=comp_label + ) + + +if __name__ == '__main__': + main() diff --git a/bin/panels_computedna2protein.py b/bin/panels_computedna2protein.py index 69630008..55e01af0 100755 --- a/bin/panels_computedna2protein.py +++ b/bin/panels_computedna2protein.py @@ -85,20 +85,46 @@ def get_transcript_gene_from_maf(path_maf, consensus_file): return gene_transcript_pairs # Generator to filter lines before they reach a dataframe -def get_gff_to_generator(release: int = 111, species: str = "homo_sapiens", genome: str = "GRCh38"): +def get_gff_to_generator(release: int = 111, species: str = "homo_sapiens", genome: str = "GRCh38", gff3_file: str | None = None): """ - Get GFF file from ensembl FTP and filter it on the fly to keep only exon and CDS lines. + Get GFF file from a local path or Ensembl FTP and filter it on the fly to + keep only exon and CDS lines. Parameters ------------ release : int The release number of the Ensembl GFF file to retrieve (default is 111). + gff3_file : str | None + Optional path to a local GFF3 or GFF3.GZ file. If provided, this file + is used instead of downloading from Ensembl. Returns ------------ generator A generator that yields lines from the GFF file that correspond to exon and CDS features. """ + if gff3_file: + LOG.info(f"Using local GFF3 file: {gff3_file}") + + open_func = gzip.open if gff3_file.endswith(".gz") else open + mode = "rt" + try: + with open_func(gff3_file, mode, encoding="utf-8") as handle: + for line_str in handle: + # Skip comments + if line_str.startswith('#'): + continue + + # Only "yield" lines that are exon or CDS + parts = line_str.split('\t') + if len(parts) > 2 and parts[2] in ["exon", "CDS"]: + yield line_str + except OSError as e: + error = RuntimeError(f"Failed to read local GFF3 file: {gff3_file}") + LOG.error(error) + raise error from e + return + url = f"https://ftp.ensembl.org/pub/release-{release}/gff3/{species}/{species.capitalize()}.{genome}.{release}.gff3.gz" # Open request @@ -118,22 +144,22 @@ def get_gff_to_generator(release: int = 111, species: str = "homo_sapiens", geno LOG.error(error) raise error - # decompress the stream on the fly + # Decompress the stream on the fly with gzip.GzipFile(fileobj=response.raw) as gz: for line in gz: # Decode bytes to string line_str = line.decode('utf-8') - + # Skip comments if line_str.startswith('#'): continue - + # Only "yield" lines that are exon or CDS parts = line_str.split('\t') - if parts[2] in ["exon", "CDS"]: + if len(parts) > 2 and parts[2] in ["exon", "CDS"]: yield line_str -def gff_to_filtered_df(gene_n_transcript: pd.DataFrame, release: int) -> pd.DataFrame: +def gff_to_filtered_df(gene_n_transcript: pd.DataFrame, release: int, species: str, genome: str, gff3_file: str | None = None) -> pd.DataFrame: """ Transforms the yields from get_gff_to_generator into a filtered DataFrame. The reading and filtering is done with polars to improve efficiency. @@ -144,6 +170,13 @@ def gff_to_filtered_df(gene_n_transcript: pd.DataFrame, release: int) -> pd.Data A DataFrame containing gene and transcript information. release : int The Ensembl release number to use for GFF file retrieval (default is 111). + species : str + The species for which to retrieve GFF data. + genome : str + The genome assembly for which to retrieve GFF data. + gff3_file : str | None + Optional path to a local GFF3 or GFF3.GZ file. If provided, this file + is used instead of downloading from Ensembl. Returns ------------ @@ -151,7 +184,7 @@ def gff_to_filtered_df(gene_n_transcript: pd.DataFrame, release: int) -> pd.Data A filtered DataFrame of GFF lines for the specified genes and release. """ # Join the generator into a single buffer for Polars to read - filtered_buffer = io.StringIO("".join(get_gff_to_generator(release=release))) + filtered_buffer = io.StringIO("".join(get_gff_to_generator(release=release, species=species, genome=genome, gff3_file=gff3_file))) # Read generator with polars, generate transcript_id column and filter by the genes in the panel df = ( @@ -163,7 +196,7 @@ def gff_to_filtered_df(gene_n_transcript: pd.DataFrame, release: int) -> pd.Data schema_overrides={"start": pl.Int64, "end": pl.Int64, "chr": pl.Utf8, "feature": pl.Utf8, "attributes": pl.Utf8} ) .with_columns([ - pl.col("attributes").str.extract(r"transcript:(ENST\d+)", 1).alias("transcript_id") + pl.col("attributes").str.extract(r"transcript:(ENS[A-Z]*T\d+)", 1).alias("transcript_id") ]) .filter( [pl.col("transcript_id").is_in(gene_n_transcript["Ens_transcript_ID"].to_list())] @@ -278,7 +311,11 @@ def get_exon_coord_wrapper(gene_n_transcript: pd.DataFrame, gff_df: pd.DataFrame final_coord_df = pd.concat(coord_df_lst) if coord_df_lst else pd.DataFrame() final_exons_df = pd.DataFrame(exons_coord_df_lst, columns=["ID", "Chr", "Start", "End", "Strand"]) - LOG.info(f"Retrieved coordinates for {len(final_coord_df['Gene'].unique())} genes and {len(final_exons_df['ID'].unique())} exons.") + if final_coord_df.empty: + LOG.error("No CDS coordinates were retrieved for the specified genes and transcripts.") + exit(1) + else: + LOG.info(f"Retrieved coordinates for {len(final_coord_df['Gene'].unique())} genes and {len(final_exons_df['ID'].unique())} exons.") return final_coord_df, final_exons_df # Coordinate parsing and DNA-to-protein mapping functions @@ -454,7 +491,7 @@ def dna2prot_depth(gene_list: list, coord_df: pd.DataFrame, dna_sites: pd.DataFr return dna_prot_df -def get_dna2prot_depth(gene_n_transcript_info: pd.DataFrame, depth_file: str, consensus_file: str, release: int, species: str, genome: str) -> tuple[pd.DataFrame, pd.DataFrame]: +def get_dna2prot_depth(gene_n_transcript_info: pd.DataFrame, depth_file: str, consensus_file: str, release: int, species: str, genome: str, gff3_file: str | None = None) -> tuple[pd.DataFrame, pd.DataFrame]: """ Function to get the DNA to protein mapping for all positions in the provided list of genes, along with coverage and depth information for each position, and the definition of all exons of @@ -468,6 +505,9 @@ def get_dna2prot_depth(gene_n_transcript_info: pd.DataFrame, depth_file: str, co Path to the file containing depth information for all positions in the genome. consensus_file : str Path to the consensus panel file containing DNA positions and coverage information. + gff3_file : str | None + Optional path to a local GFF3 or GFF3.GZ file. If provided, this file + is used instead of downloading from Ensembl. Returns ------------ @@ -481,7 +521,7 @@ def get_dna2prot_depth(gene_n_transcript_info: pd.DataFrame, depth_file: str, co consensus_df = pd.read_table(consensus_file) depth_df = pd.read_table(depth_file) - gff_df = gff_to_filtered_df(gene_n_transcript_info, release=release) + gff_df = gff_to_filtered_df(gene_n_transcript_info, release=release, species=species, genome=genome, gff3_file=gff3_file) consensus_df = consensus_df.merge(depth_df[["CHROM", "POS", "CONTEXT"]], on = ["CHROM", "POS"], how = 'left') consensus_df = consensus_df.rename(columns={"POS" : "DNA_POS"}) @@ -528,7 +568,7 @@ def plot_coverage_per_gene(depths_df: pd.DataFrame) -> None: plot_single_coverage(coverage_summary, prefix) -def plot_single_coverage(coverage_summary: pd.DataFrame, prefix: str, batch_size: int = 5): +def plot_single_coverage(coverage_summary: pd.DataFrame, prefix: str, batch_size: int = 40): """ Plot coverage for a single prefix (DNA, protein or exon) and save the results in a PDF file. @@ -539,7 +579,7 @@ def plot_single_coverage(coverage_summary: pd.DataFrame, prefix: str, batch_size prefix : str The prefix to plot (DNA, Protein or Exon). batch_size : int, optional - The number of genes to include in each batch when plotting (default is 5). + The number of genes to include in each batch when plotting (default is 40). """ # Generate copy of the coverage summary to avoid modifying the original DataFrame coverage_summary_cp = coverage_summary.copy() @@ -554,7 +594,7 @@ def plot_single_coverage(coverage_summary: pd.DataFrame, prefix: str, batch_size if len(genes_list) < batch_size: - fig, axes = plt.subplots(2, 1, figsize=(8, 6), sharex=True) + fig, axes = plt.subplots(2, 1, figsize=(10, 5), sharex=True) covered_col = True if True in coverage_pivot.columns else (1 if 1 in coverage_pivot.columns else coverage_pivot.columns[-1]) @@ -584,7 +624,7 @@ def plot_single_coverage(coverage_summary: pd.DataFrame, prefix: str, batch_size batch_coverage_pivot = coverage_pivot.loc[batch_genes] batch_coverage_perc = coverage_perc.loc[batch_genes] - fig, axes = plt.subplots(2, 1, figsize=(8, 6), sharex=True) + fig, axes = plt.subplots(2, 1, figsize=(10, 5), sharex=True) covered_col = True if True in batch_coverage_pivot.columns else (1 if 1 in batch_coverage_pivot.columns else batch_coverage_pivot.columns[-1]) @@ -613,11 +653,20 @@ def plot_single_coverage(coverage_summary: pd.DataFrame, prefix: str, batch_size @click.option('--ensembl-species', type=str, default="homo_sapiens", help='Ensembl species name to use for GFF file retrieval (default)') @click.option('--ensembl-genome', type=str, default="GRCh38", help='Ensembl genome name to use for GFF file retrieval (default)') @click.option('--ensembl-release', type=int, default=111, help='Ensembl release number to use for GFF file retrieval (default is 111)') -def main(mutations_file, consensus_file, depths_file, ensembl_species, ensembl_genome, ensembl_release): +@click.option('--gff3-file', type=click.Path(exists=True), default=None, help='Optional local GFF3(.gz) file. If provided, skips Ensembl download') +def main(mutations_file, consensus_file, depths_file, ensembl_species, ensembl_genome, ensembl_release, gff3_file): # Count each mutation only ones if it appears in multiple reads gene_n_transcript = get_transcript_gene_from_maf(mutations_file, consensus_file) - exons_depth, exons_coord_id = get_dna2prot_depth(gene_n_transcript, depths_file, consensus_file, ensembl_release, ensembl_species, ensembl_genome) + exons_depth, exons_coord_id = get_dna2prot_depth( + gene_n_transcript, + depths_file, + consensus_file, + ensembl_release, + ensembl_species, + ensembl_genome, + gff3_file, + ) LOG.info("Exons coordinates and depth computed!") exons_depth.to_csv("depths_per_position_exon_gene.tsv", header = True, index = False, sep = '\t') diff --git a/bin/plot_depths.py b/bin/plot_depths.py index 631b9c03..fc033940 100755 --- a/bin/plot_depths.py +++ b/bin/plot_depths.py @@ -106,14 +106,10 @@ def general_plotting(sample_name, samples_list, bed6_probesByGene_df, genes_list sns.boxplot(data = bed6_probesByGene_df, x = "GENE", y = "MEAN_GENE_DEPTH", ax = ax2, showfliers = False, - # order = panel, palette = colors ) sns.stripplot(data = bed6_probesByGene_df, x = "GENE", y = "MEAN_GENE_DEPTH", ax = ax2, - # order = panel, palette = colors - # hue = "PROJECT_NAME", - # palette = colors, jitter = True, alpha = 0.5, size = 4) ax2.set_title("Depth per GENE") @@ -512,14 +508,15 @@ def process_depths(sample_name, depth_file, panel_bed6_file, panel_name, plot_wi """ depth_df, samples_list, avgdepth_per_sample_names, bed6_probes_df, bed6_probesByGene_df, genes_list = load_depths_and_panel(sample_name, depth_file, panel_bed6_file) - general_plotting(sample_name, samples_list, bed6_probesByGene_df, genes_list) + if len(samples_list) < 200 and len(genes_list) < 200: + general_plotting(sample_name, samples_list, bed6_probesByGene_df, genes_list) - #### - ## Until here all the plots and all the information was at the gene or sample level. - # Now it starts to contain within gene information, exons, normalized scores, ... - #### - if plot_within_gene: - process_within_gene_depths(sample_name, depth_df, bed6_probes_df, bed6_probesByGene_df, genes_list, samples_list, avgdepth_per_sample_names) + #### + ## Until here all the plots and all the information was at the gene or sample level. + # Now it starts to contain within gene information, exons, normalized scores, ... + #### + if plot_within_gene: + process_within_gene_depths(sample_name, depth_df, bed6_probes_df, bed6_probesByGene_df, genes_list, samples_list, avgdepth_per_sample_names) diff --git a/bin/plot_explore_variability.py b/bin/plot_explore_variability.py index e089490c..5119d15f 100755 --- a/bin/plot_explore_variability.py +++ b/bin/plot_explore_variability.py @@ -21,6 +21,39 @@ def filter_data_from_config(dataa, config): # print(filtered_data.shape) return filtered_data +def plot_mutdensity_per_sample(data, samples_list, value, title, sample_column_name = "SAMPLE_ID"): + print("Plotting mutation density per sample for:", title, value) + + muts_per_sample = data[(data["SAMPLE_ID"].isin(samples_list)) + & (data["GENE"] == 'ALL_GENES') + ] + + # Calculate the length of the longest sample name + max_label_length = muts_per_sample[sample_column_name].astype(str).str.len().max() + + # Determine the rotation angle and adjust figure size accordingly + rotation_angle = 90 if (max_label_length > 7) or (samples_list > 30) else 30 # Adjust threshold as needed + # fig_height = 4 + (max_label_length / 10) * 2 # Adjust multiplier as needed + # # Calculate the figure width based on the number of samples + # fig_width = min(18, max(2, len(muts_per_sample) * 0.5)) + + # # Create the figure and axis + # fig, ax = plt.subplots(figsize=(fig_width, fig_height)) + + fig, ax = plt.subplots(figsize=(max(12, 0.1*len(samples_list)), 4)) + sns.barplot(data = muts_per_sample, x = "SAMPLE_ID", + y = value, order=samples_list, + ax = ax, palette = ["salmon"], dodge = False) + plt.xticks(rotation = rotation_angle) + plt.xlabel("") + plt.ylabel("Mutation density\n(muts per Mb)") + plt.title(title) + + plt.tight_layout() + + return fig + + def mut_density_heatmaps(data, genes_list, samples_list, outdir, prefix = '', config_datasets = { "all" : ({"MUTTYPES": 'all_types', "REGIONS": 'all'}, 'MUTDENSITY_MB'), @@ -51,8 +84,11 @@ def mut_density_heatmaps(data, genes_list, samples_list, outdir, prefix = '', for title, (config, value) in config_datasets.items(): print("Creating heatmap for:", title, config, value) - filtered_data = filter_data_from_config(data, config) - # print(filtered_data[['GENE', 'SAMPLE_ID', value]].head()) + filtered_data = filter_data_from_config(data, config) # filtered_data[['GENE', 'SAMPLE_ID', value]] + fig = plot_mutdensity_per_sample(filtered_data, samples_list, value, title) + pdf.savefig(fig) + plt.close(fig) + # Create a pivot table for the heatmap heatmap_data = filtered_data.pivot_table(index='GENE', columns='SAMPLE_ID', values=value) heatmap_data = heatmap_data.reindex(index=genes_list, columns=samples_list) @@ -82,17 +118,16 @@ def mut_density_heatmaps(data, genes_list, samples_list, outdir, prefix = '', + def adj_mut_density_heatmaps(data, genes_list, samples_list, outdir, prefix = '', config_datasets = { - "all" : ({"MUTTYPES": 'all_types', "REGIONS": 'all'}, 'MUTDENSITY_MB'), - "all protein-affecting" : ({"MUTTYPES": 'all_types', "REGIONS": 'protein_affecting'}, 'MUTDENSITY_MB'), - "all non-protein-affecting" : ({"MUTTYPES": 'all_types', "REGIONS": 'non_protein_affecting'}, 'MUTDENSITY_MB'), - "SNVs" : ({"MUTTYPES": 'SNV', "REGIONS": 'all'}, 'MUTDENSITY_MB'), - "SNVs protein-affecting" : ({"MUTTYPES": 'SNV', "REGIONS": 'protein_affecting'}, 'MUTDENSITY_MB'), - "SNVs non-protein-affecting" : ({"MUTTYPES": 'SNV', "REGIONS": 'non_protein_affecting'}, 'MUTDENSITY_MB'), - "INDELs" : ({"MUTTYPES": 'DELETION-INSERTION', "REGIONS": 'all'}, 'MUTDENSITY_MB'), - "INDELs protein-affecting" : ({"MUTTYPES": 'DELETION-INSERTION', "REGIONS": 'protein_affecting'}, 'MUTDENSITY_MB'), - "INDELs non-protein-affecting" : ({"MUTTYPES": 'DELETION-INSERTION', "REGIONS": 'non_protein_affecting'}, 'MUTDENSITY_MB') + "synonymous" : "synonymous", + "missense" : "missense", + "nonsense" : "nonsense", + "essential_splice" : "essential_splice", + "truncating" : "truncating", + "nonsynonymous_splice" : "nonsynonymous_splice", + "all_impacts" : "all_impacts", } ): """ @@ -105,13 +140,16 @@ def adj_mut_density_heatmaps(data, genes_list, samples_list, outdir, prefix = '' print("No data available for the selected samples/groups") return - pdf_filename = f"{outdir}/{prefix}mut_density_heatmaps.pdf" + pdf_filename = f"{outdir}/{prefix}_adjusted_mut_density_heatmaps.pdf" with PdfPages(pdf_filename) as pdf: - for title, (config, value) in config_datasets.items(): - print("Creating heatmap for:", title, config, value) - filtered_data = filter_data_from_config(data, config) - # print(filtered_data[['GENE', 'SAMPLE_ID', value]].head()) + for title, value in config_datasets.items(): + print("Creating heatmap for:", title) + filtered_data = data[["GENE", "SAMPLE_ID", value]] + fig = plot_mutdensity_per_sample(filtered_data, samples_list, value, title) + pdf.savefig(fig) + plt.close(fig) + # Create a pivot table for the heatmap heatmap_data = filtered_data.pivot_table(index='GENE', columns='SAMPLE_ID', values=value) heatmap_data = heatmap_data.reindex(index=genes_list, columns=samples_list) @@ -208,13 +246,12 @@ def main(outdir, panel_regions, samples_json, all_groups_json, mutdensities, adj plotting_manager(outdir, genes_list, samples_list, "samples.", data_string, data_objects) except Exception as e: print("Error in the process", e) - + try: plotting_manager(outdir, genes_list, groups_names, "groups.", data_string, data_objects) except Exception as e: print("Error in the process", e) - if __name__ == '__main__': main() diff --git a/bin/plot_gene_saturation.py b/bin/plot_gene_saturation.py index 39dad72d..9185f476 100755 --- a/bin/plot_gene_saturation.py +++ b/bin/plot_gene_saturation.py @@ -447,7 +447,7 @@ def plot_gene_selection(mut_count_df, plot_pars, title, ddg_df=None, - thr_selection=1e-5, + thr_selection=0.05, lst_tracks=["Mut_count", "Site_selection", "Res_depth", "Domain"], default_track_order=False, save=False, @@ -517,7 +517,7 @@ def plot_gene_selection(mut_count_df, ) site_selection_hits_df = pd.DataFrame({"Protein_position" : np.arange(protein_len)+1}).merge( - site_selection_df[site_selection_df["p_value"] < thr_selection].reset_index(drop=True), how="left") + site_selection_df[site_selection_df["p_value_adj"] < thr_selection].reset_index(drop=True), how="left") axes[ax].fill_between( site_selection_hits_df.Protein_position, 0, site_selection_hits_df.Selection, color=plot_pars["colors"]["site_selection_hits"], alpha=1, zorder=2, lw=1.5, label="Significant" @@ -583,7 +583,7 @@ def plot_gene_selection(mut_count_df, for impact in ["missense", "truncating"]: exon_selection_impact = exon_selection_df[exon_selection_df["impact"] == impact] - exon_selection_impact_hits = exon_selection_impact[exon_selection_impact["pvalue"] < thr_selection].reset_index(drop=True) + exon_selection_impact_hits = exon_selection_impact[exon_selection_impact["pvalue_adj"] < thr_selection].reset_index(drop=True) axes[ax].scatter( exon_selection_impact["MID_PROT_POS"], exon_selection_impact["dnds"], zorder=3, color=plot_pars["colors"]["exon_selection"][impact], s=60, lw=0.1, ec="black" @@ -984,7 +984,7 @@ def plot_domain_selection( # this dictionary is defined in the gene plot with the gene--domain name domain_selection_in_gene["color"] = domain_selection_in_gene["DOMAIN_ID"].map(color_map) - domain_selection_in_gene["edge_width"] = domain_selection_in_gene["pvalue"].apply(lambda p: 1.5 if p < pvalue_threshold else 0.2) + domain_selection_in_gene["edge_width"] = domain_selection_in_gene["pvalue_adj"].apply(lambda p: 1.5 if p < pvalue_threshold else 0.2) fig = plt.figure(figsize=(5, 2.5)) @@ -1085,9 +1085,9 @@ def get_selection_groups(df, thr_selection): df = df.copy().rename(columns={"Protein_position": "Pos"}) g0 = df[df["Selection"] == 0].copy() - g1 = df[(df["Selection"] != 0) & (df["p_value"] >= thr_selection)].copy() + g1 = df[(df["Selection"] != 0) & (df["p_value_adj"] >= thr_selection)].copy() - g2 = df[df["p_value"] < thr_selection].copy() + g2 = df[df["p_value_adj"] < thr_selection].copy() g2 = g2.sort_values("Selection").reset_index(drop=True) g0["Group"] = "G0" @@ -1159,7 +1159,7 @@ def plot_all_domain_selection(df, color_map, figsize=(10, 3), show_domain_legend custom_markers = {"missense": "o", "truncating": "D"} # 'o' = Circle, 'D' = Rotated Square custom_size = {"missense": 150, "truncating": 100} # Larger for missense df["color"] = df["domain_id"].map(color_map) - df["edge_width"] = df["pvalue"].apply(lambda p: 1.5 if p < 0.05 else 0.2) + df["edge_width"] = df["pvalue_adj"].apply(lambda p: 1.5 if p < 0.05 else 0.2) fig = plt.figure(figsize=figsize) @@ -1268,7 +1268,8 @@ def plotting_wrapper(maf, exons_depth, o3d_df, exon_selection, def plotting_single_gene(gene, maf, exons_depth, o3d_df, exon_selection, domain_selection, site_selection, output_dir, o3d_seq_df, o3d_pdb_tool_df, domain, - lst_tracks + lst_tracks, + p_value_threshold = 0.05 ): # Data @@ -1311,7 +1312,7 @@ def plotting_single_gene(gene, maf, exons_depth, o3d_df, exon_selection_gene = exon_selection[exon_selection["gene"].str.startswith(gene)] exon_selection_gene = exon_selection_gene.sort_values("exon_rank").reset_index(drop=True).rename(columns={"exon_rank" : "EXON_RANK"}) - exon_selection_gene = exon_selection_gene[["EXON_RANK", "impact", "dnds", "lower", "upper", "pvalue"]] + exon_selection_gene = exon_selection_gene[["EXON_RANK", "impact", "dnds", "lower", "upper", "pvalue_adj"]] site_selection_gene = site_selection[site_selection["GENE"] == gene].sort_values("Protein_position").reset_index(drop=True) @@ -1360,7 +1361,7 @@ def plotting_single_gene(gene, maf, exons_depth, o3d_df, plot_pars=plot_pars, title=f"{gene}", lst_tracks=lst_tracks, - thr_selection=0.00001, + thr_selection=p_value_threshold, default_track_order=False, save=True, filename=f"{output_dir}/{gene}.saturation_all.png" @@ -1397,7 +1398,7 @@ def plotting_single_gene(gene, maf, exons_depth, o3d_df, # Selection groups feat # ===================== - site_selection_gene_grouped = get_selection_groups(site_selection_gene, thr_selection=0.00001) + site_selection_gene_grouped = get_selection_groups(site_selection_gene, thr_selection=p_value_threshold) site_selection_gene_grouped = site_selection_gene_grouped.merge(pdb_tool_gene[["Pos", "pACC"]]) if ddg_gene is not None: @@ -1417,7 +1418,7 @@ def data_loading(sample_name, domain, exons_depth, track_list): # helper: check for files and call readers only when present ## Positive selection files - omega_file = f"output_mle.{sample_name}.global_loc.tsv" + omega_file = f"all_omegas_global_loc.tsv" site_selection_file = f"{sample_name}.aminoacid.comparison.tsv.gz" o3d_df_file = f"{sample_name}.3d_clustering_pos.csv" @@ -1425,6 +1426,7 @@ def data_loading(sample_name, domain, exons_depth, track_list): if os.path.isfile(omega_file): try: omega_table = pd.read_table(omega_file) + omega_table = omega_table[omega_table["sample"] == sample_name] except Exception as e: print(f"Could not read omega file {omega_file}: {e}") omega_table = pd.DataFrame() diff --git a/bin/plot_qc_mutations_vaf.py b/bin/plot_qc_mutations_vaf.py index 347cf08a..4d0365d6 100755 --- a/bin/plot_qc_mutations_vaf.py +++ b/bin/plot_qc_mutations_vaf.py @@ -388,6 +388,42 @@ def plot_vaf_vs_vafam_histogram (maf_df, output_pdf): plt.show() +def vaf_pseudocount(alt_depth, depth, weight, prior_vaf=None): + return (alt_depth + prior_vaf * weight) / (depth + weight) + +def plot_vaf_pseudocount_curve(maf_df, output_pdf, suffix=''): + """ + Plot VAF distribution compared to VAF_AM in a histogram. + + Parameters: + ----------- + maf_df : DataFrame + MAF dataframe containing VAF and VAF_AM columns + """ + + fig, axes = plt.subplots(3, 3, figsize=(8, 8)) + axes = axes.flatten() + average_depth = maf_df[f'DEPTH{suffix}'].mean() + dg = maf_df + prior_vaf = dg[f'ALT_DEPTH{suffix}'].sum() / dg[f'DEPTH{suffix}'].sum() + fig.suptitle(f'VAF{suffix} with pseudocounts, prior vaf={prior_vaf:.2e}, avg. depth={average_depth:.1f}') + for i, weight_prop in enumerate([0, 0.2, 0.5, 0.75, 1, 1.25, 1.5, 2, 3]): + weight = weight_prop * average_depth // 1 + dg['VAF_PSEUDO'] = vaf_pseudocount(dg[f'ALT_DEPTH{suffix}'], dg[f'DEPTH{suffix}'], weight, prior_vaf=prior_vaf) + axes[i].scatter(dg[f'DEPTH{suffix}'], dg['VAF_PSEUDO'], s=3, alpha=0.1) + rho = np.corrcoef(np.log(dg[f'DEPTH{suffix}']), np.log(dg['VAF_PSEUDO']))[0, 1] + axes[i].set_xlabel(f'DEPTH{suffix}') + axes[i].set_ylabel(f'VAF_PSEUDO', fontsize=6) + axes[i].set_yscale('log') + axes[i].set_xscale('log') + axes[i].set_title(f"{weight_prop}|{weight}\nrho={rho:.2f}, \nprop_mut_tissue={dg['VAF_PSEUDO'].sum():.2f}", fontsize=6) + plt.tight_layout() + output_pdf.savefig() + plt.close() + plt.show() + + + @click.command() @click.option('--sample_name', type=str, required=True, help='Name of the sample') @click.option('--maf_file', type=click.Path(exists=True), required=False, help='MAF file with mutations') @@ -405,16 +441,26 @@ def main(sample_name, maf_file, output_prefix, max_n): """ output_pdf_path = f"{output_prefix}.mutations_vaf.pdf" - with PdfPages(output_pdf_path) as pdf: - # Plot VAF vs depth per site - if maf_file: - print(f"Generating VAF vs depth plot from {maf_file}") - maf_df = pd.read_csv(maf_file, sep='\t', na_values=custom_na_values) - plot_vaf_vs_depth_per_site(maf_df, pdf, sample_name, max_n=max_n) - plot_vaf_depth_heatmap(maf_df, pdf, sample_name) - plot_vaf_vs_vafam_histogram(maf_df, pdf) - - print(f"Plots saved to {output_pdf_path}") + # Plot VAF vs depth per site + if maf_file: + print(f"Generating VAF vs depth plot from {maf_file}") + maf_df = pd.read_csv(maf_file, sep='\t', na_values=custom_na_values) + if maf_df.shape[0] < 5: + print(f"There are less than 5 mutations in MAF file {maf_file}. Skipping VAF vs depth plot.") + else: + with PdfPages(output_pdf_path) as pdf: + plot_vaf_vs_depth_per_site(maf_df, pdf, sample_name, max_n=max_n) + print(f"VAF vs depth plot complete") + plot_vaf_depth_heatmap(maf_df, pdf, sample_name) + print(f"VAF vs depth heatmap complete") + plot_vaf_vs_vafam_histogram(maf_df, pdf) + print(f"VAF vs VAF_AM histogram complete") + plot_vaf_pseudocount_curve(maf_df, pdf, suffix='') + print(f"VAF pseudocount vs depth plot complete") + plot_vaf_pseudocount_curve(maf_df, pdf, suffix='_AM') + print(f"VAF AM pseudocount vs depth plot complete") + + print(f"Plots saved to {output_pdf_path}") get_top_mutations(maf_file, output_prefix) if len(maf_df["SAMPLE_ID"].unique()) > 1: diff --git a/bin/plot_saturation_in_genes.py b/bin/plot_saturation_in_genes.py index 8e40f935..39f23caf 100755 --- a/bin/plot_saturation_in_genes.py +++ b/bin/plot_saturation_in_genes.py @@ -81,9 +81,9 @@ def group_mutations(df, mode="aminoacid", count_mutations=True): mode: 'aminoacid', 'protein_position', 'nucleotide_change', 'nucleotide_position' count_mutations: if True, count mutations (for mutation tables); if False, just group (for panel tables) - ## define these as different options within a function: - # this should be coupled with a proper processing of the consensus_enriched_expanded table - # so that it is also grouped by in the same way, there is no need for counting in there + define these as different options within a function: + this should be coupled with a proper processing of the consensus_enriched_expanded table + so that it is also grouped by in the same way, there is no need for counting in there """ if mode == "aminoacid": @@ -221,12 +221,15 @@ def plot_genes(df, mode=None): g.savefig(plot_path, bbox_inches='tight', dpi=300) plt.close(g.fig) -def plot_domains(df, mode=None): +def plot_domains(df, genes = None, mode=None): + if genes is None: + genes = df["GENE_NAME"].unique() + seg_type = 'domain' suffix = f"_{mode}" if mode else "" pdf_path = f"{plots_dir}/saturation_domains_all{suffix}.pdf" with PdfPages(pdf_path) as pdf: - for gene in df["GENE_NAME"].unique(): + for gene in genes: df_gene = df[(df["GENE_NAME"] == gene) & (df["SEGMENT_TYPE"] == "domain")] if df_gene.empty: continue @@ -254,12 +257,15 @@ def plot_domains(df, mode=None): pdf.savefig(g.fig, bbox_inches='tight', dpi=300) plt.close(g.fig) -def plot_exons(df, mode=None): +def plot_exons(df, genes = None, mode=None): + if genes is None: + genes = df["GENE_NAME"].unique() + seg_type = 'exon' suffix = f"_{mode}" if mode else "" pdf_path = f"{plots_dir}/saturation_exons_all{suffix}.pdf" with PdfPages(pdf_path) as pdf: - for gene in df["GENE_NAME"].unique(): + for gene in genes: df_gene = df[(df["GENE_NAME"] == gene) & (df["SEGMENT_TYPE"] == "exon")] if df_gene.empty: continue @@ -289,13 +295,16 @@ def plot_exons(df, mode=None): # Frequency-stratified domain plot -def plot_domains_by_freq(df, mode=None): +def plot_domains_by_freq(df, genes = None, mode=None): + if genes is None: + genes = df["GENE_NAME"].unique() + freq_bin_order = ['3+', '2', '1'] seg_type = 'domain' suffix = f"_{mode}" if mode else "" pdf_path = f"{plots_dir}/saturation_domains_byfreq_all{suffix}.pdf" with PdfPages(pdf_path) as pdf: - for gene in df["GENE_NAME"].unique(): + for gene in genes: df_gene = df[(df["GENE_NAME"] == gene) & (df["SEGMENT_TYPE"] == "domain")] if df_gene.empty: continue @@ -335,13 +344,15 @@ def plot_domains_by_freq(df, mode=None): plt.close(fig) # Frequency-stratified exon plot -def plot_exons_by_freq(df, mode=None): +def plot_exons_by_freq(df, genes = None, mode=None): + if genes is None: + genes = df["GENE_NAME"].unique() freq_bin_order = ['3+', '2', '1'] seg_type = 'exon' suffix = f"_{mode}" if mode else "" pdf_path = f"{plots_dir}/saturation_exons_byfreq_all{suffix}.pdf" with PdfPages(pdf_path) as pdf: - for gene in df["GENE_NAME"].unique(): + for gene in genes: df_gene = df[(df["GENE_NAME"] == gene) & (df["SEGMENT_TYPE"] == "exon")] if df_gene.empty: continue @@ -385,14 +396,20 @@ def plot_exons_by_freq(df, mode=None): # Call all three plotting functions from the same table def plot_all_saturation_tables(df, mode=None): plot_genes(df, mode) - plot_domains(df, mode) - plot_exons(df, mode) + + genes_list = df["GENE_NAME"].unique() + genes_list = genes_list[:200] + plot_domains(df, genes=genes_list, mode=mode) + plot_exons(df, genes=genes_list, mode=mode) # Call all three frequency-stratified plotting functions from the same table def plot_all_saturation_tables_by_freq(df, mode=None): - plot_domains_by_freq(df, mode) - plot_exons_by_freq(df, mode) + genes_list = df["GENE_NAME"].unique() + genes_list = genes_list[:200] + + plot_domains_by_freq(df, genes=genes_list, mode=mode) + plot_exons_by_freq(df, genes=genes_list, mode=mode) def generate_all_saturation_plots(consensus_enriched_expanded, somatic_maf_clean, diff --git a/bin/plot_selectionsideplots.py b/bin/plot_selectionsideplots.py index 8efc7812..db01a109 100755 --- a/bin/plot_selectionsideplots.py +++ b/bin/plot_selectionsideplots.py @@ -60,7 +60,7 @@ def generate_all_side_figures(sample, omega_data = omega_data.merge(flagged_omega_data, how='left') omega_data = omega_data[(omega_data["impact"].isin(['missense', 'truncating'])) & ~(omega_data["gene"].str.contains('--')) # select only genes - & ~(omega_data["flagged"]) # remove flagged genes + & ~(omega_data["flagged"] == True) # remove flagged genes ] if "omega_trunc" in tools : omega_truncating = omega_data[omega_data["impact"] == "truncating"].reset_index(drop = True)[["gene", "mutations", "dnds", "pvalue", "lower", "upper"]] @@ -590,7 +590,7 @@ def get_all_data(sample, outdir, omega_data = omega_data.merge(flagged_omega_data, how='left') omega_data = omega_data[(omega_data["impact"].isin(['missense', 'truncating'])) & ~(omega_data["gene"].str.contains('--')) # select only genes - & ~(omega_data["flagged"]) # remove flagged genes + & ~(omega_data["flagged"] == True) # remove flagged genes ] if "omega_trunc" in tracks: diff --git a/bin/test/test_check_contamination.py b/bin/test/test_check_contamination.py new file mode 100644 index 00000000..ab0d365d --- /dev/null +++ b/bin/test/test_check_contamination.py @@ -0,0 +1,93 @@ +#!/usr/bin/env python3 +""" +Unit tests for check_contamination.py + +Tests the functionality of checking contamination in genomic data. +""" + +import sys +import unittest +from pathlib import Path +import pandas as pd +import numpy as np + +# Add the bin directory to the path to import the module +sys.path.insert(0, str(Path(__file__).parent.parent)) +from check_contamination import compute_shared_variants + +class TestComputeSharedVariants(unittest.TestCase): + """Test suite for compute_shared_variants function.""" + + def _somatic_df(self): + """Return a minimal somatic variants DataFrame.""" + return pd.DataFrame({ + 'SAMPLE_ID': ['s1', 's1', 's2'], + 'MUT_ID': ['mut1', 'mut2', 'mut3'], + }) + + def _germline_df(self): + """Return a minimal germline variants DataFrame.""" + return pd.DataFrame({ + 'SAMPLE_ID': ['s1', 's1', 's2'], + 'MUT_ID': ['mut1', 'mut4', 'mut5'], + }) + + # ---- basic behaviour ------------------------------------------------ + + def test_returns_dataframe(self): + result = compute_shared_variants(self._somatic_df(), self._germline_df()) + self.assertIsInstance(result, pd.DataFrame) + + def test_dtype_is_int(self): + result = compute_shared_variants(self._somatic_df(), self._germline_df()) + self.assertEqual(result.dtypes['s1'], np.int64) + + def test_shared_mutations_count(self): + """s1 shares mut1 with s1 → cell [s1, s1] == 1.""" + result = compute_shared_variants(self._somatic_df(), self._germline_df()) + self.assertEqual(result.loc['s1', 's1'], 1) + + def test_no_shared_mutations(self): + """s2 has mut3 which s1/s2 don't have → cell [s2, s1] == 0.""" + result = compute_shared_variants(self._somatic_df(), self._germline_df()) + self.assertEqual(result.loc['s2', 's1'], 0) + + def test_multiple_shared_mutations(self): + som = pd.DataFrame({'SAMPLE_ID': ['s1'], 'MUT_ID': ['mut1']}) + ger = pd.DataFrame({ + 'SAMPLE_ID': ['s1', 's1', 's1'], + 'MUT_ID': ['mut1', 'mut1', 'mut2'], # mut1 appears twice + }) + result = compute_shared_variants(som, ger) + # mut1 is in both sets → intersection size is 1 (set semantics) + self.assertEqual(result.loc['s1', 's1'], 1) + + def test_empty_somatic(self): + result = compute_shared_variants( + pd.DataFrame({'SAMPLE_ID': [], 'MUT_ID': []}), + self._germline_df(), + ) + self.assertEqual(len(result), 0) + + def test_empty_germline(self): + result = compute_shared_variants( + self._somatic_df(), + pd.DataFrame({'SAMPLE_ID': [], 'MUT_ID': []}), + ) + self.assertEqual(len(result.columns), 0) + + def test_both_empty(self): + result = compute_shared_variants( + pd.DataFrame({'SAMPLE_ID': [], 'MUT_ID': []}), + pd.DataFrame({'SAMPLE_ID': [], 'MUT_ID': []}), + ) + self.assertIsInstance(result, pd.DataFrame) + self.assertEqual(result.shape, (0, 0)) + + def test_samples_sorted(self): + """Sample IDs should appear in sorted order in index/columns.""" + som = pd.DataFrame({'SAMPLE_ID': ['z', 'a'], 'MUT_ID': ['m1', 'm2']}) + ger = pd.DataFrame({'SAMPLE_ID': ['y', 'b'], 'MUT_ID': ['m3', 'm4']}) + result = compute_shared_variants(som, ger) + self.assertEqual(list(result.index), ['a', 'z']) + self.assertEqual(list(result.columns), ['b', 'y']) diff --git a/bin/test/test_utils_filter.py b/bin/test/test_utils_filter.py new file mode 100644 index 00000000..cfb630b6 --- /dev/null +++ b/bin/test/test_utils_filter.py @@ -0,0 +1,554 @@ +#!/usr/bin/env python3 +""" +Unit tests for utils_filter.py. + +Covers: + - somatic_mask / germline_mask (each and combined) + - filter_maf (all criterion branches) + - load_filter_criteria (parsing, combining, prefix stripping) + - expand_filter_column (boolean column creation + required columns) + - extract_flagged_regions_bed (empty and non-empty BED output) +""" + +import os +import sys +import tempfile +import unittest +from pathlib import Path + +import pandas as pd + +# Add the bin directory to the path to import sibling modules +sys.path.insert(0, str(Path(__file__).parent.parent)) +from utils_filter import ( + expand_filter_column, + extract_flagged_regions_bed, + filter_maf, + germline_mask, + load_filter_criteria, + somatic_mask, +) + +THRESHOLD = 0.3 + +# --------------------------------------------------------------------------- +# Helpers +# --------------------------------------------------------------------------- + + +def _make_df(rows: list[tuple[float, float, float]]) -> pd.DataFrame: + """Build a minimal MAF DataFrame with VAF, vd_VAF, and VAF_AM columns.""" + vafs, vd_vafs, vaf_ams = zip(*rows) + return pd.DataFrame({"VAF": list(vafs), "vd_VAF": list(vd_vafs), "VAF_AM": list(vaf_ams)}) + + +def _make_maf(rows: list[dict]) -> pd.DataFrame: + """Build a MAF DataFrame from a list of row dicts, preserving column order.""" + return pd.DataFrame(rows) + + +# --------------------------------------------------------------------------- +# somatic_mask +# --------------------------------------------------------------------------- + + +class TestSomaticMask(unittest.TestCase): + """Tests for somatic_mask(maf_df, threshold).""" + + def test_all_below_threshold_is_somatic(self): + """All three VAF columns strictly below threshold → somatic True.""" + df = _make_df([(0.1, 0.2, 0.05)]) + result = somatic_mask(df, THRESHOLD) + self.assertListEqual(result.tolist(), [True]) + + def test_all_above_threshold_is_not_somatic(self): + """All three VAF columns strictly above threshold → somatic False.""" + df = _make_df([(0.5, 0.6, 0.4)]) + result = somatic_mask(df, THRESHOLD) + self.assertListEqual(result.tolist(), [False]) + + def test_boundary_equality_is_somatic(self): + """All three VAF columns exactly equal to threshold → somatic True (≤ is inclusive).""" + df = _make_df([(THRESHOLD, THRESHOLD, THRESHOLD)]) + result = somatic_mask(df, THRESHOLD) + self.assertListEqual(result.tolist(), [True]) + + def test_asymmetric_is_not_somatic(self): + """Mixed columns (some ≤ threshold, some > threshold) → somatic False.""" + df = _make_df([(0.1, 0.1, 0.5)]) + result = somatic_mask(df, THRESHOLD) + self.assertListEqual(result.tolist(), [False]) + + def test_multiple_rows(self): + """Full four-case matrix in a single DataFrame.""" + df = _make_df( + [ + (0.1, 0.2, 0.05), # all-below → True + (0.5, 0.6, 0.4), # all-above → False + (THRESHOLD, THRESHOLD, THRESHOLD), # boundary → True + (0.1, 0.1, 0.5), # asymmetric → False + ] + ) + result = somatic_mask(df, THRESHOLD) + self.assertListEqual(result.tolist(), [True, False, True, False]) + + +# --------------------------------------------------------------------------- +# germline_mask +# --------------------------------------------------------------------------- + + +class TestGermlineMask(unittest.TestCase): + """Tests for germline_mask(maf_df, threshold).""" + + def test_all_below_threshold_is_not_germline(self): + """All three VAF columns strictly below threshold → germline False.""" + df = _make_df([(0.1, 0.2, 0.05)]) + result = germline_mask(df, THRESHOLD) + self.assertListEqual(result.tolist(), [False]) + + def test_all_above_threshold_is_germline(self): + """All three VAF columns strictly above threshold → germline True.""" + df = _make_df([(0.5, 0.6, 0.4)]) + result = germline_mask(df, THRESHOLD) + self.assertListEqual(result.tolist(), [True]) + + def test_boundary_equality_is_not_germline(self): + """All three VAF columns exactly equal to threshold → germline False (> is exclusive).""" + df = _make_df([(THRESHOLD, THRESHOLD, THRESHOLD)]) + result = germline_mask(df, THRESHOLD) + self.assertListEqual(result.tolist(), [False]) + + def test_asymmetric_is_not_germline(self): + """Mixed columns (some ≤ threshold, some > threshold) → germline False.""" + df = _make_df([(0.1, 0.1, 0.5)]) + result = germline_mask(df, THRESHOLD) + self.assertListEqual(result.tolist(), [False]) + + def test_multiple_rows(self): + """Full four-case matrix in a single DataFrame.""" + df = _make_df( + [ + (0.1, 0.2, 0.05), # all-below → False + (0.5, 0.6, 0.4), # all-above → True + (THRESHOLD, THRESHOLD, THRESHOLD), # boundary → False + (0.1, 0.1, 0.5), # asymmetric → False + ] + ) + result = germline_mask(df, THRESHOLD) + self.assertListEqual(result.tolist(), [False, True, False, False]) + + +# --------------------------------------------------------------------------- +# somatic + germline masks are never simultaneously True +# --------------------------------------------------------------------------- + + +class TestMasksAreNotComplements(unittest.TestCase): + """Verify that somatic and germline masks are never simultaneously True.""" + + def test_no_row_is_true_in_both_masks(self): + """No row should satisfy both somatic and germline conditions at once.""" + df = _make_df( + [ + (0.1, 0.2, 0.05), # all-below + (0.5, 0.6, 0.4), # all-above + (THRESHOLD, THRESHOLD, THRESHOLD), # boundary + (0.1, 0.1, 0.5), # asymmetric + ] + ) + somatic = somatic_mask(df, THRESHOLD) + germline = germline_mask(df, THRESHOLD) + both_true = (somatic & germline).tolist() + self.assertListEqual(both_true, [False, False, False, False]) + + def test_asymmetric_row_is_false_in_both_masks(self): + """An asymmetric row must be False in somatic AND False in germline (the 'neither' case).""" + df = _make_df([(0.1, 0.1, 0.5)]) + self.assertFalse(somatic_mask(df, THRESHOLD).iloc[0]) + self.assertFalse(germline_mask(df, THRESHOLD).iloc[0]) + + +# --------------------------------------------------------------------------- +# filter_maf +# --------------------------------------------------------------------------- + + +class TestFilterMaf(unittest.TestCase): + """Tests for filter_maf(maf_df, filter_criteria).""" + + def _base_maf(self) -> pd.DataFrame: + """Return a small MAF DataFrame exercising all criterion branches.""" + return _make_maf( + [ + {"MUT_ID": "M1", "VAF": 0.1, "DEPTH": 50, "FILTER": "PASS", "TYPE": "SNV", "FILTER.not_covered": False}, + { + "MUT_ID": "M2", + "VAF": 0.4, + "DEPTH": 30, + "FILTER": "n_rich;NM20", + "TYPE": "SNV", + "FILTER.not_covered": False, + }, + { + "MUT_ID": "M3", + "VAF": 0.2, + "DEPTH": 60, + "FILTER": "low_mappability", + "TYPE": "INDEL", + "FILTER.not_covered": True, + }, + { + "MUT_ID": "M4", + "VAF": 0.05, + "DEPTH": 80, + "FILTER": "PASS", + "TYPE": "SNV", + "FILTER.not_covered": False, + }, + ] + ) + + # --- numeric operator branch (len(operator)==2) --- + + def test_numeric_le_filters_correctly(self): + """('VAF', 'le 0.3') keeps only rows with VAF ≤ 0.3.""" + df = self._base_maf() + result = filter_maf(df, [("VAF", "le 0.3")]) + self.assertListEqual(sorted(result["MUT_ID"].tolist()), ["M1", "M3", "M4"]) + + def test_numeric_ge_filters_correctly(self): + """('DEPTH', 'ge 50') keeps only rows with DEPTH ≥ 50.""" + df = self._base_maf() + result = filter_maf(df, [("DEPTH", "ge 50")]) + self.assertListEqual(sorted(result["MUT_ID"].tolist()), ["M1", "M3", "M4"]) + + def test_numeric_lt_filters_correctly(self): + """('VAF', 'lt 0.2') keeps only rows with VAF < 0.2.""" + df = self._base_maf() + result = filter_maf(df, [("VAF", "lt 0.2")]) + self.assertListEqual(sorted(result["MUT_ID"].tolist()), ["M1", "M4"]) + + def test_multiple_numeric_criteria_are_anded(self): + """Combining two numeric criteria narrows the result set.""" + df = self._base_maf() + result = filter_maf(df, [("VAF", "le 0.3"), ("DEPTH", "ge 50")]) + self.assertListEqual(sorted(result["MUT_ID"].tolist()), ["M1", "M3", "M4"]) + + # --- notcontains / contains branch --- + + def test_notcontains_excludes_matching_filter_token(self): + """('FILTER', 'notcontains n_rich') removes rows whose FILTER cell contains 'n_rich'.""" + df = self._base_maf() + result = filter_maf(df, [("FILTER", "notcontains n_rich")]) + # M2 has 'n_rich' in its FILTER → removed + self.assertNotIn("M2", result["MUT_ID"].tolist()) + self.assertIn("M1", result["MUT_ID"].tolist()) + + def test_notcontains_respects_semicolon_split(self): + """notcontains splits on ';' so a token that is a prefix of another is handled correctly.""" + df = _make_maf( + [ + {"MUT_ID": "A", "FILTER": "NM20;PASS"}, + {"MUT_ID": "B", "FILTER": "NM200;PASS"}, # 'NM20' is NOT a token here + {"MUT_ID": "C", "FILTER": "PASS"}, + ] + ) + result = filter_maf(df, [("FILTER", "notcontains NM20")]) + self.assertNotIn("A", result["MUT_ID"].tolist()) + self.assertIn("B", result["MUT_ID"].tolist()) + self.assertIn("C", result["MUT_ID"].tolist()) + + def test_contains_keeps_only_matching_filter_token(self): + """('FILTER', 'contains n_rich') keeps only rows whose FILTER cell contains 'n_rich'.""" + df = self._base_maf() + result = filter_maf(df, [("FILTER", "contains n_rich")]) + self.assertListEqual(result["MUT_ID"].tolist(), ["M2"]) + + # --- boolean column branch --- + + def test_boolean_criterion_true_selects_matching_rows(self): + """('FILTER.not_covered', True) keeps only rows where FILTER.not_covered is True.""" + df = self._base_maf() + result = filter_maf(df, [("FILTER.not_covered", True)]) + self.assertListEqual(result["MUT_ID"].tolist(), ["M3"]) + + def test_boolean_criterion_false_selects_matching_rows(self): + """('FILTER.not_covered', False) keeps only rows where FILTER.not_covered is False.""" + df = self._base_maf() + result = filter_maf(df, [("FILTER.not_covered", False)]) + self.assertListEqual(sorted(result["MUT_ID"].tolist()), ["M1", "M2", "M4"]) + + # --- plain-value (equality) branch --- + + def test_plain_value_equality_match(self): + """('TYPE', 'SNV') keeps only rows where TYPE == 'SNV'.""" + df = self._base_maf() + result = filter_maf(df, [("TYPE", "SNV")]) + self.assertListEqual(sorted(result["MUT_ID"].tolist()), ["M1", "M2", "M4"]) + + def test_plain_value_no_match_returns_empty(self): + """('TYPE', 'NONEXISTENT') returns an empty DataFrame.""" + df = self._base_maf() + result = filter_maf(df, [("TYPE", "NONEXISTENT")]) + self.assertEqual(len(result), 0) + + # --- no criteria leaves DataFrame unchanged --- + + def test_empty_criteria_returns_all_rows(self): + """An empty criteria list leaves the DataFrame unchanged.""" + df = self._base_maf() + result = filter_maf(df, []) + self.assertEqual(len(result), len(df)) + + +# --------------------------------------------------------------------------- +# load_filter_criteria +# --------------------------------------------------------------------------- + + +class TestLoadFilterCriteria(unittest.TestCase): + """Tests for load_filter_criteria(filters, somatic_filters).""" + + def test_extracts_notcontains_entries_from_filters(self): + """Items starting with 'notcontains ' in filters are returned with prefix stripped.""" + result = load_filter_criteria("notcontains n_rich,notcontains NM20", "") + self.assertListEqual(sorted(result), ["NM20", "n_rich"]) + + def test_extracts_notcontains_entries_from_somatic_filters(self): + """Items from somatic_filters starting with 'notcontains ' are included.""" + result = load_filter_criteria("", "notcontains low_mappability") + self.assertListEqual(result, ["low_mappability"]) + + def test_combines_both_lists(self): + """Entries from both arguments are merged before filtering.""" + result = load_filter_criteria("notcontains n_rich", "notcontains NM20") + self.assertListEqual(sorted(result), ["NM20", "n_rich"]) + + def test_non_notcontains_entries_are_excluded(self): + """Entries that do not start with 'notcontains ' are silently dropped.""" + result = load_filter_criteria("notcontains n_rich,VAF le 0.3,PASS", "") + self.assertListEqual(result, ["n_rich"]) + + def test_empty_strings_return_empty_list(self): + """Both arguments being empty strings yields an empty list.""" + result = load_filter_criteria("", "") + self.assertListEqual(result, []) + + def test_whitespace_is_trimmed_around_items(self): + """Leading/trailing whitespace around comma-separated items is stripped.""" + result = load_filter_criteria(" notcontains n_rich , notcontains NM20 ", "") + self.assertListEqual(sorted(result), ["NM20", "n_rich"]) + + def test_duplicate_entries_are_preserved(self): + """Duplicates across both arguments are preserved (no deduplication contract).""" + result = load_filter_criteria("notcontains n_rich", "notcontains n_rich") + self.assertEqual(result.count("n_rich"), 2) + + +# --------------------------------------------------------------------------- +# expand_filter_column +# --------------------------------------------------------------------------- + + +class TestExpandFilterColumn(unittest.TestCase): + """Tests for expand_filter_column(maf_df).""" + + def _make_filter_df(self, filter_values: list[str]) -> pd.DataFrame: + """Build a minimal MAF DataFrame with only a FILTER column.""" + return pd.DataFrame({"FILTER": filter_values}) + + def test_creates_boolean_column_for_each_token(self): + """Each unique ';'-delimited token gets its own FILTER. boolean column.""" + df = self._make_filter_df(["n_rich;NM20", "PASS", "NM20"]) + result = expand_filter_column(df) + self.assertIn("FILTER.n_rich", result.columns) + self.assertIn("FILTER.NM20", result.columns) + self.assertIn("FILTER.PASS", result.columns) + + def test_boolean_values_are_correct(self): + """True only where the token is present in that row's FILTER value.""" + df = self._make_filter_df(["n_rich;NM20", "PASS", "NM20"]) + result = expand_filter_column(df) + # Row 0: n_rich and NM20 present + self.assertTrue(result.loc[0, "FILTER.n_rich"]) + self.assertTrue(result.loc[0, "FILTER.NM20"]) + self.assertFalse(result.loc[0, "FILTER.PASS"]) + # Row 1: only PASS present + self.assertFalse(result.loc[1, "FILTER.n_rich"]) + self.assertTrue(result.loc[1, "FILTER.PASS"]) + # Row 2: only NM20 present + self.assertFalse(result.loc[2, "FILTER.n_rich"]) + self.assertTrue(result.loc[2, "FILTER.NM20"]) + + def test_required_columns_always_exist(self): + """FILTER.not_covered and FILTER.not_in_exons are always created even if absent in data.""" + df = self._make_filter_df(["PASS", "PASS"]) + result = expand_filter_column(df) + self.assertIn("FILTER.not_covered", result.columns) + self.assertIn("FILTER.not_in_exons", result.columns) + + def test_required_columns_are_false_when_token_absent(self): + """Required columns are all False when neither token appears in the data.""" + df = self._make_filter_df(["PASS", "n_rich"]) + result = expand_filter_column(df) + self.assertFalse(result["FILTER.not_covered"].any()) + self.assertFalse(result["FILTER.not_in_exons"].any()) + + def test_single_token_per_row(self): + """A FILTER column with no semicolons creates one boolean column per distinct value.""" + df = self._make_filter_df(["alpha", "beta", "alpha"]) + result = expand_filter_column(df) + self.assertIn("FILTER.alpha", result.columns) + self.assertIn("FILTER.beta", result.columns) + self.assertListEqual(result["FILTER.alpha"].tolist(), [True, False, True]) + self.assertListEqual(result["FILTER.beta"].tolist(), [False, True, False]) + + def test_all_required_columns_true_when_token_present(self): + """FILTER.not_covered is True exactly for the rows that contain 'not_covered'.""" + df = self._make_filter_df(["not_covered;n_rich", "PASS", "not_covered"]) + result = expand_filter_column(df) + self.assertListEqual(result["FILTER.not_covered"].tolist(), [True, False, True]) + + +# --------------------------------------------------------------------------- +# extract_flagged_regions_bed +# --------------------------------------------------------------------------- + + +class TestExtractFlaggedRegionsBed(unittest.TestCase): + """Tests for extract_flagged_regions_bed(maf_df, name, filters, specification).""" + + def setUp(self): + """Switch into a fresh temporary directory for each test; restore on teardown.""" + self._tmpdir = tempfile.mkdtemp() + self._orig_dir = os.getcwd() + os.chdir(self._tmpdir) + + def tearDown(self): + """Restore original working directory.""" + os.chdir(self._orig_dir) + + def _make_expanded_maf(self, rows: list[dict]) -> pd.DataFrame: + """Build a MAF with CHROM/POS columns then run expand_filter_column.""" + df = pd.DataFrame(rows) + return expand_filter_column(df) + + # --- empty case --- + + def test_empty_case_creates_empty_bed_file(self): + """No flagged rows → an empty .bed file is touched and the function returns None.""" + df = self._make_expanded_maf( + [ + {"CHROM": "chr1", "POS": 100, "FILTER": "PASS"}, + {"CHROM": "chr1", "POS": 200, "FILTER": "PASS"}, + ] + ) + result = extract_flagged_regions_bed(df, "sample1", ["n_rich"]) + self.assertIsNone(result) + bed_path = Path("sample1.flagged-pos.bed") + self.assertTrue(bed_path.exists()) + self.assertEqual(bed_path.stat().st_size, 0) + + def test_empty_case_with_specification_uses_correct_filename(self): + """specification parameter is included in the BED file name for the empty case.""" + df = self._make_expanded_maf([{"CHROM": "chr1", "POS": 100, "FILTER": "PASS"}]) + extract_flagged_regions_bed(df, "sample1", ["n_rich"], specification="cohort-") + self.assertTrue(Path("sample1.cohort-flagged-pos.bed").exists()) + + def test_empty_case_no_matching_filter_columns(self): + """When filter names have no corresponding FILTER.* columns, result is empty BED.""" + df = self._make_expanded_maf([{"CHROM": "chr1", "POS": 100, "FILTER": "PASS"}]) + result = extract_flagged_regions_bed(df, "sampleX", ["nonexistent_filter"]) + self.assertIsNone(result) + self.assertTrue(Path("sampleX.flagged-pos.bed").exists()) + + # --- non-empty case --- + + def test_nonempty_case_writes_bed_with_correct_columns(self): + """BED file has four tab-separated columns: CHROM, START, END, FILTERS.""" + df = self._make_expanded_maf( + [ + {"CHROM": "chr1", "POS": 500, "FILTER": "n_rich"}, + {"CHROM": "chr2", "POS": 1000, "FILTER": "PASS"}, + ] + ) + extract_flagged_regions_bed(df, "sample2", ["n_rich"]) + bed_path = Path("sample2.flagged-pos.bed") + self.assertTrue(bed_path.exists()) + bed = pd.read_csv(bed_path, sep="\t", header=None, names=["CHROM", "START", "END", "FILTERS"]) + self.assertEqual(len(bed), 1) + row = bed.iloc[0] + self.assertEqual(row["CHROM"], "chr1") + self.assertEqual(row["START"], 500) + self.assertEqual(row["END"], 500) + self.assertIn("FILTER.n_rich", row["FILTERS"]) + + def test_nonempty_case_multiple_filters_joined_with_comma(self): + """When a position has two active filter flags, FILTERS column is comma-joined.""" + df = self._make_expanded_maf( + [ + {"CHROM": "chr1", "POS": 300, "FILTER": "n_rich;NM20"}, + ] + ) + extract_flagged_regions_bed(df, "sample3", ["n_rich", "NM20"]) + bed = pd.read_csv( + Path("sample3.flagged-pos.bed"), sep="\t", header=None, names=["CHROM", "START", "END", "FILTERS"] + ) + self.assertEqual(len(bed), 1) + filters_value = bed.iloc[0]["FILTERS"] + # Both filter column names should appear, joined by comma + self.assertIn("FILTER.n_rich", filters_value) + self.assertIn("FILTER.NM20", filters_value) + self.assertIn(",", filters_value) + + def test_nonempty_case_multiple_rows_all_written(self): + """Multiple flagged positions each produce a row in the BED file.""" + df = self._make_expanded_maf( + [ + {"CHROM": "chr1", "POS": 100, "FILTER": "n_rich"}, + {"CHROM": "chr1", "POS": 200, "FILTER": "NM20"}, + {"CHROM": "chr2", "POS": 50, "FILTER": "PASS"}, + ] + ) + extract_flagged_regions_bed(df, "sample4", ["n_rich", "NM20"]) + bed = pd.read_csv( + Path("sample4.flagged-pos.bed"), sep="\t", header=None, names=["CHROM", "START", "END", "FILTERS"] + ) + self.assertEqual(len(bed), 2) + self.assertSetEqual(set(bed["START"].tolist()), {100, 200}) + + def test_nonempty_case_returns_none(self): + """Function has no explicit return in the non-empty path, so returns None.""" + df = self._make_expanded_maf([{"CHROM": "chr1", "POS": 100, "FILTER": "n_rich"}]) + result = extract_flagged_regions_bed(df, "sample5", ["n_rich"]) + self.assertIsNone(result) + + def test_nonempty_case_with_specification_uses_correct_filename(self): + """specification parameter is included in the BED file name for the non-empty case.""" + df = self._make_expanded_maf([{"CHROM": "chr1", "POS": 100, "FILTER": "n_rich"}]) + extract_flagged_regions_bed(df, "sample6", ["n_rich"], specification="cohort-") + self.assertTrue(Path("sample6.cohort-flagged-pos.bed").exists()) + + def test_nonempty_case_bed_is_sorted_by_chrom_and_pos(self): + """BED rows are sorted by CHROM then POS (ascending).""" + df = self._make_expanded_maf( + [ + {"CHROM": "chr2", "POS": 800, "FILTER": "n_rich"}, + {"CHROM": "chr1", "POS": 999, "FILTER": "n_rich"}, + {"CHROM": "chr1", "POS": 100, "FILTER": "n_rich"}, + ] + ) + extract_flagged_regions_bed(df, "sample7", ["n_rich"]) + bed = pd.read_csv( + Path("sample7.flagged-pos.bed"), sep="\t", header=None, names=["CHROM", "START", "END", "FILTERS"] + ) + self.assertEqual(len(bed), 3) + self.assertEqual(bed.iloc[0]["CHROM"], "chr1") + self.assertEqual(bed.iloc[0]["START"], 100) + self.assertEqual(bed.iloc[1]["START"], 999) + self.assertEqual(bed.iloc[2]["CHROM"], "chr2") + + +if __name__ == "__main__": + unittest.main() diff --git a/bin/utils.py b/bin/utils.py index 9418d59a..c6dd2829 100755 --- a/bin/utils.py +++ b/bin/utils.py @@ -2,6 +2,8 @@ import json import pandas as pd +MIN_NONZERO_PVALUE = 1.17e-38 + def add_filter(old_filt, add_filt, filt_name): """ diff --git a/bin/utils_filter.py b/bin/utils_filter.py index bfd201be..3209ae36 100644 --- a/bin/utils_filter.py +++ b/bin/utils_filter.py @@ -1,15 +1,18 @@ #!/usr/bin/env python import logging -import pandas as pd from pathlib import Path + +import pandas as pd + """ Utility functions for extracting filters from a MAF DataFrame. """ LOG = logging.getLogger(__name__) + def filter_maf(maf_df, filter_criteria): - ''' + """ Filter a MAF dataframe with filtering information coming from a list of tuples. This can be either a dictionary transformed to list with the .items() method or by directly creating a list of tuples. [('VAF', 'le 0.3'), ('VAF_AM', 'le 0.3'), ('vd_VAF', 'le 0.3'), @@ -17,76 +20,127 @@ def filter_maf(maf_df, filter_criteria): ('FILTER', 'notcontains cohort_n_rich_uni'), ('FILTER', 'notcontains NM20'), ('FILTER', 'notcontains no_pileup_support'), ('FILTER', 'notcontains other_sample_SNP'), ('FILTER', 'notcontains low_mappability')] - ''' + """ # Define mappings for operators used in criteria operators = { - 'eq': lambda x, y: x == y, - 'ne': lambda x, y: x != y, - 'lt': lambda x, y: x < y, - 'le': lambda x, y: x <= y, - 'gt': lambda x, y: x > y, - 'ge': lambda x, y: x >= y, - 'not': lambda x, y: x != y, - 'notcontains': lambda x, y: x.apply(lambda z : y not in z.split(";")), # (~maf_df["FILTER"].str.contains("not_in_panel")) - 'contains': lambda x, y: x.apply(lambda z : y in z.split(";")) + "eq": lambda x, y: x == y, + "ne": lambda x, y: x != y, + "lt": lambda x, y: x < y, + "le": lambda x, y: x <= y, + "gt": lambda x, y: x > y, + "ge": lambda x, y: x >= y, + "not": lambda x, y: x != y, + "notcontains": lambda x, y: x.apply( + lambda z: y not in z.split(";") + ), # (~maf_df["FILTER"].str.contains("not_in_panel")) + "contains": lambda x, y: x.apply(lambda z: y in z.split(";")), } # Apply filters based on criteria from the JSON file for col, criterion in filter_criteria: - if isinstance(criterion, bool): pref_len = maf_df.shape[0] maf_df = maf_df[maf_df[col] == criterion] - print(f"Applying {col}:{criterion} filter implied going from {pref_len} mutations to {maf_df.shape[0]} mutations.") + print( + f"Applying {col}:{criterion} filter implied going from {pref_len} mutations to {maf_df.shape[0]} mutations." + ) - elif ' ' in criterion: + elif " " in criterion: operator, value = criterion.split(maxsplit=1) if len(operator) == 2 and operator in operators: # 'VAF' : 'le 0.35' pref_len = maf_df.shape[0] maf_df = maf_df[operators[operator](maf_df[col], float(value))] - print(f"Applying {col}:{criterion} filter implied going from {pref_len} mutations to {maf_df.shape[0]} mutations.") + print( + f"Applying {col}:{criterion} filter implied going from {pref_len} mutations to {maf_df.shape[0]} mutations." + ) elif operator in operators: # 'FILTER' : 'notcontains n_rich', pref_len = maf_df.shape[0] maf_df = maf_df[operators[operator](maf_df[col], value)] - print(f"Applying {col}:{criterion} filter implied going from {pref_len} mutations to {maf_df.shape[0]} mutations.") + print( + f"Applying {col}:{criterion} filter implied going from {pref_len} mutations to {maf_df.shape[0]} mutations." + ) else: print(f"We have no filtering criteria defined for {col}:{criterion} filter.") - else: # 'TYPE' : 'SNV' pref_len = maf_df.shape[0] maf_df = maf_df[maf_df[col] == criterion] - print(f"Applying {col}:{criterion} filter implied going from {pref_len} mutations to {maf_df.shape[0]} mutations.") + print( + f"Applying {col}:{criterion} filter implied going from {pref_len} mutations to {maf_df.shape[0]} mutations." + ) return maf_df + +def somatic_mask(maf_df: pd.DataFrame, threshold: float) -> pd.Series: + """ + Return a boolean mask identifying somatic variants. + + Parameters + ---------- + maf_df : pd.DataFrame + MAF dataframe containing at least the columns ``VAF``, ``vd_VAF``, and + ``VAF_AM``. + threshold : float + Upper bound (inclusive) on VAF for a variant to be called somatic. + + Returns + ------- + pd.Series + Boolean series with the same index as *maf_df*; ``True`` where the + variant is somatic. + """ + return (maf_df["VAF"] <= threshold) & (maf_df["vd_VAF"] <= threshold) & (maf_df["VAF_AM"] <= threshold) + + +def germline_mask(maf_df: pd.DataFrame, threshold: float) -> pd.Series: + """ + Return a boolean mask identifying germline variants. + + Parameters + ---------- + maf_df : pd.DataFrame + MAF dataframe containing at least the columns ``VAF``, ``vd_VAF``, and + ``VAF_AM``. + threshold : float + Lower bound (exclusive) on VAF for a variant to be called germline. + + Returns + ------- + pd.Series + Boolean series with the same index as *maf_df*; ``True`` where the + variant is germline. + """ + return (maf_df["VAF"] > threshold) & (maf_df["vd_VAF"] > threshold) & (maf_df["VAF_AM"] > threshold) + + def load_filter_criteria(filters: str, somatic_filters: str) -> list[str]: """ Parse filter criteria from comma-separated strings. - + Parameters ---------- filters : str Comma-separated list of filter criteria somatic_filters : str Comma-separated list of somatic filter criteria - + Returns ------- list[str] List of filter names to apply """ # Parse comma-separated strings into lists - filter_list = [f.strip() for f in filters.split(',') if f.strip()] - somatic_filter_list = [f.strip() for f in somatic_filters.split(',') if f.strip()] - + filter_list = [f.strip() for f in filters.split(",") if f.strip()] + somatic_filter_list = [f.strip() for f in somatic_filters.split(",") if f.strip()] + # Combine both lists all_filters = filter_list + somatic_filter_list @@ -95,21 +149,22 @@ def load_filter_criteria(filters: str, somatic_filters: str) -> list[str]: LOG.info(f"Loaded {len(result)} filter criteria: {result}") return result + def expand_filter_column(maf_df: pd.DataFrame) -> pd.DataFrame: """ Expands the FILTER column by creating new columns for each unique filter. Each new column indicates if the corresponding filter is present (True/False). """ # Split FILTER column once per row and convert to set for O(1) lookup - filter_sets = maf_df["FILTER"].str.split(";").apply(lambda x: set(x) if x != [''] else set()) - + filter_sets = maf_df["FILTER"].str.split(";").apply(lambda x: set(x) if x != [""] else set()) + # Get all unique filter values (excluding empty strings) all_filters = set( - filter_val - for filter_val in maf_df["FILTER"].str.split(";").explode().unique() - if filter_val and filter_val != '' + filter_val + for filter_val in maf_df["FILTER"].str.split(";").explode().unique() + if filter_val and filter_val != "" ) - + # Ensure "not_covered" and "not_in_exons" exist required_filters = {"not_covered", "not_in_exons"} all_filters.update(required_filters) @@ -120,7 +175,10 @@ def expand_filter_column(maf_df: pd.DataFrame) -> pd.DataFrame: return maf_df -def extract_flagged_regions_bed(maf_df: pd.DataFrame, name: str, FILTERS: list[str], specification: str = "") -> pd.DataFrame | None: + +def extract_flagged_regions_bed( + maf_df: pd.DataFrame, name: str, filters: list[str], specification: str = "" +) -> pd.DataFrame | None: """ Returns a BED file with the regions discarded, including the list of filters applied to each mutation. Creates a properly formatted BED file with 0-based coordinates and half-open intervals. @@ -131,7 +189,7 @@ def extract_flagged_regions_bed(maf_df: pd.DataFrame, name: str, FILTERS: list[s Input MAF dataframe with filter columns. POS column should contain 1-based coordinates. name : str Sample name to be used in the output BED file name. - FILTERS : list[str] + filters : list[str] List of filter criteria to check for in the MAF dataframe. specification : str, optional Additional string to include in the output BED file name (e.g., "cohort-"), by default "". @@ -143,7 +201,7 @@ def extract_flagged_regions_bed(maf_df: pd.DataFrame, name: str, FILTERS: list[s Output coordinates are 0-based with half-open intervals [start, end). """ # List of filter columns you want to check for - filter_columns = [f"FILTER.{f}" for f in FILTERS if f"FILTER.{f}" in maf_df.columns] + filter_columns = [f"FILTER.{f}" for f in filters if f"FILTER.{f}" in maf_df.columns] maf_df_filters = maf_df[maf_df[filter_columns].any(axis=1)] if filter_columns else pd.DataFrame() @@ -157,24 +215,20 @@ def extract_flagged_regions_bed(maf_df: pd.DataFrame, name: str, FILTERS: list[s bed_df = maf_df_filters[["CHROM", "POS"] + filter_columns] # Transform to long format - _bed_melt = (pd.melt(bed_df, - id_vars=["CHROM", "POS"], - value_vars=filter_columns, - var_name="FILTERS") - .query("value == True") - ) + _bed_melt = pd.melt(bed_df, id_vars=["CHROM", "POS"], value_vars=filter_columns, var_name="FILTERS").query( + "value == True" + ) LOG.info("Mutations flagged: %s", _bed_melt.shape[0]) # Aggregate filters per position bed_annotated = ( - _bed_melt - .drop_duplicates() - .sort_values(by=["CHROM", "POS"]) - .groupby(["CHROM","POS"])["FILTERS"] - .agg(','.join) - .reset_index() - .rename(columns={"POS": "START"}) + _bed_melt.drop_duplicates() + .sort_values(by=["CHROM", "POS"]) + .groupby(["CHROM", "POS"])["FILTERS"] + .agg(",".join) + .reset_index() + .rename(columns={"POS": "START"}) ) # The idea is to filter depth files at these positions, so make END = START (1-based) @@ -183,6 +237,8 @@ def extract_flagged_regions_bed(maf_df: pd.DataFrame, name: str, FILTERS: list[s LOG.info("Unique regions flagged: %s", bed_annotated.shape[0]) # Write BED file without header or index - (bed_annotated[["CHROM", "START", "END", "FILTERS"]] - .to_csv(f"{name}.{specification}flagged-pos.bed", sep="\t", header=False, index=False) - ) \ No newline at end of file + ( + bed_annotated[["CHROM", "START", "END", "FILTERS"]].to_csv( + f"{name}.{specification}flagged-pos.bed", sep="\t", header=False, index=False + ) + ) diff --git a/bin/vaf_smoothing.py b/bin/vaf_smoothing.py new file mode 100755 index 00000000..e7a034dd --- /dev/null +++ b/bin/vaf_smoothing.py @@ -0,0 +1,527 @@ +#!/usr/bin/env python + + +import click +import pandas as pd +import numpy as np +import matplotlib.pyplot as plt +import seaborn as sns +from matplotlib.backends.backend_pdf import PdfPages +from read_utils import custom_na_values +from utils_plot import plots_general_config +import matplotlib as mpl + + + + +def compute_priors(mutdensity_file, samples): + """ + From simple adjusted version + This function needs to be removed eventually when the use of the + updated adjusted mutation density is functional in omega + """ + + mutdensity_df = pd.read_csv(mutdensity_file, sep = "\t", header = 0, na_values = custom_na_values) + npa_mutdensities = mutdensity_df[(mutdensity_df["MUTTYPES"] == "all_types") + & (mutdensity_df["GENE"] == "ALL_GENES") + & (mutdensity_df["REGIONS"] == "non_protein_affecting") + & (mutdensity_df["SAMPLE_ID"].isin(samples)) + ].reset_index(drop = True) + npa_mutdensities['prior'] = npa_mutdensities['N_MUTATED'] / npa_mutdensities['DEPTH'] + npa_priors = npa_mutdensities[['SAMPLE_ID', 'prior']] + + # npa_priors.to_csv(f"sample_specific_priors.tsv.gz", + # header=True, + # index=False, + # sep="\t") + return npa_priors + + + +def apply_correction_per_sample(mutations_table, priors_per_sample, depths_per_sample, + samples, + weights_list = [0, 0.2, 0.5, 0.75, 1, 1.25, 1.5, 2, 3] + ): + """ + Apply the pseudocount correction to the VAF values in the mutations table. + """ + + mutation_tables_versions = [] + + for sample_id in samples: + prior_vaf = priors_per_sample.loc[priors_per_sample['SAMPLE_ID'] == sample_id, 'prior'].values[0] + sample_mutations = mutations_table[mutations_table['SAMPLE_ID'] == sample_id][ + ['SAMPLE_ID', 'MUT_ID', 'ALT_DEPTH', 'DEPTH', 'VAF', 'ALT_DEPTH_AM', 'DEPTH_AM', 'VAF_AM'] + ].copy() + sample_depths = depths_per_sample[depths_per_sample['SAMPLE_ID'] == sample_id] + average_depth = sample_depths['avg_depth_sample'].values[0] + for weight_prop in weights_list: + weight = weight_prop * average_depth // 1 + sample_mutations[f'VAF_PSEUDO_{weight_prop}'] = vaf_pseudocount( + sample_mutations['ALT_DEPTH'], + sample_mutations['DEPTH'], + weight, + prior_vaf=prior_vaf + ) + sample_mutations[f'VAF_AM_PSEUDO_{weight_prop}'] = vaf_pseudocount( + sample_mutations['ALT_DEPTH_AM'], + sample_mutations['DEPTH_AM'], + weight, + prior_vaf=prior_vaf + ) + + mutation_tables_versions.append(sample_mutations) + + return pd.concat(mutation_tables_versions, ignore_index=True) + + +def vaf_pseudocount(alt_depth, depth, weight, prior_vaf=None): + return (alt_depth + prior_vaf * weight) / (depth + weight) + +def plot_vaf_pseudocount_curve(maf_df, samples, prior_vafs, output_pdf, suffix='', + weights_list=[0, 0.2, 0.5, 0.75, 1, 1.25, 1.5, 2, 3]): + """ + Plot VAF distribution compared to VAF_AM in a histogram. + + Parameters: + ----------- + maf_df : DataFrame + MAF dataframe containing VAF and VAF_AM columns + """ + + n_cols = 5 if len(weights_list) > 9 else 3 + n_rows = len(weights_list) // n_cols if len(weights_list) % n_cols == 0 else len(weights_list) // n_cols + 1 + for sample in samples: + fig, axes = plt.subplots(n_rows, n_cols, + figsize=(n_cols * 2.5, n_rows * 2.5), sharex=True, sharey=True) + axes = axes.flatten() + fig.suptitle(f'{sample} : VAF{suffix} with pseudocounts') + + samples_maf_df = maf_df[maf_df['SAMPLE_ID'] == sample] + for i, weight_prop in enumerate(weights_list): + axes[i].scatter(samples_maf_df[f'DEPTH{suffix}'], samples_maf_df[f'VAF{suffix}_PSEUDO_{weight_prop}'], s=3, alpha=0.1) + rho = np.corrcoef(np.log(samples_maf_df[f'DEPTH{suffix}']), np.log(samples_maf_df[f'VAF{suffix}_PSEUDO_{weight_prop}']))[0, 1] + axes[i].set_xlabel(f'DEPTH{suffix}') + axes[i].set_ylabel(f'VAF{suffix}_PSEUDO', fontsize=6) + axes[i].set_yscale('log') + axes[i].set_xscale('log') + axes[i].set_title(f"{weight_prop}\nrho={rho:.2f}, \nprop_mut_tissue={samples_maf_df[f'VAF{suffix}_PSEUDO_{weight_prop}'].sum():.2f}", fontsize=6) + plt.tight_layout() + output_pdf.savefig() + plt.close() + plt.show() + + + + +def calc_vaf_distance_summary( + df, + weights, + vaf_col="VAF", + target_col="VAF_AM", + pseudo_prefix="VAF_PSEUDO", + depth_col="DEPTH", + depth_quantile=0.1, + large_clone_quantile=0.95, +): + """ + Calculate distance between VAF_PSEUDO and target VAF column + for low-depth clones and large clones. + """ + + depth_cutoff_low = df[depth_col].quantile(depth_quantile) + + df_low_depth = df[(df[depth_col] < depth_cutoff_low) + & (df['ALT_DEPTH'] < 2)].copy() + + # define large clones except from low depth clones + large_clone_quantile_value = df[(df[depth_col] > depth_cutoff_low) + ][target_col].quantile(large_clone_quantile) + df_large = df[(df[depth_col] > depth_cutoff_low) + & (df[target_col] > large_clone_quantile_value)].copy() + + df_large_ALTDEPTH = df[(df[depth_col] > depth_cutoff_low) + # & (df[target_col.replace("VAF", "ALT_DEPTH")] > 1) + & (df["ALT_DEPTH"] > 1) + ].copy() + + # print(df[(df[depth_col] > depth_cutoff_low)]["ALT_DEPTH"].value_counts()) + # print(df[(df[depth_col] > depth_cutoff_low)]["ALT_DEPTH_AM"].value_counts()) + print(f"Low depth clones: {df_low_depth.shape[0]}") + print(f"Large clones: {df_large.shape[0]}") + print(f"Large clones ALTDEPTH: {df_large_ALTDEPTH.shape[0]}") + + # Reference: original VAF vs target VAF + ref_dist_log_low_depth_sum = np.abs(np.log10(df_low_depth[vaf_col]) - np.log10(df_low_depth[target_col])).sum() + ref_dist_low_depth_sum = np.abs(df_low_depth[vaf_col] - df_low_depth[target_col]).sum() + ref_dist_log_low_depth_mean = np.abs(np.log10(df_low_depth[vaf_col]) - np.log10(df_low_depth[target_col])).mean() + ref_dist_low_depth_mean = np.abs(df_low_depth[vaf_col] - df_low_depth[target_col]).mean() + + + ref_dist_log_large_sum = np.abs(np.log10(df_large[vaf_col]) - np.log10(df_large[target_col])).sum() + ref_dist_log_large_mean = np.abs(np.log10(df_large[vaf_col]) - np.log10(df_large[target_col])).mean() + ref_dist_large_sum = np.abs(df_large[vaf_col] - df_large[target_col]).sum() + ref_dist_large_mean = np.abs(df_large[vaf_col] - df_large[target_col]).mean() + + ref_dist_log_largeALTDEPTH_sum = np.abs(np.log10(df_large_ALTDEPTH[vaf_col]) - np.log10(df_large_ALTDEPTH[target_col])).sum() + ref_dist_log_largeALTDEPTH_mean = np.abs(np.log10(df_large_ALTDEPTH[vaf_col]) - np.log10(df_large_ALTDEPTH[target_col])).mean() + ref_dist_largeALTDEPTH_sum = np.abs(df_large_ALTDEPTH[vaf_col] - df_large_ALTDEPTH[target_col]).sum() + ref_dist_largeALTDEPTH_mean = np.abs(df_large_ALTDEPTH[vaf_col] - df_large_ALTDEPTH[target_col]).mean() + + results = [] + + for w in weights: + + col = f"{pseudo_prefix}_{w}" + + dist_log_low_depth = np.abs(np.log10(df_low_depth[col]) - np.log10(df_low_depth[target_col])) + dist_low_depth = np.abs(df_low_depth[col] - df_low_depth[target_col]) + + dist_log_large = np.abs(np.log10(df_large[col]) - np.log10(df_large[target_col])) + dist_large = np.abs(df_large[col] - df_large[target_col]) + + dist_log_largeALTDEPTH = np.abs(np.log10(df_large_ALTDEPTH[col]) - np.log10(df_large_ALTDEPTH[target_col])) + dist_largeALTDEPTH = np.abs(df_large_ALTDEPTH[col] - df_large_ALTDEPTH[target_col]) + + results.append({ + "w": w, + "sum_dist_log_low_depth": dist_log_low_depth.sum(), + "sum_dist_low_depth": dist_low_depth.sum(), + "mean_dist_log_low_depth": dist_log_low_depth.mean(), + "sum_dist_log_large_clones": dist_log_large.sum(), + "sum_dist_large_clones": dist_large.sum(), + "mean_dist_log_large_clones": dist_log_large.mean(), + "sum_dist_log_largeALTDEPTH": dist_log_largeALTDEPTH.sum(), + "sum_dist_largeALTDEPTH": dist_largeALTDEPTH.sum(), + "mean_dist_log_largeALTDEPTH": dist_log_largeALTDEPTH.mean(), + "n_total": len(df), + "n_low_depth": len(df_low_depth), + "n_large_clones": len(df_large), + "n_largeALTDEPTH": len(df_large_ALTDEPTH), + "depth_cutoff_low": depth_cutoff_low, + "large_clone_definition": large_clone_quantile_value, + "ref_dist_log_low_depth_sum": ref_dist_log_low_depth_sum, + "ref_dist_log_large_clones_sum": ref_dist_log_large_sum, + "ref_dist_low_depth_sum": ref_dist_low_depth_sum, + "ref_dist_large_clones_sum": ref_dist_large_sum, + "ref_dist_log_largeALTDEPTH_sum": ref_dist_log_largeALTDEPTH_sum, + "ref_dist_largeALTDEPTH_sum": ref_dist_largeALTDEPTH_sum, + "ref_dist_log_low_depth_mean": ref_dist_log_low_depth_mean, + "ref_dist_log_large_clones_mean": ref_dist_log_large_mean, + "ref_dist_low_depth_mean": ref_dist_low_depth_mean, + "ref_dist_large_clones_mean": ref_dist_large_mean, + "ref_dist_log_largeALTDEPTH_mean": ref_dist_log_largeALTDEPTH_mean, + "ref_dist_largeALTDEPTH_mean": ref_dist_largeALTDEPTH_mean + }) + + return pd.DataFrame(results) + + + +def plot_vaf_distance_summary( + dist_df, + use_mean=False, + figsize=(7, 4), + log_dist=False, + log_y=False, + sample_name = "sample" +): + """ + Plot pseudocount distance curves and reference VAF vs VAF_AM distances. + """ + + if use_mean: + if log_dist: + ref_low = dist_df["ref_dist_log_low_depth_mean"].iloc[0] + ref_large = dist_df["ref_dist_log_large_clones_mean"].iloc[0] + ref_large_alt = dist_df["ref_dist_log_largeALTDEPTH_mean"].iloc[0] + y_low = "mean_dist_log_low_depth" + y_large = "mean_dist_log_large_clones" + y_large_alt = "mean_dist_log_largeALTDEPTH" + ylabel = "Mean log-distance" + else: + ref_low = dist_df["ref_dist_low_depth_mean"].iloc[0] + ref_large = dist_df["ref_dist_large_clones_mean"].iloc[0] + ref_large_alt = dist_df["ref_dist_largeALTDEPTH_mean"].iloc[0] + y_low = "mean_dist_low_depth" + y_large = "mean_dist_large_clones" + y_large_alt = "mean_dist_log_largeALTDEPTH" + ylabel = "Mean distance" + else: + if log_dist: + ref_low = dist_df["ref_dist_log_low_depth_sum"].iloc[0] + ref_large = dist_df["ref_dist_log_large_clones_sum"].iloc[0] + ref_large_alt = dist_df["ref_dist_log_largeALTDEPTH_sum"].iloc[0] + y_low = "sum_dist_log_low_depth" + y_large = "sum_dist_log_large_clones" + y_large_alt = "sum_dist_log_largeALTDEPTH" + ylabel = "Sum log distance" + else: + ref_low = dist_df["ref_dist_low_depth_sum"].iloc[0] + ref_large = dist_df["ref_dist_large_clones_sum"].iloc[0] + ref_large_alt = dist_df["ref_dist_largeALTDEPTH_sum"].iloc[0] + y_low = "sum_dist_low_depth" + y_large = "sum_dist_large_clones" + y_large_alt = "sum_dist_largeALTDEPTH" + ylabel = "Sum distance" + + + + fig, ax = plt.subplots(figsize=figsize) + + ax.plot( + dist_df["w"], + dist_df[y_low], + "-o", + lw=2, + ms=5, + label=f"Low-depth clones (n={dist_df['n_low_depth'].iloc[0]})" + ) + + ax.plot( + dist_df["w"], + dist_df[y_large], + "-o", + lw=2, + ms=5, + label=f"Large clones (n={dist_df['n_large_clones'].iloc[0]})" + ) + + if dist_df['n_largeALTDEPTH'].iloc[0] > 0: + ax.plot( + dist_df["w"], + dist_df[y_large_alt], + "-o", + lw=2, + ms=5, + label=f"Large clones ALTDEPTH (n={dist_df['n_largeALTDEPTH'].iloc[0]})" + ) + + if not use_mean: + ax.axhline( + ref_low, + color="C0", + linestyle="--", + linewidth=2, + label="VAF vs VAF_AM (low-depth)" + ) + + ax.axhline( + ref_large, + color="C1", + linestyle="--", + linewidth=2, + label="VAF vs VAF_AM (large)" + ) + + if dist_df['n_largeALTDEPTH'].iloc[0] > 0: + ax.axhline( + ref_large_alt, + color="C2", + linestyle="--", + linewidth=2, + label="VAF vs VAF_AM (largeALTDEPTH)" + ) + + ax.set_xlabel("Pseudocount weight (w)") + ax.set_ylabel(ylabel) + + if log_y: + ax.set_yscale("log") + + plt.title(f"{sample_name} (n={dist_df['n_total'].iloc[0]})\nlow_depth cutoff: {dist_df['depth_cutoff_low'].iloc[0]:.2f}, large clone definition: {dist_df['large_clone_definition'].iloc[0]:.2e}") + ax.legend(frameon=False) + sns.despine() + plt.tight_layout() + + return fig, ax + + + +def select_vafpseudo_per_sample(selected_weights_df, mutations_table): + """ + Select the best VAF_PSEUDO column for each sample based on the selected weights. + """ + + selected_vafpseudo_columns = [] + columns_to_keep = ['SAMPLE_ID', 'MUT_ID', 'ALT_DEPTH', 'DEPTH', 'VAF', 'ALT_DEPTH_AM', 'DEPTH_AM', 'VAF_AM', 'prior', 'avg_depth_sample'] + for _, row in selected_weights_df.iterrows(): + sample_id = row['SAMPLE_ID'] + selected_weight = row['selected_weight'] + selected_column = f'VAF_PSEUDO_{selected_weight}' + selected_vafpseudo_columns.append((sample_id, selected_column, selected_weight)) + + # Create a new DataFrame with the selected VAF_PSEUDO columns + selected_vafpseudo_df = pd.DataFrame(selected_vafpseudo_columns, columns=['SAMPLE_ID', 'selected_column', 'selected_weight']) + + # Merge with the original mutations table to get the selected VAF_PSEUDO values + merged_df = mutations_table.merge(selected_vafpseudo_df, on='SAMPLE_ID', how='left') + + # Create a new column for the final selected VAF_PSEUDO values + merged_df['selected_VAF_PSEUDO'] = merged_df.apply(lambda x: x[x['selected_column']], axis=1) + merged_df["artificial_depth"] = merged_df["avg_depth_sample"] * merged_df["selected_weight"] + merged_df = merged_df[columns_to_keep + ['selected_weight', 'artificial_depth', 'selected_VAF_PSEUDO']] + + # Save the final DataFrame with selected VAF_PSEUDO values + merged_df.to_csv("mutations_with_final_selected_VAF_PSEUDO.tsv.gz", sep="\t", index=False) + + corrections_summary_df = merged_df[['SAMPLE_ID', 'prior', 'avg_depth_sample', 'selected_weight', 'artificial_depth']].drop_duplicates() + + # Save the final DataFrame with a summary of the corrections applied per sample + corrections_summary_df.to_csv("corrections_per_sample_summary.tsv.gz", sep="\t", index=False) + + +@click.command() +@click.option('--mutdensities', type=click.Path(exists=True), help='Input mutation density file') +@click.option('--mutations', type=click.Path(), help='Mutations file') +@click.option('--depth-sample', type=click.Path(), help='Depth per sample file') + +def main(mutdensities, mutations, depth_sample): + """ + Main function to execute the VAF smoothing pipeline. + """ + weights = [x.round(1) for x in np.arange(0, 1.6, 0.1)] + + click.echo("Loading data...") + mutations_table = pd.read_csv(mutations, sep = "\t", header = 0, na_values = custom_na_values) + mutations_0_vaf = mutations_table[~(mutations_table["VAF"] > 0) + | ~(mutations_table["VAF_AM"] > 0) + ] + if not mutations_0_vaf.empty: + click.echo(f"Warning: {mutations_0_vaf.shape[0]} mutations have VAF or VAF_AM <= 0 and will be excluded from the analysis.") + click.echo("These mutations are:") + click.echo(mutations_0_vaf[['SAMPLE_ID', 'MUT_ID', 'VAF', 'VAF_AM']]) + mutations_table = mutations_table[(mutations_table["VAF"] > 0) + & (mutations_table["VAF_AM"] > 0) + ].reset_index(drop = True) + + depths_per_sample = pd.read_csv(depth_sample, sep = "\t", header = 0, na_values = custom_na_values) + + samples_list = sorted(mutations_table['SAMPLE_ID'].unique()) + + click.echo("Computing priors...") + priors_per_sample = compute_priors(mutdensities, samples_list) + + click.echo("Applying corrections...") + all_mutations_table = apply_correction_per_sample(mutations_table, priors_per_sample, depths_per_sample, samples_list, weights_list=weights) + priors_n_depth = priors_per_sample.merge(depths_per_sample, on='SAMPLE_ID', how='left') + + mutations_plus_info = all_mutations_table.merge(priors_n_depth, on='SAMPLE_ID', how='left') + mutations_plus_info = mutations_plus_info[['SAMPLE_ID', 'MUT_ID', 'prior', 'avg_depth_sample'] + [x for x in all_mutations_table.columns if x not in ['SAMPLE_ID', 'MUT_ID', 'prior', 'avg_depth_sample']]] + mutations_plus_info.to_csv(f"mutations_with_smoothed_VAF.tsv.gz", + header=True, + index=False, + sep="\t") + + click.echo("Plotting VAF pseudocount curves...") + with PdfPages("vaf_pseudocounts_curves.pdf") as pdf: + plt.figure(figsize=(8, 6)) + plt.hist(priors_per_sample['prior'], bins=10) + plt.title('Distribution of sample-specific priors') + plt.xlabel('Prior VAF') + plt.ylabel('Frequency') + plt.tight_layout() + pdf.savefig() + plt.close() + + plot_vaf_pseudocount_curve(all_mutations_table, samples_list, priors_per_sample, pdf, suffix='', weights_list=weights) + plot_vaf_pseudocount_curve(all_mutations_table, samples_list, priors_per_sample, pdf, suffix='_AM', weights_list=weights) + + click.echo("Plotting pseudocount curves weight comparison...") + with PdfPages("vaf_pseudocount_weights_comparison.pdf") as pdf: + + selected_weights_dict = {} + try : + for sample in samples_list: + sample_mutations = all_mutations_table[all_mutations_table['SAMPLE_ID'] == sample] + + dist_df = calc_vaf_distance_summary( + sample_mutations, + weights=weights, + vaf_col="VAF", + target_col="VAF_AM", + pseudo_prefix="VAF_PSEUDO", + depth_col="DEPTH", + depth_quantile=0.2, + large_clone_quantile=0.95, + ) + + #get minimum distance weight for low depth clones + min_dist_low_depth = dist_df.loc[dist_df['sum_dist_low_depth'].idxmin(), 'w'] + selected_weights_dict[sample] = min_dist_low_depth + + print(f"Sample {sample}: Minimum distance weight for low depth clones: {min_dist_low_depth}") + + fig, ax = plot_vaf_distance_summary( + dist_df, + log_dist=False, + use_mean=False, + log_y=False, + sample_name = sample + ) + pdf.savefig() + plt.close() + + fig, ax = plot_vaf_distance_summary( + dist_df, + log_dist=True, + use_mean=True, + log_y=False, + sample_name = f'{sample}_mean' + ) + pdf.savefig() + plt.close() + + + dist_df = calc_vaf_distance_summary( + all_mutations_table, + weights=weights, + vaf_col="VAF", + target_col="VAF_AM", + pseudo_prefix="VAF_PSEUDO", + depth_col="DEPTH", + depth_quantile=0.2, + large_clone_quantile=0.95, + ) + #get minimum distance weight for low depth clones + min_dist_low_depth = dist_df.loc[dist_df['sum_dist_low_depth'].idxmin(), 'w'] + selected_weights_dict['all_samples'] = min_dist_low_depth + + fig, ax = plot_vaf_distance_summary( + dist_df, + log_dist=False, + use_mean=False, + log_y=False, + sample_name = 'all_samples' + ) + pdf.savefig() + plt.close() + + fig, ax = plot_vaf_distance_summary( + dist_df, + log_dist=True, + use_mean=True, + log_y=False, + sample_name = 'all_samples_mean' + ) + pdf.savefig() + plt.close() + + + # store selected weights in a tsv file + selected_weights_df = pd.DataFrame(list(selected_weights_dict.items()), columns=['SAMPLE_ID', 'selected_weight']) + # selected_weights_df.to_csv("selected_weights_per_sample.tsv.gz", sep="\t", index=False) + + except Exception as e: + click.echo(f"Error in plotting VAF distance summary: {e}") + + click.echo("Storing summary of priors and weights used.") + select_vafpseudo_per_sample(selected_weights_df, mutations_plus_info) + + click.echo("VAF smoothing pipeline completed successfully.") + + + +if __name__ == '__main__': + main() + diff --git a/bin/vcf2maf.py b/bin/vcf2maf.py index f4528d03..beecec66 100755 --- a/bin/vcf2maf.py +++ b/bin/vcf2maf.py @@ -18,7 +18,8 @@ def read_from_vardict_VCF_all(sample, name, columns_to_keep = ['CHROM', 'POS', 'REF', 'ALT', 'DEPTH', 'REF_DEPTH', 'ALT_DEPTH', 'VAF', - 'vd_DEPTH', 'vd_REF_DEPTH', 'vd_ALT_DEPTH']): + 'vd_DEPTH', 'vd_REF_DEPTH', 'vd_ALT_DEPTH'], + vaf_distortion_threshold = 3): """ Read VCF file coming from Vardict2 Note that the file can only contain one sample @@ -33,6 +34,7 @@ def read_from_vardict_VCF_all(sample, Optional arguments: columns_to_keep = ['CHROM', 'POS', 'REF', 'ALT', 'DEPTH', 'ALT_DEPTH', 'VAF'], # add 'PID' for phased mutations + vaf_distortion_threshold = 3 """ print(f"Processing {sample}") @@ -180,7 +182,7 @@ def read_from_vardict_VCF_all(sample, dat_full["VAF_distortion"] = dat_full["VAF_AM"] / dat_full["VAF"] dat_full["VAF_distortion_sq"] = np.log10(dat_full["VAF"]) / np.log10(dat_full["VAF_AM"]) - dat_full["VAF_distorted_expanded"] = dat_full["VAF_distortion"] > 3 + dat_full["VAF_distorted_expanded"] = dat_full["VAF_distortion"] > vaf_distortion_threshold dat_full["VAF_distorted_expanded_sq"] = dat_full["VAF"] < ( dat_full["VAF_AM"] ** 1.5 ) dat_full["VAF_distorted_expanded"] = dat_full["VAF_distorted_expanded"].fillna(True) @@ -331,7 +333,8 @@ def update_indel_info(df): @click.option('--sampleid', type=str, required=True, help='Sample ID.') @click.option('--level', type=str, default = 'med', help='Level of confidence of the mutations.') @click.option('--annotation_file', type=click.Path(exists=True), required=True, help='Path to the annotation file.') -def main(vcf, sampleid, level, annotation_file): +@click.option('--vaf_distortion_threshold', type=float, default=3, show_default=True, help='Threshold for defining VAF distortion outliers.') +def main(vcf, sampleid, level, annotation_file, vaf_distortion_threshold): keep_all_columns = [ "CHROM", "POS", "REF", "ALT", "FILTER", "INFO", "FORMAT", "SAMPLE", "DEPTH", "ALT_DEPTH", "REF_DEPTH", "VAF", @@ -353,7 +356,8 @@ def main(vcf, sampleid, level, annotation_file): sample_muts = read_from_vardict_VCF_all( sampleid, vcf, - columns_to_keep=keep_all_columns + columns_to_keep=keep_all_columns, + vaf_distortion_threshold=vaf_distortion_threshold ) sample_muts["SAMPLE_ID"] = sampleid @@ -376,4 +380,3 @@ def main(vcf, sampleid, level, annotation_file): if __name__ == '__main__': main() - diff --git a/conf/base.config b/conf/base.config index c5a2d9d8..fbacfc64 100644 --- a/conf/base.config +++ b/conf/base.config @@ -44,7 +44,7 @@ process { memory = { 6.GB * task.attempt } } - withName: 'BBGTOOLS:DEEPCSA:CREATEPANELS:CUSTOMPROCESSING.*' { + withName: 'CUSTOMPROCESSING.*' { memory = { 10.GB * task.attempt } } @@ -89,39 +89,18 @@ process { memory = { 1.GB * task.attempt } } - withName: 'BBGTOOLS:DEEPCSA:OMEGANONPROT.*:SUBSETPANEL*' { - cpus = { 2 * task.attempt } - memory = { 4.GB * task.attempt } - } - withName: 'BBGTOOLS:DEEPCSA:CREATEPANELS:POSTPROCESSVEPPANEL*' { cpus = { 2 * task.attempt } memory = { 8.GB * task.attempt } } - withName: 'BBGTOOLS:DEEPCSA:MUTRATE.*:MUTRATE*' { - memory = { 8.GB * task.attempt } - } - - withName: 'BBGTOOLS:DEEPCSA:OMEGA.*:(PREPROCESSING|ESTIMATOR).*' { - memory = { 4.GB * task.attempt } - } - - withName: 'BBGTOOLS:DEEPCSA:OMEGA.*:ESTIMATOR_DNDSCV*' { - memory = { 2.GB * task.attempt } - } - - withName: 'BBGTOOLS:DEEPCSA:OMEGA.*:ESTIMATOR_DRIVERMUTRATE*' { - memory = { 1.GB * task.attempt } - } - - // Catch-all for other OMEGA estimators not explicitly configured - // Keep conservative settings until we have usage data - withName: 'BBGTOOLS:DEEPCSA:OMEGA.*:ESTIMATOR_(?!DNDSCV|DRIVERMUTRATE).*' { + // BBGTOOLS:DEEPCSA:OMEGA:ESTIMATOR and BBGTOOLS:DEEPCSA:OMEGA:PREPROCESSING + withName: 'PREPROCESSING|ESTIMATOR' { memory = { 4.GB * task.attempt } } - withName: 'BBGTOOLS:DEEPCSA:SIGNATURESNONPROT:SIGPROFILERASSIGNMENT*' { + // BBGTOOLS:DEEPCSA:SIGNATURESNONPROT:SIGPROFILERASSIGNMENT + withName: 'SIGPROFILERASSIGNMENT' { memory = { 2.GB * task.attempt } } @@ -135,7 +114,7 @@ process { memory = { 4.GB * task.attempt } } - withName: 'BBGTOOLS:DEEPCSA:MUTATEDCELLSVAF:MUTATEDGENOMESFROMVAFAM' { + withName: 'MUTATEDGENOMESFROMVAFAM' { errorStrategy = 'ignore' maxRetries = 1 } @@ -153,7 +132,7 @@ process { memory = { 30.GB * task.attempt } } - withName: 'BBGTOOLS:DEEPCSA:MUT_PREPROCESSING:FILTERNANOSEQSNP' { + withName: 'FILTERNANOSEQSNP' { memory = { 8.GB * task.attempt } } diff --git a/conf/exome.config b/conf/exome.config index ce7dbac5..a74e952a 100644 --- a/conf/exome.config +++ b/conf/exome.config @@ -93,17 +93,14 @@ process { - withName: '(BBGTOOLS:DEEPCSA:MUT_PREPROCESSING:SOMATICMUTATIONS*|BBGTOOLS:DEEPCSA:OMEGANONPROT.*:SUBSETPANEL*)' { + withName: 'BBGTOOLS:DEEPCSA:MUT_PREPROCESSING:SOMATICMUTATIONS*' { cpus = { 2 * task.attempt } memory = { 4.GB * task.attempt } time = { 360.min * task.attempt } } - - withName: 'BBGTOOLS:DEEPCSA:MUTRATE.*:MUTRATE*' { - memory = { 8.GB * task.attempt } - } - - withName: '(BBGTOOLS:DEEPCSA:OMEGA.*:(PREPROCESSING|ESTIMATOR).*|BBGTOOLS:DEEPCSA:MULTIQC)' { + + // BBGTOOLS:DEEPCSA:OMEGA:ESTIMATOR and BBGTOOLS:DEEPCSA:OMEGA:PREPROCESSING + withName: 'PREPROCESSING|ESTIMATOR|BBGTOOLS:DEEPCSA:MULTIQC' { memory = { 4.GB * task.attempt } } diff --git a/conf/general_files_IRB.config b/conf/general_files_IRB.config index ad1fb044..7d973664 100644 --- a/conf/general_files_IRB.config +++ b/conf/general_files_IRB.config @@ -9,17 +9,17 @@ params { cosmic_ref_signatures = "/data/bbg/datasets/COSMIC_signatures/COSMIC_v3.5_SBS_GRCh38.txt" indel_ref_signatures = "/data/bbg/datasets/COSMIC_signatures/COSMIC_v3.5_ID_GRCh37.txt" - wgs_trinuc_counts = "/data/bbg/datasets/transfer/ferriol_deepcsa/trinucleotide_counts/trinuc_counts.homo_sapiens.tsv" + wgs_trinuc_counts = "/data/bbg/datasets/pipelines/deepCSA/reference_datasets/trinuc_counts.homo_sapiens.tsv" cadd_scores = "/data/bbg/datasets/CADD/v1.7/hg38/whole_genome_SNVs.tsv.gz" cadd_scores_ind = "/data/bbg/datasets/CADD/v1.7/hg38/whole_genome_SNVs.tsv.gz.tbi" - dnds_ref_transcripts = "/data/bbg/projects/prominent/analysis/dNdScv/data/reference_files/RefCDS_human_latest_intogen.rda" - dnds_covariates = "/data/bbg/projects/prominent/analysis/dNdScv/data/reference_files/covariates_hg19_hg38_epigenome_pcawg.rda" + dnds_biomart_ref = "/data/bbg/datasets/pipelines/deepCSA/reference_datasets/dNdScv_reference/homo_sapiens.v111.MANE.biomart.tsv" + dnds_covariates = "/data/bbg/datasets/pipelines/deepCSA/reference_datasets/dNdScv_covariates/covariates_hg19_hg38_epigenome_pcawg.rda" // oncodrive3d - datasets3d = "/data/bbg/nobackup/scratch/oncodrive3d/datasets_240506" - annotations3d = "/data/bbg/nobackup/scratch/oncodrive3d/annotations_240506" - domains_file = "/data/bbg/projects/prominent/dev/internal_development/domains/o3d_pfam_parsed.tsv" + datasets3d = "/data/bbg/datasets/oncodrive3d/datasets/datasets-260603" + annotations3d = "/data/bbg/datasets/oncodrive3d/annotations/annotations-260603" + domains_file = "/data/bbg/projects/prominent/dev/internal_development/domains/bbgdomains.v2025.annotated_deepCSA.tsv" // Nanoseq masks nanoseq_snp = "/data/bbg/datasets/genomes/GRCh38/masking_bedfiles/nanoseq_masks/SNP_GRCh38.wgns.bed.gz" diff --git a/conf/modules.config b/conf/modules.config index 9d8c16b3..1c8f3bee 100644 --- a/conf/modules.config +++ b/conf/modules.config @@ -41,6 +41,7 @@ process { withName: COMPUTEDEPTHS { ext.restrict_panel = params.use_custom_bedfile ext.minimum_depth = params.use_custom_minimum_depth + ext.remove_chrM = params.remove_chrM ext.args = "-H" publishDir = [ enabled: false @@ -83,6 +84,7 @@ process { ext.ensembl_release = params.vep_cache_version ?: 111 ext.ensembl_species = params.vep_species ?: 'homo_sapiens' ext.ensembl_genome = params.vep_genome ?: 'GRCh38' + ext.gff3_provided = params.gff3_file ? true : false } withName: QUERYDEPTHS { @@ -131,22 +133,11 @@ process { ] } - withName: SUBSETPILEUP { - ext.prefix = { ".subset_pileup" } - ext.args = '' - ext.args2 = '-s 1 -b 2 -e 2' - ext.args3 = '-h' - ext.extension = 'tsv' - ext.header = 'pile' - publishDir = [ - enabled: false - ] - } - withName: 'BBGTOOLS:DEEPCSA:.*ALL:.*' { ext.prefix = { ".all" } } + // Gated by params.profileintrons; "no matching selector" warning is expected when the feature is off withName: 'BBGTOOLS:DEEPCSA:.*INTRONS:.*' { ext.prefix = { ".introns" } } @@ -168,7 +159,7 @@ process { // } withName: VCF2MAF { - ext.args = "--level ${params.confidence_level}" + ext.args = "--level ${params.confidence_level} --vaf_distortion_threshold ${params.vaf_distortion_threshold}" } withName: COMPUTEPROFILE { @@ -356,6 +347,10 @@ process { ext.prop_samples_nrich = params.prop_samples_nrich } + withName: CONTAMINATION { + ext.germline_threshold = params.germline_threshold + } + withName: "TABLE2GROUP" { ext.unique_identifier = params.features_unique_identifier ext.feature_groups = params.features_groups_list @@ -414,6 +409,7 @@ process { ] } + // Gated by params.profilenonprot; "no matching selector" warning is expected when the feature is off withName: '.*NONPROT:SUBSETMUTPROFILE' { ext.filters = { [ @@ -490,7 +486,7 @@ process { ext.output_fmt = { [ '"header": false', - '"columns": ["SAMPLE_ID", "CHROM_ensembl", "POS", "REF", "ALT"]', + '"columns": ["SAMPLE_ID", "CHROM", "POS", "REF", "ALT"]', ].join(',\t').trim() } publishDir = [ @@ -530,6 +526,10 @@ process { ext.recode_list = "${params.mutepi_genes_to_recode}" } + withName: PLOTMETRICSVSDEPTHQC { + ext.all_adjusted_mutdensities = params.omega | (params.profileall && params.mutationdensity) + ext.all_omegas_globalloc = (params.omega && params.omega_globalloc) + } withName: SUBSETONCODRIVECLUSTL { ext.filters = { "" } @@ -564,13 +564,6 @@ process { ] } - withName: READSPOSBED { - ext.tool = "readsxposition" - publishDir = [ - enabled: false - ] - } - withName: ONCODRIVECLUSTL { ext.args = "-sim region_restricted \ -kmer 3 \ @@ -645,7 +638,7 @@ process { } // BBGRegressions output configuration - withName: '.*:REGRESSIONS.*:EDITCONFIG' { + withName: 'EDITCONFIG' { publishDir = [ [ mode: params.publish_dir_mode, @@ -656,7 +649,8 @@ process { ] } - withName: '.*:REGRESSIONS.*:CREATE_INPUT' { + // REGRESSIONS:CREATE_INPUT + withName: 'CREATE_INPUT' { publishDir = [ [ mode: params.publish_dir_mode, @@ -667,7 +661,8 @@ process { ] } - withName: '.*:REGRESSIONS.*:MODELS' { + // REGRESSIONS:MODELS + withName: 'MODELS' { publishDir = [ [ mode: params.publish_dir_mode, @@ -677,7 +672,8 @@ process { ] } - withName: '.*:REGRESSIONS.*:PLOT' { + // REGRESSIONS:PLOT + withName: 'PLOT' { publishDir = [ [ mode: params.publish_dir_mode, diff --git a/conf/results_outputs.config b/conf/results_outputs.config index 7e76515c..a821bd13 100644 --- a/conf/results_outputs.config +++ b/conf/results_outputs.config @@ -35,14 +35,14 @@ process { } withName: PLOTMAF{ publishDir = [ - path: { "${params.outdir}/plots/mutations_summary_all" }, + path: { "${params.outdir}/plots/mutations_summary" }, mode: params.publish_dir_mode, pattern: '**{tsv,pdf,png}', ] } withName: PLOTSOMATICMAF{ publishDir = [ - path: { "${params.outdir}/plots/mutations_summary_somatic" }, + path: { "${params.outdir}/plots/mutations_summary" }, mode: params.publish_dir_mode, pattern: '**{tsv,pdf,png}', ] @@ -96,6 +96,13 @@ process { pattern: '**{tsv,pdf,png}', ] } + withName: VAFSMOOTHING{ + publishDir = [ + path: { "${params.outdir}/qc/vafsmoothing" }, + mode: params.publish_dir_mode, + pattern: '**{tsv.gz,pdf}', + ] + } withName: PLOTMUTATIONSPECIFIC{ publishDir = [ path: { "${params.outdir}/qc/mutationspecific" }, @@ -110,6 +117,21 @@ process { pattern: '**{tsv,csv,pdf,png}', ] } + withName: PLOTMETRICSVSDEPTHQC{ + publishDir = [ + path: { "${params.outdir}/qc/metrics_vs_depth" }, + mode: params.publish_dir_mode, + pattern: '**{tsv,pdf,png}', + ] + } + + withName: PLOTOMEGAVSGLOBALVSDNDSCV { + publishDir = [ + path: { "${params.outdir}/qc/omegaqc/compare_dnds_values" }, + mode: params.publish_dir_mode, + pattern: '**{tsv,pdf,png}', + ] + } withName: APPLYOMEGAQC { publishDir = [ @@ -119,7 +141,7 @@ process { pattern: 'omega.flagged_annotated.tsv' ], [ - path: { "${params.outdir}/qc/omega_flagged" }, + path: { "${params.outdir}/qc/omegaqc/omega_flagged" }, mode: params.publish_dir_mode, pattern: '*.{tsv,pdf,png}', saveAs: { filename -> @@ -153,7 +175,7 @@ process { ] } - withName: 'SORTPANELRICH|SORTPANELRICHALL' { + withName: 'SORTPANELRICH|SORTPANELCOMPACT' { publishDir = [ path: { "${params.outdir}/regions/annotations" }, mode: params.publish_dir_mode, @@ -185,7 +207,7 @@ process { withName: SITESFROMPOSITIONS { publishDir = [ - path: { "${params.outdir}/processing_files/sitesfrompositions" }, + path: { "${params.outdir}/processing_files/all_possible_sites" }, mode: params.publish_dir_mode, pattern: '**{tsv}', ] @@ -215,11 +237,27 @@ process { withName: COMPUTEMATRIX { publishDir = [ - path: { "${params.outdir}/processing_files/computematrix" }, + path: { "${params.outdir}/processing_files/mutations_matrix/per_sample" }, mode: params.publish_dir_mode, pattern: '**{tsv,per_sample,sigprofiler}', ] } + + withName: DNDSPROXY { + publishDir = [ + path: { "${params.outdir}/selection/dndsproxy" }, + mode: params.publish_dir_mode, + pattern: '**{tsv,log}', + ] + } + + withName: MATRIXCONCATWGS { + publishDir = [ + path: { "${params.outdir}/processing_files/mutations_matrix" }, + mode: params.publish_dir_mode, + pattern: '**{tsv,csv}', + ] + } withName: COMPUTETRINUC { publishDir = [ @@ -300,6 +338,47 @@ process { pattern: '**{tsv.gz}', ] } + withName: 'BBGTOOLS:DEEPCSA:OMEGA:HOTSPOTSSELECTION' { + publishDir = [ + path: { "${params.outdir}/selection/sitecomparison/hotspot_selection" }, + mode: params.publish_dir_mode, + pattern: '**{tsv.gz}', + ] + } + withName: 'BBGTOOLS:DEEPCSA:OMEGAMULTI:HOTSPOTSSELECTION' { + publishDir = [ + path: { "${params.outdir}/selection/sitecomparison/hotspot_selection/multi" }, + mode: params.publish_dir_mode, + pattern: '**{tsv.gz}', + ] + } + + withName: 'DNDSRUN' { + publishDir = [ + [ + path: { "${params.outdir}/selection/dndscv/cv" }, + mode: params.publish_dir_mode, + pattern: '**{.cv.tsv}', + ], + [ + path: { "${params.outdir}/selection/dndscv/persample" }, + mode: params.publish_dir_mode, + pattern: '**{.globaldnds.tsv}', + ], + [ + path: { "${params.outdir}/selection/dndscv/local" }, + mode: params.publish_dir_mode, + pattern: '**{.loc.tsv}', + ] + ] + } + withName: 'BIOMARTPANEL4REFCDS' { + publishDir = [ + path: { "${params.outdir}/regions/dndscv" }, + mode: params.publish_dir_mode, + pattern: '**{tsv,pdf}', + ] + } withName: 'MUTDENSITY.*' { publishDir = [ path: { "${params.outdir}/mutdensity/individual_vals" }, @@ -307,6 +386,13 @@ process { pattern: '**{tsv}', ] } + withName: 'WGSCALEDMUTDENSITY' { + publishDir = [ + path: { "${params.outdir}/mutdensity/individual_vals_wgs" }, + mode: params.publish_dir_mode, + pattern: '**{tsv}', + ] + } withName: 'MUTDENSITYADJ' { publishDir = [ path: { "${params.outdir}/mutdensity_adjusted/individual_vals" }, diff --git a/conf/tmp_quick_fixes.config b/conf/tmp_quick_fixes.config index 4a473f0b..84b75089 100644 --- a/conf/tmp_quick_fixes.config +++ b/conf/tmp_quick_fixes.config @@ -15,11 +15,8 @@ process { errorStrategy = 'ignore' } - withName: 'BBGTOOLS:DEEPCSA:MUTATEDCELLSVAF:MUTATEDGENOMESFROMVAFAM' { - errorStrategy = 'ignore' - maxRetries = 1 - } + // Gated by params.expected_mutated_cells; "no matching selector" warning is expected when the feature is off withName: "BBGTOOLS:DEEPCSA:EXPECTEDMUTATEDCELLS:.*" { errorStrategy = 'ignore' maxRetries = 1 diff --git a/conf/tools/mutdensity.config b/conf/tools/mutdensity.config index 05f5501d..4a7f31b0 100644 --- a/conf/tools/mutdensity.config +++ b/conf/tools/mutdensity.config @@ -94,4 +94,12 @@ process { withName: SYNMUTREADSDENSITY { ext.mode = 'mutated_reads' } + + withName: UPDSYNMUTDENSITY { + ext.mode = 'new' + } + withName: UPDSYNMUTREADSDENSITY { + ext.mode = 'new' + } + } diff --git a/conf/tools/omega.config b/conf/tools/omega.config index 0d28c498..9f8cee7b 100644 --- a/conf/tools/omega.config +++ b/conf/tools/omega.config @@ -14,6 +14,7 @@ process { withName: 'SITECOMPARISON.*' { ext.size = params.site_comparison_grouping + ext.genes_subset = params.selected_genes } @@ -72,6 +73,7 @@ process { ext.prefix = { ".multi" } } + // Gated by params.profilenonprot && params.positive_selection_non_protein_affecting && params.omega_multi; "no matching selector" warning is expected when the feature is off withName: 'BBGTOOLS:DEEPCSA:OMEGANONPROTMULTI:.*' { ext.prefix = { ".non_prot_aff.multi" } } @@ -85,6 +87,7 @@ process { ext.prefix = { ".multi.global_loc" } } + // Gated by params.omega_globalloc (+ params.profilenonprot && params.positive_selection_non_protein_affecting); "no matching selector" warning is expected when the feature is off withName: 'BBGTOOLS:DEEPCSA:OMEGANONPROT:.*GLOBALLOC' { ext.prefix = { ".non_prot_aff.global_loc" } } @@ -105,6 +108,7 @@ process { ] } + // BBGTOOLS:DEEPCSA:OMEGA:ESTIMATOR withName: ESTIMATOR { ext.option = 'mle' ext.args = "" @@ -128,6 +132,16 @@ process { ] } + withName: OMEGAMULTIPLETEST { + publishDir = [ + [ + mode: params.publish_dir_mode, + path: { "${params.outdir}/selection/omega" }, + pattern: "*{tsv}", + ] + ] + } + withName: PREPROCESSINGGLOBALLOC { ext.args = "" @@ -164,8 +178,18 @@ process { ] } + withName: OMEGAMULTIPLETESTGLOBALLOC { + publishDir = [ + [ + mode: params.publish_dir_mode, + path: { "${params.outdir}/selection/omegagloballoc" }, + pattern: "*{tsv}", + ] + ] + } - withName: 'PREPROCESSING.*|ESTIMATOR.*' { + // BBGTOOLS:DEEPCSA:OMEGA:ESTIMATOR and BBGTOOLS:DEEPCSA:OMEGA:PREPROCESSING + withName: 'PREPROCESSING|ESTIMATOR' { ext.assembly = params.vep_genome == 'GRCh38' ? 'hg38' : params.vep_genome == 'GRCm39' diff --git a/conf/tools/panels.config b/conf/tools/panels.config index bd681691..16681a58 100644 --- a/conf/tools/panels.config +++ b/conf/tools/panels.config @@ -47,7 +47,7 @@ process { withName: 'BBGTOOLS:DEEPCSA:CREATEPANELS:CREATECONSENSUSPANELS.*' { ext.args = "--consensus_min_depth ${params.consensus_panel_min_depth} \ --compliance_threshold ${params.consensus_compliance} " - ext.genes_subset = "${params.selected_genes}" + ext.genes_subset = params.selected_genes publishDir = [ [ mode: params.publish_dir_mode, @@ -57,67 +57,7 @@ process { ] } - withName: 'BBGTOOLS:DEEPCSA:CREATEPANELS:CREATESAMPLEPANELSALL' { - publishDir = [ - [ - mode: params.publish_dir_mode, - path: { "${params.outdir}/regions/samplepanels/createsamplepanelsall" }, - pattern: "*{tsv,bed}", - ] - ] - } - - withName: 'BBGTOOLS:DEEPCSA:CREATEPANELS:CREATESAMPLEPANELSPROTAFFECT' { - publishDir = [ - [ - mode: params.publish_dir_mode, - path: { "${params.outdir}/regions/samplepanels/createsamplepanelsprotaffect" }, - pattern: "*{tsv,bed}", - ] - ] - } - - withName: 'BBGTOOLS:DEEPCSA:CREATEPANELS:CREATESAMPLEPANELSNONPROTAFFECT' { - publishDir = [ - [ - mode: params.publish_dir_mode, - path: { "${params.outdir}/regions/samplepanels/createsamplepanelsnonprotaffect" }, - pattern: "*{tsv,bed}", - ] - ] - } - - withName: 'BBGTOOLS:DEEPCSA:CREATEPANELS:CREATESAMPLEPANELSEXONS' { - publishDir = [ - [ - mode: params.publish_dir_mode, - path: { "${params.outdir}/regions/samplepanels/createsamplepanelsexons" }, - pattern: "*{tsv,bed}", - ] - ] - } - - withName: 'BBGTOOLS:DEEPCSA:CREATEPANELS:CREATESAMPLEPANELSINTRONS' { - publishDir = [ - [ - mode: params.publish_dir_mode, - path: { "${params.outdir}/regions/samplepanels/createsamplepanelsintrons" }, - pattern: "*{tsv,bed}", - ] - ] - } - - withName: 'BBGTOOLS:DEEPCSA:CREATEPANELS:CREATESAMPLEPANELSSYNONYMOUS' { - publishDir = [ - [ - mode: params.publish_dir_mode, - path: { "${params.outdir}/regions/samplepanels/createsamplepanelssynonymous" }, - pattern: "*{tsv,bed}", - ] - ] - } - - withName: 'BBGTOOLS:DEEPCSA:ENRICHPANELS:EXPANDREGIONSALL' { + withName: 'EXPANDREGIONSALL|EXPANDREGIONSNONPROT|EXPANDREGIONSPROT|EXPANDREGIONSSYNONYMOUS|EXPANDREGIONSEXONS' { publishDir = [ [ mode: params.publish_dir_mode, @@ -127,43 +67,4 @@ process { ] } - withName: 'BBGTOOLS:DEEPCSA:ENRICHPANELS:EXPANDREGIONSNONPROT' { - publishDir = [ - [ - mode: params.publish_dir_mode, - path: { "${params.outdir}/regions/expandedregions/" }, - pattern: "*{tsv,json}", - ] - ] - } - - withName: 'BBGTOOLS:DEEPCSA:ENRICHPANELS:EXPANDREGIONSPROT' { - publishDir = [ - [ - mode: params.publish_dir_mode, - path: { "${params.outdir}/regions/expandedregions/" }, - pattern: "*{tsv,json}", - ] - ] - } - - withName: 'BBGTOOLS:DEEPCSA:ENRICHPANELS:EXPANDREGIONSSYNONYMOUS' { - publishDir = [ - [ - mode: params.publish_dir_mode, - path: { "${params.outdir}/regions/expandedregions/" }, - pattern: "*{tsv,json}", - ] - ] - } - - withName: 'BBGTOOLS:DEEPCSA:ENRICHPANELS:EXPANDREGIONSEXONS' { - publishDir = [ - [ - mode: params.publish_dir_mode, - path: { "${params.outdir}/regions/expandedregions/" }, - pattern: "*{tsv,json}", - ] - ] - } } diff --git a/docs/README.md b/docs/README.md index 8e0c0b90..19103546 100644 --- a/docs/README.md +++ b/docs/README.md @@ -4,9 +4,15 @@ The bbglab/deepCSA documentation is split into the following pages: - [Usage](usage.md) - An overview of how the pipeline works and how to run it. +- [Input scenarios](input_scenarios.md) + - The three supported input modes (VCF + BAM, VCF + precomputed depths, cohort MAF + precomputed depths) and when to use each. - [File formatting](file_formatting.md) - An overview of the specific formats required for each of the custom mandatory or optional files. - [Output](output.md) - An overview of the different results produced by the pipeline and how to interpret them. - [Tools](tools.md) - An overview of the explanation of the tools used in deepCSA and the rationale behind some of the decisions or computations. +- [Test data](test_data.md) + - Where the test data lives, what it contains, and how it is consumed by the nf-test suite. +- [Issue resolution](issue_resolution.md) + - Known issues encountered during development and how they were resolved. diff --git a/docs/file_formatting.md b/docs/file_formatting.md index 98458877..da53ff1e 100644 --- a/docs/file_formatting.md +++ b/docs/file_formatting.md @@ -216,16 +216,105 @@ params { ### cosmic_ref_signatures +Path to the COSMIC SBS signature reference file used by SigProfilerAssignment. Use the SBS 96 context file for your genome build (e.g., `COSMIC_v3.4_SBS_GRCh38.txt`). The file is a tab-delimited matrix where the first column encodes mutation context and each additional column corresponds to a signature. + ### wgs_trinuc_counts +Tab-delimited file with two columns: + +```text +CONTEXT COUNT +ACA 118979126 +ACC 67570313 +... +``` + +The file represents the **total number of occurrences of each trinucleotide** in the reference genome. The pipeline provides a default example in `assets/trinucleotide_counts/`. + ### cadd_scores +Path to the CADD "All possible SNVs" file (BGZIP-compressed TSV). This file is used for OncodriveFML scoring. + +Recommended download: [CADD downloads](https://cadd.gs.washington.edu/download) → "All possible SNVs of GRCh38/hg38". + ### cadd_scores_ind -### dnds_ref_transcripts +Tabix index (`.tbi`) for the `cadd_scores` file. If you need to generate it: + +```bash +bgzip -c whole_genome_SNVs.tsv > whole_genome_SNVs.tsv.gz +tabix -s 1 -b 2 -e 2 whole_genome_SNVs.tsv.gz +``` + +### dnds_biomart_ref + +https://github.com/bbglab/deepCSA/tree/dev/assets/build_datasets/dndscv ### dnds_covariates -### datasets3d +dNdScv covariates file, usually `covariates_hg19_hg38_epigenome_pcawg.rda`. This provides covariate regression terms for mutation rate modeling. + +### Oncodrive3D datasets and annotations + +Directory containing precomputed Oncodrive3D datasets (structure and mutation mapping information). + +Directory containing Oncodrive3D annotation datasets (protein annotations, stability data, etc.) + +Use the same build process for datasets and annotations to ensure compatibility. + +Build using the [Oncodrive3D dataset builder](https://github.com/bbglab/oncodrive3d?tab=readme-ov-file#building-datasets). + +These datasets can be downloaded from [Zenodo](https://zenodo.org/records/21031511). +See explanation here: [Oncodrive3D pre-built datasets](https://github.com/bbglab/oncodrive3d?tab=readme-ov-file#pre-built-datasets-download-instead-of-building) + +Files for human: + +```sh +o3d_annotations_human.v1.0.9.tar.zst +o3d_datasets_human.v1.0.9.tar.zst +``` + +Files for mouse: + +```sh +o3d_annotations_mouse.v1.0.9.tar.zst +o3d_datasets_mouse.v1.0.9.tar.zst +``` + + + +### gff3_file + +Optional local GFF3 file used by the DNA2PROTEINMAPPING step. If not provided, the pipeline downloads the GFF3 from Ensembl. If provided, it must match the Ensembl release, species, and genome build you are using (compressed `.gff3.gz` files are supported). + +## Examples + +### Blacklist mutations + +``` +chr1:11107296_C>CA +chr1:11107450_C>A +chr1:11108379_T>A +``` + +### Gene grouping + +``` +chr15q chr15q IDH2 SIN3A +chr17p chr17p MAP2K4 NCOR1 TP53 USP6 +``` + +### Custom annotation + +See `assets/example_inputs/custom_regions.example.tsv` for a full example of a custom region annotation file. + +### Omega hotspots / subgenic regions + +Provide a BED file with 3 or 4 columns (`CHROM`, `START`, `END`, optional `NAME`): + +``` +chr7 55191765 55191840 EGFR_L858R_region +chr12 25245300 25245380 KRAS_G12_region +``` -### annotations3d +You can expand these regions with `hotspot_expansion` and optionally generate complements with `subgenic_regions_complement`. diff --git a/docs/images/ClustermapProfileSimilarity.png b/docs/images/ClustermapProfileSimilarity.png new file mode 100644 index 00000000..6e1d38f9 Binary files /dev/null and b/docs/images/ClustermapProfileSimilarity.png differ diff --git a/docs/images/DepthsPerSampleGene.png b/docs/images/DepthsPerSampleGene.png new file mode 100644 index 00000000..b1bced53 Binary files /dev/null and b/docs/images/DepthsPerSampleGene.png differ diff --git a/docs/images/DomainSelection.png b/docs/images/DomainSelection.png new file mode 100644 index 00000000..e190e8f7 Binary files /dev/null and b/docs/images/DomainSelection.png differ diff --git a/docs/images/EvalOmegaGloc.png b/docs/images/EvalOmegaGloc.png new file mode 100644 index 00000000..2d569e06 Binary files /dev/null and b/docs/images/EvalOmegaGloc.png differ diff --git a/docs/images/MutDensityNdepth.png b/docs/images/MutDensityNdepth.png new file mode 100644 index 00000000..0da0380a Binary files /dev/null and b/docs/images/MutDensityNdepth.png differ diff --git a/docs/images/MutDensityQCgene.png b/docs/images/MutDensityQCgene.png new file mode 100644 index 00000000..fbc628a0 Binary files /dev/null and b/docs/images/MutDensityQCgene.png differ diff --git a/docs/images/MutatedReads.png b/docs/images/MutatedReads.png new file mode 100644 index 00000000..cb2fa6f6 Binary files /dev/null and b/docs/images/MutatedReads.png differ diff --git a/docs/images/MutdensityQCsample.png b/docs/images/MutdensityQCsample.png new file mode 100644 index 00000000..b46ddfbf Binary files /dev/null and b/docs/images/MutdensityQCsample.png differ diff --git a/docs/images/TrinucleotideProportions.png b/docs/images/TrinucleotideProportions.png new file mode 100644 index 00000000..24519714 Binary files /dev/null and b/docs/images/TrinucleotideProportions.png differ diff --git a/docs/images/VAF_qc.png b/docs/images/VAF_qc.png new file mode 100644 index 00000000..f67cbd00 Binary files /dev/null and b/docs/images/VAF_qc.png differ diff --git a/docs/images/mutational_signatures.png b/docs/images/mutational_signatures.png new file mode 100644 index 00000000..f2c5720a Binary files /dev/null and b/docs/images/mutational_signatures.png differ diff --git a/docs/images/needle_plots.png b/docs/images/needle_plots.png new file mode 100644 index 00000000..e562e3b3 Binary files /dev/null and b/docs/images/needle_plots.png differ diff --git a/docs/images/saturation_plots.png b/docs/images/saturation_plots.png new file mode 100644 index 00000000..05e2ae30 Binary files /dev/null and b/docs/images/saturation_plots.png differ diff --git a/docs/images/saturation_proportions.png b/docs/images/saturation_proportions.png new file mode 100644 index 00000000..449d8e69 Binary files /dev/null and b/docs/images/saturation_proportions.png differ diff --git a/docs/images/selection_summary.png b/docs/images/selection_summary.png new file mode 100644 index 00000000..ed4d596f Binary files /dev/null and b/docs/images/selection_summary.png differ diff --git a/docs/input_scenarios.md b/docs/input_scenarios.md new file mode 100644 index 00000000..0ddeda63 --- /dev/null +++ b/docs/input_scenarios.md @@ -0,0 +1,96 @@ +# bbglab/deepCSA: Input scenarios + +deepCSA supports three input scenarios depending on what you already have available (BAMs, mutations as VCFs, or a cohort-level MAF together with a precomputed depths table). All scenarios still require the standard samplesheet CSV passed via `--input`. + +Sample naming rules apply to every scenario: avoid `.` in sample names and prefer text-like names instead of purely numeric ones. See [File formatting](file_formatting.md) for details on each file. + +## Scenario summary + +| Scenario | `--input` columns | Depth source | Extra flags | +|---|---|---|---| +| 1. VCF + BAM (default) | `sample,vcf,bam` | Computed from BAMs | — | +| 2. VCF + precomputed depths | `sample,vcf` | `--custom_depths_table` | `--use_custom_depths true` | +| 3. Cohort MAF + precomputed depths | `sample,vcf` (metadata only) | `--custom_depths_table` | `--input_maf ` + `--use_custom_depths true` | + +The pipeline validates these combinations at start-up and stops with an explicit error if `--input_maf` is set without `--use_custom_depths true` (see [workflows/deepcsa.nf](../workflows/deepcsa.nf)). + +## Scenario 1 — VCF + BAM (default) + +Use this scenario when you have per-sample variant calls and the BAM files that were used to produce them. + +```csv +sample,vcf,bam +sample1,sample1.filtered.vcf,sample1.sorted.bam +sample2,sample2.filtered.vcf,sample2.sorted.bam +``` + +The pipeline derives per-position sequencing depth directly from the BAMs (subworkflow `depthanalysis`). No extra flag is needed. + +## Scenario 2 — VCF + precomputed depths + +Use this scenario when you already have a depths table (for example produced by a previous deepCSA run, or by an external tool) and you want to skip BAM-based pileup. + +```csv +sample,vcf +sample1,sample1.filtered.vcf +sample2,sample2.filtered.vcf +``` + +```console +params { + use_custom_depths = true + custom_depths_table = '/path/to/precomputed_depths_table.tsv' +} +``` + +Notes: + +- The depths-table column names must match the sample names declared in the `sample` column of the input CSV. +- `custom_depths_table` may be TSV or CSV but must follow the per-position depth layout that deepCSA expects. +- If the file is missing or unreadable the pipeline fails immediately. + +See [Usage — Using a precomputed depths table](usage.md#using-a-precomputed-depths-table) for additional notes on how columns are matched and on preparing the file from a previous deepCSA run. + +## Scenario 3 — Cohort MAF + precomputed depths + +Use this scenario when all mutations for the cohort are already consolidated in a single MAF/TSV file and you also have the matching precomputed depths table. + +```console +params { + input = "samplesheet.csv" + input_maf = "cohort_mutations.maf" + use_custom_depths = true + custom_depths_table = "precomputed_depths.tsv" +} +``` + +```bash +nextflow run bbglab/deepCSA \ + --input samplesheet.csv \ + --outdir results/ \ + --input_maf cohort_mutations.maf \ + --use_custom_depths true \ + --custom_depths_table precomputed_depths.tsv \ + -profile +``` + +What happens under the hood: + +1. The MAF file is split into one VCF per unique `SAMPLE_ID` by `INPUTMAF2VCF` (script [assets/useful_scripts/deepcsa_maf2samplevcfs.py](../assets/useful_scripts/deepcsa_maf2samplevcfs.py)). +2. The per-sample VCFs are published under `/processing_files/input_vcfs/`. +3. The rest of the pipeline runs as in Scenario 2. + +The standard `--input` samplesheet is still required, because it provides the sample metadata used by other pipeline steps. The `SAMPLE_ID` values in the MAF must match the `sample` column of the samplesheet. + +For the expected MAF columns (deepCSA-generated MAF vs external MAF) see [Usage — MAF file format](usage.md#maf-file-format). + +## Related parameters + +| Parameter | Purpose | +|---|---| +| `input` | Samplesheet CSV with `sample,vcf[,bam]` columns. Always required. | +| `input_maf` | Cohort-level MAF file (Scenario 3). Requires `use_custom_depths = true`. | +| `use_custom_depths` | Skip BAM-based depth computation. Required for Scenarios 2 and 3. | +| `custom_depths_table` | Path to the precomputed per-position depths table. Required when `use_custom_depths = true`. | + +Custom-mutation workflows (e.g. forcing your own filter list) layer on top of these scenarios. See [Usage — Custom mutation calls](usage.md#custom-mutation-calls----option-1-building-input-vcfs-and-providing-them-via-normal-input) for the advanced options. diff --git a/docs/issue_resolution.md b/docs/issue_resolution.md index c57a74fd..9133926b 100644 --- a/docs/issue_resolution.md +++ b/docs/issue_resolution.md @@ -100,8 +100,27 @@ This becomes a problem when subgenic elements that are not fully covered by the The solution is to include the covered portions of a region in the expanded panel even if the region is only partially covered. This ensures representation of all elements in downstream analyses, but it does not guarantee full comparability between runs because different sample groups may yield different panel and subgenic-region definitions. +## 4. CDKN2A missing in dNdScv covariates file -## 4. Issue 4: Description of the issue +When trying to match CDKN2A to the gene names in the dNdScv covariates file it is missing. This is because there are the two isoforms of it corresponding to two different proteins: + +- "ENSP00000462950" ~ "CDKN2A.p14arf", + +- "ENSP00000307101" ~ "CDKN2A.p16INK4a" <---- This one! + +### Processes affected: + +- `DNDSRUN` + +### GitHub tracking + +- **GitHub PR (Resolution):** [https://github.com/bbglab/deepCSA/pull/465](https://github.com/bbglab/deepCSA/pull/465) + +### Resolution + +What we did in this PR is to rename the row of `CDKN2A.p16INK4a` in the covariates file to CDKN2A and we removed the row for `CDKN2A.p14arf` this way we avoid the ambiguity, but at the expenses of losing information from these second protein version of CDKN2A. + +## 5. Issue 5: Description of the issue Description of the issue, including its impact and context. @@ -116,7 +135,6 @@ Description of the issue, including its impact and context. ### Resolution - ## Conclusion In this document, we have documented the issues identified during development and their corresponding resolutions. diff --git a/docs/output.md b/docs/output.md index 574d2d27..b3c5dd00 100644 --- a/docs/output.md +++ b/docs/output.md @@ -19,239 +19,247 @@ The pipeline is built using [Nextflow](https://www.nextflow.io/) and processes d - [Additional clonal structure metrics](#additional-clonal-structure-metrics) - [Mutational signatures](#mutational-signatures) - [Plotting functionalities](#plotting-functionalities) +- [QC outputs](#qc-outputs) - [Additional outputs](#additional-outputs) ## Directory Structure -The directory structure listed below will be created in the results directory after the pipeline has finished. -The structure captures the maximum diversity of created outputs, but when only certain run options are turned on, not all directories will be generated. -All paths are relative to the top-level results directory. +The directory tree below shows the maximum diversity of outputs the pipeline can publish. When only some run options are turned on, only the corresponding subdirectories will be generated. All paths are relative to the top-level results directory. ```{console} {outdir} -├──absolutemutabilities -├──absolutemutabilitiesgloballoc -├──annotatedepths -├──clean_germline_somatic -├──clean_somatic -├──computematrix -├──computeprofile -├──createpanels -│ ├── consensus -│ │ └── .consensus.bed -│ │ └── .consensus.tsv -│ ├── captured -│ │ └── .captured.bed -│ │ └── .captured.tsv -│ └── sample -│ └── ..bed -│ └── ..tsv -├──customannotation -├──customprocessing -├──customprocessingrich -├──depthssummary -├──dna2proteinmapping -├──domainannotation -├──expandregions -├──filterexons -├──germline_somatic -├──groupgenes -├──indels -├──matrixconcatwgs -├──multiqc -├──mutability -├──mutatedcellsfromvafam -├──mutatedgenomesfromvafam -├──mutrate -├──muts2sigs -├──omega -│ ├── preprocessing -│ │ └── syn_muts. -│ │ └── mutabilities. -│ └── output_mle..tsv -├──omegagloballoc -│ ├── preprocessing -│ │ └── syn_muts. -│ │ └── mutabilities. -│ └── output_mle..tsv -├──oncodrive3d -│ ├── run -│ └── -│ └── plot -│ └── -├──oncodrivefmlsnvs -├──pipeline_info -├──plotmaf -├──plotneedles -│ └── -│ └── -├──plotselection -├──plotsomaticmaf -├──postprocessveppanel -├──signatures_hdp -│ └── output. -│ └── -├──sigprobs -├──sigprofilerassignment -│ └── output. -│ └── -├──sitecomparison -├──sitecomparisongloballoc -├──sitecomparisongloballocmulti -├──sitecomparisonmulti -├──sitesfrompositions -├──sumannotation -├──synmutrate -├──synmutreadsdensity -└──table2group -work/ -.nextflow.log +├── depths +│ ├── individual # per-sample depth tables +│ ├── plots_per_group # depth plots split by sample groupings +│ └── summary # exons / exons_cons / all_cons depth summaries +├── group_definition +│ ├── genes +│ └── samples +├── mutations +│ ├── germline_somatic # all calls labelled germline + somatic +│ ├── clean_somatic # somatic calls after filtering +│ └── clean_germline_somatic # cleaned germline + somatic +├── mutational_profile # trinucleotide profiles (all / exons / introns / non-prot / synonymous) +├── mutdensity +│ └── individual_vals # flat mutation density per sample/group +├── mutdensity_adjusted +│ └── individual_vals # trinucleotide-adjusted mutation density +├── regions +│ ├── allsites # captured positions ready for VEP +│ ├── annotations # panel annotation tables and plots +│ ├── capturedpanels # per-region captured panels +│ ├── consensuspanels # consensus panels (cohort-level) +│ ├── samplepanels # per-sample panels per region type +│ │ ├── createsamplepanelsall +│ │ ├── createsamplepanelsexons +│ │ ├── createsamplepanelsintrons +│ │ ├── createsamplepanelsnonprotaffect +│ │ ├── createsamplepanelsprotaffect +│ │ └── createsamplepanelssynonymous +│ ├── expandedregions # subgenic / domain / exon expansions +│ ├── panelannotation +│ └── dndscv # biomart filtered by panel BED (dynamic RefCDS input) +├── selection +│ ├── omega +│ │ ├── preprocessing # syn_muts., mutabilities. +│ │ └── estimator # all_omegas.tsv, output_mle..tsv +│ ├── omegagloballoc +│ │ ├── preprocessing +│ │ └── estimator +│ ├── sitecomparison # background × count combinations +│ │ ├── bckg_single_count_single +│ │ ├── bckg_single_count_multi +│ │ ├── bckg_multi_count_single +│ │ ├── bckg_multi_count_multi +│ │ ├── bckg_glocsingle_count_single +│ │ ├── bckg_glocsingle_count_multi +│ │ ├── bckg_glocmulti_count_single +│ │ └── bckg_glocmulti_count_multi +│ ├── oncodrivefml +│ ├── oncodrive3d +│ │ └── run # per-sample Oncodrive3D results +│ ├── dndscv # dNdScv (R) outputs +│ │ ├── cv # *.cv.tsv +│ │ ├── persample # *.globaldnds.tsv +│ │ └── local # *.loc.tsv +│ └── dndsproxy # dN/dS proxy from adjusted vs synonymous densities +├── signatures +│ ├── sigprofilerassignment +│ ├── sigprofilerassignment_indels +│ ├── sigprofilermatrixgenerator +│ ├── signatures_hdp +│ └── hdp_decomposition_spa +├── plots +│ ├── mutations_summary # plot_maf / plot_somatic_maf +│ ├── needle_plots # per-sample, per-gene needles +│ ├── selection_summary +│ ├── selection +│ │ ├── omega +│ │ ├── omegagloballoc +│ │ └── oncodrive3d +│ │ └── chimerax +│ ├── gene_subgenic_selection +│ ├── saturation_proportions +│ └── interindividual_variability +├── qc +│ ├── trinucleotide_proportions +│ ├── mutational_profiles_comparison +│ ├── mutdensityqc +│ ├── metrics_vs_depth +│ ├── mutationspecific +│ ├── omega_flagged +│ ├── evaluate_omega_globalloc +│ └── contamination +├── processing_files +│ ├── input_vcfs # per-sample VCFs (when --input_maf is used) +│ ├── all_possible_sites +│ ├── sumannotation +│ ├── synmutdensity +│ ├── synmutreadsdensity +│ ├── mutations_matrix +│ │ └── per_sample # SBS matrices for signature analysis +│ ├── relativemutability +│ ├── flagged_positions +│ └── multiqc +├── regressions +├── pipeline_info +└── multiqc ``` ## Input and configuration -See Usage docs for extensive explanation on required inputs and format. Including documentation on parameters to run on for 4 different suggested running modes. +See [Usage](usage.md) and [Input scenarios](input_scenarios.md) for an explanation of the required inputs, the three supported input modes, and the parameter presets for the four suggested run profiles. ## Depth analysis ### Key role -- Computation of depth per sample for each specific position - Most analysis may be influenced by sequencing depth, it is essential to correct for these values. +- Computation of depth per sample for each specific position. + Most analyses are influenced by sequencing depth, so it is essential to correct for these values. -- Definition of regions to analyze - Only genomic areas that have been properly covered across samples will be used for the analysis. +- Definition of regions to analyse. + Only genomic areas that have been properly covered across samples will be used. -**Note 1:** There is a depth difference between the depth reported in the files in the annotated depths directory and the values of depth reported in each of the mutations. This difference is because we do not count Ns when computing th depth of specific mutations. This means that the values of VAF are computed with N-discounted depth, while other metrics are not. +**Note 1:** There is a depth difference between the depth reported in the files under `depths/individual/` and the values reported per mutation. This difference is because Ns are not counted when computing the depth at the specific mutation position. Therefore VAF values are computed with N-discounted depth while other metrics are not. -### Detailed explanation of depthssummary depths versions +### Detailed explanation of `depths/summary/` versions -In this directory you will find different versions of TSVs and PDFs summarizing the depths of the samples/genes sequenced. +In this directory you will find different versions of TSVs and PDFs summarising the depths of the samples/genes sequenced. -Each of the versions provides slightly different information, as you can see in the image below: +Each version provides slightly different information, as shown below: ![depths summary slide](images/deepCSA_depths_summary.png) -- exons contains the average depth in all the exonic regions sequenced in the genome no matter which minimum consensus coverage was reached. -- exons_cons contains the average depth in the exonic regions sequenced in the genome to a minimum consensus depth threshold. (only exons in the well covered regions) -- all_cons contains the average depth of all sequenced regions of the genome that are well covered across the samples in the cohort, without any distinction of exons/introns/others. - -We will work on a better representation of the different metrics of depth so that is it more understandable, but for now we include this schematic and brief explanations. - -Reach out if you have more questions! +- `exons` — average depth in all the exonic regions sequenced, regardless of consensus coverage. +- `exons_cons` — average depth in exonic regions reaching the minimum consensus depth threshold (i.e. exons within the well-covered regions). +- `all_cons` — average depth of all well-covered sequenced regions across the cohort, with no exonic/intronic distinction. ### Outputs -- sitesfrompositions -- postprocessveppanel -- createpanels -- annotatedepths -- depthssummary +- `depths/` (individual, summary, plots_per_group) +- `regions/allsites/`, `regions/annotations/`, `regions/panelannotation/` +- `regions/capturedpanels/`, `regions/consensuspanels/`, `regions/samplepanels/` + +Optional (subgenic / domain expansion): -Optional: +- `regions/expandedregions/` +- `regions/annotations/` (domain and DNA-to-protein mapping outputs) -- dna2proteinmapping -- domainannotation -- customprocessing -- customprocessingrich +![DepthsVariability](images/DepthsPerSampleGene.png) ## Mutation preprocessing ### Key role -- VCF annotation: Annotate mutations with Ensembl VEP. -- VCF to MAF conversion: Convert VCFs to MAF, define VAF, and merge with annotation. -- Custom region annotation: Allow user to define different consequence types for specific regions. -- Hotspot annotation: Add known hotspots to mutation annotation. +- VCF annotation with Ensembl VEP. +- VCF → MAF conversion, VAF computation, merge with annotation. +- Custom region annotation: user-defined consequence types for specific regions. +- Hotspot annotation: add known hotspots to the mutation table. - Filtering: - - Filter mutations at the sample level (e.g., VAF distortion). - - Filter at the cohort level (e.g., other_sample_SNP, repetitive_variant, not_covered, not_in_exons). -- Blacklist mutations if activated (see assets for example). -- Downsample mutations if activated. + - Sample-level filters (e.g. VAF distortion via `vaf_distortion_threshold`). + - Cohort-level filters (e.g. `other_sample_SNP`, `repetitive_variant`, `not_covered`, `not_in_exons`). +- Optional blacklist of mutations (see assets for example). +- Optional downsampling of mutations. ### Outputs -- sumannotation -- customannotation -- germline_somatic -- clean_somatic -- clean_germline_somatic +- `mutations/germline_somatic/` +- `mutations/clean_somatic/` +- `mutations/clean_germline_somatic/` +- `processing_files/sumannotation/`, `processing_files/flagged_positions/` ## Basic analysis ### Key role -- Mutation density computation - Correct the number of mutations observed by the number of sequenced nucleotides. - -- Mutational profile computation - Capture the mutation probability of each trinucleotide. Represent it in three different normalization conditions. +- Mutation density computation — corrects the number of observed mutations by the number of sequenced nucleotides. +- Mutational profile computation — captures the mutation probability of each trinucleotide, in three different normalisation conditions. ### Outputs -- computematrix -- computeprofile -- mutrate +- `mutdensity/individual_vals/` +- `mutdensity_adjusted/individual_vals/` (trinucleotide-adjusted; see [Tools — Adjusted mutation density](tools.md#adjusted-mutation-density)) +- `mutational_profile/` +- `processing_files/mutations_matrix/` (per-sample SBS matrix) ## Intermediate outputs ### Key role -- Matrix concatenation - Combine WGS-renomralized matrices for mutational signature analysis. - -- Mutability calculation - Compute relative mutabilities using depths and mutational profile. - -- Choose synonymous mutation rates for downstream analysis. +- Matrix concatenation — combine WGS-renormalised matrices for mutational signature analysis. +- Mutability calculation — compute relative mutabilities using depths and the mutational profile. +- Selection of the synonymous mutation rate used downstream. ### Outputs -- matrixconcatwgs -- mutability -- synmutrate -- synmutreadsdensity +- `processing_files/mutations_matrix/` (cohort-level concatenated matrix) +- `processing_files/relativemutability/` +- `processing_files/synmutdensity/` +- `processing_files/synmutreadsdensity/` ## Positive selection ### Key role -- Compute multiple positive selection metrics - This is done at the cohort-level, but also for each sample or group of samples. - -- OncodriveFML: Detects functional impact bias in observed mutations. - -- Oncodrive3D: Identifies 3D protein regions with mutation clustering, using relative mutabilities and raw VEP annotation. - -- Omega: dN/dS-based, quantifies selection pressure in defined regions (genes, exons, domains, hotspots, etc.). - -- Indels: Analysis of indel selection. +- Compute several positive selection metrics at the cohort level and per sample/group: + - **OncodriveFML** — functional-impact bias. + - **Oncodrive3D** — 3D protein clustering, optionally on raw VEP annotation. + - **Omega** — dN/dS-based selection in defined regions (genes, exons, domains, hotspots, ...). + - **dNdScv** — R implementation, run with a per-run RefCDS built dynamically from the panel BED + a biomart export (`dnds_biomart_ref`) + the genome FASTA. See [Tools — dNdScv](tools.md#dndscv). + - **dN/dS proxy** — quick ratio of adjusted vs synonymous mutation densities, output as `*.gene_mutdensities_n_dnds.tsv`. + - **Indels** — indel selection analysis. ### Outputs -- omega -- omegagloballoc -- oncodrive3d -- oncodrivefmlsnvs -- indels +- `selection/omega/{preprocessing,estimator}/` +- `selection/omegagloballoc/{preprocessing,estimator}/` +- `selection/oncodrive3d/run/` +- `selection/oncodrivefml/` +- `selection/dndscv/{cv,persample,local}/` +- `selection/dndsproxy/` ## Site selection metrics ### Key role - Compute absolute mutabilities for each position. - -- Compare the observed number of mutations per site to the expected number of mutations and estimate a site selection value. +- Compare the observed number of mutations per site to the expected number and estimate a site-selection value. ### Outputs -- absolutemutabilities -- absolutemutabilitiesgloballoc +- The recommended ones to use are: + + - For reporting selection at a cohort-level: + + `selection/sitecomparison/bckg_single_count_single` + + - For estimating selection accounting for the expansions or multiple occurrences of specific mutations: + + `selection/sitecomparison/bckg_single_count_multi` -- sitecomparison -- sitecomparisongloballoc -- sitecomparisongloballocmulti -- sitecomparisonmulti + `selection/sitecomparison/bckg_multi_count_multi` + +- But all possible combinations are available: `selection/sitecomparison/` (8 background × count combinations: `bckg_{single,multi,glocsingle,glocmulti}_count_{single,multi}/`) ## Additional clonal structure metrics @@ -261,50 +269,105 @@ Optional: ### Outputs -- mutatedcellsfromvafam -- mutatedgenomesfromvafam +- Subdirectories of the mutated-cells analyses are published under `selection/` and `mutations/` according to the configured grouping; the corresponding processes are `mutated_cells_from_vaf` and `mutated_genomes_from_vaf` (controlled by `params.mutated_cells_vaf`). ## Mutational signatures ### Key role -- Signature assignment: Use SigProfilerAssignment with optional custom signatures. -- HDP: Hierarchical Dirichlet Process for signature extraction. -- (Pending) Signature extraction: SigProfilerExtractor support. +- Signature assignment with SigProfilerAssignment (optional custom signatures). +- HDP — Hierarchical Dirichlet Process signature extraction. +- SigProfilerExtractor is supported but must be run externally. ### Outputs -- signatures_hdp -- sigprofilerassignment -- sigprobs -- muts2sigs +- `signatures/sigprofilerassignment/` +- `signatures/sigprofilerassignment_indels/` +- `signatures/sigprofilermatrixgenerator/` +- `signatures/signatures_hdp/` +- `signatures/hdp_decomposition_spa/` + +### Examples + +![MutationalSignatures](images/mutational_signatures.png) ## Plotting functionalities ### Key role -- Plotting basic statistics of numbers and distribution of mutations in genes. +- Plot basic statistics on numbers and distribution of mutations in genes. +- Plot selection results (omega, OncodriveFML, Oncodrive3D, gene/subgenic saturation, interindividual variability). + +Plotting scope can be controlled with `plot_only_allsamples`: when `true`, only cohort-level plots are generated; when `false`, plots are also produced for each defined subgroup. + +### Outputs + +- `plots/mutations_summary/` +- `plots/needle_plots/` +- `plots/selection_summary/` +- `plots/selection/{omega,omegagloballoc,oncodrive3d}/` +- `plots/gene_subgenic_selection/` +- `plots/saturation_proportions/` +- `plots/interindividual_variability/` + +### Examples + +![NeedlePlots](images/needle_plots.png) + +![SaturationPlots](images/saturation_plots.png) + +![DomainSelection](images/DomainSelection.png) + +![SelectionSummary](images/selection_summary.png) + +![SaturationProportions](images/saturation_proportions.png) + +![MutationsSummary](images/MutatedReads.png) + +## QC outputs -- Optionally think on adding more plots. +### Key role + +A `qc/` umbrella collects all the quality-control views; `qc/metrics_vs_depth/` always runs and produces depth-vs-metric scatter plots for raw and adjusted mutation densities and omega-globalloc. ### Outputs -- plotmaf -- plotneedles -- plotselection -- plotsomaticmaf +- `qc/trinucleotide_proportions/` +- `qc/mutational_profiles_comparison/` +- `qc/mutdensityqc/` +- `qc/metrics_vs_depth/` +- `qc/mutationspecific/` +- `qc/omega_flagged/` +- `qc/evaluate_omega_globalloc/` +- `qc/contamination/` + +### Examples + +![TrinucleotideProportionsInPanel](images/TrinucleotideProportions.png) + +![MutProfileComparisons](images/ClustermapProfileSimilarity.png) + +![MutDensityQCSample](images/MutdensityQCsample.png) + +![MutDensityQCGenes](images/MutDensityQCgene.png) + +![EvalOmegaGloc](images/EvalOmegaGloc.png) + +![MutationDensityVSdepth](images/MutDensityNdepth.png) + +![VAFmutationQC](images/VAF_qc.png) ## Additional outputs ### Key role -- Definition of groups, expanded regions and other metrics related with the full pipeline execution. +- Definition of sample/gene groups, expanded regions, regression configs, and pipeline-level reports. ### Outputs -- table2group -- groupgenes -- expandregions -- filterexons -- multiqc -- pipeline_info +- `group_definition/{samples,genes}/` +- `regions/expandedregions/` +- `regressions/` +- `multiqc/` +- `pipeline_info/` +- `processing_files/input_vcfs/` (when `--input_maf` is used) diff --git a/docs/test_data.md b/docs/test_data.md new file mode 100644 index 00000000..f07b81f8 --- /dev/null +++ b/docs/test_data.md @@ -0,0 +1,77 @@ +# bbglab/deepCSA: Test data + +deepCSA ships a minimal nf-test suite ([tests/deepcsa.nf.test](../tests/deepcsa.nf.test)) that exercises the main input scenarios and validation paths. This document describes where the test data lives, what it contains, and how it is consumed by the tests. + +## Where the test data lives + +The reference test datasets are hosted in the [bbglab/DeepClone_protocol](https://github.com/bbglab/DeepClone_protocol) repository, under `test_datasets/deepCSA/testdata/`: + +``` +test_datasets/deepCSA/testdata/ +├── maf/ +│ └── all_samples.somatic.mutations.maf # cohort-level MAF (3 samples) +├── depth/ +│ └── all_samples_indv.depths.tsv.gz # precomputed per-position depths table +└── input_vcfs/ + ├── P19_0002_BDO_01.vcf + ├── P19_0002_BTR_01.vcf + └── P19_0003_BDO_01.vcf +``` + +The three test samples (`P19_0002_BDO_01`, `P19_0002_BTR_01`, `P19_0003_BDO_01`) come from a bladder duplex-sequencing experiment and are large enough to exercise the panel, mutational-profile, depth, and omega code paths while keeping runtimes short. + +Locally committed inputs under [tests/test_data/](../tests/test_data/) only contain the small CSV samplesheets and one toy MAF used by the validation-failure tests: + +| File | Purpose | +|---|---| +| `input.csv` | Samplesheet with `sample,vcf,bam` columns referring to internal IRB paths (not used by the public CI tests). | +| `input_maf.csv` | Samplesheet with `sample,vcf` columns pointing to remote VCFs from `bbglab/DeepClone_protocol`. Used by the MAF-input test. | +| `input_no_bam.csv` | Same as `input_maf.csv`, used by the VCF-without-BAM tests. | +| `test_mutations.maf` | Tiny MAF used only by the parameter-validation failure tests (3, 4, 5). | + +## Remote-fetching convention + +Following the convention used by nf-core pipelines (e.g. [nf-core/fastquorum](https://github.com/nf-core/fastquorum)), the MAF and depths files are fetched **at runtime** directly from `bbglab/DeepClone_protocol` rather than pre-downloaded: + +```groovy +input_maf = 'https://raw.githubusercontent.com/bbglab/DeepClone_protocol/main/test_datasets/deepCSA/testdata/maf/all_samples.somatic.mutations.maf' +use_custom_depths = true +custom_depths_table = 'https://raw.githubusercontent.com/bbglab/DeepClone_protocol/main/test_datasets/deepCSA/testdata/depth/all_samples_indv.depths.tsv.gz' +``` + +> ⚠️ The `bbglab/DeepClone_protocol` repository must remain **publicly accessible** for Nextflow to fetch these files at runtime. If access is restricted the tests fail with a "No such file or directory" error. + +Because nf-schema 2.x validates `file-path` parameters for local existence, [tests/nextflow.config](../tests/nextflow.config) excludes the remote-URL parameters from that check: + +```groovy +validation { + ignoreParams = ['input_maf', 'custom_depths_table'] +} +``` + +## How tests map to input scenarios + +The five nf-test cases cover all three [input scenarios](input_scenarios.md) plus three validation-failure paths: + +| Test | Scenario covered | Inputs | +|---|---|---| +| TEST 1 — basic MAF processing | Scenario 3 (cohort MAF + depths) | `input_maf.csv` + remote MAF + remote depths | +| TEST 1b — VCF + depths | Scenario 2 (VCF + precomputed depths) | `input_no_bam.csv` + remote depths | +| TEST 2 — omega run | Scenario 3 with `omega = true` | same as TEST 1 | +| TEST 3 — `--input_maf` without `--use_custom_depths` | Validation failure | `input_maf.csv` + local toy MAF | +| TEST 4 — VCF samplesheet without BAMs and `use_custom_depths = false` | Validation failure | `input_no_bam.csv` | +| TEST 5 — `use_custom_depths = true` without a depths table | Validation failure | `input_no_bam.csv` | + +Snapshots (MD5 of selected outputs) are stored in [tests/deepcsa.nf.test.snap](../tests/deepcsa.nf.test.snap). + +## Running and updating the tests + +For the full execution-environment notes (SLURM, Singularity, `DEEPCSA_TEST_WORKDIR`, configuring for a non-IRB site) see [tests/README.md](../tests/README.md). The short version: + +```bash +nf-test test tests/deepcsa.nf.test # run the whole suite +nf-test test tests/deepcsa.nf.test --tag omega # run a single test +nf-test test tests/deepcsa.nf.test --update-snapshot # regenerate snapshots +``` + +The tests must be run on a SLURM cluster with Singularity; local execution is not supported because resource limits won't be met. Snapshots must be regenerated whenever default pipeline parameters change. diff --git a/docs/tools.md b/docs/tools.md index b74a4805..e862155e 100644 --- a/docs/tools.md +++ b/docs/tools.md @@ -8,6 +8,51 @@ Here, you can find an explanation of the different computations, tools or metrics implemented in deepCSA. +## Interpreting outputs (sanity checks and key metrics) + +### Sanity checks / QC + +Use these outputs to assess overall data quality before interpreting biological signals: + +- **Depth summaries** (`depthssummary/`): verify consistent coverage across samples and genes. +- **Mutation density vs depth** (`qc/metrics_vs_depth/`): check that mutation density does not collapse in low-depth samples. +- **Omega QC** (`qc/metrics_vs_depth/` + `qc/annotated_omegas`): highlights genes/samples with unstable omega estimates. +- **Mutational profile stability** (`computeprofile/*.profile_stability.tsv`): higher deviations indicate unstable mutational profiles (see below). + +### Omega vs omegagloballoc + +- **`omega/`** uses **per-sample mutational profiles** and per-sample synonymous rates to estimate selection. +- **`omegagloballoc/`** uses a **global cohort mutational profile** and global synonymous rates (shared across samples), which stabilizes estimates in low-burden samples and facilitates cohort-level comparisons. + +Use `omega` for sample-specific selection signals and `omegagloballoc` for conservative cohort-level estimates. + +### Site selection values + +Outputs in `sitecomparison/` and `sitecomparisongloballoc/` compare observed vs expected mutations per site or residue: + +- `OBSERVED_MUTS`: number of observed mutations. +- `EXPECTED_MUTS`: expected mutations from mutability models. +- `OBS/EXP`: selection enrichment ratio. +- `p_value`: Poisson p-value for observing at least `OBSERVED_MUTS` given `EXPECTED_MUTS`. + +The resolution is controlled by `site_comparison_grouping` (`site`, `aminoacid`, or `aminoacid_change`). + +### Mutational signatures + +- **`sigprofilerassignment/`**: assignments of known COSMIC signatures; includes activity tables and plots. +- **`signatures_hdp/`**: extracted signatures using a hierarchical Dirichlet process. +- **`sigprobs/` / `muts2sigs/`**: per-mutation signature probabilities (useful for downstream stratification). + +Interpret signature results alongside mutation counts and profile stability to avoid over-interpreting low-burden samples. + +### Mutational profile stability + +The file `*.profile_stability.tsv` is generated by adding a single mutation to each of the 96 SBS channels and measuring the L1 deviation from the original profile. Reported statistics include: + +- `mean_deviation`, `min_deviation`, `max_deviation`, `std_deviation` + +Lower deviations indicate a more stable (less noisy) profile. + ## Publications with detailed explanation We are in the process of completing the documentation, but in the meantime you can check the recently published [paper and its supplementary material for more details](https://www.nature.com/articles/s41586-025-09521-x). @@ -110,10 +155,44 @@ Defining the mutation density as $m/L$, according to the preceding explanation, For more explanations on omega go to the [corresponding repo](https://github.com/bbglab/omega). +deepCSA applies Benjamini-Hochberg multiple-testing correction separately for each of the following +comparison sets: all-samples (cohort), sample groups, and per-sample results, and it does this +independently for gene-level and subgenic regions. The corrected values are reported in the +`pvalue_adj` column. P-values equal to 0 are set to `1.17e-38` (minimum non-zero float32) before +correction to avoid underflow. + ## Site comparison The site comparison step takes advantage of the computation of mutabilities in [omega](https://github.com/bbglab/omega), and then compares these mutabilities either by residue, residue change or nucleotide change. +## dNdScv + +deepCSA wraps the [dNdScv](https://github.com/im3sanger/dndscv) R package and runs it with a **dynamically built** `RefCDS` reference instead of relying on a pre-baked `.rda` transcripts file. + +The `dnds` subworkflow performs three steps for every run: + +1. `ADAPT_PANEL_REFCDS` (`dNdScv_panel_prep.py`) — filter the biomart export referenced by `params.dnds_biomart_ref` to the transcripts overlapping the panel BED. +2. `BUILD_REFCDS` — call `dndscv::buildref` using `params.fasta` to produce a fresh `RefCDS_custom.rda`. +3. `DNDSRUN` (`dNdS_run.R`) — run dNdScv on the cohort, producing `*.cv.tsv`, `*.globaldnds.tsv` and `*.loc.tsv` under `selection/dndscv/{cv,persample,local}/`. + +Instructions for regenerating the biomart TSV are in [assets/build_datasets/dndscv/instructions.txt](../assets/build_datasets/dndscv/instructions.txt). The previously required `dnds_ref_transcripts` parameter has been removed. + +## dN/dS proxy + +When mutation density and the all-regions profile are computed, deepCSA also generates a quick **dN/dS proxy** per gene by taking the ratio of non-synonymous vs synonymous adjusted mutation densities. The implementation is in `mut_density_adjusted_dnds.py` and the results are published to `selection/dndsproxy/` as `*.gene_mutdensities_n_dnds.tsv`. + +This metric is intended as a fast sanity check and is independent of the R-based dNdScv run and of omega, both of which provide dN/dS estimates with significance testing. It is gated by `run_mutdensity` (which itself is enabled by either `mutationdensity` or `omega`) combined with `profileall`. + +## Depth-vs-metric QC + +The `qc/metrics_vs_depth/` directory is produced by `PLOT_METRICS_VS_DEPTH_QC` (in the `plotting_qc` subworkflow) and is always generated. It joins per-gene/sample average depth (from the `PLOTDEPTHSEXONSCONS` step) with: + +- raw mutation densities (`mutdensity/`) +- adjusted mutation densities (`mutdensity_adjusted/`) +- omega-globalloc estimates (`selection/omegagloballoc/`) + +Each combination yields a scatter PDF and a status TSV under `*.metrics_depth_qc/`, used to flag samples/genes whose metric values may be confounded by sequencing depth. + ## Mutational signatures We provide two different strategies for signature analysis. @@ -122,4 +201,32 @@ We provide two different strategies for signature analysis. - Using a Hierarchical Dirichlet Process algorithm developed by Nicola Robets and compacted by the McGranahan lab into a wrapped version. + - The outputs of the signature extraction process are then further processed downstream using SigProfilerAssignment to decompose the de novo signature and reassign mutational processes to samples. + +- Additionally we also output mutation count matrices that are ready to be run through [MSA](https://gitlab.com/s.senkin/MSA) which is another method for mutational signature attribution. + Additionally one could run SigProfilerExtractor on the data but this needs to be done externally. + +## Containers and reproducibility + +deepCSA defines container images directly in module files and `conf/modules.config`. For bbglab-maintained images (`bbglab/*`), Dockerfile recipes are tracked in the lab repository: https://github.com/bbglab/containers-recipes. External images (e.g., `ferriolcalvet/*`, `rblancomi/*`, `biocontainers/*`) should be mirrored locally if strict reproducibility is required. + +Key images used by the pipeline: + +| Component | Image | +| --- | --- | +| Core utilities | `docker.io/bbglab/deepcsa-core:0.1.0` | +| Panel BED tools | `docker.io/bbglab/deepcsa_bed:latest` | +| Omega | `docker.io/bbglab/omega:0.2.1` | +| Oncodrive3D | `docker.io/bbglab/oncodrive3d:1.0.5` | +| Oncodrive3D (ChimeraX plots) | `docker.io/spellegrini87/oncodrive3d_chimerax:latest` | +| OncodriveFML | `docker.io/ferriolcalvet/oncodrivefml:latest` | +| OncodriveCLUSTL | `docker.io/ferriolcalvet/oncodriveclustl:latest` | +| SigProfilerAssignment | `docker.io/ferriolcalvet/sigprofiler_assignment:1.1.3` | +| SigProfilerMatrixGenerator | `docker.io/ferriolcalvet/sigprofilermatrixgenerator:1.3.5` | +| mSigHdp (HDP) | `docker.io/ferriolcalvet/msighdp:latest` | +| bbgregressions | `docker.io/rblancomi/bbgregressions:dev` | +| Ensembl VEP | `biocontainers/ensembl-vep:111.0--pl5321h2a3209d_0` (version depends on `vep_cache_version`) | +| SAMtools | `biocontainers/samtools:1.18--h50ea8bc_1` | + +To override any image, set `process.container` or the relevant module label in your `nextflow.config`. diff --git a/docs/usage.md b/docs/usage.md index a2c69a1d..329f9056 100644 --- a/docs/usage.md +++ b/docs/usage.md @@ -6,6 +6,7 @@ - [Introduction](#introduction) - [How to run the pipeline](#how-to-run-the-pipeline) +- [Input scenarios](#input-scenarios) - [Samplesheet input](#samplesheet-input) - [Available genomes](#available-genomes) - [Proposed run modes](#proposed-run-modes) @@ -32,6 +33,10 @@ nextflow run bbglab/deepCSA --outdir -profile --input For more information on how to run Nextflow pipelines check a more detailed explanation [below](#running-the-pipeline) in this same document or check the [Nextflow](https://www.nextflow.io/docs/latest/index.html) or [nf-core](https://nf-co.re) community documentations. +## Input scenarios + +deepCSA accepts three different input combinations: per-sample VCF + BAM (default), per-sample VCF + a precomputed depths table, or a cohort-level MAF + a precomputed depths table. The sections below describe each piece in detail; for a concise summary of the three modes and when to use each, see [Input scenarios](input_scenarios.md). + ## Samplesheet input You will need to create a samplesheet with information about the samples you would like to analyse before running the pipeline. Use this parameter to specify its location. It has to be a comma-separated file with 3 columns, and a header row as shown in the examples below. @@ -56,6 +61,18 @@ sample2,sample2.high.filtered.vcf,sample2.sorted.bam An [example samplesheet](../assets/example_inputs/input_example.csv) has been provided with the pipeline. +### Input files from deepUMIcaller + +If your mutations were called with [deepUMIcaller](https://github.com/bbglab/deepUMIcaller), use the **final** per-sample outputs produced by that pipeline: + +- **VCF**: the filtered duplex VCF for each sample (commonly named `*.duplex.filtered.vcf` or `*.high.filtered.vcf`), typically found under the `mutations_vcf/` output directory. The VCF should be uncompressed. +- **BAM**: the **duplex-consensus** BAM used for calling those variants (commonly named `*.sorted.bam` in `sortbamduplexcons/`). This BAM must be aligned to the same reference genome as the VCF. + +Batch/sample guidance: + +- **One row per library/run**: if you have multiple sequencing libraries for the same biological sample and want to **aggregate** them, keep the same `sample` name across rows and list each matching VCF/BAM pair. +- **Separate batches**: if you want to compare batches or runs, keep distinct sample names and (optionally) add a `BATCH` or `RUN_ID` column in the feature groups table to stratify the analysis. + ## Available genomes deepCSA pipeline heavily relies on bgreference and bgdata tools so the use of this pipeline is limited to those genomes available in these packages. In particular, the default containers that are being used already have the hg38 and mm39 genomes cached, if you want to use any other genome, open an issue and we will address it as soon as we can. @@ -183,19 +200,49 @@ params { } ``` +### Feature toggles (turning steps on/off) + +All pipeline parameters are defined in `nextflow_schema.json` and exposed via `--help`. Most analysis steps can be enabled/disabled with boolean flags such as `mutationdensity`, `omega`, `oncodrive3d`, `signatures`, and `regressions`. If a step is turned off, its output directories are not produced. For a full list of parameters run: + +```bash +nextflow run bbglab/deepCSA --help +``` + ## Definition of structural parameters -- Container pulling (either prior to running the pipeline or directly as the pipeline runs) -- Generation of Oncodrive3D datasets (see: [Oncodrive3D repo datasets building process](https://github.com/bbglab/oncodrive3d?tab=readme-ov-file#building-datasets)) +Before running, ensure container images can be pulled (or are already available) and that all reference datasets are accessible from your execution environment. -- Download of additional specific datasets - - Ensembl VEP (see: [Ensembl VEP docs](https://www.ensembl.org/info/docs/tools/vep/script/vep_cache.html#cache)). Modify accordingly your `nextflow.config` vep parameters, `vep_cache`, `vep_cache_version`, etc. - - - CADD scores (see: [CADD downloads page](https://cadd.gs.washington.edu/download) "All possible SNVs of GRCh38/hg38" file) - - COSMIC signatures (i.e. [COSMIC signatures downloads page](https://cancer.sanger.ac.uk/signatures/downloads/) (select context size = 96 and your desired species of interest)) +### Data sources and reference downloads -- Provide custom domain definition file. - +- **Ensembl VEP cache** + Download from the [Ensembl VEP cache](https://www.ensembl.org/info/docs/tools/vep/script/vep_cache.html#cache) page and set `vep_cache`, `vep_cache_version`, `vep_species`, and `vep_genome` accordingly. + +- **CADD scores** + Use the ["All possible SNVs" GRCh38/hg38 file](https://cadd.gs.washington.edu/download). Provide both the compressed TSV (`cadd_scores`) and its tabix index (`cadd_scores_ind`). + +- **COSMIC signatures** + Download the SBS signatures (context size = 96) for your genome build from the [COSMIC signatures downloads page](https://cancer.sanger.ac.uk/signatures/downloads/). Set `cosmic_ref_signatures`. + +- **dNdScv reference data** + Download `covariates_hg19_hg38_epigenome_pcawg.rda` from the [dNdScv](https://github.com/im3sanger/dndscv) reference data (also mirrored by [IntOGen](https://intogen.org/download)). Set `dnds_covariates`. + Run the scripts for the generation of the dNdScv required reference input that you can find here: https://github.com/bbglab/deepCSA/tree/dev/assets/build_datasets/dndscv + +- **Oncodrive3D datasets** + Build datasets and annotations following the [Oncodrive3D dataset instructions](https://github.com/bbglab/oncodrive3d?tab=readme-ov-file#building-datasets). Provide the resulting `datasets3d` and `annotations3d` directories. + +- **Trinucleotide counts** + Provide a `wgs_trinuc_counts` file with the total count of each trinucleotide in your reference genome (see `assets/trinucleotide_counts/` for the expected format). + +- **DNA2PROTEINMAPPING GFF3 (optional)** + By default, the pipeline fetches a GFF3 file from Ensembl FTP at runtime. For local/offline use, download the matching file from + `https://ftp.ensembl.org/pub/release-/gff3//` + (e.g., `Homo_sapiens.GRCh38.111.gff3.gz`) and provide it via `gff3_file`. + +- **Domain definitions** + Supply a Pfam/InterPro domain file (see [file formatting](file_formatting.md#domain-definition-file)). + +- **NanoSeq masks (optional)** + See [Nanoseq genomic masks](#nanoseq-genomic-masks) below. ### Mandatory parameter configuration @@ -214,9 +261,16 @@ params { cadd_scores_ind = "CADD/v1.7/hg38/whole_genome_SNVs.tsv.gz.tbi" // dnds - dnds_ref_transcripts = "RefCDS_human_latest_intogen.rda" + // dnds_biomart_ref is a biomart TSV; deepCSA dynamically builds a per-run RefCDS_custom.rda + // by intersecting it with the panel BED (replaces the previously required static + // RefCDS_*.rda transcripts file). See assets/build_datasets/dndscv/instructions.txt + // for how to regenerate the biomart export. + dnds_biomart_ref = "biomart_export.tsv" dnds_covariates = "covariates_hg19_hg38_epigenome_pcawg.rda" + // GFF3 annotation for the genome assembly, consumed when building exon/domain panels + gff3_file = "Homo_sapiens.GRCh38.111.gff3.gz" + // oncodrive3d + fancy plots datasets3d = "oncodrive3d/datasets" annotations3d = "oncodrive3d/annotations" @@ -275,6 +329,15 @@ These files identify sites overlapping common SNPs and noisy or variable genomic - Nanoseq SNP: Common SNP positions that should be excluded from analysis - Nanoseq Noise: Regions with high noise or variability +Enable them with: + +```console +params { + nanoseq_snp = "SNP_GRCh38.wgns.bed.gz" + nanoseq_noise = "NOISE_GRCh38.wgns.bed.gz" +} +``` + Both files are available for GRCh38 at the [shared folder](https://drive.google.com/drive/folders/1wqkgpRTuf4EUhqCGSLA4fIg9qEEw3ZcL) from Iñigo Martincorena's group, at the Wellcome Sanger Institute. ## Additional customizable parameters @@ -301,6 +364,12 @@ This value is used for filtering the mutations by depth. Meaning that if a mutat This value is the less stringent depth threshold and is used in the first step of computing the positions that may be part of the so called "panels". This value indicates the minimum average depth at a given position for this position to be kept for the posterior depth analysis and definition on panels. The main use of this value should be to reduce the size of the files that are being processed afterwards. This can be set to 20 or more very safely. +### VAF-distortion filter + +- vaf_distortion_threshold = 3 + +Mutations whose ratio `VAF_AM / VAF` (all-molecules VAF over duplex VAF) exceeds this threshold are flagged as VAF-distorted during mutation filtering. Lower values are more conservative. + ### Using a precomputed depths table If you already have a precomputed table with per-position depths for your cohort (for example produced by a previous run or an external tool), you can instruct the pipeline to use that table instead of re-computing depths from the BAM files. This can save time and compute resources when depth computation has been performed once and re-used. @@ -322,6 +391,30 @@ Notes and requirements: - If your input.csv file contains `sample`, vcf and bam columns, the columns of the depths table have to be the same as the name of the BAM files of each sample in the input.csv file. - Make sure that you remove the column CONTEXT from the table in case you are starting with the all_samples individual depths table that is outputted by deepCSA. Check out the assets/useful_scripts/downsample_depths.ipynb file for an example on how to prepare the input for this parameter. +### Plotting controls + +If you are running with sample groups (see [Feature groups](file_formatting.md#feature-groups)), you can control whether plotting steps generate cohort-only outputs or include all group-level plots: + +```console +params { + plot_only_allsamples = true // only cohort-level plots +} +``` + +Set `plot_only_allsamples = false` to generate per-group plots alongside the cohort summaries. + +### Positive selection with non-protein-affecting profiles + +deepCSA can compute positive selection metrics using a **non-protein-affecting** mutational profile (synonymous + intronic + intergenic mutations). This is useful to compare selection metrics against a background that excludes protein-altering events. + +```console +params { + positive_selection_non_protein_affecting = true +} +``` + +When enabled, outputs will be labeled with the `.non_prot_aff` suffix in the corresponding selection directories (e.g., `omega/`, `omegagloballoc/`). + ## Custom mutation calls -- option 1 (building input VCFs and providing them via normal input) If you want to run deepCSA with your own mutation calls, this is also possible. Reasons behind this would be: diff --git a/modules/local/bbgtools/omega/estimator/main.nf b/modules/local/bbgtools/omega/estimator/main.nf index 9f85057b..44ebf696 100644 --- a/modules/local/bbgtools/omega/estimator/main.nf +++ b/modules/local/bbgtools/omega/estimator/main.nf @@ -11,6 +11,7 @@ process OMEGA_ESTIMATOR { tuple val(meta) , path(mutabilities_table), path(mutations_table), path(depths) tuple val(meta2), path(annotated_panel) path (genes_json) + path (impacts_json) output: tuple val(meta), path("output_*.tsv"), emit: results @@ -22,21 +23,10 @@ process OMEGA_ESTIMATOR { def prefix = task.ext.prefix ?: "" prefix = "${meta.id}${prefix}" """ - mkdir groups; mv ${genes_json} groups/group_genes.json - - cat > groups/group_impacts.json << EOF - { - "missense": ["missense"], - "nonsense": ["nonsense"], - "essential_splice": ["essential_splice"], - "truncating": ["nonsense", "essential_splice"], - "nonsynonymous_splice": ["missense", "nonsense", "essential_splice"] - } - EOF - + mv ${impacts_json} groups/group_impacts.json cat > groups/group_samples.json << EOF { "${meta.id}" : ["${meta.id}"] diff --git a/modules/local/bbgtools/oncodrive3d/plot/main.nf b/modules/local/bbgtools/oncodrive3d/plot/main.nf index 6337a514..ee8d79e3 100644 --- a/modules/local/bbgtools/oncodrive3d/plot/main.nf +++ b/modules/local/bbgtools/oncodrive3d/plot/main.nf @@ -2,7 +2,7 @@ process ONCODRIVE3D_PLOT { tag "$meta.id" label 'process_medium' - container 'docker.io/bbglab/oncodrive3d:1.0.5' + container 'docker.io/spellegrini87/oncodrive3d:1.0.9-light' input: @@ -11,13 +11,13 @@ process ONCODRIVE3D_PLOT { path(annotations) output: - tuple val(meta), path("**.summary_plot.png") , emit: summary_plot, optional: true - tuple val(meta), path("**.genes_plots/**.png") , emit: genes_plot, optional: true - tuple val(meta), path("**.associations_plots/**.logodds_plot.png") , emit: logodds_plot, optional: true - tuple val(meta), path("**.associations_plots/**.volcano_plot.png") , emit: volcano_plot, optional: true - tuple val(meta), path("**.associations_plots/**.volcano_plot_gene.png") , emit: volcano_plot_gene, optional: true - tuple val(meta), path("**.3d_clustering_pos.annotated.csv") , emit: pos_annotated_csv, optional: true - tuple val(meta), path("**plot_*.log") , emit: log + tuple val(meta), path("${meta.id}**.summary_plot.png") , emit: summary_plot, optional: true + tuple val(meta), path("${meta.id}**.genes_plots/**.png") , emit: genes_plot, optional: true + tuple val(meta), path("${meta.id}**.associations_plots/**.logodds_plot.png") , emit: logodds_plot, optional: true + tuple val(meta), path("${meta.id}**.associations_plots/**.volcano_plot.png") , emit: volcano_plot, optional: true + tuple val(meta), path("${meta.id}**.associations_plots/**.volcano_plot_gene.png") , emit: volcano_plot_gene, optional: true + tuple val(meta), path("${meta.id}**.3d_clustering_pos.annotated.csv") , emit: pos_annotated_csv, optional: true + tuple val(meta), path("${meta.id}**plot_*.log") , emit: log path "versions.yml" , topic: versions diff --git a/modules/local/bbgtools/oncodrive3d/plot_chimerax/main.nf b/modules/local/bbgtools/oncodrive3d/plot_chimerax/main.nf index 1cc74028..bdc24664 100644 --- a/modules/local/bbgtools/oncodrive3d/plot_chimerax/main.nf +++ b/modules/local/bbgtools/oncodrive3d/plot_chimerax/main.nf @@ -2,8 +2,7 @@ process ONCODRIVE3D_PLOT_CHIMERAX { tag "$meta.id" label 'process_medium' - // TODO pending to push the container somewhere and be able to retrieve it - container 'docker.io/spellegrini87/oncodrive3d_chimerax:latest' + container 'docker.io/spellegrini87/oncodrive3d:1.0.9-chimerax' input: @@ -11,10 +10,10 @@ process ONCODRIVE3D_PLOT_CHIMERAX { path(datasets) output: - tuple val(meta), path("**.chimerax/attributes/**.defattr") , emit: chimerax_defattr, optional: true - tuple val(meta), path("**.chimerax/plots/**.png") , emit: chimerax_plot, optional: true - tuple val(meta), path("**.log") , emit: log - path "versions.yml" , topic: versions + tuple val(meta), path("${meta.id}**.chimerax/attributes/**.defattr") , emit: chimerax_defattr, optional: true + tuple val(meta), path("${meta.id}**.chimerax/plots/**.png") , emit: chimerax_plot, optional: true + tuple val(meta), path("${meta.id}**.log") , emit: log + path "versions.yml" , topic: versions script: diff --git a/modules/local/bbgtools/oncodrive3d/run/main.nf b/modules/local/bbgtools/oncodrive3d/run/main.nf index f1f7f0e9..43752d2e 100644 --- a/modules/local/bbgtools/oncodrive3d/run/main.nf +++ b/modules/local/bbgtools/oncodrive3d/run/main.nf @@ -2,7 +2,7 @@ process ONCODRIVE3D_RUN { tag "$meta.id" label 'process_high' - container 'docker.io/bbglab/oncodrive3d:1.0.5' + container 'docker.io/spellegrini87/oncodrive3d:1.0.9-light' input: @@ -11,12 +11,12 @@ process ONCODRIVE3D_RUN { output: - tuple val(meta), path("**genes.csv") , emit: csv_genes - tuple val(meta), path("**pos.csv") , emit: csv_pos - tuple val(meta), path("**mutations.processed.tsv") , emit: mut_processed, optional: true - tuple val(meta), path("**miss_prob.processed.json") , emit: prob_processed, optional: true - tuple val(meta), path("**seq_df.processed.tsv") , emit: seq_processed, optional: true - tuple val(meta), path("**run_*.log") , emit: log + tuple val(meta), path("${meta.id}**genes.csv") , emit: csv_genes + tuple val(meta), path("${meta.id}**pos.csv") , emit: csv_pos + tuple val(meta), path("${meta.id}**mutations.processed.tsv") , emit: mut_processed, optional: true + tuple val(meta), path("${meta.id}**miss_prob.processed.json") , emit: prob_processed, optional: true + tuple val(meta), path("${meta.id}**seq_df.processed.tsv") , emit: seq_processed, optional: true + tuple val(meta), path("${meta.id}**run_*.log") , emit: log path "versions.yml" , topic: versions diff --git a/modules/local/bbgtools/sitecomparison/main.nf b/modules/local/bbgtools/sitecomparison/main.nf index 6b86839b..5813002d 100644 --- a/modules/local/bbgtools/sitecomparison/main.nf +++ b/modules/local/bbgtools/sitecomparison/main.nf @@ -19,12 +19,15 @@ process SITE_COMPARISON { def prefix = task.ext.prefix ?: "" prefix = "${meta.id}${prefix}" def size = task.ext.size ?: "all" // other options are 'site', 'aa_change', 'aa', '3aa', '3aa_rolling' // think if is worth having 'Naa', 'Naa_rolling' + def genes_subset = task.ext.genes_subset ?: "" + target_genes = genes_subset != "" ? "--genes ${genes_subset}": "" """ omega_comparison_per_site.py --mutations-file ${mutations} \\ --panel-file ${annotated_panel_richer} \\ --mutabilities-file ${mutabilities_per_site} \\ --size ${size} \\ - --output-prefix ${prefix} + --output-prefix ${prefix} \\ + ${target_genes} cat <<-END_VERSIONS > versions.yml "${task.process}": diff --git a/modules/local/computedepths/main.nf b/modules/local/computedepths/main.nf index 6279a636..efffa339 100644 --- a/modules/local/computedepths/main.nf +++ b/modules/local/computedepths/main.nf @@ -26,6 +26,7 @@ process COMPUTEDEPTHS { // positions with a mean depth above a given value // if the provided value is 0 this is not used def minimum_depth = task.ext.minimum_depth ? "| awk 'NR == 1 {print; next} {sum = 0; for (i=3; i<=NF; i++) sum += \$i; mean = sum / (NF - 2); if (mean >= ${task.ext.minimum_depth} ) print }'": "" + def remove_chrM = task.ext.remove_chrM ? "| egrep -v '^chrM'" : "" """ ls -1 *.bam > bam_files_list.txt; samtools \\ @@ -35,6 +36,7 @@ process COMPUTEDEPTHS { -@ $task.cpus \\ -f bam_files_list.txt \\ | tail -c +2 \\ + ${remove_chrM} \\ ${minimum_depth} \\ | gzip -c > ${prefix}.depths.tsv.gz; diff --git a/modules/local/concatprofiles/main.nf b/modules/local/concatprofiles/main.nf index 87bcad6c..3cb6c5f8 100644 --- a/modules/local/concatprofiles/main.nf +++ b/modules/local/concatprofiles/main.nf @@ -8,10 +8,10 @@ process CONCAT_PROFILES { path (all_groups) output: - path("*_heatmap*.png") , emit: heatmap - path("*_clustermap*.png") , emit: clustermap - path("*.cosine_similarity*.tsv") , emit: cosine_similarity - path("*.compiled_profiles.tsv") , emit: compiled_profiles + path("*_heatmap*.png") , optional:true, emit: heatmap + path("*_clustermap*.png") , optional:true, emit: clustermap + path("*.cosine_similarity*.tsv") , optional:true, emit: cosine_similarity + tuple val(meta), path("*.compiled_profiles.tsv") , emit: compiled_profiles path "versions.yml" , topic: versions script: diff --git a/modules/local/contamination/main.nf b/modules/local/contamination/main.nf index 7185b5d4..84dc1de0 100644 --- a/modules/local/contamination/main.nf +++ b/modules/local/contamination/main.nf @@ -16,10 +16,12 @@ process COMPUTE_CONTAMINATION { path "versions.yml" , topic: versions script: + def somatic_vaf_boundary = task.ext.germline_threshold ? "--somatic-vaf-boundary ${task.ext.germline_threshold}" : "" """ check_contamination.py \\ --maf_path ${maf} \\ - --somatic_maf ${somatic_maf} + --somatic_maf ${somatic_maf} \\ + ${somatic_vaf_boundary} cat <<-END_VERSIONS > versions.yml "${task.process}": diff --git a/modules/local/createpanels/sample/main.nf b/modules/local/createpanels/sample/main.nf deleted file mode 100644 index 02db32ef..00000000 --- a/modules/local/createpanels/sample/main.nf +++ /dev/null @@ -1,59 +0,0 @@ -process CREATESAMPLEPANELS { - tag "$meta.id" - label 'process_single' - label 'time_low' - - conda "bioconda::pybedtools=0.9.1--py38he0f268d_0" - container "${ workflow.containerEngine == 'singularity' && !task.ext.singularity_pull_docker_container ? - 'https://depot.galaxyproject.org/singularity/pybedtools:0.9.1--py38he0f268d_0' : - 'biocontainers/pybedtools:0.9.1--py38he0f268d_0' }" - - - input: - tuple val(meta) , path(compact_captured_panel_annotation) - tuple val(meta2), path(depths) - val(min_depth) - - output: - path("*.tsv") , emit: sample_specific_panel - path("*.bed") , emit: sample_specific_panel_bed - path "versions.yml" , topic: versions - - - script: - def prefix = task.ext.prefix ?: "" - prefix = "${meta.id}${prefix}" - // TODO min_depth should be provided from modules.config - """ - create_panel4sample.py \\ - --compact-annot-panel-path ${compact_captured_panel_annotation} \\ - --depths-path all_samples.depths.tsv.gz \\ - --panel-name ${prefix} \\ - --min-depth ${min_depth} - - for sample_panel in \$(ls *${prefix}.tsv ); do - bedtools merge \\ - -i <( - tail -n +2 \$sample_panel | \\ - awk -F'\\t' '{print \$1, \$2-1, \$2}' OFS='\\t' | uniq - ) > \${sample_panel%.tsv}.bed; - done - - cat <<-END_VERSIONS > versions.yml - "${task.process}": - python: \$(python --version | sed 's/Python //g') - END_VERSIONS - """ - - stub: - def prefix = task.ext.prefix ?: "TargetRegions" - """ - touch ${prefix}.tsv - touch ${prefix}.bed - - cat <<-END_VERSIONS > versions.yml - "${task.process}": - python: \$(python --version | sed 's/Python //g') - END_VERSIONS - """ -} diff --git a/modules/local/dna2protein/main.nf b/modules/local/dna2protein/main.nf index 9e67d616..f49d4fea 100644 --- a/modules/local/dna2protein/main.nf +++ b/modules/local/dna2protein/main.nf @@ -8,6 +8,7 @@ process DNA_2_PROTEIN_MAPPING { tuple val(meta) , path(mutations_file) tuple val(meta2), path(panel_file) tuple val(meta3), path(all_samples_depths) + path(gff3_file) output: @@ -22,6 +23,7 @@ process DNA_2_PROTEIN_MAPPING { def ensembl_release = "--ensembl-release \"${task.ext.ensembl_release}\"" def ensembl_species = "--ensembl-species \"${task.ext.ensembl_species}\"" def ensembl_genome = "--ensembl-genome \"${task.ext.ensembl_genome}\"" + def gff3_local = task.ext.gff3_provided ? "--gff3-file \"${gff3_file}\"" : "" """ cut -f 1,2,6 ${panel_file} | uniq > ${meta2.id}.panel.unique.tsv panels_computedna2protein.py \\ @@ -30,7 +32,9 @@ process DNA_2_PROTEIN_MAPPING { --depths-file ${all_samples_depths} \\ ${ensembl_release} \\ ${ensembl_species} \\ - ${ensembl_genome} + ${ensembl_genome} \\ + ${gff3_local} + cat <<-END_VERSIONS > versions.yml "${task.process}": diff --git a/modules/local/dnds/adaptpanelrefcds/main.nf b/modules/local/dnds/adaptpanelrefcds/main.nf new file mode 100644 index 00000000..3e3972c1 --- /dev/null +++ b/modules/local/dnds/adaptpanelrefcds/main.nf @@ -0,0 +1,43 @@ +process ADAPT_PANEL_REFCDS { + + tag "$meta.id" + label 'cpu_single_fixed' + label 'time_low' + label 'process_high_memory' + + label 'deepcsa_core' + + input: + path (biomart) + tuple val(meta), path(bedfile) + + output: + tuple val(meta), path("custom_filtered_biomart.tsv"), emit: filtered_biomart + tuple val(meta), path("splice_sites.tsv") , emit: splice_sites + path "versions.yml" , topic: versions + + script: + """ + dNdScv_panel_prep.py --bed ${bedfile} \\ + --genes ${biomart} \\ + --output custom_filtered_biomart.tsv \\ + --verbose \\ + -s splice_sites.tsv + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + dNdScv_panel_prep : \$(dNdScv_panel_prep.py --version) + END_VERSIONS + """ + + stub: + """ + touch custom_filtered_biomart.tsv + touch splice_sites.tsv + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + dNdScv_panel_prep : \$(dNdScv_panel_prep.py --version) + END_VERSIONS + """ +} \ No newline at end of file diff --git a/modules/local/dnds/buildref/main.nf b/modules/local/dnds/buildref/main.nf new file mode 100644 index 00000000..bc38b99e --- /dev/null +++ b/modules/local/dnds/buildref/main.nf @@ -0,0 +1,48 @@ +process BUILD_REFCDS { + + tag "$meta.id" + label 'cpu_single_fixed' + label 'time_low' + label 'process_high_memory' + + + container 'docker.io/ferriolcalvet/dnds:latest' + + input: + tuple val(meta) , path(biomart_cds) + path(reference_genome) + + + output: + tuple val(meta), path("RefCDS_custom.rda") , emit: ref_cds + path "versions.yml" , topic: versions + + when: + task.ext.when == null || task.ext.when + + script: + def args = task.ext.args ?: "" + def prefix = task.ext.prefix ?: "${meta.id}" + """ + Rscript -e "library(dndscv); buildref('${biomart_cds}', '${reference_genome}', outfile = 'RefCDS_custom.rda')" + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + R : 4.4.2 + dNdScv: 0.1.0 + END_VERSIONS + """ + + stub: + def args = task.ext.args ?: '' + def prefix = task.ext.prefix ?: "${meta.id}" + """ + touch RefCDS_custom.rda + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + R : 4.4.2 + dNdScv: 0.1.0 + END_VERSIONS + """ +} \ No newline at end of file diff --git a/modules/local/dnds/run/main.nf b/modules/local/dnds/run/main.nf index 8441f30e..5ea84dfb 100644 --- a/modules/local/dnds/run/main.nf +++ b/modules/local/dnds/run/main.nf @@ -13,20 +13,24 @@ process RUN_DNDS { path (covariates) output: - tuple val(meta), path("*.out.tsv*") , emit: results - path "versions.yml" , topic: versions + tuple val(meta), path("*.cv.tsv*") , emit: results_cv + tuple val(meta), path("*.loc.tsv*") , emit: results_local + tuple val(meta), path("*.globaldnds.tsv*") , emit: results_global + path "versions.yml" , topic: versions script: + def args = task.ext.args ?: "" def prefix = task.ext.prefix ?: "" prefix = "${meta.id}${prefix}" """ dNdS_run.R --inputfile ${mutations_table} \\ - --outputfile ${prefix}.out.tsv \\ + --outputprefix ${prefix} \\ --samplename ${prefix} \\ --covariates ${covariates} \\ --referencetranscripts ${ref_cds} \\ - --genedepth ${depths} + --genedepth ${depths} \\ + ${args} # --cores ${task.cpus} cat <<-END_VERSIONS > versions.yml "${task.process}": @@ -38,7 +42,9 @@ process RUN_DNDS { def prefix = task.ext.prefix ?: "" prefix = "${meta.id}${prefix}" """ - touch ${prefix}.out.tsv + touch ${prefix}.cv.tsv + touch ${prefix}.loc.tsv + touch ${prefix}.globaldnds.tsv cat <<-END_VERSIONS > versions.yml "${task.process}": @@ -47,17 +53,7 @@ process RUN_DNDS { """ } -// "--referencetranscripts" -// default="/workspace/projects/prominent/analysis/dNdScv/data/reference_files/RefCDS_human_latest_intogen.rda", -// --covariates -// "/workspace/projects/prominent/analysis/dNdScv/data/reference_files/covariates_hg19_hg38_epigenome_pcawg.rda", -// help="Human GRCh38 covariates file [default= %default]", metavar="character"), -// --genelist"), type="character", -// default=NULL, -// help="Gene list file [default= %default]", metavar="character"), -// --genedepth"), type="character", -// default=NULL, -// help="Gene depth file (2 columns: GENE\tAVG_DEPTH) [default= %default]", metavar="character"), + // --snvsonly"), type="logical", // default=FALSE, // help="Only use SNVs for the analysis [default= %default]", metavar="logical") diff --git a/modules/local/dnds_proxy/main.nf b/modules/local/dnds_proxy/main.nf new file mode 100644 index 00000000..5179ad4a --- /dev/null +++ b/modules/local/dnds_proxy/main.nf @@ -0,0 +1,45 @@ +process DNDS_PROXY { + tag "$meta.id" + label 'process_single' + + label 'deepcsa_core' + + input: + path(all_mutation_densities) + tuple val(meta), path(cohort_synonymous_mutdensities) + + output: + tuple val(meta), path("*.gene_mutdensities_n_dnds.tsv") , emit: mutdensity_with_dnds + path "versions.yml" , topic: versions + + + + script: + def prefix = task.ext.prefix ?: "" + prefix = "${meta.id}${prefix}" + def mode = task.ext.mode ?: "mutations" + """ + mut_density_adjusted_dnds.py \\ + --mutdensities ${all_mutation_densities} \\ + --cohort-syn-mutdensities ${cohort_synonymous_mutdensities} \\ + --output ${prefix}.gene_mutdensities_n_dnds.tsv \\ + --mode ${mode}; + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + python: \$(python --version | sed 's/Python //g') + END_VERSIONS + """ + + stub: + def prefix = task.ext.prefix ?: "all_samples" + """ + touch ${prefix}.gene_mutdensities_n_dnds.tsv + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + python: \$(python --version | sed 's/Python //g') + END_VERSIONS + """ + +} diff --git a/modules/local/group_genes/main.nf b/modules/local/group_genes/main.nf index cf1aee15..ec21bfa6 100644 --- a/modules/local/group_genes/main.nf +++ b/modules/local/group_genes/main.nf @@ -20,7 +20,7 @@ process GROUP_GENES { def grouping_info = "--table-file ${features_table} --separator ${separator} --output-json-groups pathway_groups_out.json" def custom_groups = task.ext.custom ? "${grouping_info}" : "" """ - awk 'NR>1 {print \$1}' ${mutations_table} | sort -u > gene_list.txt + awk 'NR>1 {print \$6}' ${mutations_table} | sort -u > gene_list.txt features_2group_genes.py \\ --panel-genes-file gene_list.txt \\ diff --git a/modules/local/hotspots_selection/main.nf b/modules/local/hotspots_selection/main.nf new file mode 100644 index 00000000..31ce3e3f --- /dev/null +++ b/modules/local/hotspots_selection/main.nf @@ -0,0 +1,46 @@ +process HOTSPOTS_SELECTION { + tag "$meta.id" + label 'cpu_single_fixed' + label 'time_low' + label 'process_high_memory' + + label 'deepcsa_core' + + input: + tuple val(meta), path(comparisons) + tuple val(meta2), path(annotated_panel_richer) + path(hotspots_file) + + output: + tuple val(meta), path("*.hotspots_selection.tsv.gz") , emit: selections + path "versions.yml" , topic: versions + + script: + def prefix = task.ext.prefix ?: "" + prefix = "${meta.id}${prefix}" + def comparison_args = comparisons.collect { "--comparisons ${it}" }.join(' ') + """ + compute_hotspots_selection.py \\ + ${comparison_args} \\ + --panel-file ${annotated_panel_richer} \\ + --hotspots-file ${hotspots_file} \\ + --output-prefix ${prefix} + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + python: \$(python --version | sed 's/Python //g') + END_VERSIONS + """ + + stub: + def prefix = task.ext.prefix ?: "" + prefix = "${meta.id}${prefix}" + """ + touch ${prefix}.site.hotspots_selection.tsv.gz + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + python: \$(python --version | sed 's/Python //g') + END_VERSIONS + """ +} diff --git a/modules/local/mut_density/main.nf b/modules/local/mut_density/adjusted/main.nf similarity index 95% rename from modules/local/mut_density/main.nf rename to modules/local/mut_density/adjusted/main.nf index 890b506f..eb855631 100644 --- a/modules/local/mut_density/main.nf +++ b/modules/local/mut_density/adjusted/main.nf @@ -8,7 +8,7 @@ process MUTATION_DENSITY { input: tuple val(meta), path(somatic_mutations_file), path(depths_file), path(mutability_file) tuple val(meta2), path(panel_file) - path(trinucleotide_counts_file) + path (trinucleotide_counts_file) output: @@ -20,7 +20,7 @@ process MUTATION_DENSITY { script: def sample_name = "${meta.id}" """ - mut_density.py \\ + mut_density_adjusted.py \\ --sample_name ${sample_name} \\ --depths_file ${depths_file} \\ --somatic_mutations_file ${somatic_mutations_file} \\ diff --git a/modules/local/computemutdensity/main.nf b/modules/local/mut_density/simple/main.nf similarity index 97% rename from modules/local/computemutdensity/main.nf rename to modules/local/mut_density/simple/main.nf index 2e13ee83..0fd0e9a5 100644 --- a/modules/local/computemutdensity/main.nf +++ b/modules/local/mut_density/simple/main.nf @@ -15,7 +15,7 @@ process MUTATION_DENSITY { def sample_name = "${meta.id}" def panel_version = task.ext.panel_version ?: "${meta2.id}" """ - compute_mutdensity.py \\ + mut_density_simple.py \\ --maf_path ${mutations} \\ --depths_path ${depth} \\ --annot_panel_path ${consensus_panel} \\ diff --git a/modules/local/mut_density/wgscaled/main.nf b/modules/local/mut_density/wgscaled/main.nf new file mode 100644 index 00000000..cc4b0c87 --- /dev/null +++ b/modules/local/mut_density/wgscaled/main.nf @@ -0,0 +1,56 @@ +process WG_SCALED_MUTATION_DENSITY { + + tag "${meta.id}" + + label 'cpu_single_fixed' + label 'time_low' + label 'process_high_memory' + + + container 'docker.io/ferriolcalvet/runningr:v1' + + input: + tuple val(meta), path(mutations_file), path(depths_file) + tuple val(meta2), path(consensus_bed_all) + path (wgs_counts) + + output: + tuple val(meta), path("*.tsv") , emit: adjusted_mutrate + path "versions.yml" , topic: versions + + + script: + def args = task.ext.args ?: "" + def prefix = task.ext.prefix ?: "" + prefix = "${meta.id}${prefix}" + def panel_version = task.ext.panel_version ?: "${meta2.id}" + """ + mutrate_genome_trinuc_corrected.R \\ + --samplename ${meta.id} \\ + --mutations ${mutations_file} \\ + --depths ${depths_file} \\ + --consensus_bed ${consensus_bed_all} \\ + --wgs_counts ${wgs_counts} \\ + --outputprefix ${prefix} \\ + --panel_version ${panel_version} + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + R: 4.3.1 + Rscript: 4.3.1 + END_VERSIONS + """ + + stub: + def prefix = task.ext.prefix ?: "" + prefix = "${meta.id}${prefix}" + """ + touch ${prefix}_mutrates_results.tsv + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + R: 4.3.1 + Rscript: 4.3.1 + END_VERSIONS + """ +} \ No newline at end of file diff --git a/modules/local/omega_multipletesting/main.nf b/modules/local/omega_multipletesting/main.nf new file mode 100644 index 00000000..2a0b9397 --- /dev/null +++ b/modules/local/omega_multipletesting/main.nf @@ -0,0 +1,42 @@ +process OMEGA_MULTITEST { + tag "all" + label 'process_low' + + label 'deepcsa_core' + + input: + path(omegas) + tuple path(samples_json), path(groups_json), path(all_groups_json) + + output: + path("${omegas.name}") , emit: corrected + path "versions.yml" , topic: versions + + script: + def output_name = "${omegas.name}" + def temp_name = "${omegas.baseName}.corrected.tsv" // avoid overwriting staged input + """ + omega_multiple_testing.py \\ + --omegas-file ${omegas} \\ + --samples-json ${samples_json} \\ + --groups-json ${groups_json} \\ + --output ${temp_name} + + mv ${temp_name} ${output_name} + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + python: \$(python --version | sed 's/Python //g') + END_VERSIONS + """ + + stub: + """ + touch ${omegas.name} + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + python: \$(python --version | sed 's/Python //g') + END_VERSIONS + """ +} diff --git a/modules/local/plot/depths_summary/main.nf b/modules/local/plot/depths_summary/main.nf index f9c475c4..7c7584aa 100644 --- a/modules/local/plot/depths_summary/main.nf +++ b/modules/local/plot/depths_summary/main.nf @@ -9,12 +9,12 @@ process PLOT_DEPTHS { tuple val(meta2), path(panel) output: - tuple val(meta), path("*.pdf") , emit: plots - tuple val(meta), path("*.avgdepth_per_sample.tsv") , emit: average_per_sample - tuple val(meta), path("*.avgdepth_per_gene.tsv") , emit: average_per_gene + tuple val(meta), path("*.pdf") , optional : true, emit: plots + tuple val(meta), path("*.avgdepth_per_sample.tsv") , emit: average_per_sample + tuple val(meta), path("*.avgdepth_per_gene.tsv") , emit: average_per_gene tuple val(meta), path("*.depth_per_gene_per_sample.tsv") , emit: average_per_gene_sample - tuple val(meta), path("*depth*.tsv") , emit: depths - path "versions.yml" , topic: versions + tuple val(meta), path("*depth*.tsv") , optional : true, emit: depths + path "versions.yml" , topic: versions diff --git a/modules/local/plot/interindividual_variability/main.nf b/modules/local/plot/interindividual_variability/main.nf index 34c801e2..50c83031 100644 --- a/modules/local/plot/interindividual_variability/main.nf +++ b/modules/local/plot/interindividual_variability/main.nf @@ -6,10 +6,11 @@ process PLOT_INTERINDIVIDUAL_VARIABILITY { label 'deepcsa_core' input: - path(samples_json) - path(all_groups_json) + path (samples_json) + path (all_groups_json) tuple val(meta), path(panel_file) - path(mutdensities_file) + path (mutdensities_file) + path (adjusted_mutdensities_file) output: path("**.pdf") , emit: plots @@ -23,6 +24,7 @@ process PLOT_INTERINDIVIDUAL_VARIABILITY { mkdir ${prefix}.variability_plots plot_explore_variability.py \\ --mutdensities ${mutdensities_file} \\ + --adjusted-mutdensities ${adjusted_mutdensities_file} \\ --panel-regions ${panel_file} \\ --outdir ${prefix}.variability_plots \\ --samples-json ${samples_json} \\ diff --git a/modules/local/plot/qc/annotate_omega/main.nf b/modules/local/plot/qc/annotate_omega/main.nf index c1aa8932..d3a60270 100644 --- a/modules/local/plot/qc/annotate_omega/main.nf +++ b/modules/local/plot/qc/annotate_omega/main.nf @@ -10,11 +10,12 @@ process ANNOTATE_OMEGA_QC { path (compiled_flagged_cases) output: - path("*flagged_annotated.tsv") , emit: all_omegas_annotated - path("*flagged.tsv") , optional: true, emit: flagged_cases - path("*.tsv") , optional: true, emit: files - path("*.png") , optional: true, emit: plots - path "versions.yml" , topic: versions + path("*flagged_annotated.tsv") , emit: all_omegas_annotated + path("debug.*flagged*.tsv") , optional: true, emit: flagged_cases + path("debug.syn_flagged_gene.tsv") , optional: true, emit: flagged_synonymous_cases + path("*.tsv") , optional: true, emit: files + path("*.png") , optional: true, emit: plots + path "versions.yml" , topic: versions script: """ @@ -35,6 +36,7 @@ process ANNOTATE_OMEGA_QC { def prefix = task.ext.prefix ?: "all_samples" """ touch ${prefix}.pdf + touch omega.flagged_annotated.tsv cat <<-END_VERSIONS > versions.yml "${task.process}": diff --git a/modules/local/plot/qc/metrics_vs_depth/main.nf b/modules/local/plot/qc/metrics_vs_depth/main.nf new file mode 100644 index 00000000..89229935 --- /dev/null +++ b/modules/local/plot/qc/metrics_vs_depth/main.nf @@ -0,0 +1,53 @@ +process PLOT_METRICS_VS_DEPTH_QC { + + tag "${group_name}" + label 'process_low' + + label 'deepcsa_core' + + input: + path (all_mutdensities) + path (depth_gene_sample) + path (groups_json) + val (group_name) + path (all_adjusted_mutdensities, stageAs: 'adjusted_mutdensities.tsv') + path (all_omegas_globalloc, stageAs: 'omega_gloc.tsv') + + output: + path("${group_name}.metrics_depth_qc/*.pdf"), optional: true , emit: plots + path("${group_name}.metrics_depth_qc/*.tsv"), optional: true , emit: tables + path "versions.yml" , topic: versions + + script: + def adjusted_arg = task.ext.all_adjusted_mutdensities ? "--adjusted-mutdensity-file ${all_adjusted_mutdensities}" : "" + def omega_arg = task.ext.all_omegas_globalloc ? "--omegas-file ${all_omegas_globalloc}" : "" + """ + mkdir ${group_name}.metrics_depth_qc + metrics_vs_depth_qc.py \\ + --mutdensity-file ${all_mutdensities} \\ + --depth-gene-sample-file ${depth_gene_sample} \\ + --group-definition ${groups_json} \\ + --group-name ${group_name} \\ + --output-dir ${group_name}.metrics_depth_qc \\ + ${adjusted_arg} \\ + ${omega_arg} + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + python: \$(python --version | sed 's/Python //g') + END_VERSIONS + """ + + stub: + """ + mkdir -p ${group_name}.metrics_depth_qc + touch ${group_name}.metrics_depth_qc/${group_name}.mutdensity.depth_scatter_per_sample.pdf + touch ${group_name}.metrics_depth_qc/${group_name}.mutdensity.depth_effect_summary.tsv + touch ${group_name}.metrics_depth_qc/${group_name}.metrics_vs_depth_qc.status.tsv + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + python: \$(python --version | sed 's/Python //g') + END_VERSIONS + """ +} diff --git a/modules/local/plot/qc/omega_vs_global_vs_dndscv/main.nf b/modules/local/plot/qc/omega_vs_global_vs_dndscv/main.nf new file mode 100644 index 00000000..59cfede1 --- /dev/null +++ b/modules/local/plot/qc/omega_vs_global_vs_dndscv/main.nf @@ -0,0 +1,47 @@ +process PLOT_OMEGA_VS_GLOBAL_VS_DNDSCV { + + tag "all_samples" + label 'process_low' + label 'deepcsa_core' + + input: + path (all_omegas) + path (all_omegas_globalloc) + path (dndscv_cv) + path (compiled_flagged) + path (groups_json) + + output: + path("**.pdf") , optional: true , emit: plots + path("**.tsv") , optional: true , emit: tables + path "versions.yml" , topic: versions + + script: + def dndscv_flag = params.dnds ? "--input-dndscv-file ${dndscv_cv}" : "" + """ + omega_vs_global_vs_dndscv_qc.py \\ + --input-omega-file ${all_omegas} \\ + --input-omegaglobal-file ${all_omegas_globalloc} \\ + ${dndscv_flag} \\ + --output-dir . \\ + --flagged-genes-omega ${compiled_flagged} \\ + --defined-groups ${groups_json} + + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + python: \$(python --version | sed 's/Python //g') + END_VERSIONS + """ + + stub: + def prefix = task.ext.prefix ?: "all_samples" + """ + touch ${prefix}.pdf + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + python: \$(python --version | sed 's/Python //g') + END_VERSIONS + """ +} diff --git a/modules/local/vaf_smoothing/main.nf b/modules/local/vaf_smoothing/main.nf new file mode 100644 index 00000000..28ddfbe9 --- /dev/null +++ b/modules/local/vaf_smoothing/main.nf @@ -0,0 +1,40 @@ +process VAF_SMOOTHING { + tag "groups" + label 'process_low' + + label 'deepcsa_core' + + input: + tuple val(meta) , path(all_mutations) + path (all_mutdensities) + tuple val(meta2), path(average_depth_sample) + + output: + path("*.tsv.gz") , emit: smoothed_vaf_tables + path("*.pdf") , optional: true, emit: plots + path "versions.yml" , topic: versions + + script: + """ + vaf_smoothing.py \\ + --mutations ${all_mutations} \\ + --mutdensities ${all_mutdensities} \\ + --depth-sample ${average_depth_sample} + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + python: \$(python --version | sed 's/Python //g') + END_VERSIONS + """ + + stub: + """ + touch vaf_pseudocounts_curves.pdf + touch mutations_with_smoothed_VAF.tsv.gz + + cat <<-END_VERSIONS > versions.yml + "${task.process}": + python: \$(python --version | sed 's/Python //g') + END_VERSIONS + """ +} \ No newline at end of file diff --git a/nextflow.config b/nextflow.config index c485f68b..6ef3ee92 100644 --- a/nextflow.config +++ b/nextflow.config @@ -30,6 +30,7 @@ params { use_custom_bedfile = false custom_bedfile = null use_custom_minimum_depth = 30 + remove_chrM = false hotspots_annotation = false hotspots_definition_file = '' @@ -47,7 +48,7 @@ params { oncodriveclustl = false oncodrive3d = false - o3d_raw_vep = false + o3d_raw_vep = true o3d_plot = false o3d_plot_chimerax = false @@ -96,7 +97,7 @@ params { bbgr_mode = "default" - filter_criteria = ["notcontains low_mappability", "notcontains not_covered", "notcontains no_pileup_support", "notcontains nanoseq_noise", "notcontains NM20", "notcontains cohort_n_rich", "notcontains n_rich", "notcontains cohort_n_rich_threshold"] + filter_criteria = ["notcontains low_mappability", "notcontains not_covered", "notcontains no_pileup_support", "notcontains AM_no_pileup_support", "notcontains nanoseq_noise", "notcontains NM20", "notcontains cohort_n_rich", "notcontains n_rich", "notcontains cohort_n_rich_threshold"] filter_criteria_somatic = ["notcontains nanoseq_snp", "notcontains gnomAD_SNP"] no_filter = false @@ -106,11 +107,12 @@ params { consensus_panel_min_depth = 200 consensus_compliance = 0.8 min_muts_per_sample = 0 - selected_genes = '' + selected_genes = null panel_with_canonical = true panel_sites_chunk_size = 1000000 // 0 means no chunking (default), set to positive integer to enable chunking germline_threshold = 0.3 + vaf_distortion_threshold = 3 mutation_depth_threshold = 100 gnomad_af_threshold = 1e-3 @@ -126,7 +128,7 @@ params { cadd_scores_ind = "CADD/v1.7/hg38/whole_genome_SNVs.tsv.gz.tbi" // dnds - dnds_ref_transcripts = "RefCDS_human_latest_intogen.rda" + dnds_biomart_ref = "homo_sapiens.v111.canonical.biomart.tsv" dnds_covariates = "covariates_hg19_hg38_epigenome_pcawg.rda" // oncodrive3d @@ -151,6 +153,7 @@ params { vep_out_format = "tab" vep_params = "--no_stats --cache --offline --symbol --protein --canonical --af_gnomadg --af_gnomade" vep_params_panel = "--no_stats --cache --offline --symbol --protein --canonical" + gff3_file = null } // Load default regressions parameters diff --git a/nextflow_schema.json b/nextflow_schema.json index 750561e6..e0541ff0 100644 --- a/nextflow_schema.json +++ b/nextflow_schema.json @@ -178,6 +178,13 @@ "default": 0, "help_text": "", "fa_icon": "far fa-file-code" + }, + "remove_chrM": { + "type": "boolean", + "description": "Whether to remove chrM from the depths table.", + "default": true, + "help_text": "If true, any positions on chrM will be removed from the computed depths table. This is useful if you want to exclude mitochondrial DNA from downstream analyses.", + "fa_icon": "fas fa-ban" } } }, @@ -337,6 +344,8 @@ }, "o3d_raw_vep": { "type": "boolean", + "default": true, + "hidden": true, "description": "Do you want to run oncodrive3d with the raw vep output as input?", "fa_icon": "fas fa-book" }, @@ -642,6 +651,13 @@ "fa_icon": "fas fa-book", "help_text": "" }, + "vaf_distortion_threshold": { + "type": "number", + "description": "VAF distortion threshold used to flag VAF-distorted mutations (VAF_AM / VAF).", + "default": 3, + "fa_icon": "fas fa-book", + "help_text": "" + }, "mutation_depth_threshold": { "type": "integer", "description": "Depth threshold for discarding mutations. Minimum to have a reliable estimation of the VAF.", @@ -776,7 +792,7 @@ "fa_icon": "far fa-file-code", "default": "CADD/v1.7/hg38/whole_genome_SNVs.tsv.gz.tbi" }, - "dnds_ref_transcripts": { + "dnds_biomart_ref": { "type": "string", "description": "Path to the dNdScv RefCDS file for all genes. See https://github.com/im3sanger/dndscv", "help_text": "Path to the dNdScv RefCDS file for all genes. See https://github.com/im3sanger/dndscv", @@ -860,6 +876,13 @@ "fa_icon": "fas fa-book", "help_text": "Define a place to store the Ensembl VEP cache.", "default": ".vep" + }, + "gff3_file": { + "type": "string", + "description": "Path to the GFF3 file for the genome assembly being used.", + "fa_icon": "far fa-file-code", + "help_text": "Path to the GFF3 file for the genome assembly being used. This is used for building the exon and domain panels. It should be compatible with the genome assembly and species being used for VEP annotation.", + "default": null } } }, diff --git a/subworkflows/local/adjmutdensity/main.nf b/subworkflows/local/adjmutdensity/main.nf index 1da8bcee..613ead87 100644 --- a/subworkflows/local/adjmutdensity/main.nf +++ b/subworkflows/local/adjmutdensity/main.nf @@ -2,7 +2,7 @@ include { TABIX_BGZIPTABIX_QUERY as QUERYMUTATIONS } from '../../.. include { SUBSET_MAF as SUBSETMUTDENSITYADJUSTED } from '../../../modules/local/subsetmaf/main' -include { MUTATION_DENSITY as MUTDENSITYADJ } from '../../../modules/local/mut_density/main' +include { MUTATION_DENSITY as MUTDENSITYADJ } from '../../../modules/local/mut_density/adjusted/main' workflow MUTATION_DENSITY { diff --git a/subworkflows/local/dnds/main.nf b/subworkflows/local/dnds/main.nf index edf8e534..0f137f5c 100644 --- a/subworkflows/local/dnds/main.nf +++ b/subworkflows/local/dnds/main.nf @@ -1,6 +1,9 @@ include { SUBSET_MAF as SUBSET_DNDS } from '../../../modules/local/subsetmaf/main' include { PREPROCESS_DNDS as PREPROCESSDEPTHS } from '../../../modules/local/dnds/preprocess/main' +include { ADAPT_PANEL_REFCDS as BIOMARTPANEL4REFCDS } from '../../../modules/local/dnds/adaptpanelrefcds/main' +include { BUILD_REFCDS as BUILDREFCDS } from '../../../modules/local/dnds/buildref/main' + include { RUN_DNDS as DNDSRUN } from '../../../modules/local/dnds/run/main' @@ -8,12 +11,15 @@ workflow DNDS { take: mutations depth + panel_bedfile panel + fasta + main: covariates = params.dnds_covariates ? channel.fromPath( params.dnds_covariates, checkIfExists: true).first() : channel.empty() - ref_trans = params.dnds_ref_transcripts ? channel.fromPath( params.dnds_ref_transcripts, checkIfExists: true).first() : channel.empty() + full_biomart_output = params.dnds_biomart_ref ? channel.fromPath( params.dnds_biomart_ref, checkIfExists: true).first() : channel.empty() SUBSET_DNDS(mutations) @@ -23,13 +29,24 @@ workflow DNDS { .join(PREPROCESSDEPTHS.out.depths) .set{ mutations_n_depth } - ref_trans.map{it -> [ ["id" : "global_RefCDS"] , it ] } - .set{ refcds_global } - DNDSRUN(mutations_n_depth, refcds_global, covariates) + BIOMARTPANEL4REFCDS(full_biomart_output, panel_bedfile) + + BUILDREFCDS(BIOMARTPANEL4REFCDS.out.filtered_biomart, fasta) + + DNDSRUN(mutations_n_depth, BUILDREFCDS.out.ref_cds, covariates) + + DNDSRUN.out.results_cv.map{ it -> it[1]}.flatten().set{ cv_results } + cv_results.collectFile(name: "all_dNdScv.cv.tsv", storeDir:"${params.outdir}/selection/dndscv/cv", skip: 1, keepHeader: true).set{ all_dndscv_results } + + DNDSRUN.out.results_global.map{ it -> it[1]}.flatten().set{ global_results } + global_results.collectFile(name: "all_dNdScv.global.tsv", storeDir:"${params.outdir}/selection/dndscv/persample", skip: 1, keepHeader: true).set{ all_dndscv_global_results } + + DNDSRUN.out.results_local.map{ it -> it[1]}.flatten().set{ local_results } + local_results.collectFile(name: "all_dNdScv.local.tsv", storeDir:"${params.outdir}/selection/dndscv/local", skip: 1, keepHeader: true).set{ all_dndscv_local_results } - // // uncomment whenever we can use the custom RefCDS file - // DNDSRUN(mutations_n_depth, ref_trans, covariates) + emit: + all_dndscv_results + all_dndscv_global_results + all_dndscv_local_results - // emit: - // dnds_values = DNDSRUN.out.dnds_values } diff --git a/subworkflows/local/enrichpanels/main.nf b/subworkflows/local/enrichpanels/main.nf index 076ff346..de56bcf1 100644 --- a/subworkflows/local/enrichpanels/main.nf +++ b/subworkflows/local/enrichpanels/main.nf @@ -26,7 +26,8 @@ workflow ENRICHPANELS { main: - DNA2PROTEINMAPPING(mutations, exons_consensus_panel, all_samples_depths) + gff3_channel = params.gff3_file ? file(params.gff3_file, checkIfExists: true) : file(params.input, checkIfExists: true) + DNA2PROTEINMAPPING(mutations, exons_consensus_panel, all_samples_depths, gff3_channel) // Create a channel for the domains file if autodomains is true domains_ch = params.autodomains ? domains_file : [] // .map{ it -> it[1]} : [] @@ -37,25 +38,25 @@ workflow ENRICHPANELS { if (params.create_subgenic_regions){ EXPANDREGIONSALL(all_consensus_panel, domains_ch, exons_ch, subgenic_ch) - all_expanded_panel = EXPANDREGIONSALL.out.panel_increased + all_expanded_panel = EXPANDREGIONSALL.out.panel_increased.first() EXPANDREGIONSNONPROT(nonprot_consensus_panel, domains_ch, exons_ch, subgenic_ch) - nonprot_expanded_panel = EXPANDREGIONSNONPROT.out.panel_increased + nonprot_expanded_panel = EXPANDREGIONSNONPROT.out.panel_increased.first() EXPANDREGIONSPROT(prot_consensus_panel, domains_ch, exons_ch, subgenic_ch) - prot_expanded_panel = EXPANDREGIONSPROT.out.panel_increased + prot_expanded_panel = EXPANDREGIONSPROT.out.panel_increased.first() EXPANDREGIONSSYNONYMOUS(synonymous_consensus_panel, domains_ch, exons_ch, subgenic_ch) - synonymous_expanded_panel = EXPANDREGIONSSYNONYMOUS.out.panel_increased + synonymous_expanded_panel = EXPANDREGIONSSYNONYMOUS.out.panel_increased.first() EXPANDREGIONSEXONS(exons_consensus_panel, domains_ch, exons_ch, subgenic_ch) - exons_expanded_panel = EXPANDREGIONSEXONS.out.panel_increased + exons_expanded_panel = EXPANDREGIONSEXONS.out.panel_increased.first() // all_json_subgenic = EXPANDREGIONSALL.out.new_regions_json // nonprot_json_subgenic = EXPANDREGIONSNONPROT.out.new_regions_json // prot_json_subgenic = EXPANDREGIONSPROT.out.new_regions_json // synonymous_json_subgenic = EXPANDREGIONSSYNONYMOUS.out.new_regions_json - exons_json_subgenic = EXPANDREGIONSEXONS.out.new_regions_json + exons_json_subgenic = EXPANDREGIONSEXONS.out.new_regions_json.first() } else { all_expanded_panel = all_consensus_panel @@ -69,13 +70,13 @@ workflow ENRICHPANELS { emit: - all_consensus_expanded_panel = all_expanded_panel.first() - nonprot_consensus_expanded_panel = nonprot_expanded_panel.first() - prot_consensus_expanded_panel = prot_expanded_panel.first() - synonymous_consensus_expanded_panel = synonymous_expanded_panel.first() - exons_consensus_expanded_panel = exons_expanded_panel.first() + all_consensus_expanded_panel = all_expanded_panel + nonprot_consensus_expanded_panel = nonprot_expanded_panel + prot_consensus_expanded_panel = prot_expanded_panel + synonymous_consensus_expanded_panel = synonymous_expanded_panel + exons_consensus_expanded_panel = exons_expanded_panel - exons_json_subgenic = exons_json_subgenic.first() + exons_json_subgenic = exons_json_subgenic dna2protein_mapping_depth_exons = DNA2PROTEINMAPPING.out.depths_exons_positions.first() dna2protein_mapping_panel_exons = DNA2PROTEINMAPPING.out.panel_exons_bed.first() diff --git a/subworkflows/local/mutationdensity/main.nf b/subworkflows/local/mutationdensity/main.nf index e9569ad9..a0f08f6b 100644 --- a/subworkflows/local/mutationdensity/main.nf +++ b/subworkflows/local/mutationdensity/main.nf @@ -1,8 +1,9 @@ -include { TABIX_BGZIPTABIX_QUERY as QUERYMUTATIONS } from '../../../modules/nf-core/tabix/bgziptabixquery/main' +include { TABIX_BGZIPTABIX_QUERY as QUERYMUTATIONS } from '../../../modules/nf-core/tabix/bgziptabixquery/main' -include { SUBSET_MAF as SUBSETMUTDENSITY } from '../../../modules/local/subsetmaf/main' +include { SUBSET_MAF as SUBSETMUTDENSITY } from '../../../modules/local/subsetmaf/main' -include { MUTATION_DENSITY as MUTDENSITY } from '../../../modules/local/computemutdensity/main' +include { MUTATION_DENSITY as MUTDENSITY } from '../../../modules/local/mut_density/simple/main' +include { WG_SCALED_MUTATION_DENSITY as WGSCALEDMUTDENSITY } from '../../../modules/local/mut_density/wgscaled/main' workflow MUTATION_DENSITY{ @@ -11,6 +12,9 @@ workflow MUTATION_DENSITY{ depth bedfile panel + samples_ch + wgs_trinucs + raw_depth main: @@ -25,7 +29,16 @@ workflow MUTATION_DENSITY{ MUTDENSITY(mutations_n_depth, panel) + QUERYMUTATIONS.out.subset + .map { mut -> tuple(mut[0].id, mut) } + .join(samples_ch) + .map { it -> it[1] } + .join(raw_depth) + .set{ mutations_n_depths_samples} + WGSCALEDMUTDENSITY(mutations_n_depths_samples, bedfile, wgs_trinucs) + emit: mutdensities = MUTDENSITY.out.mutdensities + mutdensities_wgs = WGSCALEDMUTDENSITY.out.adjusted_mutrate } diff --git a/subworkflows/local/mutationpreprocessing/main.nf b/subworkflows/local/mutationpreprocessing/main.nf index d5265230..789903e2 100644 --- a/subworkflows/local/mutationpreprocessing/main.nf +++ b/subworkflows/local/mutationpreprocessing/main.nf @@ -62,7 +62,7 @@ workflow MUTATION_PREPROCESSING { CUSTOMANNOTATION(SUMANNOTATION.out.tab, custom_annotation_tsv) summary_of_mutations = CUSTOMANNOTATION.out.mutations.first() } else { - summary_of_mutations = SUMANNOTATION.out.tab.first() + summary_of_mutations = SUMANNOTATION.out.tab } VCF2MAF(vcfs, summary_of_mutations) diff --git a/subworkflows/local/mutationprofile/main.nf b/subworkflows/local/mutationprofile/main.nf index 877fabb9..ee957be1 100644 --- a/subworkflows/local/mutationprofile/main.nf +++ b/subworkflows/local/mutationprofile/main.nf @@ -60,7 +60,7 @@ workflow MUTATIONAL_PROFILE { .concat(COMPUTEPROFILE.out.wgs_sigprofiler) .set{ sigprofiler_wgs } - compile_all_profiles = COMPUTEPROFILE.out.profile.map{ it -> it[1] }.collect().map { files -> [ [id:'all_profiles'], files ] } + compile_all_profiles = COMPUTEPROFILE.out.profile.map{ it -> it[1] }.collect().map { files -> [ [id:'all_samples'], files ] } CONCATPROFILES(compile_all_profiles, all_groups) compile_stabilities = COMPUTEPROFILE.out.profile_stability.map{ it -> it[1] }.collect().map { files -> [ [id:'all_stabilities'], files ] } diff --git a/subworkflows/local/omega/main.nf b/subworkflows/local/omega/main.nf index 95024b8c..9fa767b9 100644 --- a/subworkflows/local/omega/main.nf +++ b/subworkflows/local/omega/main.nf @@ -1,23 +1,26 @@ -include { SUBSET_MAF as SUBSETOMEGA } from '../../../modules/local/subsetmaf/main' -include { SUBSET_MAF as SUBSETOMEGAMULTI } from '../../../modules/local/subsetmaf/main' - -include { TABIX_BGZIPTABIX_QUERY as QUERYPANEL } from '../../../modules/nf-core/tabix/bgziptabixquery/main' - -include { OMEGA_PREPROCESS as PREPROCESSING } from '../../../modules/local/bbgtools/omega/preprocess/main' -include { GROUP_GENES as GROUPGENES } from '../../../modules/local/group_genes/main' -include { OMEGA_ESTIMATOR as ESTIMATOR } from '../../../modules/local/bbgtools/omega/estimator/main' -include { OMEGA_MUTABILITIES as ABSOLUTEMUTABILITIES } from '../../../modules/local/bbgtools/omega/mutabilities/main' -include { PLOT_OMEGA as PLOTOMEGA } from '../../../modules/local/plot/omega/main' -include { SITE_COMPARISON as SITECOMPARISON } from '../../../modules/local/bbgtools/sitecomparison/main' -include { SITE_COMPARISON as SITECOMPARISONMULTI } from '../../../modules/local/bbgtools/sitecomparison/main' -include { PLOT_OMEGASYN_QC as EVALOMEGAGLOCESTIMATION } from '../../../modules/local/plot/qc/globalloc_synonymous/main' - -include { OMEGA_PREPROCESS as PREPROCESSINGGLOBALLOC } from '../../../modules/local/bbgtools/omega/preprocess/main' -include { OMEGA_ESTIMATOR as ESTIMATORGLOBALLOC } from '../../../modules/local/bbgtools/omega/estimator/main' +include { SUBSET_MAF as SUBSETOMEGA } from '../../../modules/local/subsetmaf/main' +include { SUBSET_MAF as SUBSETOMEGAMULTI } from '../../../modules/local/subsetmaf/main' + +include { TABIX_BGZIPTABIX_QUERY as QUERYPANEL } from '../../../modules/nf-core/tabix/bgziptabixquery/main' + +include { OMEGA_PREPROCESS as PREPROCESSING } from '../../../modules/local/bbgtools/omega/preprocess/main' +include { GROUP_GENES as GROUPGENES } from '../../../modules/local/group_genes/main' +include { OMEGA_ESTIMATOR as ESTIMATOR } from '../../../modules/local/bbgtools/omega/estimator/main' +include { OMEGA_MUTABILITIES as ABSOLUTEMUTABILITIES } from '../../../modules/local/bbgtools/omega/mutabilities/main' +include { PLOT_OMEGA as PLOTOMEGA } from '../../../modules/local/plot/omega/main' +include { SITE_COMPARISON as SITECOMPARISON } from '../../../modules/local/bbgtools/sitecomparison/main' +include { SITE_COMPARISON as SITECOMPARISONMULTI } from '../../../modules/local/bbgtools/sitecomparison/main' +include { PLOT_OMEGASYN_QC as EVALOMEGAGLOCESTIMATION } from '../../../modules/local/plot/qc/globalloc_synonymous/main' +include { OMEGA_MULTITEST as OMEGAMULTIPLETEST } from '../../../modules/local/omega_multipletesting/main' + +include { OMEGA_PREPROCESS as PREPROCESSINGGLOBALLOC } from '../../../modules/local/bbgtools/omega/preprocess/main' +include { OMEGA_ESTIMATOR as ESTIMATORGLOBALLOC } from '../../../modules/local/bbgtools/omega/estimator/main' include { OMEGA_MUTABILITIES as ABSOLUTEMUTABILITIESGLOBALLOC } from '../../../modules/local/bbgtools/omega/mutabilities/main' include { PLOT_OMEGA as PLOTOMEGAGLOBALLOC } from '../../../modules/local/plot/omega/main' include { SITE_COMPARISON as SITECOMPARISONGLOBALLOC } from '../../../modules/local/bbgtools/sitecomparison/main' include { SITE_COMPARISON as SITECOMPARISONGLOBALLOCMULTI } from '../../../modules/local/bbgtools/sitecomparison/main' +include { OMEGA_MULTITEST as OMEGAMULTIPLETESTGLOBALLOC } from '../../../modules/local/omega_multipletesting/main' +include { HOTSPOTS_SELECTION as HOTSPOTSSELECTION } from '../../../modules/local/hotspots_selection/main' workflow OMEGA_ANALYSIS{ @@ -67,13 +70,10 @@ workflow OMEGA_ANALYSIS{ .join( depth ) .set{ preprocess_n_depths } - channel.of([ [ id: "all_samples" ] ]) - .join( PREPROCESSING.out.syn_muts_tsv ) - .set{ all_samples_muts } - - GROUPGENES(all_samples_muts, custom_gene_groups, json_subgenic) + GROUPGENES(expanded_panel, custom_gene_groups, json_subgenic) - ESTIMATOR( preprocess_n_depths, expanded_panel, GROUPGENES.out.json_genes.first()) + ESTIMATOR( preprocess_n_depths, expanded_panel, + GROUPGENES.out.json_genes.first(), "${projectDir}/assets/omega_consequences_groupings.json") if (params.omega_plot){ mutations @@ -110,7 +110,7 @@ workflow OMEGA_ANALYSIS{ PREPROCESSINGGLOBALLOC(muts_n_depths_n_profile, expanded_panel, - mutationdensities.first(), + mutationdensities, all_samples_mut_profile) PREPROCESSINGGLOBALLOC.out.mutabs_n_mutations_tsv @@ -119,12 +119,18 @@ workflow OMEGA_ANALYSIS{ ESTIMATORGLOBALLOC(preprocess_globalloc_n_depths, expanded_panel, - GROUPGENES.out.json_genes.first()) + GROUPGENES.out.json_genes.first(), + "${projectDir}/assets/omega_consequences_groupings.json") global_loc_results = ESTIMATORGLOBALLOC.out.results global_loc_results.map{ it -> it[1]}.flatten().set{ all_gloc_indv_results } - all_gloc_indv_results.collectFile(name: "all_omegas${suffix}_global_loc.tsv", storeDir:"${params.outdir}/selection/omegagloballoc", skip: 1, keepHeader: true).set{ all_gloc_results } + // Keep the concatenated file in the work directory; the corrected output is published below. + all_gloc_indv_results + .collectFile(name: "all_omegas${suffix}_global_loc.tsv", skip: 1, keepHeader: true) + .set{ all_gloc_results_raw } + OMEGAMULTIPLETESTGLOBALLOC(all_gloc_results_raw, grouping_defs) + all_gloc_results = OMEGAMULTIPLETESTGLOBALLOC.out.corrected PREPROCESSING.out.syn_muts_tsv.map{ it -> it[1]}.flatten().collect().set{ all_syn_muts } PREPROCESSINGGLOBALLOC.out.syn_muts_tsv.map{ it -> it[1]}.flatten().collect().set{ all_syn_muts_gloc } @@ -168,9 +174,25 @@ workflow OMEGA_ANALYSIS{ [meta, all_files] }.set{ site_comparison_results_flattened } + if (params.hotspots_annotation && params.hotspots_definition_file) { + hotspots_file = channel.fromPath(params.hotspots_definition_file, checkIfExists: true).first() + + HOTSPOTSSELECTION( + site_comparison_results, + QUERYPANEL.out.subset.first(), + hotspots_file + ) + // If needed, we can also collect or emit these results + } + ESTIMATOR.out.results.map{ it -> it[1]}.flatten().set{ all_indv_results } - all_indv_results.collectFile(name: "all_omegas${suffix}.tsv", storeDir:"${params.outdir}/selection/omega", skip: 1, keepHeader: true).set{ all_results } + // Keep the concatenated file in the work directory; the corrected output is published below. + all_indv_results + .collectFile(name: "all_omegas${suffix}.tsv", skip: 1, keepHeader: true) + .set{ all_results_raw } + OMEGAMULTIPLETEST(all_results_raw, grouping_defs) + all_results = OMEGAMULTIPLETEST.out.corrected emit: diff --git a/subworkflows/local/plotting_qc/main.nf b/subworkflows/local/plotting_qc/main.nf index 8c3fd8f5..d705cc17 100644 --- a/subworkflows/local/plotting_qc/main.nf +++ b/subworkflows/local/plotting_qc/main.nf @@ -1,21 +1,27 @@ -include { PLOT_MUTDENSITY_QC as PLOTMUTDENSITYQC } from '../../../modules/local/plot/qc/mutation_densities/main' -include { ANNOTATE_OMEGA_QC as APPLYOMEGAQC } from '../../../modules/local/plot/qc/annotate_omega/main' -include { PLOT_MUTATION_SPECIFIC as PLOTMUTATIONSPECIFIC } from '../../../modules/local/plot/qc/mutation_specific/main' +include { PLOT_MUTDENSITY_QC as PLOTMUTDENSITYQC } from '../../../modules/local/plot/qc/mutation_densities/main' +include { PLOT_METRICS_VS_DEPTH_QC as PLOTMETRICSVSDEPTHQC } from '../../../modules/local/plot/qc/metrics_vs_depth/main' +include { ANNOTATE_OMEGA_QC as APPLYOMEGAQC } from '../../../modules/local/plot/qc/annotate_omega/main' +include { PLOT_MUTATION_SPECIFIC as PLOTMUTATIONSPECIFIC } from '../../../modules/local/plot/qc/mutation_specific/main' +include { PLOT_OMEGA_VS_GLOBAL_VS_DNDSCV as PLOTOMEGAVSGLOBALVSDNDSCV } from '../../../modules/local/plot/qc/omega_vs_global_vs_dndscv/main' workflow PLOTTING_QC { take: all_mutations - // positive_selection_results_ready all_mutdensities - // all_samples_depth - // all_groups + all_adjusted_mutdensities + all_omegas_globalloc + average_depth_gene_sample all_omegas panel groups_definition group_name + dndscv_cv + groups_only_definition + // all_samples_depth + // all_groups // full_panel_rich // seqinfo_df // domain_df @@ -24,36 +30,46 @@ workflow PLOTTING_QC { main: + dndscv_channel = params.dnds ? dndscv_cv : channel.value(file("${projectDir}/assets/placeholder_no_file.tsv", checkIfExists: true)) + // Channel.of([ [ id: "all_samples" ] ]) // .join( all_mutations ) // .set{ mutations } - PLOTMUTATIONSPECIFIC(all_mutations) - - - // pdb_tool_df = params.annotations3d - // ? channel.fromPath( "${params.annotations3d}/pdb_tool_df.tsv", checkIfExists: true).first() - // : channel.empty() - // plotting only for the entire cohort group // channel.of([ [ id: "all_samples" ] ]) // .join( positive_selection_results_ready ) // .set{ all_samples_results } + PLOTMUTATIONSPECIFIC(all_mutations) + PLOTMUTDENSITYQC(all_mutdensities, panel, groups_definition, group_name) - // mutation density per gene cohort-level - // mutation density per gene & sample - // synonymous - // non-protein-affecting - // pending: - // protein-affecting - // truncating - // missense + + PLOTMETRICSVSDEPTHQC( + all_mutdensities, + average_depth_gene_sample.map { it -> it[1] }, + groups_definition, + group_name, + all_adjusted_mutdensities, + all_omegas_globalloc + ) + APPLYOMEGAQC(all_omegas, PLOTMUTDENSITYQC.out.compiled_flagged.collect()) + // Run omega qc script independently of dndscv output (handled in the script) + PLOTOMEGAVSGLOBALVSDNDSCV( + all_omegas, + all_omegas_globalloc, + dndscv_channel, + APPLYOMEGAQC.out.flagged_synonymous_cases, + groups_only_definition) + emit: - mutdensity_plots = PLOTMUTDENSITYQC.out.plots - flagged_omegas = APPLYOMEGAQC.out.all_omegas_annotated + mutdensity_plots = PLOTMUTDENSITYQC.out.plots + metrics_vs_depth_plots = PLOTMETRICSVSDEPTHQC.out.plots + metrics_vs_depth_tables = PLOTMETRICSVSDEPTHQC.out.tables + flagged_omegas = APPLYOMEGAQC.out.all_omegas_annotated + omega_vs_global_vs_dndscv_plots = PLOTOMEGAVSGLOBALVSDNDSCV.out.plots } diff --git a/subworkflows/local/plottingsummary/main.nf b/subworkflows/local/plottingsummary/main.nf index afb4d9dc..70c3f3be 100644 --- a/subworkflows/local/plottingsummary/main.nf +++ b/subworkflows/local/plottingsummary/main.nf @@ -14,6 +14,7 @@ workflow PLOTTING_SUMMARY { positive_selection_results_ready all_mutations all_mutdensities + all_mutdensities_adjusted site_comparison all_samples_depth samples @@ -80,7 +81,7 @@ workflow PLOTTING_SUMMARY { // ? plot saturation kinetics curves - PLOTINTERINDIVIDUALVARIABILITY(samples, all_groups, panel, all_mutdensities) + PLOTINTERINDIVIDUALVARIABILITY(samples, all_groups, panel, all_mutdensities, all_mutdensities_adjusted) // heatmaps: // mutations per gene/sample (total, SNV only, INDEL only, per consequence type) // driver mutations per gene/sample diff --git a/tests/README.md b/tests/README.md index bacb8a20..40115af4 100644 --- a/tests/README.md +++ b/tests/README.md @@ -68,7 +68,7 @@ nf-test test tests/deepcsa.nf.test --tag omega --update-snapshot - `mutational_profile/` and `omega/` directories exist - `omega/all_omegas.tsv` exists - `oncodrivefml/`, `oncodrive3d/` directories do **not** exist - - Structural checks on `all_omegas.tsv`: header contains `gene`, `sample`, `dnds`; all rows contain same columns, all samples are present. +- Structural checks on `all_omegas.tsv`: header contains `gene`, `sample`, `dnds`, `pvalue_adj`; all rows contain same columns, all samples are present. - Snapshot of `mutational_profile/all_samples.all.profile.tsv` (MD5) - `input_maf_validation` (Parameter validation): @@ -142,7 +142,7 @@ The parameters you will need to provide are: | `cadd_scores` / `cadd_scores_ind` | CADD scores TSV + index | | `cosmic_ref_signatures` | COSMIC SBS signatures file | | `nanoseq_snp` / `nanoseq_noise` | NanoSeq masking BED files | -| `dnds_ref_transcripts` / `dnds_covariates` | dNdScv reference files | +| `dnds_biomart_ref` / `dnds_covariates` | dNdScv reference files | | `datasets3d` / `annotations3d` | Oncodrive3D datasets | | `singularity.cacheDir` / `singularity.libraryDir` | Singularity image cache | diff --git a/tests/deepcsa.nf.test b/tests/deepcsa.nf.test index 1d34c44e..023f83f9 100644 --- a/tests/deepcsa.nf.test +++ b/tests/deepcsa.nf.test @@ -183,6 +183,7 @@ nextflow_pipeline { assert header.contains("gene") : "Omega output should contain 'gene' column" assert header.contains("sample") : "Omega output should contain 'sample' column" assert header.contains("dnds") : "Omega output should contain 'dnds' column" + assert header.contains("pvalue_adj") : "Omega output should contain 'pvalue_adj' column" assert lines.size() > 1 : "Omega output should have at least one data row" // Structural integrity: every row must have the same number of columns @@ -228,21 +229,28 @@ nextflow_pipeline { //TODO Include omega output snapshot when stable // Filter out empty lines and the header (collectFile keepHeader is non-deterministic) - def expectedHeader = "gene\tsample\timpact\tmutations\tdnds\tpvalue\tlower\tupper" - def dataLines = lines.findAll { it && it != expectedHeader } + def dataLines = lines.drop(1).findAll { it } assert dataLines.size() == 252 : "Omega output should contain 252 data rows, found ${dataLines.size()}" // Sort by key columns (gene, sample, impact) to avoid floating-point differences affecting order // Round numeric columns (dnds, pvalue, lower, upper) to 2 decimals for deterministic comparison + def snapshotColumns = ["gene", "sample", "impact", "mutations", "dnds", "pvalue", "lower", "upper"] + def snapshotIndices = snapshotColumns.collect { column -> + header.findIndexOf { it == column } + } + assert snapshotIndices.every { it >= 0 } : "Omega output should contain snapshot columns: ${snapshotColumns}" + def sortedRounded = filteredRows.sort { line -> def cols = line.split('\t') "${cols[0]}\t${cols[1]}\t${cols[2]}" }.collect { line -> - def cols = line.split('\t') - (4..7).each { i -> cols[i] = String.format("%.2f", cols[i] as Double) } - cols.join('\t') + // Use -1 to preserve trailing empty columns. + def cols = line.split('\t', -1) + def snapshotCols = snapshotIndices.collect { idx -> cols[idx] } + (4..7).each { i -> snapshotCols[i] = String.format("%.2f", snapshotCols[i] as Double) } + snapshotCols.join('\t') } - def sortedContent = [expectedHeader] + sortedRounded + def sortedContent = [snapshotColumns.join('\t')] + sortedRounded // Snapshot both files assert snapshot(path("${params.outdir}/mutational_profile/all_samples.all.profile.tsv")).match("mutational_profile") diff --git a/workflows/deepcsa.nf b/workflows/deepcsa.nf index a180ea2d..6d12580a 100644 --- a/workflows/deepcsa.nf +++ b/workflows/deepcsa.nf @@ -96,25 +96,29 @@ include { CUSTOM_DUMPSOFTWAREVERSIONS } from '../modules/n ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ */ -include { MAF_2_VCF as INPUTMAF2VCF } from '../modules/local/maf2vcf/main' -include { TABLE_2_GROUP as TABLE2GROUP } from '../modules/local/table2groups/main' -include { ANNOTATE_DEPTHS as ANNOTATEDEPTHS } from '../modules/local/annotatedepth/main' -include { DOWNSAMPLE_DEPTHS as DOWNSAMPLEDEPTHS } from '../modules/local/downsample/depths/main' +include { MAF_2_VCF as INPUTMAF2VCF } from '../modules/local/maf2vcf/main' +include { TABLE_2_GROUP as TABLE2GROUP } from '../modules/local/table2groups/main' +include { ANNOTATE_DEPTHS as ANNOTATEDEPTHS } from '../modules/local/annotatedepth/main' +include { DOWNSAMPLE_DEPTHS as DOWNSAMPLEDEPTHS } from '../modules/local/downsample/depths/main' +include { DOWNSAMPLE_DEPTHS as DOWNSAMPLEDEPTHSALLSAMPLES } from '../modules/local/downsample/depths/main' -include { TABIX_BGZIPTABIX_QUERY as QUERYMUTATIONSEXONS } from '../modules/nf-core/tabix/bgziptabixquery/main' +include { TABIX_BGZIPTABIX_QUERY as QUERYMUTATIONSEXONS } from '../modules/nf-core/tabix/bgziptabixquery/main' -include { ANALYZE_DEPTHS_GROUPS as ANALYZEDEPTHSGROUPS } from '../modules/local/analyzedepths/main' +include { ANALYZE_DEPTHS_GROUPS as ANALYZEDEPTHSGROUPS } from '../modules/local/analyzedepths/main' -include { SELECT_MUTDENSITIES as SYNMUTDENSITY } from '../modules/local/select_mutdensity/main' -include { SELECT_MUTDENSITIES as SYNMUTREADSDENSITY } from '../modules/local/select_mutdensity/main' +include { VAF_SMOOTHING as VAFSMOOTHING } from '../modules/local/vaf_smoothing/main' -include { DNA_2_PROTEIN_MAPPING as DNA2PROTEINMAPPING } from '../modules/local/dna2protein/main' +include { SELECT_MUTDENSITIES as SYNMUTDENSITY } from '../modules/local/select_mutdensity/main' +include { SELECT_MUTDENSITIES as SYNMUTREADSDENSITY } from '../modules/local/select_mutdensity/main' +include { SELECT_MUTDENSITIES as UPDSYNMUTDENSITY } from '../modules/local/select_mutdensity/main' +include { SELECT_MUTDENSITIES as UPDSYNMUTREADSDENSITY } from '../modules/local/select_mutdensity/main' +include { DNDS_PROXY as DNDSPROXY } from '../modules/local/dnds_proxy/main' -include { MAF_2_VCF as MAF2VCF } from '../modules/local/maf2vcf/main' -include { SIGPROFILER_MATRIXGENERATOR as SIGPROMATRIXGENERATOR } from '../modules/local/signatures/sigprofiler/matrixgenerator/main' -include { SIGPROFILERASSIGNMENT_COSMIC_FIT as SIGPROFILERASSIGNMENTINDELS } from '../modules/local/signatures/sigprofiler/assignment/cosmic_fit/main' +include { MAF_2_VCF as MAF2VCF } from '../modules/local/maf2vcf/main' +include { SIGPROFILER_MATRIXGENERATOR as SIGPROMATRIXGENERATOR } from '../modules/local/signatures/sigprofiler/matrixgenerator/main' +include { SIGPROFILERASSIGNMENT_COSMIC_FIT as SIGPROFILERASSIGNMENTINDELS } from '../modules/local/signatures/sigprofiler/assignment/cosmic_fit/main' -include { MUTATIONS_2_SIGNATURES as MUTS2SIGS } from '../modules/local/mutations2sbs/main' +include { MUTATIONS_2_SIGNATURES as MUTS2SIGS } from '../modules/local/mutations2sbs/main' /* ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ @@ -156,9 +160,11 @@ workflow DEEPCSA { site_comparison_results = channel.empty() all_compiled_omegas = channel.empty() - all_compiled_omegasgloballoc = channel.empty() + all_compiled_omegasgloballoc = channel.value(file("${projectDir}/assets/placeholder_no_file.tsv", checkIfExists: true)) all_mutdensities_file = channel.empty() + all_adjusted_mutdensities_file = channel.value(file("${projectDir}/assets/placeholder_no_file.tsv", checkIfExists: true)) all_compiled_stabilities = channel.empty() + dndscv_table = channel.empty() // if the user wants to use custom gene groups, import the gene groups table // otherwise I am using the input csv as a dummy value channel @@ -175,6 +181,7 @@ workflow DEEPCSA { // Initialize booleans based on user params def run_mutabilities = (params.oncodrivefml || params.oncodriveclustl || params.oncodrive3d) def run_mutdensity = (params.mutationdensity || params.omega) + def run_profile_all = (params.profileall || run_mutabilities || run_mutdensity || params.omega) // Validate input_maf usage: it requires use_custom_depths to be enabled if ( params.input_maf && !params.use_custom_depths ) { @@ -218,6 +225,12 @@ workflow DEEPCSA { grouping_definitions = TABLE2GROUP.out.json_samples.concat(TABLE2GROUP.out.json_groups).concat(TABLE2GROUP.out.json_allgroups).collect() // Load group keys from JSON file in 'groups' channel + TABLE2GROUP.out.json_samples.map { json_path -> + def json = file(json_path).text + groovy.json.JsonSlurper.newInstance().parseText(json).keySet() + }.flatten().unique() + .set { samples_keys_ch } // this is a channel that contains only the group names as elements of the channel + TABLE2GROUP.out.json_groups.map { json_path -> def json = file(json_path).text groovy.json.JsonSlurper.newInstance().parseText(json).keySet() @@ -258,8 +271,12 @@ workflow DEEPCSA { if (params.downsample ){ DOWNSAMPLEDEPTHS(annotated_depths_full) annotated_depths = DOWNSAMPLEDEPTHS.out.downsampled_depths + + DOWNSAMPLEDEPTHSALLSAMPLES(ANNOTATEDEPTHS.out.all_samples_depths) + all_samples_indv_annotated_depths = DOWNSAMPLEDEPTHSALLSAMPLES.out.downsampled_depths } else { annotated_depths = annotated_depths_full + all_samples_indv_annotated_depths = ANNOTATEDEPTHS.out.all_samples_depths } if (params.plot_depths){ @@ -297,40 +314,13 @@ workflow DEEPCSA { DEPTHSSYNONYMOUSCONS(annotated_depths, CREATEPANELS.out.synonymous_consensus_bed) } - if (run_mutdensity){ - // Mutation Density - MUTDENSITYALL(somatic_mutations, DEPTHSALLCONS.out.subset, CREATEPANELS.out.all_consensus_bed, ENRICHPANELS.out.all_consensus_expanded_panel) - MUTDENSITYPROT(somatic_mutations, DEPTHSPROTCONS.out.subset, CREATEPANELS.out.prot_consensus_bed, ENRICHPANELS.out.prot_consensus_expanded_panel) - MUTDENSITYNONPROT(somatic_mutations, DEPTHSNONPROTCONS.out.subset, CREATEPANELS.out.nonprot_consensus_bed, ENRICHPANELS.out.nonprot_consensus_expanded_panel) - MUTDENSITYSYNONYMOUS(somatic_mutations, DEPTHSSYNONYMOUSCONS.out.subset, CREATEPANELS.out.synonymous_consensus_bed, ENRICHPANELS.out.synonymous_consensus_expanded_panel) - - channel.of([ [ id: "all_samples" ] ]) - .join( MUTDENSITYSYNONYMOUS.out.mutdensities ) - .set{ all_samples_syn_mutdensity } - - SYNMUTDENSITY(all_samples_syn_mutdensity) - - SYNMUTREADSDENSITY(all_samples_syn_mutdensity) - - - // Concatenate all outputs into a single file - channel.empty() - .concat(MUTDENSITYALL.out.mutdensities.map{ it -> it[1]}.flatten()) - .concat(MUTDENSITYPROT.out.mutdensities.map{ it -> it[1]}.flatten()) - .concat(MUTDENSITYNONPROT.out.mutdensities.map{ it -> it[1]}.flatten()) - .concat(MUTDENSITYSYNONYMOUS.out.mutdensities.map{ it -> it[1]}.flatten()) - .set{ all_mutdensities } - all_mutdensities.collectFile(name: "all_mutdensities.tsv", storeDir:"${params.outdir}/mutdensity", skip: 1, keepHeader: true).set{ all_mutdensities_file } - - } - // Intersect BED of all sites with somatic mutations to keep only those mutations in the exons consensus panel QUERYMUTATIONSEXONS(somatic_mutations, CREATEPANELS.out.exons_consensus_bed) mutations_in_exons = QUERYMUTATIONSEXONS.out.subset // Mutational profile - if ( params.profileall || run_mutabilities || params.omega ){ + if ( run_profile_all ){ MUTPROFILEALL(somatic_mutations, DEPTHSALLCONS.out.subset, CREATEPANELS.out.all_consensus_bed, wgs_trinucs, TABLE2GROUP.out.json_allgroups) all_compiled_stabilities = all_compiled_stabilities.concat(MUTPROFILEALL.out.profile_stabilities.map{ it -> it[1] }) if (run_mutdensity){ @@ -342,11 +332,21 @@ workflow DEEPCSA { // Concatenate all outputs into a single file MUTDENSITYADJUSTED.out.mutdensities.map{ it -> it[1]}.flatten() .set{ all_adjusted_mutdensities } - all_adjusted_mutdensities.collectFile(name: "all_adjusted_mutdensities.tsv", storeDir:"${params.outdir}/mutdensity_adjusted", skip: 1, keepHeader: true) + all_adjusted_mutdensities.collectFile(name: "all_adjusted_mutdensities.tsv", storeDir:"${params.outdir}/mutdensity_adjusted", skip: 1, keepHeader: true).first().set{ all_adjusted_mutdensities_file } MUTDENSITYADJUSTED.out.mutdensities_flat.map{ it -> it[1]}.flatten() .set{ all_adjusted_mutdensities_flat } all_adjusted_mutdensities_flat.collectFile(name: "all_adjusted_mutdensities_flat.tsv", storeDir:"${params.outdir}/mutdensity_adjusted", skip: 1, keepHeader: true) + + channel.of([ [ id: "all_samples" ] ]) + .join( MUTDENSITYADJUSTED.out.mutdensities ) + .set{ all_samples_adj_mutdensity } + + UPDSYNMUTDENSITY(all_samples_adj_mutdensity) + + // UPDSYNMUTREADSDENSITY(all_samples_adj_mutdensity) + + DNDSPROXY(all_adjusted_mutdensities_file, UPDSYNMUTDENSITY.out.mutdensity.first()) } } if (params.profilenonprot){ @@ -365,14 +365,54 @@ workflow DEEPCSA { all_compiled_stabilities.flatten().collectFile(name: "all_profile_stabilities.tsv", storeDir:"${params.outdir}/mutational_profile", skip: 1, keepHeader: true) + if (run_mutdensity){ + // Mutation Density + MUTDENSITYALL(somatic_mutations, DEPTHSALLCONS.out.subset, CREATEPANELS.out.all_consensus_bed, ENRICHPANELS.out.all_consensus_expanded_panel, samples_keys_ch, wgs_trinucs, annotated_depths) + MUTDENSITYPROT(somatic_mutations, DEPTHSPROTCONS.out.subset, CREATEPANELS.out.prot_consensus_bed, ENRICHPANELS.out.prot_consensus_expanded_panel, samples_keys_ch, wgs_trinucs, annotated_depths) + MUTDENSITYNONPROT(somatic_mutations, DEPTHSNONPROTCONS.out.subset, CREATEPANELS.out.nonprot_consensus_bed, ENRICHPANELS.out.nonprot_consensus_expanded_panel, samples_keys_ch, wgs_trinucs, annotated_depths) + MUTDENSITYSYNONYMOUS(somatic_mutations, DEPTHSSYNONYMOUSCONS.out.subset, CREATEPANELS.out.synonymous_consensus_bed, ENRICHPANELS.out.synonymous_consensus_expanded_panel, samples_keys_ch, wgs_trinucs, annotated_depths) + + // Concatenate all outputs into a single file + channel.empty() + .concat(MUTDENSITYALL.out.mutdensities.map{ it -> it[1]}.flatten()) + .concat(MUTDENSITYPROT.out.mutdensities.map{ it -> it[1]}.flatten()) + .concat(MUTDENSITYNONPROT.out.mutdensities.map{ it -> it[1]}.flatten()) + .concat(MUTDENSITYSYNONYMOUS.out.mutdensities.map{ it -> it[1]}.flatten()) + .set{ all_mutdensities } + all_mutdensities.collectFile(name: "all_mutdensities.tsv", storeDir:"${params.outdir}/mutdensity", skip: 1, keepHeader: true).set{ all_mutdensities_file } + + // Concatenate all outputs into a single file + channel.empty() + .concat(MUTDENSITYALL.out.mutdensities_wgs.map{ it -> it[1]}.flatten()) + .concat(MUTDENSITYPROT.out.mutdensities_wgs.map{ it -> it[1]}.flatten()) + .concat(MUTDENSITYNONPROT.out.mutdensities_wgs.map{ it -> it[1]}.flatten()) + .concat(MUTDENSITYSYNONYMOUS.out.mutdensities_wgs.map{ it -> it[1]}.flatten()) + .set{ all_mutdensities_wgs } + all_mutdensities_wgs.collectFile(name: "all_mutdensities_wgs.tsv", storeDir:"${params.outdir}/mutdensity", skip: 1, keepHeader: true).set{ all_mutdensities_wgs_file } + + + channel.of([ [ id: "all_samples" ] ]) + .join( MUTDENSITYSYNONYMOUS.out.mutdensities ) + .set{ all_samples_syn_mutdensity } + + SYNMUTDENSITY(all_samples_syn_mutdensity) + + SYNMUTREADSDENSITY(all_samples_syn_mutdensity) + + channel.of([ [ id: "all_samples" ] ]) + .join( somatic_mutations ) + .set{ mutations_all_samples } + VAFSMOOTHING(mutations_all_samples, all_mutdensities_file, PLOTDEPTHSEXONSCONS.out.average_depth_sample) + + } + + if (run_mutabilities) { - if (params.profileall){ - MUTABILITYALL(mutations_in_exons, - annotated_depths, - MUTPROFILEALL.out.profile, - CREATEPANELS.out.exons_consensus_panel - ) - } + MUTABILITYALL(mutations_in_exons, + annotated_depths, + MUTPROFILEALL.out.profile, + CREATEPANELS.out.exons_consensus_panel + ) if (params.profilenonprot){ MUTABILITYNONPROT(mutations_in_exons, annotated_depths, @@ -405,40 +445,37 @@ workflow DEEPCSA { // OncodriveFML if (params.oncodrivefml){ - if (params.profileall){ - mode = "all" - ONCODRIVEFMLALL(mutations_in_exons, MUTABILITYALL.out.mutability, - CREATEPANELS.out.exons_consensus_panel, - cadd_scores, mode - ) - positive_selection_results = positive_selection_results.join(ONCODRIVEFMLALL.out.results_snvs, remainder: true) - } + ONCODRIVEFMLALL(mutations_in_exons, MUTABILITYALL.out.mutability, + CREATEPANELS.out.exons_consensus_panel, + cadd_scores, "all" + ) + positive_selection_results = positive_selection_results.join(ONCODRIVEFMLALL.out.results_snvs, remainder: true) + if (params.profilenonprot && params.positive_selection_non_protein_affecting){ - mode = "non_prot_aff" ONCODRIVEFMLNONPROT(mutations_in_exons, MUTABILITYNONPROT.out.mutability, CREATEPANELS.out.exons_consensus_panel, - cadd_scores, mode + cadd_scores, "non_prot_aff" ) } } if (params.oncodrive3d){ - if (params.profileall){ - // Oncodrive3D - ONCODRIVE3D(mutations_in_exons, MUTABILITYALL.out.mutability, - datasets3d, annotations3d, MUT_PREPROCESSING.out.all_raw_vep_annotation) - positive_selection_results = positive_selection_results.join(ONCODRIVE3D.out.results, remainder: true) - positive_selection_results = positive_selection_results.join(ONCODRIVE3D.out.results_pos, remainder: true) - - } + // Oncodrive3D + ONCODRIVE3D(mutations_in_exons, MUTABILITYALL.out.mutability, + datasets3d, annotations3d, MUT_PREPROCESSING.out.all_raw_vep_annotation) + positive_selection_results = positive_selection_results.join(ONCODRIVE3D.out.results, remainder: true) + positive_selection_results = positive_selection_results.join(ONCODRIVE3D.out.results_pos, remainder: true) } // if (params.expected_mutated_cells & params.dnds){ if (params.dnds){ DNDS(mutations_in_exons, DEPTHSEXONSCONS.out.subset, - CREATEPANELS.out.exons_consensus_panel + CREATEPANELS.out.exons_consensus_bed, + CREATEPANELS.out.exons_consensus_panel, + params.fasta ) + dndscv_table = DNDS.out.all_dndscv_results } if (params.omega){ @@ -446,65 +483,63 @@ workflow DEEPCSA { omega_regressions_files_gloc = channel.empty() // Omega - if (params.profileall){ - OMEGA(mutations_in_exons, - DEPTHSEXONSCONS.out.subset, - MUTPROFILEALL.out.profile, - CREATEPANELS.out.exons_consensus_bed.first(), - ENRICHPANELS.out.exons_consensus_expanded_panel.first(), - custom_groups_table, - SYNMUTDENSITY.out.mutdensity.first(), - CREATEPANELS.out.panel_annotated_rich, - "", - grouping_definitions, - ENRICHPANELS.out.exons_json_subgenic - ) - positive_selection_results = positive_selection_results.join(OMEGA.out.results, remainder: true) - all_compiled_omegas = OMEGA.out.all_compiled - if (params.omega_mutabilities){ - site_comparison_results = OMEGA.out.site_comparison - } - if (params.omega_globalloc){ - positive_selection_results = positive_selection_results.join(OMEGA.out.results_global, remainder: true) - all_compiled_omegasgloballoc = OMEGA.out.all_globalloc_compiled - } + OMEGA(mutations_in_exons, + DEPTHSEXONSCONS.out.subset, + MUTPROFILEALL.out.profile, + CREATEPANELS.out.exons_consensus_bed, + ENRICHPANELS.out.exons_consensus_expanded_panel, + custom_groups_table, + SYNMUTDENSITY.out.mutdensity.first(), + CREATEPANELS.out.panel_annotated_rich, + "", + grouping_definitions, + ENRICHPANELS.out.exons_json_subgenic + ) + positive_selection_results = positive_selection_results.join(OMEGA.out.results, remainder: true) + all_compiled_omegas = OMEGA.out.all_compiled + if (params.omega_mutabilities){ + site_comparison_results = OMEGA.out.site_comparison + } + if (params.omega_globalloc){ + positive_selection_results = positive_selection_results.join(OMEGA.out.results_global, remainder: true) + all_compiled_omegasgloballoc = OMEGA.out.all_globalloc_compiled.first() + } - if (params.regressions){ - omega_regressions_files = omega_regressions_files.mix(OMEGA.out.results.map{ it -> it[1] }) - omega_regressions_files_gloc = omega_regressions_files_gloc.mix(OMEGA.out.results_global.map{ it -> it[1] }) - } + if (params.regressions){ + omega_regressions_files = omega_regressions_files.mix(OMEGA.out.results.map{ it -> it[1] }) + omega_regressions_files_gloc = omega_regressions_files_gloc.mix(OMEGA.out.results_global.map{ it -> it[1] }) + } - if (params.omega_multi){ - // Omega multi - OMEGAMULTI(mutations_in_exons, - DEPTHSEXONSCONS.out.subset, - MUTPROFILEALL.out.profile, - CREATEPANELS.out.exons_consensus_bed.first(), - ENRICHPANELS.out.exons_consensus_expanded_panel.first(), - custom_groups_table, - SYNMUTREADSDENSITY.out.mutdensity.first(), - CREATEPANELS.out.panel_annotated_rich, - ".multi", - grouping_definitions, - ENRICHPANELS.out.exons_json_subgenic - ) - positive_selection_results = positive_selection_results.join(OMEGAMULTI.out.results, remainder: true) - if (params.omega_globalloc){ - positive_selection_results = positive_selection_results.join(OMEGAMULTI.out.results_global, remainder: true) - } - if (params.regressions){ - omega_regressions_files = omega_regressions_files.mix(OMEGAMULTI.out.results.map{ it -> it[1] }) - omega_regressions_files_gloc = omega_regressions_files_gloc.mix(OMEGAMULTI.out.results_global.map{ it -> it[1] }) - } - } + if (params.omega_multi){ + // Omega multi + OMEGAMULTI(mutations_in_exons, + DEPTHSEXONSCONS.out.subset, + MUTPROFILEALL.out.profile, + CREATEPANELS.out.exons_consensus_bed, + ENRICHPANELS.out.exons_consensus_expanded_panel, + custom_groups_table, + SYNMUTREADSDENSITY.out.mutdensity.first(), + CREATEPANELS.out.panel_annotated_rich, + ".multi", + grouping_definitions, + ENRICHPANELS.out.exons_json_subgenic + ) + positive_selection_results = positive_selection_results.join(OMEGAMULTI.out.results, remainder: true) + if (params.omega_globalloc){ + positive_selection_results = positive_selection_results.join(OMEGAMULTI.out.results_global, remainder: true) + } + if (params.regressions){ + omega_regressions_files = omega_regressions_files.mix(OMEGAMULTI.out.results.map{ it -> it[1] }) + omega_regressions_files_gloc = omega_regressions_files_gloc.mix(OMEGAMULTI.out.results_global.map{ it -> it[1] }) + } } if (params.profilenonprot && params.positive_selection_non_protein_affecting){ OMEGANONPROT(mutations_in_exons, DEPTHSEXONSCONS.out.subset, MUTPROFILENONPROT.out.profile, - CREATEPANELS.out.exons_consensus_bed.first(), - ENRICHPANELS.out.exons_consensus_expanded_panel.first(), + CREATEPANELS.out.exons_consensus_bed, + ENRICHPANELS.out.exons_consensus_expanded_panel, custom_groups_table, SYNMUTDENSITY.out.mutdensity.first(), CREATEPANELS.out.panel_annotated_rich, @@ -517,8 +552,8 @@ workflow DEEPCSA { OMEGANONPROTMULTI(mutations_in_exons, DEPTHSEXONSCONS.out.subset, MUTPROFILENONPROT.out.profile, - CREATEPANELS.out.exons_consensus_bed.first(), - ENRICHPANELS.out.exons_consensus_expanded_panel.first(), + CREATEPANELS.out.exons_consensus_bed, + ENRICHPANELS.out.exons_consensus_expanded_panel, custom_groups_table, SYNMUTREADSDENSITY.out.mutdensity.first(), CREATEPANELS.out.panel_annotated_rich, @@ -532,7 +567,7 @@ workflow DEEPCSA { } - if (params.mutated_cells_vaf){ + if (params.mutated_cells_vaf && params.omega && params.omega_globalloc){ MUT_PREPROCESSING.out.somatic_mafs .join(meta_samples_alone) .set{ sample_mutations_only } @@ -594,13 +629,18 @@ workflow DEEPCSA { PLOTTINGQC( somatic_mutations, all_mutdensities_file.first(), + all_adjusted_mutdensities_file, + all_compiled_omegasgloballoc, + PLOTDEPTHSEXONSCONS.out.average_depth_gene_sample.first(), all_compiled_omegas, // site_comparison_results, // ANNOTATEDEPTHS.out.all_samples_depths, // TABLE2GROUP.out.json_allgroups, CREATEPANELS.out.exons_consensus_panel, TABLE2GROUP.out.json_allgroups.first(), - group_keys_ch + group_keys_ch, + dndscv_table, + TABLE2GROUP.out.json_groups.first(), // CREATEPANELS.out.panel_annotated_rich, // seqinfo_df, // CREATEPANELS.out.domains_in_panel, @@ -610,24 +650,34 @@ workflow DEEPCSA { if (params.omega || params.oncodrive3d || params.oncodrivefml || params.indels || run_mutdensity) { if (params.omega){ positive_selection_results = positive_selection_results.combine(PLOTTINGQC.out.flagged_omegas) + if (params.omega_globalloc){ + positive_selection_results = positive_selection_results.combine(all_compiled_omegasgloballoc) + } } - positive_selection_results_ready = positive_selection_results.map { element -> [element[0], element[1..-1]] } + positive_selection_results_ready = positive_selection_results + .map { element -> + def meta = element[0] + def files = element[1..-1].findAll { it -> it != null } + [meta, files] + } + .filter { _meta, files -> files.size() > 0 } + PLOTTINGSUMMARY(positive_selection_results_ready, somatic_mutations, all_mutdensities_file.first(), - + all_adjusted_mutdensities_file, site_comparison_results, ANNOTATEDEPTHS.out.all_samples_depths.first(), TABLE2GROUP.out.json_samples.first(), TABLE2GROUP.out.json_allgroups.first(), - CREATEPANELS.out.exons_consensus_panel.first(), - ENRICHPANELS.out.exons_consensus_expanded_panel.first(), - CREATEPANELS.out.panel_annotated_rich.first(), + CREATEPANELS.out.exons_consensus_panel, + ENRICHPANELS.out.exons_consensus_expanded_panel, + CREATEPANELS.out.panel_annotated_rich, seqinfo_df, - CREATEPANELS.out.domains_in_panel.first(), - ENRICHPANELS.out.dna2protein_mapping_depth_exons.first(), + CREATEPANELS.out.domains_in_panel, + ENRICHPANELS.out.dna2protein_mapping_depth_exons, group_keys_ch ) }