svaha2 is designed for high-performance variation graph construction, specifically optimized for cancer genomics workflows where multiple variant types and formats (VCF, MAF, SV, BEDPE) must be integrated into a unified graph.
The build command is the primary entry point for graph construction.
./svaha build -r <reference.fa> [options]-r <file>: The reference genome in FASTA format. A corresponding.faiindex file must exist in the same directory.
You can provide multiple variant files, even of different formats, in a single run:
-v | --vcf <file>: Standard VCF 4.0+ file. Handles SNPs, Indels, and SVs (DEL, INS, INV, BND).--maf <file>: cBioPortal/TCGA Mutation Annotation Format. Optimized for somatic mutations (SNP, DNP, TNP, ONP, INS, DEL).--sv <file>: cBioPortal structural variant format (data_sv.txt). Parses connection types (3to5,5to5, etc.) to set correct graph orientations.--bedpe <file>: Standard BEDPE format for structural variants.
-m <int>: Maximum Node Size (default: 32). This controls the discretization of the reference backbone. Smaller values increase graph resolution but result in more nodes/edges. Larger values reduce memory but might "hide" variants if they fall within the same node.-R <chrom:start-end>: Region Filtering. Only variants and reference sequences within this coordinate range will be processed. Coordinates are 1-based.- Example:
-R chr17:7500000-8000000
- Example:
--relax-chrom: Chromosome Naming Normalization. By default,svaha2requires exact matches between variant files and the reference genome. Use this flag to allow automated mapping between plain numbers (e.g.,17) andchr-prefixed names (e.g.,chr17).
svaha2 parses the following columns from cBioPortal MAFs:
Chromosome,Start_Position,End_PositionReference_Allele,Tumor_Seq_Allele2Variant_Type(SNP, DNP, TNP, ONP, INS, DEL)
Small variants are automatically normalized:
- Leading anchor bases are stripped to ensure clean graph bubbles.
- Coordinates are adjusted to represent only the novel sequence.
Designed for data_sv.txt files, this parser interprets:
Site1_Chromosome,Site1_PositionSite2_Chromosome,Site2_PositionClass(Inversion, Deletion, Translocation, etc.)Connection_Type(Used to determine head-to-head, tail-to-tail, or standard orientations).
The VCF parser handles:
- Small Variants: Automatically detects SNPs and Indels.
- Structural Variants: Parses
SVTYPEandENDtags. - Breakends: Supports
BNDandTRAtypes for translocations. - Multi-allelic Sites: Correctly creates separate alternate paths for each allele at a single position.
For large genomes (e.g., human hg38), construction can take several minutes. svaha2 provides structured reporting to stderr:
- Phase 0 (Parsing): Reports the file being read and the number of variants successfully loaded.
- Discretization: Informs you when the reference backbone is being calculated.
- Phase 1 (Topology): Summarizes the number of backbone nodes, variant nodes, and edges created.
- Phase 2 (Materialization): For full-genome builds, reports progress every 1,000,000 nodes.
# 1. Compile
make
# 2. Build graph for a specific cancer study
./svaha build -r /path/to/hg38.fa \
--maf study_data/data_mutations.txt \
--sv study_data/data_sv.txt \
-v common_germline.vcf \
--relax-chrom \
-R chr17 > my_graph.gfa
# 3. Check graph stats
./svaha stats my_graph.gfa
# 4. Visualize the graph (regional only)
./svaha view my_graph.gfa > my_graph.dot
dot -Tpdf my_graph.dot -o my_graph.pdf- Built-in
viewcommand: Best for small regional graphs (e.g., < 100 nodes). Converts GFA to DOT format for use with Graphviz. - Bandage: The gold standard for large-scale GFA visualization. It provides an interactive GUI for exploring complex variation graphs.
- vg view: If you have the
vgtoolkit, you can usevg view -dto generate more complex SVG/DOT visualizations.