Thank you for the incredible work on minigraph-cactus and the vg toolkit!
We are planning a large association study (GWAS) of structural variants and would greatly appreciate your feedback: are we overlooking any obvious technical problems that explain why so few population-scale SV-GWAS studies use this approach?
Study design
Cohort: 1,500–10,000 short-read WGS samples from a single ethnicity (underrepresented in current reference panels).
Pangenome: Minigraph-Cactus graph built from 40 high-quality haplotypes of the same ethnicity.
Goal: Genotype SVs >50 bp for GWAS.
Pipeline: vg giraffe → vg call → truvari collapse.
We benchmarked our pangenome against a linear reference (GRCh38) using long reads from a subset of samples. The pangenome showed substantially better sensitivity and specificity for SVs, especially in reference-gap regions and poorly resolved loci. This convinced us that a graph reference is the right choice for our population.
We understand that vg call effectively genotypes known SVs already embedded in the graph - including insertions that close reference gaps, alternate haplotypes, and complex rearrangements present in our 40 founder haplotypes. We are not expecting de novo SV discovery from short reads alone. For GWAS, genotyping a representative catalog is acceptable, provided the catalog is sufficiently complete for common and low-frequency SVs in our ethnicity.
Despite the promising benchmark, we can find almost no examples of population-scale SV-GWAS specifically using short reads + minigraph-cactus / vg giraffe. This is concerning: we worry that we may be missing a fundamental limitation that makes this design impractical at scale.
We are happy to benchmark any suggested parameters or workflows and share the results with the community. Thank you for any guidance you can provide!
Thank you for the incredible work on minigraph-cactus and the vg toolkit!
We are planning a large association study (GWAS) of structural variants and would greatly appreciate your feedback: are we overlooking any obvious technical problems that explain why so few population-scale SV-GWAS studies use this approach?
Study design
We benchmarked our pangenome against a linear reference (GRCh38) using long reads from a subset of samples. The pangenome showed substantially better sensitivity and specificity for SVs, especially in reference-gap regions and poorly resolved loci. This convinced us that a graph reference is the right choice for our population.
We understand that vg call effectively genotypes known SVs already embedded in the graph - including insertions that close reference gaps, alternate haplotypes, and complex rearrangements present in our 40 founder haplotypes. We are not expecting de novo SV discovery from short reads alone. For GWAS, genotyping a representative catalog is acceptable, provided the catalog is sufficiently complete for common and low-frequency SVs in our ethnicity.
Despite the promising benchmark, we can find almost no examples of population-scale SV-GWAS specifically using short reads + minigraph-cactus / vg giraffe. This is concerning: we worry that we may be missing a fundamental limitation that makes this design impractical at scale.
We are happy to benchmark any suggested parameters or workflows and share the results with the community. Thank you for any guidance you can provide!