We are excited to announce a new Long-read Structural Variants container track on the human assemblies GRCh38/hg38 and T2T-CHM13/hs1. The container track brings together structural variant (SV) callsets from 14 long-read sequencing studies worldwide into a single place where you can compare large genomic rearrangements (deletions, insertions, duplications, inversions, and complex events) across populations, cohorts, and calling strategies. Long-read technologies span repetitive regions and resolve complex loci that are difficult to detect with short reads, giving more precise breakpoints and better sensitivity for large variants.
At the center of the container track is an All long-read SVs merged track that unifies every source callset on identical position, type, and length into roughly 2.3 million distinct SV sites, each carrying per-database allele counts so you can see at a glance how many studies support a given variant and at what frequency. This merged track is the best place to start: we recommend using it to survey structural variation across all studies at a locus, then turning to the individual dataset tracks for cohort-specific allele frequencies, genotypes, and annotations on the variants you want to follow up.
The Long-read Structural Variants container track at UGT2B17 (chr4). The merged track (top) unifies a common whole-gene deletion of about 117 kb, which recurs across the individual dataset tracks below. The mouseover reports a 79% allele frequency in the Chinese Pangenome Consortium, consistent with the higher frequency of the UGT2B17 deletion in East Asian populations.
Each variant is colored by SV type (deletions, insertions, duplications, inversions, and complex or multi-allele events), and carries allele frequencies, sample counts, and per-study annotations. Cross-database filters let you narrow by SV type, source study, variant length, allele frequency, or the number of supporting studies, in any combination.
The container track pulls together cohorts from across the world. A high-level summary is shown below; a complete table with per-study sample counts, cohorts, sequencing coverage, and SV counts is on the track description page.
| Region / theme | Contributing studies |
|---|---|
| East Asia | Han Chinese 945, ToMMo Japan, Chinese Pangenome Consortium |
| 1000 Genomes (global) | 1KG ONT 100 (Gustafson), 1KG ONT Vienna (1,019) |
| Americas | All of Us, GA4K (pediatric rare disease), SVatalog (cystic fibrosis) |
| Europe | deCODE (Iceland) |
| Middle East | Arab Pangenome Reference |
| Global reference & pangenome callsets | CoLoRSdb, HPRC v2.1, HGSVC2, HGSVC3 |
Six of these datasets (CoLoRSdb, 1KG ONT Vienna, HGSVC3, HPRC v2.1, Arab Pangenome Reference, and the Chinese Pangenome Consortium) are also released on the T2T-CHM13/hs1 assembly in native coordinates.
License restrictions on some sources limit redistribution; see the track description page for per-study details.
We plan to keep adding long-read SV callsets as more become available. If you are involved with a project that publishes long-read structural variants and would like to contribute, please reach out.
We would like to thank the many participants who donated samples, and the consortia and investigators who generated and shared these callsets, including Glenn Hickey and the Human Pangenome Reference Consortium graph team for the HPRC v2.1 callset; Evan Eichler and Jiadong Lin; and Kwanho Kim, Joshua Levin, and colleagues in the Aligning Science Across Parkinson's (ASAP) network.