LD Score regression distinguishes confounding from polygenicity
quantifies the contribution of each by examining the relationship between test statistics and linkage disequilibrium (LD).Read the original source
Numbers & models
What can we learn about the genetics of a trait from studies that already exist?
Interactive study
Explore the exampleLDSC makes published genetic studies useful for more than their original question. It estimates how much variation common genetic variants explain, and whether two traits share genetic influences. The clever part is that correlated variants help separate many small genetic effects from study bias: a variant that tags more neighbors tends to collect more real genetic signal. You need published genome-wide association study (GWAS) summaries and a suitable reference panel, but you don't need the participants' individual genomes. I'm interested in that reuse, and in making the expensive reference calculation fast enough to compare more datasets and settings.
The Rust engine calculates LD scores, heritability, and genetic correlation, with exact and sketch paths. In a recorded comparison on the same workstation, the exact f64 LD-score step took about 85 seconds instead of nearly 26 minutes, an 18.1× speedup. The page below connects that work to real 1000G output and a published BMI study. The next question is how much time sketching can save without losing accuracy in the final genetic estimates.
1000 Genomes tells us how genetic variants correlate with each other. Published GWAS summaries tell us how those variants associate with a trait. LDSC uses the first to interpret the second.
That's what makes the method so useful for reuse. One suitable reference can serve many studies, and their published summaries let us ask new questions without access to participants' individual genomes.
18.1× faster LD scores
In the recorded comparison, the Rust f64 run took about 85 seconds where Python took 25 minutes 49 seconds. That's 24.4 minutes saved on one reference calculation.
1,664,852 variants · 2,490 reference samples · 1000 kb window · Ryzen 5 5600X (6 cores / 12 threads), 32 GB RAM. Recorded 2026-03-07 with hyperfine: one Python run and three Rust runs. ± values show standard deviations.
A shorter wait makes it easier to try another window, compare reference panels, or repeat an accuracy check. The implementation reuses memory and gives matrix kernels blocks of work to distribute across CPU threads.
Explore the computational methods
The log reports zero maximum absolute difference in this f64 comparison. That result applies to the recorded build and settings.
The f32 path changes precision, with a recorded maximum difference of 0.008 LD-score units. Sketch modes introduce a separate approximation tradeoff.
These are results from a native CPU run. The speed depends on the machine and settings. The browser displays the saved results.
Read the benchmark extract or inspect the pinned performance log.
Each variant needs correlations with its neighbors, and every correlation reads across the reference samples. Those comparisons add up quickly. The Rust engine reuses data between windows and turns the correlations into larger matrix operations that the CPU can handle efficiently.
The 18.1× benchmark measures the complete f64 run. These small diagrams explain the implementation, but they don't measure each optimization's share of the speedup.
Move to the next chunk, and much of its neighborhood is still the same. A ring buffer keeps those normalized genotype columns, so the engine can reuse them.
Call the new chunk B and the retained neighbors A. Two matrix products calculate the correlations:
This gives faer blocks of arithmetic to distribute across CPU threads. Keeping the columns together in memory also helps the CPU reuse cached data.
The engine keeps its scratch matrices between chunks, too. Each product overwrites them, so it skips repeated allocation and zero fill.
The default keeps Python's chunk-level window convention. The optional --snp-level-masking flag removes pairs outside each variant's actual window.
With c new variants, w retained neighbors, and N samples, the arithmetic work is roughly N × c × (w + c). Denser regions increase w.
Read the implementation of the ring buffer and matrix products.
Blue: BᵀBGreen: AᵀB and its symmetric contribution
The --fast-f32 option uses the same windows and matrix algorithm, with 32-bit floats for genotype storage and matrix products.
Each normalized value takes four bytes instead of eight. That means less data to move through memory and more room in the CPU cache.
For 200 variants and 2,490 reference samples, one genotype chunk takes 3.80 MiB in f64 or 1.90 MiB in f32. Those sizes cover the chunk, not the whole process.
There's a precision tradeoff. Normalization statistics and corrected LD-score totals stay in f64, but the matrix products use f32.
The recorded f32 run took 66.3 seconds, with a maximum difference of 0.008 LD-score units. That gives us a useful comparison for this dataset. Other datasets still need their own accuracy checks.
CountSketch reduces N sample rows to d bucket rows. Each sample gets a bucket and a random plus or minus sign. The engine adds those signed, normalized values together.
The matrix products then work on the smaller sketch. With 2,490 samples and 500 buckets, each dot product has about one-fifth as many row terms.
The input still needs a full read. On the contiguous path, one pass calculates statistics from packed BED bytes. A second pass decodes, normalizes, and adds each value directly to its bucket.
Combining those steps skips a full decoded matrix write before the sketch. Irregular variant extraction still uses a full-size fallback buffer.
One-fifth as many row terms is a work estimate, not a fivefold runtime result. The 18.1× benchmark above uses the exact f64 path.
The diagram uses one synthetic variant: six normalized values enter three buckets, then rescale to squared norm N = 6. The bucket assignment is illustrative.
When samples land in the same bucket, their contributions mix and add noise. More buckets usually reduce that noise, but they also increase matrix work.
The engine rescales columns and applies a quadratic bias correction before the finite-sample correction. Those steps reduce bias, but they don't make the sketch exact. We still need to compare both LD scores and final heritability estimates with the exact results.
For c new variants and w retained neighbors, the work per chunk is roughly N × c + d × c × (w + c). The input scan still grows with N.
Plotters 0.3.7 draws these SVGs directly from Rust. The controls switch between the diagrams.
For an animated walkthrough, MotionGfx offers procedural animation in Rust with a Bevy integration. That's an option for a future version. The diagrams here use SVG.
Nearby variants often travel together through inheritance. A variant can therefore carry information about its neighbors. Its LD score adds up squared correlations across the window, including the correlation with itself.
The chart shows the Rust engine's saved chromosome 22 results. Most variants have modest scores, while a smaller group has much stronger LD with its neighbors.
This helps explain why an association can point to a region with several correlated variants. A high LD score alone doesn't tell us which variant causes a disease.
| LD score interval | Variants |
|---|---|
| 0–10 | 6,173 |
| 10–20 | 6,640 |
| 20–30 | 2,742 |
| 30–40 | 1,227 |
| 40–50 | 892 |
| 50–60 | 356 |
| 60–70 | 309 |
| 70–80 | 181 |
| 80–90 | 10 |
| 90–100 | 9 |
| 100+ | 88 |
Intervals include the lower bound and exclude the upper bound.
Here's a complete example. The native pipeline used a 503-person 1000G European reference to calculate LD scores, then paired them with Yengo et al.'s published BMI study of roughly 700,000 people. The reference supplies the correlations. The study supplies the association statistics.
Correlations across 1.66 million variants produce LD scores.
29 sAlleles, sample counts, and association statistics enter a common format.
10 sRegress the study’s association signal against the reference LD scores.
2 sRecorded pipeline total: 41 seconds on an Apple M5 Pro. This is separate from the Ryzen benchmark above.
20.32% of BMI variation tagged by common genetic variants in this analysis
Standard error: 0.55 percentage points. The regression merged 1,030,015 variants.
The estimate describes variation across the study population. It doesn't mean that 20.32% of one person's weight is genetic, and it doesn't include every genetic effect.
The BMI intercept was 1.0972 ± 0.0193. The confounding ratio was 3.30% ± 0.65 percentage points, with one standard error.
Under the LDSC model, about 3.3% of the mean association statistic's excess over 1 comes from confounding. The rest reflects the model's genetic signal.
That's an estimate under the model. Population structure, relatedness, reference choice, and study corrections all affect its interpretation. The intercept helps distinguish those effects from many small genetic contributions.
With two GWAS summaries and a suitable LD reference, cross-trait LDSC estimates genetic correlation. This lets us ask whether genetic influences overlap across metabolic, psychiatric, or other traits.
A positive correlation indicates shared influences in the same direction. A negative one indicates opposing directions. Neither proves that one trait causes the other.
The original cross-trait paper explains the method. The BMI example here uses one trait, so it doesn't calculate a genetic correlation.
The charts use saved results from the native engine. Rust calculates the displayed ratios and draws the histogram from the saved counts.
The small example below uses synthetic genotypes so you can follow the LD-score formula separately from the recorded runs.
Start with six samples and three variants. Compare the raw sum, the finite-sample correction, and a shorter window to see why each choice changes the LD score.
Each score sums squared correlations across all three variants, with the self-correlation included.
A: 0 0 1 1 2 2
B: 0 1 0 2 1 2
C: 2 2 1 1 0 0Variant 1: 2.250000
Variant 2: 1.500000
Variant 3: 2.250000
The finite-sample correction uses r² − (1 − r²) / (N − 2), with N = 6. Negative contributions remain possible.
A: 0 0 1 1 2 2
B: 0 1 0 2 1 2
C: 2 2 1 1 0 0Variant 1: 2.062500
Variant 2: 1.125000
Variant 3: 2.062500
The short window omits the distant pair. The input stays the same, but the scores change.
A: 0 0 1 1 2 2
B: 0 1 0 2 1 2
C: 2 2 1 1 0 0Variant 1: 1.062500
Variant 2: 1.125000
Variant 3: 1.062500
Rust calculates this small example from synthetic genotypes. It illustrates the LD-score formula, while the full regression and sketch modes run in the native engine.
quantifies the contribution of each by examining the relationship between test statistics and linkage disequilibrium (LD).Read the original source
estimating genetic correlation that requires only GWAS summary statistics and is not biased by sample overlap.Read the original source
Repeat the full pipeline with fixed inputs and builds, then compare exact and sketch modes on larger panels. Measure the error in heritability and genetic correlation alongside runtime, so the speed tradeoff stays useful for the actual analysis.
How much can published genetic studies reveal without access to individual genomes?
The display combines recorded native benchmarks, aggregate chromosome 22 LD scores, and a retained BMI regression result. A separate three-variant example uses synthetic data.
Compare runtime, examine the 1000G LD-score distribution, and trace the BMI result back to its reference and study inputs.
Rust calculates display ratios and SVG charts from saved results at export. The full genotype and regression pipeline stays in the native CLI.
A Rust reimplementation of the LDSC method and reference software.