All posts
Research2026-06-227 min readAustin Talbot

Detecting Parent-of-Origin Effects Without Family Data: Introducing POISE

The Problem: POEs Are Pervasive but Hard to Detect at Scale

Surprisingly, not all alleles behave the same way depending on which parent they came from. Parent-of-origin effects (POEs), where the phenotypic impact of an allele differs based on maternal versus paternal inheritance, are a well-established but less characterized layer of the genome. The canonical mechanism is genomic imprinting, in which only one parental copy of a gene is expressed. Imprinted genes play critical roles in placentation, growth, and neurodevelopment, and their dysregulation underlies disorders such as Prader–Willi and Beckwith–Wiedemann syndromes.

POEs have also been proposed as one explanation for the “missing heritability” problem, the gap between phenotypic variance explained by standard GWAS and total estimated heritability. The trouble is that detecting POEs directly requires parental genotypes or trio data, which are expensive to collect and poorly represented in modern biobank cohorts. In the first UK Biobank release, for example, fewer than 5,000 of over 400,000 genotyped participants had directly available parental genotypes.

Variance-based methods offer a workaround. The key insight is that a POE inflates the phenotypic variance of heterozygotes relative to homozygotes: because heterozygous individuals carry either the maternal or paternal copy but we cannot observe which, their phenotypes are drawn from a mixture of two shifted distributions, producing excess variance. Hoggart et al. introduced a univariate test for this signal, and POIROT extended it to the multivariate setting. But POIROT has practical gaps: it returns a single p-value per variant with no per-trait effect estimates, no uncertainty quantification, and critically it cannot distinguish a true POE from a variance QTL (vQTL) that happens to inflate heterozygote variance through a different mechanism.

The Approach: Spectral Estimation of the POE Vector

POISE (Parent of Origin Inference via Spectral Estimation) addresses these limitations by modeling the problem as a latent mixture and using spectral decomposition to recover the POE signal directly.

The core observation is that when a POE exists, the covariance matrix of the heterozygous population gains a rank-one perturbation along the direction of the POE vector Δ = βM − βP. This “bump” stretches the covariance ellipse in exactly one direction separating the two hidden inheritance groups. In a setting with uncorrelated phenotypes, this means the leading eigenvalue of the heterozygous covariance matrix exceeds 1 by an amount that encodes the POE magnitude, while all other eigenvalues remain at 1. POISE recovers Δ by extracting the leading eigenvalue and eigenvector and inverting the eigenvalue relation.

Simple case

In practice, phenotypes are correlated. Metabolic traits co-vary, expression levels within pathways move together, so a naive application of the above would conflate pre-existing correlations with the POE signal. POISE handles this via a whitening step: we first estimate the baseline covariance from homozygotes (who carry no POE signal), use it to decorrelate the heterozygous observations, apply the spectral estimator in the whitened space, and then transform back to the original phenotype coordinates. The result is a per-trait vector of POE coefficients that is directly interpretable: large entries identify phenotypes strongly affected by parental origin; entries near zero indicate phenotypes that are not.

General case

Three Additional Guarantees

Beyond the point estimate, POISE provides three features that are absent from existing GWAS-based POE methods.

Per-trait confidence intervals. POISE computes bias-corrected and accelerated (BCa) bootstrap confidence intervals for each phenotype’s POE coefficient. These account for both the bias and skewness of the bootstrap distribution, providing reliable uncertainty quantification even when spectral inflation (the tendency of leading sample eigenvalues to exceed their population counterparts at finite sample sizes) would otherwise distort naive interval estimates.

An exact permutation test. Under the null hypothesis (no POE), the genotype group labels are exchangeable after mean-centering. POISE exploits this exchangeability to construct a permutation-based p-value that has exact Type I error control without distributional assumptions and that remains valid even when point estimates are biased in magnitude, because the same bias affects the observed and permuted test statistics identically.

A minimum detectable effect size filter. Detecting a POE from unlabeled heterozygotes is fundamentally harder than a standard mean-shift test: the signal enters only through the covariance perturbation, so the sample size requirement scales with the fourth power of the effect size rather than the second. POISE derives a closed-form information-theoretic detectability floor from the Chernoff information between the null and alternative distributions, and only reports a POE when the lower bootstrap confidence interval endpoint exceeds this threshold. This simultaneously guards against practically insignificant effects and against false positives from vQTLs that inflate heterozygote variance through non-additive mechanisms.

Simulation Results: Better Calibration, Higher Power, Zero vQTL False Positives

We evaluated POISE across three simulation scenarios.

Type I error. Under 5,000 null replicates with both Gaussian and heavy-tailed (t9) noise, the permutation p-values were well-calibrated against a Uniform(0,1) distribution (Kolmogorov–Smirnov p = 0.33 and 0.09, respectively). Proposition 1 in the paper guarantees that this calibration extends to any non-singular baseline covariance structure.

QQ plot

Power. Under Gaussian errors, POISE showed uniformly higher power than POIROT across signal strengths. The advantage was most pronounced at moderate SNR: at SNR = 0.19, POISE achieved 70.4% power versus 46.4% for POIROT—a difference of 24 percentage points. POISE reached 80% power at SNR ≈ 0.24, while POIROT required SNR ≈ 0.29. Under heavy-tailed (t9) errors, the two methods performed comparably, confirming that the spectral approach does not sacrifice robustness under distributional misspecification.

Power

Robustness to vQTLs. We simulated a pure variance-shift null—no true POE, but the heterozygote covariance scaled uniformly relative to homozygotes—and varied the scaling factor κ from 0.1 to 1.1. POIROT rejected the null consistently whenever κ deviated from 1, regardless of whether the deviation was compatible with any additive POE. POISE produced zero false positives for all κ < 1. This follows directly from the eigenvalue truncation in the estimator: when the heterozygote covariance is deflated relative to the homozygote baseline, the leading eigenvalue of the whitened covariance falls below 1 and the POE estimate is forced to zero.

UK Biobank Application: A Strict Superset of POIROT

We applied POISE to 602,000 UK Biobank participants, estimating POEs jointly for BMI, LDL cholesterol, and HDL cholesterol. These three traits have established or plausible links to parent-of-origin biology. Variants were prescreened using Bonferroni-corrected POIROT p-values, and all phenotypes were inverse-normal transformed prior to analysis.

At a Bonferroni threshold of 1.5 × 10⁻⁷, POIROT identified 338 variants, of which 186 passed the effect size criterion. POISE identified 893 variants, of which 320 passed the same criterion. Critically, the POISE-significant set was a strict superset of the POIROT-significant set: every one of the 186 POIROT variants was recovered. POISE additionally identified 134 variants that POIROT did not detect.

Manhattan Plot

The concordant set (variants detected by both methods) was dominated by the extended MHC class I region on chromosome 6, which accounted for 152 of 186 variants. Outside chromosome 6, the concordant set encompassed several canonical lipid-metabolism loci with established POE evidence, including variants at APOB (chromosome 2), the APOA5/APOC3/BUD13/ZPR1 apolipoprotein cluster (chromosome 11), and loci at APOE/TOMM40, CETP, and LIPC. The concordant variants were enriched for coding variants (10.8% exonic) and carried larger absolute effect sizes on BMI than the POISE-exclusive set (mean |β| = 0.312 vs. 0.208).

The 134 POISE-exclusive variants spanned 19 chromosomes and had higher average allele frequencies (mean gnomAD global AF = 0.218 vs. 0.098 in the concordant set), consistent with POISE capturing common-variant signals that lack sufficient heterozygote-variance signature for POIROT to detect. Biologically, the exclusive set included rs12243326 in TCF7L2—the most strongly replicated type 2 diabetes susceptibility gene and a genome-wide significant BMI locus—as well as variants in ABCG8 (sterol transporter), SCARB1 (HDL scavenger receptor SR-BI), and the classic CETP TaqIB polymorphism (rs708272).

Limitations and Next Steps

POISE is complementary to, not a replacement for, phasing-based methods that directly assign parental origin. Phased methods provide directional evidence (maternal vs. paternal) and are not susceptible to non-POE variance heterogeneity. POISE, operating on the full unrelated cohort, trades directional resolution for scale: it can screen hundreds of thousands of variants across the entire biobank population rather than the fraction for which phasing or relatedness inference is feasible.

Two limitations are worth noting. First, like all variance-based methods, POISE cannot fully distinguish a true POE from an inflation-type vQTL arising from gene–environment interaction or epistasis (the eigenvalue truncation protects against deflation-type vQTLs but not inflation). Second, the direction of the parental effect is fundamentally unidentifiable from unlabeled heterozygotes. Hits identified exclusively by POISE will therefore require follow-up with phasing-based or family-based methods to determine whether the effect is maternal or paternal in origin.

Code and Paper

POISE is implemented in Python and openly available at the Bystro GitHub repository. The preprint describes the full method, proofs, and UK Biobank results in detail.