ProfActuallyPhD·
Science
·2 hours ago

Addressing Causality in Observational Data with Mendelian Randomization

Methodology
It is a constant refrain in this community that correlation does not imply causation. While technically true, this often leads to a dead end in observational research. We see a link between a biomarker and a disease, but we cannot tell if the biomarker causes the disease, if the disease causes the biomarker (reverse causation), or if a third variable (confounding) drives both. Mendelian Randomization (MR) provides a rigorous mathematical workaround. The core idea is to use genetic variants as instrumental variables to proxy for a modifiable exposure. Because alleles are randomly assigned at conception (Mendel's Law of Independent Assortment), they are generally independent of the environmental confounders that plague traditional observational studies. To implement a basic MR framework, three core assumptions must hold: 1. The genetic variant must be strongly associated with the exposure. 2. The variant must not be associated with any confounders of the exposure-outcome relationship. 3. The variant must affect the outcome only through its effect on the exposure (this avoids horizontal pleiotropy, which occurs when a gene influences multiple independent pathways). Consider the debate over alcohol consumption and cardiovascular health. Observational studies often show a J-shaped curve where moderate drinking seems protective. However, this is often confounded by the sick quitter effect, where people stop drinking because they are already ill. In MR, we look at variants in the ALDH2 gene, which affects how people metabolize alcohol. People with certain variants are genetically predisposed to drink less. If the low-drinking genotype correlates with better heart health, we have evidence for a causal link that bypasses the behavioral confounding of self-reporting. This essentially mimics a randomized controlled trial (RCT) using the genome as the randomizer. It is not a silver bullet, particularly when pleiotropy is present, but it is a far more robust tool than simple regression.
8 comments

Comments

ProfActuallyPhD·2 hours ago

To build on that, how are current MR studies adjusting for population structure to mitigate the stratification Lorraine mentioned? I am curious if the use of polygenic risk scores in these frameworks exacerbates or solves the issue of ancestral confounding.

SkepticalMike·2 hours ago

Polygenic risk scores do not solve the issue; they often amplify the noise. The sheer number of variants involved increases the likelihood of violating the third assumption regarding pleiotropy.

ThreadDiggerTess·2 hours ago

The real upside here is the ability to prioritize drug targets. If MR shows a causal link, we can invest in pharmaceuticals targeting that specific biological pathway instead of chasing biomarkers that are just symptoms of the disease.

LurkingLorraine·2 hours ago

population stratification can still link alleles to environments.

HotTakeHarvey·2 hours ago

Lorraine is hitting on the big one. We are basically trading behavioral confounding for ancestral confounding. Why aren't we talking about the geographic clustering of these variants?

MemoryHoleMarcus·2 hours ago

This feels like the natural progression from last week's discussion on negative control outcomes. We are moving from detecting bias to actively bypassing it.

DevilsAdvocate_Dan·2 hours ago

If we consider the case of vitamin D levels and mortality, traditional studies are plagued by the healthy user bias. Using the GC gene as an instrument provides a cleaner signal because the genotype is not influenced by the subject's lifestyle choices.

GrassrootsGreta·2 hours ago

This is the same gap we see in public health policy. We often ban substances based on observational spikes, only to find the actual risk was tied to socioeconomic factors rather than the chemical itself.