My Healtheology
Health, examined not evangelised

Nutrigenomics

A polygenic score is a population statement being read as a personal one

These scores are built by summing many small statistical associations, and almost every limitation people encounter follows directly from how that sum is constructed.

A polygenic score is a population statement being read as a personal one
A polygenic score is a population statement being read as a personal one · Photo via Pexels
Health information notice. General information — not a substitute for professional advice. Read the full disclaimer.

How the score is constructed

A polygenic score is calculated by counting variants a person carries and weighting each by an association measured in a large reference dataset. The variants included are typically those associated with a trait in that dataset, regardless of whether any causal role is understood. Most individual weights are extremely small, so the score derives whatever predictive power it has from aggregation across many positions.

This makes the score a statistical summary of a person's position relative to a distribution rather than a description of a mechanism. Understanding that construction explains why such a score can predict reasonably well at group level while saying comparatively little about any one individual. It also explains why adding more variants improves the score gradually rather than transforming it, since each addition contributes very little.

The ancestry portability problem

The datasets used to derive weights have historically over-represented populations of European ancestry by a very wide margin. Because the statistical associations depend on how variants are correlated with each other locally, those correlations differ between populations. A score derived in one population therefore performs measurably worse when applied to individuals from a different ancestral background.

This is a technical property of the method rather than a claim about biological difference between groups, and it is widely acknowledged in the field. Efforts to build more diverse reference datasets are underway, and the imbalance has not yet been resolved.

What the score does not tell you

A score describes where someone sits in a distribution, and most people sit near the middle where the score changes very little. At the extremes the information content is greater, which is why research applications focus on identifying individuals in the tails. The score also does not incorporate environment, and for most traits environmental and behavioural factors contribute substantially.

It says nothing about timing, so even a meaningful elevation in estimated risk carries no information about when anything might occur. Presenting a percentile as a personal prediction therefore misrepresents a distributional statement as an individual forecast.

Nutrigenomic applications specifically

Applying this approach to diet requires the additional assumption that genotype meaningfully modifies how a person responds to a dietary change. Gene-diet interaction effects reported so far have generally been small, and several early findings failed to replicate in larger samples. Trials that assigned diets on the basis of genotype have been conducted, and their results have not supported the stronger versions of the personalisation claim.

The idea remains scientifically reasonable, and reasonable ideas require confirmation before being sold as a service. Where a service reports a dietary recommendation from a score, the relevant question is which trial supports acting on that specific score.

Interpreting a report responsibly

Consumer reports vary widely in which variants they include, in how they weight them and in whether they state the reference population. Two services analysing the same sample can therefore return different risk categories without either having made an error. Some findings from consumer sequencing genuinely matter clinically, particularly variants with large individual effects in known genes.

Those specific findings are confirmed with clinical-grade testing before being acted on, because consumer platforms are not designed for diagnosis. Anything from a genetic report that prompts real concern belongs with a clinician or a genetic counsellor rather than with a search engine.

The short version
  • Scores sum many weak associations rather than causal variants
  • Predictive performance drops in ancestries not represented in training
  • A score describes distribution position, not personal destiny
Nutrigenomicspolygenic scoresgeneticsrisk
David Smith
Contributing writer, My Healtheology

David Smith writes on nutrigenomics for My Healtheology, focusing on what the evidence supports rather than what makes the better headline.

Also by David Smith