Sleep Biohacking
How Sleep Stages Are Scored And Why Wearables Approximate
Formal sleep staging relies on electrical recordings from the scalp scored in fixed intervals, and wrist devices infer the same categories from motion and heart signals instead.

Sleep stage summaries appear on almost every tracking device, presented as measured quantities. The reference method they approximate is quite different, and the gap explains most disagreements between devices.
The reference method records three signals
Laboratory sleep studies record electrical activity from the scalp, movement of the eyes, and muscle tone from beneath the chin. These three together define the stages.
Each signal contributes something distinct. Brain activity distinguishes depth by wave frequency and amplitude, eye movement identifies one stage specifically, and muscle tone drops to near zero in that same stage.
No single signal is sufficient. The stages are defined by combinations, which is why removing any one of the three degrades the classification substantially.
Scoring proceeds in fixed windows
The recording is divided into intervals of thirty seconds, and each interval is assigned a single stage according to published rules based on what predominates within it.
This creates a coarse output by design. Transitions occurring mid-interval are not represented, and brief intrusions of another state are absorbed into the surrounding assignment.
Trained scorers applying the same rules to the same recording agree most of the time but not always, with disagreement concentrated at transitions and in lighter stages.
Wearables have none of these signals
A wrist device records movement and an optical pulse signal. It cannot detect brain activity, eye movement or muscle tone, so it cannot apply the reference rules at all.
Instead it uses a model trained on data where wrist signals were recorded alongside laboratory scoring, learning which patterns of motion and heart variation accompanied each stage.
The model reproduces the labels statistically rather than measuring the underlying states. Its output is a best guess about what a scorer would have written.
Some distinctions survive better than others
Separating sleep from wakefulness works reasonably well, since movement differs markedly and long still periods are strong evidence of sleep.
The main weakness is quiet wakefulness. Someone lying still while awake resembles sleep in both movement and heart signals, so devices tend to overestimate sleep duration.
Distinguishing stages within sleep is harder still, and validation studies consistently show lower agreement for stage classification than for the simple sleep-wake distinction.
What the numbers support
Duration and timing are the most reliable outputs, and both are genuinely useful for observing patterns across weeks rather than judging a single night.
Stage percentages should be read as estimates carrying substantial uncertainty, and comparing them between devices or against published norms is not meaningful.
Persistent poor sleep, loud snoring with pauses, or daytime sleepiness are matters for clinical assessment, since several sleep disorders require diagnostic testing a wearable cannot perform.
Also by Dr. Francis Collins
- Science-Backed Strategies: Refining nad precursors synthesis for Everyday FocusAdvanced Therapies
- Science-Backed Strategies: Refining nad precursors synthesis for Everyday Focus (Insights)Advanced Therapies
- Science-Backed Strategies: Refining nad precursors synthesis for Everyday Focus (Overview)Advanced Therapies
- Science-Backed Strategies: Refining nad precursors synthesis for Everyday Focus (Tactical Update)Advanced Therapies




