Field note
A sleep score is a single morning number that folds several overnight signals into one grade. Most consumer scores mix how long you slept, how continuous that sleep looked, stage estimates, and sometimes heart or breathing summaries. It is a summary, not a lab result.
People treat the number like a teacher marked the night. It is closer to a weighted average of noisy estimates. Useful when you watch the week. Misleading when one morning becomes a verdict.
If you only remember one sentence: duration and continuity should lead the story, and stage ribbons should not.
Ohayon and colleagues (2017) published the National Sleep Foundation's first sleep quality recommendations. Quality in that framing is not one magic metric. It covers continuity, how long it takes to fall asleep, how often you wake, and how the night feels. Consumer scores try to compress a similar idea into a 0-100 style grade so a busy morning has something to glance at.
Carskadon and Dement (2017) described normal human sleep as a structured night with cycles of lighter sleep, deeper sleep, and REM. Diekelmann and Born (2010) reviewed how those stages support memory functions. That science is real. The leap from lab stages to a wrist chart is where the error bars widen.
| Ingredient many scores use | How sturdy it usually is on a wrist |
|---|---|
| Duration (time asleep) | Often the sturdiest overnight estimate |
| Continuity (wake-ups, fragmentation) | Useful when motion and heart signals agree |
| Stage labels (deep, REM, light) | Estimates; weaker than sleep-versus-wake |
| Heart or breathing summaries | Context, not a sleep diagnosis |
Different apps weight those ingredients differently. There is no universal recipe that every device shares. Treating any brand's formula as medical truth is a mistake. Read the score as "how this device summarised last night," not as "how healthy your sleep is forever."
Chinoy and colleagues (2021) compared seven consumer sleep-tracking devices with polysomnography, the lab standard. Lee and colleagues (2023) tested eleven wearable, nearable, and airable trackers in a multicenter validation study. Across that validation work, devices are generally stronger at telling sleep from wake than at naming deep versus REM versus light with lab precision.
That is the limit worth knowing once. Stage charts look clinical. They are still estimates. Are Apple Watch sleep stages accurate? is the full note on that gap. If a score leans hard on stage minutes, treat the grade more gently than a score led by duration and continuity.
I am citing Chinoy (2021) and Lee (2023) as they sit in the Field Guide source list. I have not opened the full texts for device-by-device percentage tables here, so I am not inventing a specific accuracy number for your watch model.
One low score after a late meal, a drink, travel, or a hard training day is ordinary noise. A run of softer nights is the pattern worth noticing. That is the same logic as sleep debt. What is sleep debt? explains how short nights stack even when you stop feeling as tired.
A readiness or training score often borrows from the same overnight window. What does a readiness score mean? is how devices turn recent load and recovery signals into a training call. Do not confuse a sleep grade with a training permission slip, or the reverse.
When the morning number and how you feel disagree, believe the week and your own head more than a single composite. If duration looked fine and continuity looked fine, a middling score driven by stage guesses is a weak reason to panic.
Glance at duration first. Then continuity. Then the trend across several nights. Use stages as colour, not as the judge. Ask whether bedtime and wake drifted. Regularity belongs beside the score even when the score itself is quiet about timing.
If you change one thing after a soft week, make it earlier and steadier nights before you chase a deeper-sleep target the wrist cannot measure well. The sturdy levers are the ones the sensors estimate best.
A sleep score is a compressed overnight summary. Duration and continuity deserve the most trust. Stage labels deserve the least.
Watch the week. Let one morning stay a morning.
I write these because I build Aera, an iPhone app that reads Apple Health against your own baseline and puts the research and the error bars underneath every number it shows you, including this one. You do not need it to use anything above.