Your ring says your readiness is 62. Your watch says your body battery is 41. Put a third device on the other wrist and you get a third verdict on the same night, and not one of those companies will show you the arithmetic that produced any of it.

Readiness, recovery, strain and sleep score are rollups rather than measurements. Each one takes a handful of signals a sensor can plausibly capture, applies weights to them, compares the result against some window of your own history, and returns a single number on a scale the company invented. The signals underneath are real enough. The number on top is a product decision.

What is a readiness score actually adding up?

In 2025 a research group at University College Dublin did the tedious thing nobody else had bothered to do. They went through the consumer wearable market and catalogued every composite score they could find, then tried to work out what each one was built from. They identified fourteen composite health scores across ten manufacturers: Fitbit Daily Readiness, Garmin Body Battery and Training Readiness, Oura Readiness and Resilience, WHOOP Strain, Recovery and Stress Monitor, Polar Nightly Recharge, Samsung Energy Score, Suunto Body Resources, Ultrahuman Dynamic Recovery, Coros Daily Stress, and the Withings Health Improvement Score.1

The ingredient list turned out to be fairly consistent. Heart rate variability appeared in 86 percent of the scores, resting heart rate in 79 percent, physical activity in 71 percent, and sleep duration in 71 percent.1 So four familiar inputs, recombined fourteen different ways and sold under fourteen different names.

The public descriptions match that. Oura says its Readiness Score reflects how balanced your recovery and activity are, and lists what feeds it: the lowest resting heart rate overnight and when it occurred, average body temperature, sleep quality, and movement from the previous day, plus longer-run HRV, sleep, and activity balance. The contributors with the word balance in them run on fourteen-day weighted averages, with the past two to five days weighted slightly heavier, compared against your own average over the previous two months. The scale is 0 to 100 and the bands are 85 to 100 Optimal, 70 to 84 Good, 60 to 69 Fair, and 0 to 59 Pay Attention.2

WHOOP Strain is a different shape of thing. It runs 0 to 21, it accumulates through the day and resets at midnight, and it is built from cardiovascular load derived from time spent in personalized heart rate zones plus a muscular load component that gets layered in when strength work is detected or logged. WHOOP states that the two are combined through a nonlinear formula that produces the single number, and that the zones are Light at 0 to 9, Moderate at 10 to 13, High at 14 to 17, and All Out at 18 to 21.3

Both of those are weighted composites and neither publishes the weights. I spent years reconciling other people's summary figures, and the first question is always the same one: what is in the total, and who decided how much each piece counts? Here the answer to the second half is the manufacturer, and it is not written down anywhere you can check it.

Has anyone published a validation of these scores?

The Dublin review found significant discrepancies between manufacturers in data collection timeframes, in how metrics were weighted, and in the proprietary scoring methods themselves. No manufacturer disclosed its exact algorithmic formula, and very few offered empirical validation that the score reflects anything physiologically meaningful. The authors concluded that the scientific validity, transparency, and clinical applicability of these scores remain uncertain.1

That is the finding, and it is worth being precise about it. The scores were not tested and found wanting. In most cases there is nothing published to test.

Sleep clinicians hit the same wall years earlier from a different direction. The American Academy of Sleep Medicine's 2018 position statement on consumer sleep technology noted that the data these devices generate is not standardized, and that raw data and proprietary algorithms are typically unavailable to clinicians.4 The 2017 paper that first named the anxiety problem put it more bluntly still: lack of transparency in the device algorithms makes it impossible to know how accurate they are even under the best circumstances.5

So a person can compare their 62 to yesterday's 71 and feel something about the difference, and there is no published basis for knowing what, if anything, those nine points represent. A metric nobody can audit is a mood with a number stapled to it.

What does heart rate variability actually tell you?

HRV is the single most common ingredient in these scores, so it is worth knowing what it actually indexes. It indexes neurocardiac function and is generated by heart and brain interactions and by dynamic, non-linear autonomic nervous system processes. It is a window onto autonomic regulation, and onto vagal activity in particular.6

It is also one of the most context-dependent numbers in physiology. Time-domain HRV measurements decline with age, with the sharpest drop between the second and third decades, though some measures show a U-shaped pattern across the lifespan rather than a straight decline. Women show higher mean heart rate and lower SDNN than men, along with a different distribution of frequency-domain power. Greater tidal volumes and lower respiration rates increase respiratory sinus arrhythmia. Posture matters. So does recording length, which affects both time-domain and frequency-domain values enough that the authors of the standard reference review state it is inappropriate to compare metrics like SDNN when they are calculated from epochs of different length, and that 24-hour, short-term, and ultra-short-term normative values are not interchangeable.6

Which means an HRV number is close to meaningless as a comparison between two people. Yours against a friend's compares two ages, two sexes, two body sizes, two breathing patterns, two measurement windows, and two proprietary sensor pipelines, and only incidentally two nervous systems.

Within a single person it moves a great deal too. A fourteen-day study of 41 healthy adults taking daily morning recordings found the coefficient of variation of RMSSD averaged 0.37, with individual values running from 0.14 to 0.71. Higher morning RMSSD tracked with better self-reported sleep, lower fatigue, and lower stress, but the day-to-day swing around a person's own mean was roughly a third of that mean.7 The sample was small and skewed young, and the authors say so.

A much larger analysis of roughly two million nocturnal HRV readings from more than 21,000 wearable users asked how many nights it takes to get a stable read on a person's HRV variability. The answer was at least five of seven nights to reach acceptable agreement with the full-week value.8

Five nights to characterize one week is the sample size the underlying signal demands. Your morning score is drawn from one.

Why do two devices disagree about the same night?

Because they disagree about the inputs before they ever reach the weights. In a 2025 validation study, 62 adults slept one night under polysomnography while wearing two to four consumer devices at the same time. Six devices were evaluated in total. Agreement with the lab standard on sleep staging ran from fair to moderate, with kappa coefficients between 0.21 and 0.53. Every device detected sleep well, with sensitivity above 90 percent, and detected wake poorly, with specificity between 29 and 52 percent. Wake after sleep onset was underestimated by 12 to 48 minutes depending on the device, and the authors observed a strong tendency to treat light sleep as a fallback classification when the algorithm was uncertain.9

An earlier study ran 34 healthy young adults across three consecutive nights, including one deliberately disrupted night, against polysomnography with seven consumer devices and research actigraphy. Most devices performed as well as or better than actigraphy on sleep and wake detection, with high sensitivity and low-to-medium specificity, and performance degraded on the disrupted night.10

Pooled across the literature, a 2025 meta-analysis of 24 studies and 798 participants found statistically significant differences from polysomnography of roughly 17 minutes in total sleep time, roughly 4.7 percentage points in sleep efficiency, roughly 2.6 minutes in sleep latency, and roughly 13 minutes in wake after sleep onset. Heterogeneity between studies was high, with I-squared values between 57.9 and 84.1 percent, and the authors conclude these devices are not yet valid substitutes for polysomnography while remaining useful for tracking general sleep patterns.11

Now stack the composite score on top of that. Two devices measure the same night and land twenty minutes apart on how much of it you were awake. Then each feeds its own answer into its own undisclosed formula, weighted its own way, compared against its own baseline window, and prints a verdict on its own invented scale. Of course the two verdicts differ. The only surprising outcome would be if they matched.

What happens when someone starts optimizing for the number?

There is a documented pattern here, and it is easy to overstate in either direction.

In 2017 Kelly Glazer Baron and colleagues published a three-patient case series describing people who arrived at a sleep clinic preoccupied with improving or perfecting their wearable sleep data. They coined the term orthosomnia, built on the same root as orthorexia. One patient reported daytime problems only on mornings when his tracker showed under eight hours. Another pasted tracker printouts in place of a sleep diary and was eventually given polysomnography that showed normal deep sleep while her tracker had been reporting poor sleep. The authors recommended that clinicians walk patients through what these devices can and cannot detect, and steer them toward tracking their sleep pattern and time in bed rather than the minute-by-minute split between wake and sleep.5

Three patients is three patients. The authors said so themselves, noting that they did not know about the patients' sleep before the tracker and therefore could not say whether the tracker caused the problem.5

A 2024 cross-sectional study tried to put a number on how common the pattern is. Researchers surveyed 523 adults recruited online and through community advertising, median age 21 and 81 percent female, then applied a four-part definition: owning and regularly using a sleep tracker, an Athens Insomnia Scale score at or above 6, a generalized anxiety score below the severe threshold, and sleep preoccupation above a cutoff. Depending on how strictly that last cutoff was set, prevalence came out at 3.0 percent, 8.6 percent, or 14.0 percent. Across every cutoff, people meeting the definition had higher insomnia scores than those who did not.12

The authors are direct about what that does and does not establish. The cross-sectional design prevents drawing causal conclusions. The identifying algorithm has not been validated against clinical diagnosis. The sample may not represent the general population.12

The honest summary is that clinicians have described the pattern, a survey found it associated with worse insomnia symptoms in a young and self-selected sample, and nobody has demonstrated that tracking causes it. That is weaker than the headlines suggest and considerably stronger than nothing. If sleep problems persist, the next step is a clinician rather than a different ring.

What is the data genuinely good for?

Quite a lot, once you stop asking it for a verdict.

The thing these devices count most reliably is when you were in bed and roughly how long you slept, and it turns out consistency of timing carries real weight. A study of 60,977 UK Biobank participants, each wearing a wrist accelerometer for a week, calculated a Sleep Regularity Index, which scores how closely your sleep and wake states line up from one 24-hour period to the next. The four most regular quintiles showed 20 to 48 percent lower all-cause mortality risk than the least regular quintile, and sleep regularity was a stronger predictor of all-cause mortality than sleep duration was.13 The study is correlational, drawn from a single week of recording in an older and largely homogeneous cohort, and the authors flag all of that.13

Timing regularity is exactly the kind of thing a wrist sensor can count without needing to guess at sleep stages. So use the device for what it can actually count:

  • Duration. How many hours you are getting, tracked across weeks rather than judged one night at a time.
  • Timing. Whether your bedtime and wake time hold steady from day to day, which is the metric with the strongest outcome evidence behind it.
  • Direction. Whether your own resting heart rate or HRV baseline has moved over several weeks, not whether this morning came in under yesterday.
  • Change. A sustained departure from your personal baseline is worth noticing and worth mentioning to a clinician. A single reading is not.

Why is none of this diagnostic?

Because it was never built to be, and the regulator has drawn that line explicitly. The FDA's general wellness guidance, reissued on 6 January 2026, sets out the compliance policy for low-risk products that promote a healthy lifestyle. It says the agency may treat products that use non-invasive sensing to estimate or infer physiologic parameters, naming heart rate variability among them, as general wellness products when the outputs are intended solely for wellness use. Such products may display values, ranges, trends, baselines, or longitudinal summaries, and may contextualize those outputs in relation to sleep, activity, stress, or recovery.14

The conditions attached are the interesting part. To stay inside the policy, a product must not be intended for diagnosis or treatment, must not substitute for an authorized medical device, must not include outputs that prompt or guide specific clinical action, and must not include values that mimic those used clinically unless validated to reflect those values.14

The guidance then says something that ought to be printed on every onboarding screen in the category: a product's inclusion under the general wellness policy does not establish that it has been shown to be safe or effective for its intended use.14

Sitting inside the wellness lane is not a finding of accuracy; it is a statement that the agency has chosen not to look. Sleep medicine's professional body says the same thing from the clinical side: given the lack of validation and FDA clearance, consumer sleep technologies cannot be used for the diagnosis or treatment of sleep disorders, and the data they generate should be considered within a full clinical evaluation rather than replacing validated diagnostic testing.4

So how should you read tomorrow morning's number?

Trace what it is before you react to it. Tomorrow's score is a weighted combination of about four inputs, at least one of which the device measured imperfectly and one of which swings roughly a third around your own mean from day to day. Those inputs pass through a formula the manufacturer has not published, over a comparison window the manufacturer has not fully specified, against a baseline assembled from your own past readings, and the result arrives on a scale the manufacturer invented and named.

Exactly one link in that chain is auditable by you, and it is the last one. You can see whether your own numbers are drifting in one direction over weeks. You cannot see, and no published source will tell you, what any single morning's composite is supposed to mean.

None of which is an argument for throwing the device in a drawer. Duration and timing are worth having, and a sustained change from your own baseline is a real signal worth raising with someone qualified to interpret it. It is an argument for reading the score as what it structurally is: a vendor's opinion about your body, delivered with all the confidence of a measurement and without the disclosure that would let anyone check the work.