Every morning, several million people reach for a phone before they are properly awake and read a verdict on the night they just had: minutes in deep sleep, minutes in REM, a score out of one hundred. The number arrives with the composure of a lab result, and most people treat it as one.
On August 20, 2026, a proposed class action was filed in the United States District Court for the Northern District of California, captioned Surber v. Oura, Inc., case number 3:26-cv-08686. The Clarkson Law Firm brought it for a named plaintiff, Madison Surber, who says she paid about $513.68 for an Oura Ring 4 in May 2025. Nothing in it has been proven. Oura disputes the allegations.1
The filing is worth reading for the arithmetic rather than the outrage.
What does the complaint actually allege?
The complaint does not argue that the ring is useless, only that one advertised figure cannot be supported by the hardware. Sleep stages are defined clinically by polysomnography, which records brain waves, eye movement, cardiac signals and muscle activity. REM is named for the eye movement that defines it. Deep sleep is defined by large amplitude slow wave activity in the brain. A ring on your finger records none of that. It records pulse through the skin, motion, peripheral temperature and, in newer models, blood oxygen trends, and a model turns those proxies into a stage label for every thirty second window of the night.1
The products named are the Oura Ring 5, the Oura Ring 4 and the Oura Ring 4 Ceramic. The advertising quoted in the filing includes "Built for accuracy", "Unparalleled Accuracy", "79% agreement with gold-standard polysomnography (PSG) for classifying the four stages of sleep" and "95% Sleep Staging Accuracy Compared to clinical sleep lab". The relief sought is restitution and an injunction against continuing to advertise that way. Jurisdiction is pleaded under the Class Action Fairness Act, which requires at least one hundred class members and more than five million dollars in controversy.1
What number sits under the footnote?
The claim is still on the product page. On the Oura Ring 5 page, three statistics sit in a row under the heading "Rebuilt for greater accuracy". The third reads 95%, labelled "Sleep Staging Accuracy", with the subtext "Compared to clinical sleep lab" and a superscript that jumps to the legal footnotes at the bottom of the page. The same page states elsewhere that the Oura Ring is not a medical device and is not intended to diagnose, treat, cure, monitor or prevent medical conditions.2
Footnote two names exactly one paper. "Accuracy of Oura Ring is validated by comparisons to polysomnography (PSG) based sleep staging, with participants measured simultaneously by PSG and the Oura Ring over multiple nights. Altini M, Kinnunen H. The Promise of Sleep: A Multi-Sensor Approach for Accurate Sleep Stage Detection Using the Oura Ring. Sensors. 2021; 21(13):4302."2
So read the paper. It is open access, it covers 440 nights from 106 people across three continents, and it was written by an Oura Health employee and an Oura Health adviser, both disclosed. The figure 95 percent appears in it once, in the section on two stage classification: accuracy for telling sleep from wake was 94 percent using the accelerometer alone, 95 percent once skin temperature was added, and 96 percent with heart rate variability and circadian features included. The four stage number, the one that separates wake from light, deep and REM, is 79 percent for the full model and 57 percent for the accelerometer alone.3
That is the gap. The label above the number says staging, while the number itself, in the only paper the footnote cites, is the sleep and wake figure from a study whose staging figure is 79.
Oura does not really contest that reading. In a post published on August 23, 2026, the company set out how it reports accuracy across different dimensions and said that sleep versus wake detection reaches 90 to 96 percent agreement with polysomnography, while four stage classification runs around 76 to 79 percent. The company said it disputes the allegations and stands behind its research.4
Why does sleep versus wake accuracy flatter almost any device?
Because of the denominator. On a scored night in a sleep lab, the overwhelming majority of thirty second epochs are sleep. A model that leans toward guessing "asleep" will be right most of the time without knowing anything interesting. The accuracy percentage inherits the base rate, and then gets printed as though it were a measurement of skill.
You can see the effect in the numbers themselves. A 2024 validation of the Oura Ring Generation 3 led by researchers at the University of Tokyo, covering 96 participants and 421,045 epochs of ambulatory polysomnography, reported sensitivity for detecting sleep of 94.4 to 94.5 percent and overall accuracy of 91.7 to 91.8 percent. In the same study, specificity, the ability to correctly call wake, was 73.0 to 74.6 percent, and the predictive value for wake was 66.6 to 67.0 percent.7
The pattern is not unique to rings. An independent Belgian validation of six wrist worn devices against polysomnography in 62 adults found that every device detected more than 90 percent of sleep epochs, while specificity ranged from 29.39 to 52.15 percent, and agreement measured by Cohen's kappa ranged from 0.21 to 0.53, which is fair to moderate.10
The clearest statement of the problem is in Oura's own footnoted paper. Discussing a bed sensor that reported 79 percent total accuracy on a kappa of 0.43, the authors wrote that this highlights "how accuracy is not an ideal metric for a classification problem with highly imbalanced data". That sentence is a company scientist telling you, in print, not to read the headline percentage the way the product page invites you to read it.3
What do the better studies say about stages?
There is a real literature here, and it does not all point one way.
- Company funded, favourable, and disclosed. A 2024 single night inpatient study at Brigham and Women's Hospital put 35 healthy adults on polysomnography while they wore an Oura Ring Gen3, a Fitbit Sense 2 and an Apple Watch Series 8. The Oura Ring agreed with polysomnography on 92.0 percent of epochs for sleep versus wake and 76.3 percent in the four stage comparison. Per stage sensitivity across the three devices ranged from about 50 to 86 percent. The paper states that the research was funded by Oura Ring Inc., and its first author discloses membership of the Oura Ring Medical Advisory Board and consulting fees from the company.5
- The same study, as marketed. Oura's own write up of that paper leads with 79 percent agreement in four stage classification and a comparison showing the ring ahead of the Apple Watch and the Fitbit. To the company's credit, the post says outright that the study was funded by Oura.6
- Independent, in a real clinic, much worse. A 2025 study at Charité in Berlin tested three finger ring trackers against polysomnography across 45 measurement nights on 45 patients drawn from a clinical population. The Oura Ring reached 85.03 percent accuracy for sleep versus wake, and 53.18 percent for four stage classification, on a kappa of 0.31. The authors declared no competing interests. This is the study the complaint leans on.8
- Independent, and worse with age. A 2026 study funded through a National Institutes of Health centre grant tested consumer devices including the Oura Ring in healthy young adults and adults aged 56 to 80. Accuracy fell in the older group, with total sleep time and wake after sleep onset underestimated and deep sleep overestimated. The authors recommend caution in interpreting consumer device output, particularly for older adults.9
Set against all of that, a 2025 systematic review and meta analysis pooling six studies and 388 participants found no statistically significant differences between the Oura Ring and polysomnography or actigraphy for total sleep time, sleep efficiency, wake after sleep onset, sleep onset latency, light sleep, deep sleep or REM duration, and concluded the device is useful as a self monitoring tool.11
Both of those things are true at once. Totals can agree while the minute by minute assignment is close to a coin toss. If a device calls REM when you were in light sleep for twenty minutes, and calls light sleep when you were in REM for twenty minutes somewhere else in the night, the nightly totals come out fine and the picture of your night is wrong. Aggregate agreement is not evidence that the timeline is right.
Does group level agreement tell you anything about your own night?
Less than you would hope. The Berlin authors put it directly: group level measurements showed modest differences, but individual level differences often remained large, and that spread was their reason for saying these devices cannot yet be used in clinical sleep medicine. A mean bias near zero can sit on top of very wide limits of agreement. The average customer is fine. You are not the average customer, and you are the one reading the score.8
The company funded Brigham study is unusually honest about a second problem. Its authors note that they only compared devices to polysomnography during the scheduled sleep episode rather than across a full 24 hours, and that this "overestimates concordance, since wearables erroneously score quiet wakefulness" such as reading or watching a film. The test window was chosen to be the window in which sleep was nearly certain. That is a favourable denominator, disclosed in the paper and absent from the marketing.5
One more piece of context rarely makes it onto a product page. Two qualified human technicians scoring the same polysomnography record agree with each other only about 80 percent of the time. The gold standard has its own error bars, which cuts both ways: it makes a ring's 76 percent less embarrassing, and it makes any claim of near total accuracy against that standard harder to justify, not easier.5
Who is supposed to check this before it reaches a product page?
Nobody, in the way most buyers assume.
The American Academy of Sleep Medicine settled its position in 2018: consumer sleep technologies cannot be used for the diagnosis or treatment of sleep disorders, because they have not been validated against polysomnography to the required standard and have not been cleared by the Food and Drug Administration. The academy allowed that such devices may help a conversation between patient and clinician, and asked for validation, access to raw data, algorithm transparency and regulatory oversight. Eight years later, that list is still a list of requests.12
The regulatory position will not surprise a careful reader. The FDA's general wellness guidance, reissued on January 6, 2026, says the Center for Devices and Radiological Health does not intend to examine low risk general wellness products to determine whether they are devices at all, or whether they comply with premarket review, labelling, quality system or adverse event reporting requirements. "Claims to promote sleep management, such as to track sleep trends" are listed in the guidance as an example of a general wellness claim. And the guidance contains a sentence that ought to be printed on every wearable box: "A product's inclusion under the general wellness policy in this guidance does not establish that it has been shown to be safe and/or effective for its intended use." The document also states on its face that it is nonbinding and creates no legally enforceable responsibilities.13
So the accuracy number is not audited by anyone before it goes on the page. It is a marketing claim with a footnote, and the footnote is where the work either was or was not done.
What is the cost of believing a number that is softer than it looks?
It is not only money. Clinicians named the failure mode nearly a decade ago. Orthosomnia describes patients who arrive seeking treatment for sleep problems they diagnosed from tracker output, chasing a perfect night as measured by a device, and who trust the tracker's account of their sleep more than validated measurement or their own experience. The clinical challenge, the authors wrote, is balancing education about what these devices can actually establish against a patient's enthusiasm for objective data.14
That is the cost of a precise looking estimate. Precision reads as authority, and where a range would invite the reader to exercise judgment, a single integer forecloses it.
What would an auditor ask before signing off on this number?
The same three questions you would ask about any figure that has been put in front of you as a result rather than an estimate. What is the denominator? What exactly was being classified? Who paid for the count?
Run those here and the answers are on the record. The denominator is a scored night in which most epochs are sleep, which is why a two state score climbs into the nineties. The thing classified in the 95 percent figure is sleep against wake rather than stage against stage, and the paper it points to reports 79 percent for stages. The paper in the footnote was written by the manufacturer's own staff, the favourable Brigham study was paid for by the manufacturer, and both say so; the two unfavourable studies came from a university hospital and a federal grant, and say so as well. None of that makes the ring a fraud, and none of it has been decided by any court. It does mean the number has been carrying more weight in the marketing than the citation underneath it can hold.
The commercial stakes explain why the number stays. Oura announced on May 21, 2026 that it had confidentially submitted a draft registration statement to the Securities and Exchange Commission for a proposed initial public offering, and introduced the Ring 5 a week later.15 The complaint alleges the company reached a valuation of roughly $11 billion after a 2025 funding round, with more than 5.5 million rings sold.1 Asked about the suit, a spokesperson said the company stands behind its science, research and accuracy claims, and that it will defend against the allegations in the appropriate legal forum.16
The honest version of what a finger ring can know is narrower than the product page suggests. It can tell, very reliably, whether you were asleep. It can tell you roughly how long, and it can track that consistently enough night after night that the trend line means something. What it cannot do is watch your brain change state, and every stage number it gives you is an inference from your pulse, your movement and your skin, run through a model that was trained to reproduce a scorer's judgment it cannot observe.
None of that is a scandal. Inference is a legitimate way to produce a useful number, and the trend across weeks is probably the most valuable thing the device gives you. The defect is the presentation. A model output that is right about three times in four gets rendered as minutes and seconds, sitting under a percentage borrowed from an easier question, and the buyer is not given the denominator that would let them discount it. That is a disclosure choice rather than a technology problem, and it survives because no regulator inspects it, no auditor signs it, and the only party with an incentive to print the harder number is the party selling the ring.

