Walk through a department store beauty hall and someone will offer to scan your face. Open a skincare brand's website and a panel will invite you to upload a selfie for a free skin analysis. Download the app and it will grade your pores, your wrinkles, your redness and your evenness, then hand you a number. Then it will hand you a basket.
I have not used these tools, and this is not a review of any of them. What I want is the thing an accountant always wants when a number appears on a screen: what was measured, against what standard, verified by whom, and who profits from the answer. All four questions have public answers. Most of those answers are not in the brochure.
What is the scan actually measuring?
The consumer version of skin AI works from a photograph. You point a phone at your face in whatever light you happen to be standing in. Software finds the face, maps landmarks onto it, and scores regions for a fixed list of surface features: apparent wrinkle depth, pore visibility, redness, dark spots, evenness of tone. Sometimes it collapses all of that into a single figure and calls it your skin age.
Every one of those is a description of a picture rather than a finding about tissue. The distinction sounds pedantic until you notice that the commercial pitch depends entirely on blurring it. A wrinkle score measures contrast and shadow in an image. It moves with the angle of your head, the color temperature of the bulb above you, the sensor in your phone, and whether you were smiling. Nothing in that pipeline reaches below the surface of the photograph, and nothing in it needs to, because the destination was always a product page.
Where does the law draw the line?
The Food and Drug Administration publishes the test it applies, and it repays reading in the original rather than in a brand's paraphrase. A general wellness product, in FDA's definition, is one with "an intended use that relates to maintaining or encouraging a general state of health or a healthy activity," or one that connects a healthy lifestyle to reduced risk of a chronic disease where that link is already well accepted.1 The qualifying claim categories in that guidance include, in FDA's own parenthetical, "devices with a cosmetic function that make claims related only to self-esteem."1 That is the doorway a beauty scanner walks through. It scores how your skin looks and suggests a serum. It never claims to find disease, so it never has to ask FDA for anything.
The door on the other side of that hallway is much narrower than most people assume. Congress excluded certain clinical decision support software from the device definition in 2016, and FDA's guidance sets out four criteria, all of which must be met.2 The first is the one that governs here. The software must not be "intended to acquire, process, or analyze a medical image or a signal from an in vitro diagnostic device or a pattern or signal from a signal acquisition system."2 FDA then defines a medical image far more broadly than the phrase implies. It includes "images acquired for a medical purpose (e.g., pathology, dermatology)," and adds that "images that were not originally acquired for a medical purpose but are being processed or analyzed for a medical purpose are also considered medical images."2
A selfie is not a medical image because of the camera. It becomes one because of the purpose. The identical photograph is a wellness input while the software grades your pores and a medical image the instant the software assesses it for disease. The exclusion also has a second wall. FDA states that software functions supporting or providing recommendations "to patients or caregivers," rather than to health care professionals, "meet the definition of a device."2 Consumer-facing software that assesses disease is a device. There is no carve-out waiting for it.
So the line is not drawn by the technology. It is drawn by the claim. Two products can run near-identical computer vision over near-identical photographs and sit in entirely different regulatory universes, because one says "here is your pore score" and the other says "this looks suspicious." The first has to prove nothing. The second has to prove a great deal.
Which dermatology AI tools has the FDA actually authorized?
FDA maintains a public list of AI-enabled medical devices it has cleared, granted or approved. In the version published on 16 June 2026, that list runs to 1,516 entries. Radiology accounts for 1,162 of them. The advisory panel that covers skin, General and Plastic Surgery, accounts for eight, and most of those eight are surgical systems and tissue probes. Exactly three assess a skin lesion.3
- MelaFind. Approved 1 November 2011 under premarket approval P090012, for use by dermatologists on clinically atypical pigmented lesions when deciding whether to biopsy, and the approval order states it should not be used to confirm a clinical diagnosis of melanoma.4
- Nevisense. Approved 28 June 2017 under P150046, an electrical impedance spectrometer indicated for use "when a dermatologist chooses to obtain additional information when considering biopsy."5
- DermaSensor. Granted 12 January 2024 under De Novo request DEN230008, a handheld spectroscopy device for physicians who are not dermatologists, used on lesions already assessed as suspicious.6
Three devices in thirteen years. All three are prescription devices, operated by a clinician, in a clinical setting, on a lesion a human being has already decided looks wrong. Each one begins after a professional is already worried, which is a different job from screening. FDA cautions that the list is assembled from AI-related terms in authorization documents and "is not a comprehensive resource," so treat the count as a floor rather than a census.3 The shape of it still holds. The number of consumer beauty scanners on that list is zero.
What did the newest authorized device actually score?
DermaSensor is the most instructive of the three, because FDA published the full decision summary and the numbers inside it are more sobering than the coverage was.7 The pivotal study, DERM-SUCCESS, enrolled 1,005 participants with 1,579 lesions across 22 sites, and every enrolled lesion went to biopsy.7 Overall sensitivity for malignancy was 95.5 percent, against 83.0 percent for the primary care physicians assessing the same lesions. That is the figure the headlines carried.
Specificity was 20.7 percent. The physicians' specificity was 54.2 percent.7 In plain terms: the device caught cancers the doctors were missing, and it also flagged roughly four out of five benign lesions as worth investigating further. FDA accepted that trade deliberately, on the reasoning that a missed melanoma is far worse than an unnecessary referral, and because the device is adjunctive, one input to a physician who is already concerned.7 That is a defensible design. It is also the precise opposite of what a consumer wants from a scan, which is reassurance.
The decision summary carries one more line no marketing department would volunteer. In repeatability testing, FDA found that "the device shows low repeatability and reproducibility of negative binary decisions," low repeatability of the similarity score it displays, and "high variability of the underlying continuous model output."7 Scan the same lesion twice and a negative result may not stay negative. FDA judged this acceptable in context and required it in the labeling.7 Now imagine the equivalent disclosure printed under a beauty app that tells you your skin age fell four years.
What happened when researchers tested the consumer apps?
The strongest evidence on phone apps that assess lesions is a systematic review published in The BMJ in February 2020, covering nine studies of six identifiable apps.8 SkinScan was evaluated in a single study of 15 lesions containing five melanomas, and detected none of them: sensitivity of 0 percent, specificity of 100 percent.8 SkinVision, evaluated across two studies, reached 80 percent sensitivity (95% CI 63 to 92) and 78 percent specificity (67 to 87) for malignant or premalignant lesions.8
How those studies were run matters as much as the numbers. Lesion selection and image capture were performed by clinicians, not by the people who would actually be using the app on themselves.8 Up to 45 percent of images in one study were unevaluable.8 The reviewers concluded that current algorithm-based apps "cannot be relied on to detect all cases of melanoma or other skin cancers," that real-world performance is "likely to be poorer than reported here," and that the European CE marking process for these apps "does not provide adequate protection to the public."8
The review also records what happened to two of the six. MelApp and Mole Detective were withdrawn from the market following Federal Trade Commission investigations into what the reviewers describe as "deceptively claiming the apps accurately analysed melanoma risk."8 Consumer protection reached those products before any medical regulator did, and it reached them on the advertising, not on the algorithm.
Whose skin was in the training data?
The founding result in this field is a 2017 paper in Nature, in which a convolutional network trained on 129,450 clinical images spanning 2,032 diseases matched 21 board-certified dermatologists on two binary classification tasks.9 It is a genuine achievement, and its closing observation, that mobile devices could extend the reach of dermatologists outside the clinic, is the sentence the industry took and ran with for a decade.9
Now look at what the field actually trains on. HAM10000, among the most widely used public benchmarks, consists of 10,015 dermatoscopic images gathered over twenty years from two sites: the Department of Dermatology at the Medical University of Vienna and a skin cancer practice in Queensland, Australia.10 Those are images taken through a dermatoscope, by clinicians, in two of the fairest-skinned patient populations on earth. Nothing about that is misconduct. It is simply what was available to collect.
A 2022 systematic review in The Lancet Digital Health surveyed the whole field. It identified 21 open access datasets holding 106,950 skin lesion images. Fitzpatrick skin type was recorded for 2,236 of those images, which is 2.1 percent. Ethnicity was recorded for 1,415 images, which is 1.3 percent. Of the 14 datasets that reported a country of origin, 11 held images from Europe, North America or Oceania only. The reviewers found "substantial under-representation of darker skin types."11
The 2.1 percent is not a statement that dark skin is under-represented. It is a statement that for 97.9 percent of the images in the public record, nobody wrote down whose skin it was. You cannot measure a gap you never recorded. That is a data integrity failure before it is an equity failure, and in this case they are the same one.
Researchers eventually built the missing benchmark, and the published scores moved a long way. The Diverse Dermatology Images dataset, published in Science Advances in 2022, assembled 656 pathologically confirmed images balanced across skin tones: 208 at Fitzpatrick I to II, 241 at III to IV, and 207 at V to VI.12 Three well-known dermatology AI models were run against it. ModelDerm, previously reported at an area under the curve of 0.93 to 0.94, scored 0.65. A model trained on HAM10000 fell from 0.92 to 0.67. DeepDerm fell from 0.88 to 0.56.12
An area under the curve of 0.56 is very close to a coin toss. The same study found that the dermatologists who label these datasets also performed worse on dark skin, which means the bias is baked into the ground truth before a model ever sees it.12 Fine-tuning on the diverse images closed the performance gap, which is the encouraging half of the finding and the half that requires somebody to actually go and do it.12
None of this is confined to consumer software. Return to the DermaSensor pivotal study, the one FDA authorized. Of 1,005 participants, 976 were White, or 97.1 percent. Eighteen participants, 1.8 percent, were Fitzpatrick phototype VI.7 FDA wrote the consequence into the device limitations: "Consistent with the lower prevalence of skin cancer in Fitzpatrick skin phototypes IV-VI, less data is available for sensitivity of the DermaSensor device for melanoma in these patients," and the decision to refer such a patient "should be primarily based on clinical concern."7 That is a regulator writing down in public the thing a product page never will.
Lower prevalence is real, and it is not the same thing as lower stakes. An analysis of 96,953 cutaneous melanoma cases from the SEER database found White patients had the longest survival, that Black patients had significantly lower survival at stages I and III, and that a greater proportion of cases in Black patients were diagnosed at later stages.13 A technology validated almost entirely on the population least likely to die of the disease is a strange instrument to hand to everybody.
Who is selling the answer?
Set accuracy aside for a moment and look at the structure. A diagnostic operated by the party selling the remedy is a conflict of interest. That is not an accusation about any company's conduct. It is a description of the arrangement. The output space is bounded by the catalogue. A scanner that can only return conditions the brand happens to formulate for is filling in an order form.
When the honest result is that your skin is fine, that answer sells nothing, so no part of the system has any reason to make it appear. Nobody has to rig anything for the outcome to tilt. You simply build a scoring scale where zero is not a place the needle usually lands, and let the arithmetic do the rest.
Two examples, quoted from the companies' own pages rather than from anyone's coverage of them. L'Oreal Paris says its Skin Genius tool "compares your selfie with more than 10,000 clinically graded images," is "up to 95% accurate compared with a live dermatologist consultation," was "validated in women of a wide range of ethnic origins and skin types," and "makes personal recommendations by choosing from more than 30 products." The same page offers three dollars off a skincare purchase once the analysis is done.15 Vichy, also part of the L'Oreal group, says its SkinConsult AI works "with more than 98% accuracy," drawing on "a skin strength database with 25,000 graded photos" to analyse "7 different skin concerns." That page offers twenty percent off two or more products.16
Neither page cites a study. Not a journal, not an author, not a year. Searching the full text of the Skin Genius page for the phrases "medical device" and "diagnosis" returns nothing at all.15 The same search on the Vichy page returns nothing either.16 Those absences are doing real work. They are what keeps the tool on the wellness side of the line described earlier, and they are also why the accuracy figures have no visible referent. Ninety-five percent of what, measured against whom, is not a question the page invites you to ask.
The research does exist, and it repays reading alongside the marketing. In 2022 the Journal of the European Academy of Dermatology and Venereology published a cross-sectional study of selfie images from 1,041 US women, in which seven facial signs were graded both by an automated algorithm and by 50 US dermatologists working from the same reference atlas.14 Eight of the thirteen authors were employed by L'Oreal or by ModiFace, a L'Oreal group company.14 For five of the seven signs, the algorithm correlated strongly with the dermatologists, at r of 0.75 or above. Cheek pores were moderate, at r of 0.63. Pigmentation signs, "especially for the darkest skin tones," were weakly correlated, at r of 0.40.14
The authors' own conclusion is that the procedure is accurate and clinically relevant "although skin tone requires further improvement."14 Now hold that against the sentence on the product page: validated in women of a wide range of ethnic origins and skin types. Neither statement is false. The weakest agreement in the entire study was on pigmentation in the darkest skin, the company's own scientists wrote it down in a peer-reviewed journal, and the customer never sees it. Two audiences got two documents, and only one of them carried the number.
I want to be exact about what that paper does and does not establish, because neither brand page links to it. It validates agreement between an algorithm and expert graders on the appearance of facial signs. Not disease. Not outcome. Not whether following the resulting product recommendation changes anything about your skin. And since neither page cites any study whatsoever, there is no public document connecting those published correlations to the 95 percent and 98 percent figures being advertised. The evidence and the advertisement are not joined to each other anywhere a customer can read.
That gap is the finding. When a brand says its analyzer was built with dermatologists and trained on tens of thousands of images, neither of those is a result. A training set size is an input. A dermatologist on the payroll is a credential. Neither one tells you how often the thing is right, on whom, measured against what. If a company had that number and the number was good, the number would be the advertisement.
What can a clinician do that a photograph cannot?
A dermatologist examining a lesion is doing several things at once, and only one of them is looking at it.
- Dermoscopy. A handheld lens and light source that cuts surface reflection and reveals subsurface pigment structures no phone camera can see.
- Palpation. Touch reports depth, firmness and whether a lesion is fixed to the tissue underneath, none of which survives into a photograph.
- History. How long, how fast, has it bled, has it changed, what runs in the family, what has already been treated.
- Full-body examination. Looking at the places you cannot photograph and would never think to check.
- Biopsy. Removing tissue and putting it under a microscope, which is the only step in the whole sequence that produces a diagnosis.17
The Cochrane review of dermoscopy pooled 104 study publications covering 42,788 lesions and found that accuracy was higher for diagnosis conducted in person than for image-based evaluation, with a relative diagnostic odds ratio of 4.6 (95% CI 2.4 to 9.0).18 Being in the room beats looking at a picture, and the size of that advantage has been measured rather than asserted.
I am a breast cancer survivor, and it leaves you with one narrow, permanent piece of knowledge about screening: a screening tool is only ever as good as what it was validated against. People who have been through real diagnostic medicine know the difference between a scan and an answer, and no interface has ever moved that line.
So what is the scan actually for?
Here is what I would say to someone standing at a counter with a tablet pointed at their face. The scan is not lying to you, exactly. It is measuring something real about a photograph, consistently enough to be interesting. What it is not doing is examining you. It has no access to depth, no access to history, no access to touch, and no capacity to be wrong in a way anyone has to answer for.
Separately, and having nothing to do with pores: a mole or patch of skin that is changing in size, shape or color, that has irregular borders or more than one color, that itches, oozes or bleeds, or that is simply new and behaving oddly, is a reason to see a clinician. The National Cancer Institute lists those signs and says to check with your doctor.17 It also notes that while fair complexion raises risk, "anyone can have melanoma, including people with dark skin."17 No app on the market is authorized to make that call for you, and all three of the devices that are authorized require a physician to be standing there when they are used.
A brand skin analyzer lives in the wellness space because it declines to make a medical claim, and it declines to make a medical claim because making one would require evidence it does not have and does not need. The evidence gap is the product's operating condition rather than an oversight in it. Stay on the cosmetic side of that line and you never have to prove anything, and the tool performs flawlessly, because its job was never to find out what is on your face. Its job was to find out what belongs in your cart.



