Every figure on this page is recomputed, not remembered. Luni's catalog database is committed alongside the code that scores it, so the verdicts before and after each change are both in git and the impact is a diff rather than a recollection — git show <commit>^:seed/skinlogic.db against git show <commit>:seed/skinlogic.db. The script is seed/measure_corrections.py and it prints the table below. Where a number here differs from the one in the original commit message, this one is the number — and one of them did: the pass of 6 September cleared 21 junk rows from our ingredient dictionary, not the sixteen its own commit message claims, because a second person cleared five more that were never added to the count.
Two kinds of figure appear below, and they are not equally checkable. Everything about products — scores moved, bands changed, which product went from what to what, how many carried an ingredient — is a diff of two committed databases, and the script reproduces it. A handful of figures describe the engine's internal behaviour instead (how many times a rule fired, how many products a flag applied to). Those cannot be recovered from a database diff, because the flags are computed at runtime and never stored; they were measured against the engine when the change was made, and they are marked measured at the time where they appear.
The ledger
| Date | What was wrong | Scores moved | Bands changed |
|---|---|---|---|
| 2026-07-06 | A declared dose is a dose we can read | 66 | 55 |
| 2026-07-15 | Cleansers were scored as if they stayed on the skin | 666 | 498 |
| 2026-08-05 | Invisible characters made ingredients unreadable | 1 | 1 |
| 2026-08-06 | Growth factors were not scored as growth factors | 71 | 5 |
| 2026-08-08 | An asterisk made organic labels unreadable | 3 | 6 |
| 2026-08-11 | Every peptide was graded as though equally evidenced | 545 | 57 |
| 2026-08-12 | A molecule class split by spelling | 74 | 7 |
| 2026-08-24 | A comma inside a number, and a dose in ppm | 163 | 56 |
| 2026-09-03 | A pH adjuster was read as a buried hero | 364 | 105 |
| 2026-09-05 | Naming a fragrance allergen was scored as a defect | 0 | 61 |
| 2026-09-06 | The scorer could not place salicylic acid | 1,721 | 275 |
| 2026-09-10 | A declared colourant was graded as a UV filter | 90 | 32 |
| 2026-09-10 | Moisturisers were judged on actives they are not built to have | 1,825 | 1,485 |
| 2026-09-10 | Eye creams had the same blind spot | 377 | 278 |
| 2026-09-10 | Masks did too, except the ones doing the opposite job | 403 | 301 |
“Scores moved” counts products we already held a verdict on; products added by the same change are excluded, so this measures what the correction did rather than how much the catalog grew.
We were charging brands for obeying the law
5 September 2026. EU cosmetics law requires a brand to name fragrance allergens above a threshold. Our scorer took a point of Tolerability for each disclosure, on top of the point it already took for fragrance itself. So a pack that printed “Parfum, Limonene, Linalool” scored worse than an identical pack that hid the same molecules inside “Parfum” alone. We were charging for compliance and rewarding concealment, in the product whose whole position is evidence over marketing. 1,429 products were paying that penalty and 2,636 that disclose nothing were not — measured at the time, because the flag is computed when a product is scored and never written down.
The arithmetic was small and the consequence was not. Tolerability is not an input to the Formula Score, so removing the penalty moved no score at all — the measured figure is zero. But a Tolerability under 4 caps a band at Fair, and that one point was pushing products under the line. Sixty-one products were being held at Fair for naming their allergens. The flag itself is unchanged and still shows: a reader with a known allergy needs to see it. Only the arithmetic went.
We called 17 products marketing-led over an acid the brand never mentions
3 September 2026. The owner asked whether Q+A's Caffeine Eye Serum really deserved 2.60 and the label “Mostly Marketing”. It did not. The only active our engine had detected was lactic acid, listed last of 23 ingredients, after sodium levulinate and sodium anisate — the acid that sets the pH of a clean-label preservative system. The scorer read it as a fairy-dusted AHA, and that single reading selected an accusation reserved for buried marketed heroes, for an ingredient the brand never names anywhere.
The guard is narrow on purpose: an AHA or PHA in the last two positions, below the estimated 1% line, in a product whose name claims no exfoliation and does not name the acid, is not an active. BHA is deliberately excluded — salicylic acid works at around 0.5% and sits legitimately low on a list, and a first draft of this guard stripped the real active off genuine acne treatments. A missed accusation costs less than a hidden active.
Seventeen products left the label, including Weleda's Sensitive Care Facial Cream, both Simple Water Boost moisturisers, four Babor products and Yadah's Pure Green emulsion. Ten scored higher afterwards, six scored lower, and one did not move at all — because “Mostly Marketing” is an accusation, not a rung on the scale. Removing a phantom active also removes what it was contributing. The Rosehip Bioregenerate Radiance Mask scored 3.3 before and 3.3 after; all that changed was that we stopped saying it was selling something it did not contain.
| Band change | Products |
|---|---|
| Weak → Fair | 58 |
| Mostly Marketing → Weak | 17 |
| Mostly Marketing → Fair | 12 |
| Good → Excellent | 10 |
| Excellent → Good | 4 |
Our scorer could not read a 2% BHA exfoliant
6 September 2026. Our engine's one free check is whether an active sits above or below the estimated 1% line, and it located that line at the first preservative-coded ingredient. Salicylic acid is a listed EU preservative and our only strong-tier BHA. So on a 2% BHA exfoliant the line was drawn on the product's own active, which then read as buried beneath it. It fired on 383 of 936 salicylic acid detections (measured at the time). The engine could not say a BHA was dosed.
The fix is written on the property rather than the molecule: an ingredient the scorer itself grades as an active can no longer mark the line, and the scan continues so a real preservative further down still sets it. The products that gained most are the ones the defect hit hardest — a 2% BHA exfoliant liquid went from 3.80 and Weak to 8.70 and Excellent, two more BHA exfoliants moved identically, and a fourth went from 4.10 to 8.80. Not one of their formulas changed. Only our ability to read them did.
The same pass corrected two more of our own errors: 1,241 products printed vitamin E only as the ester spelling and earned no credit for it, and 21 dictionary rows held a chemical identifier where a restriction should have been — an EC number filed as though it were a legal restriction. Both figures are recomputed here: 1,241 scorable products print tocopheryl and never tocopherol, and 21 rows lost an eu_restriction value in that commit.
The correction that lowered 545 products and raised none
11 August 2026. Every peptide in our dictionary sat at the same evidence tier, which said Palmitoyl Pentapeptide-4 and Progeline are equally well supported. They are not. Regrading them against the evidence moved 545 products and every single move was down. Thirty products fell from Good to Fair, sixteen from Excellent to Good, eleven from Fair to Weak.
Nothing entered the top tier: no topical peptide earns one on the evidence we could find. This is the correction we would least like to have made and the one that most needed making — an instrument that only ever revises upward is not measuring anything.
Rinse-off products were scored as if they stayed on the skin
15 July 2026. A cleanser was judged by the same model as a serum, which asks whether well-evidenced actives are present at a working dose. That question is close to meaningless for something you rinse off in forty seconds: a cleanser's job is to clean without stripping. Judged as leave-on products, good cleansers read as weak ones. Introducing a mildness model for rinse-off products moved 666 scores and 498 bands, the largest revision in this log.
Four times, we simply could not read the label
Four corrections were not judgements at all. They were the engine failing to parse what was printed, and each one silently produced a verdict anyway.
- A comma inside a number (24 August 2026). We split ingredient lists on every comma, which cut 1,2-Hexanediol — present in 3,425 products on the day, recomputed here from the stored lists — into “1” and “2-Hexanediol”, neither of which is a real ingredient. Each phantom token also lengthened the list, which moves the 1% line and every position that depends on it. Fixing that, plus reading a dose in ppm as a dose, moved 163 scores and 56 bands.
- An asterisk (8 August 2026). Organic certification is marked with an asterisk on a pack. Our parser could not see past it, so certified-organic ingredients read as unknown ones. Three scores and six bands, five of them out of Mostly Marketing.
- Invisible characters (5 August 2026). Zero-width characters pasted into an ingredient list made words unrecognisable to a matcher that is exact. One product moved — from Weak to Excellent.
- A spelling (12 August 2026). Korean brands write PDRN as Sodium DNA. Our dictionary knew the acronym and not the salt, so one product scored for the ingredient and an identical one did not. 74 scores moved, in both directions.
A pigment was scoring as sun protection
10 September 2026. Our actives table grades titanium dioxide on the strongest tier, because it is a UV filter. That is right in a sunscreen. It is wrong in the 554 products in our catalog that are not sunscreens and list it as CI 77891 / TITANIUM DIOXIDE — the colour-index notation a manufacturer uses to declare a pigment, the thing that makes a cream look white.
So the engine was crediting products for a colourant as though it were sun protection. Lancôme's Absolue Revitalizing Eye Serum scored 7.8 and Good on exactly one graded active: the pigment. Nothing else in it was something we grade. It now reads 2.7.
The fix reads the notation rather than the molecule: any active whose own ingredient token carries a CI 7xxxx number is a declared colourant and is never graded. It is deliberately narrow — it does not fire on a UV filter written plainly in a non-sun product, because a tinted day cream really can carry filters, and the broader version of this rule would have stripped the genuine actives off 990 products including two retinal serums. 90 scores moved, 32 bands.
Seven products became “Mostly Marketing” as a result, and we let them. That is our strongest label and none of them carried it before. The mechanism is uncomfortable and worth stating plainly: the pigment sat above the estimated 1% line, and its presence there suppressed the check that looks for a marketed hero buried further down. Remove the pigment and the engine sees what was always true. We could have added a guard to prevent a removal from ever creating an accusation. We decided that would be hiding the finding twice.
This one was found because a reader asked why a well-regarded barrier balm was missing from the catalog. Adding it produced a score of 3.0, and 1.75 of that came from its titanium dioxide, sitting last of thirty-three ingredients.
We were judging a moisturiser on actives it is not built to have
10 September 2026. Sixty-two per cent of a product's Formula Score is its Evidence axis, and that axis counts one thing: graded treatment actives — retinoids, acids, vitamin C. A barrier cream has none of those, by design. Its job is to hold water in, put lipids back and calm the skin. We were asking it a question about someone else's job and then marking it down for the answer.
The scale was not marginal. Of 3,334 leave-on hydrating and barrier products, 2,511 carried at least two of the families that actually define the category — humectants, barrier lipids, soothing agents — and 52% of those were rated Weak or “Mostly Marketing”. 421 carried all three and were still Weak or worse. CeraVe Moisturizing Cream, about as close to a consensus barrier cream as dermatology has, scored 2.5.
The fix reads the product for what it is: four families — humectants draw water, barrier lipids rebuild, occlusives hold it in, soothers make it comfortable — weighted by where they sit in the list, because a ceramide fourth is the formula and the same ceramide thirty-eighth is a claim. It is the same move we made for cleansers in July 2026, which had the identical problem: judged as leave-on products, good cleansers read as weak ones.
It is taken as the better of two readings, never a replacement. A moisturiser that also contains a real active keeps its active score. That is what makes the change safe rather than merely large, and the numbers show it held: 1,825 products moved and every single one went up. Nothing was downgraded. Across the three categories, Weak fell from 1,210 to 250 and “Mostly Marketing” from 178 to 5 — 173 products stopped being described as marketing-led. CeraVe Moisturizing Cream now reads 7.6.
It also exposed something we had been missing entirely. Our irritant check matched only the denatured spellings of alcohol, so a plain Alcohol in an ingredient list — which is ethanol — was invisible to it. It could not simply be added to the list, because “alcohol” also sits inside cetyl, stearyl and cetearyl alcohol, which are emollients and the opposite of drying. A night cream with alcohol third had just risen from 3.3 to 8.6 with its tolerability untouched. It now reads 8.6 Fair: the number still describes how the formula is built, and the band carries the ethanol. That capped 131 products.
We would rather say plainly that this was wrong for as long as the app has existed than present the new numbers as an improvement. Every one of those 421 products was being shown to someone as a poor formula, and it was our reading that was poor.
The same mistake, three more times
10 September 2026. Once we had fixed moisturisers, the obvious question was where else the same thing was true. The answer was: everywhere the category is not built around a treatment active.
Eye creams, 377 products. An eye cream is a leave-on whose job is to hold water in and keep the thinnest skin on the face comfortable. Judged on treatment actives alone, 37% of 618 read Weak or “Mostly Marketing” — including several whose names are their own function: a Moisturizing Eye Bomb at 3.5, a Soothing Eye Contour Cream at 3.1, Chanel's Sublimage La Crème Yeux at 2.5. Weak fell from 192 to 50 and “Mostly Marketing” from 34 to one.
Masks, 403 products — and this one we got wrong first. We had refused masks twice on the grounds that the category is not one thing: clay, sheet, exfoliating and overnight products do not share a job. Simulating it without a guard proved the refusal right. 23 clay-led masks reached Good or Excellent — one detox clay mask went from 3.1 to 8.0 — rewarded for the glycerin and shea they carry alongside the kaolin. That is this same correction pointing the other way: crediting a product for a job it is not doing.
So a mask is judged as a leave-on only when it is not built to absorb. The test is an adsorbent mineral in the top ten ingredients, because position separates structure from a trace — four masks in our catalog carry kaolin at positions 11, 12, 17 and 18, and they are cream masks, not clay ones. Two details decided by measurement rather than assumption: the bare word “clay” counts, because one mask lists Manicouagan (Sea Mineral) Clay second and no mineral name catches it; and silica does not, because it is a common gel-mask thickener and treating it as an absorbent wrongly excluded thirteen masks.
An absorbing mask keeps its old score, and we are not going to pretend that is a good answer. We have no way to measure how well something draws oil out of skin, so a clay mask is still being judged on a question it was never trying to answer. What changed is that we now know which products those are. Inventing a number for it would be worse than saying nothing.
Across the four categories — cleansers in July, moisturisers, eye creams and masks — 2,605 products moved and not one moved down. Each axis is taken as the better of two readings, never a replacement, so a product carrying a genuine active keeps the score that active earned.
What this page cannot show
It only shows corrections we found. Every entry here exists because someone noticed. The reason four of the fifteen are parsing failures is not that parsing is where our errors concentrate — it is that a mis-parsed ingredient is visible once you look at the row, whereas a wrong evidence grade looks exactly like a right one. The corrections we have not made yet are not on this page, and there is no honest way to count them.
A count of moved products is not a count of wronged users. These figures are catalog-wide. We do not know, and deliberately cannot know, how many people had any of these products on their shelf when the verdict was wrong — the shelf lives on the phone and is never sent to us. Nobody was told their product had been re-judged, because we have no way to reach the people it applied to.
The dates are when we shipped the fix, not when the error began. Several of these defects were present from the first version of the scorer. The cleanser model was wrong for every day the app existed before 15 July 2026.
Bands moved by more than one correction are counted more than once. The 3,222 total is band changes, not distinct products. We have not deduplicated across corrections because the catalog itself changed shape between them — products were added and, once, 99 were removed — so a product id is not a stable subject across the whole period.
Every band here is our engine's opinion about a formula, produced with no user profile attached, and it is a cosmetic-suitability signal rather than a safety judgement or a prediction that anyone will react to anything.
How to check this
Clone the repository and run python3 seed/measure_corrections.py. It extracts both catalog databases for each commit listed above straight out of git, joins them on product id, and prints the table at the top of this page. Nothing on this page is typed in by hand from a commit message; where the two disagree, the recomputed figure is the one published. Named products are quoted from the same two snapshots, and each carries its own source URL and fetch date in the catalog. Counts are as of 8 September 2026.