Reference: Measured accuracy
On 15 August I ran a structured review over this product: twelve lenses, from conversion to accessibility to the honesty of the copy, across the repository and the live site. The first pass produced 121 findings. A subtraction pass kept 106 of them, 20,092 words, preserved verbatim. Then I ran a second pass, hostile on purpose. It is the only part of the method I would defend in public: a finding survived only if the cutter opened the file it cites. Thirty-five survived. The rule caught reviewers inventing evidence. That is worth sitting with, because the reviewers were models and the inventions were plausible, specific and wrong.
One honesty constraint bounds everything the review claims. Nobody in it could reach production analytics. So every conversion finding is structural: an argument about what a page makes easy or impossible, never a measured funnel. The review says so in its own §5, rather than letting the reader assume otherwise.
Three defects were verified by hand before anything was cut
The most expensive finding was invisible locally by construction.
pool.apExtract() threw on every call in production. The worker
posted a field, the pool rebuilt its result from a hand-written list that
never gained it, and only the deployed box could fail. npm run
dev runs inspections inline rather than through the worker pool.
The second: the pricing page's credit-pack cards linked to a bare
/account. So the one path where a Brazilian buyer says yes
dropped the pack on the floor. ?pack= now mirrors
?plan=.
The third is my favourite for what it says about fixture discipline. The Portuguese pay-stub checker demonstrated with a utility bill, while this repo generates a fictional holerite for exactly that purpose. The English twin has the same mismatch and I deliberately did not fix it. No US payslip fixture exists, and faking one to close a review row would be the review failing at its own game.
The pattern in the survivors: published numbers had drifted
Read as a list, the 35 rows look miscellaneous. Read for a pattern, the expensive ones are all one thing: a number or a claim published on a page, while the source of truth moved underneath it.
- The evasion post published 12 signal families and 5 suppressed; the engine had grown to 18 and 8.
- The flagship guide's social card, a PNG, said ten families while the page under it said 18. A text lint can never reach a number baked into an image, which is why that one got its own test.
-
The pricing page sold a per-key quota while the code meters per account.
The API page opened with "1. Get a key", while an anonymous
POST /inspectalready returns a full report. -
The report header counted
signals.lengthwhere the API'ssignalCountexcludesinfo, so the site's own "Clean statement" demo read "1 signal" while the API said 0.
Each fix followed the house rule rather than re-typing the number: derive it.
The pricing figures come out of plans.ts, and the tests derive
the served pages from that same object. That is the only reason the repricing
three days later moved 54 hard-coded mentions without missing one. A
published number that no test derives is a claim with a shelf life nobody set.
The accessibility rows were re-verified, then measured
Rows 24 to 31 were closed a week later, and the method matters more than the list. Every row was re-verified against the tree before I touched it. Every fix was measured in Chromium at 390x844 and 1440x900, not reasoned about.
Two were level-A failures under WCAG, the baseline tier of the web
accessibility standard, and both sat on the primary journey. Every drop zone
was a container with role="button" and an aria-label. That
combination deletes the zone's own text from the accessibility tree, the
model of the page a screen reader or a voice tool actually works from. So a
Voice Control user saying "click browse" could not operate the site's main
control. 22 zones across 20 files now put a real button on the visible word.
And the hero animation looped forever while a comment four lines above it
specified that it must not.
The smaller rows are the kind only a checklist catches. The risk verdict was the smallest heading in the report. A 40 MB file uploaded in full before the server's 413, its "too large" answer, arrived. The phone nav's caret had been overridden into a grey bar by its own tap-target rule.
Not one new feature. Reconciling the 35 rows against the plan index concluded that no nineteenth plan was needed: every finding was wiring, correcting or publishing something already built. For a product at this stage, that is the diagnosis worth having. It is also the cheapest possible outcome of a week of review: the engine was fine, and the claims about it were rotting.
The method, if you want to steal it
Two passes, not one. The first pass optimises for recall and is allowed to be credulous. Keep its output verbatim, because the cut list is evidence too. The second pass optimises for precision under one mechanical rule: open the file, or the finding dies.
That rule is what converts a review from prose into a work queue. It is also the only defence I know against a reviewer, human or model, that writes what a defect plausibly would be instead of what is there. This product's entire pitch is that claims should be checkable against bytes. The review only worked when I held it to the same standard.
The page the review kept pointing back at is the evidence page: every accuracy figure this product publishes, dated and versioned, including the unflattering ones. That page is what the 35 rows were protecting.