Cover: White Whale Score 719 of 1000, rated “Gut”, evidence quotient 81 per cent, no critical failures, €437 lost monthly revenue, rank 1 of 7.
White Whale Score

One website, recalculated across 1000 points.

Ten categories, 106 criteria, and beside every point the artefact it came from. Whatever could not be measured leaves the calculation and is printed as its own number. We started with our own site and we show the whole result.

Why this exists

Everyone sells web design. How do you spot the good kind?

You collect three quotes. All three show handsome screenshots and use the same words: modern, fast, search optimised. Two years later the site is live and the phone does not ring. What were you supposed to have noticed beforehand?

The usual tools do not help. They honestly check what they can check and leave the real question alone. A site can pass every technical audit and still be invisible for every term it was built for. That was exactly our situation, and the distance was 281 points.

So we built a rating that asks the question directly: does this website do its job? Ten categories with weights published in advance, an artefact behind every point, seven disqualifying states that cannot be averaged away. No favours, no impressions, no marks for effort.

The gap

The same page, the same day measured twice.

Lighthouse, 29 August 2026
100 / 100 / 100

Accessibility, SEO, best practices. 52 of 52 audits passed.

281points apart
Search Console, the same 28 days
2 clicks

12 impressions in total. For the four terms the site was built for: none.

Lighthouse does not lie. It carefully checks 52 things it knows how to check, and those were fine. None of them answers whether the site does its job. Between “correctly marked up” and “serves its purpose” lay 281 points that day. That distance is the reason for everything else on this page.

Structure

Three layers instead of one number.

A single grade cannot be argued with, which means it cannot be checked either. The report returns three results of different hardness and says which is which.

01Evidence
719 / 1000

What was measured

The score. Every point hangs on an artefact that can be fetched again: a response header, an API result, a computed contrast ratio. Sentences like “looks tidy” earn nothing.

02Price
€437 / month

What it costs

Lost revenue from the visibility gap alone. Nine lines of arithmetic are printed separately, three of them assumptions and marked as such. Set the close rate to 15 per cent and the figure halves. You are meant to be able to redo the sum.

03Rank
1 of 7

Where you stand

Position in the field across eleven signals that can be collected identically from someone else's site without access to their data. This is explicitly not the whole Score, only the part of it that compares fairly from outside.

Evidence quotient

A report must not look better for having checked less.

Every rating has the same weak spot: what happens to a criterion that could not be measured? Deduct points and the grade punishes gaps in the auditor's toolkit. Deduct nothing and the result inflates.

The White Whale Score removes the unmeasured from the denominator and prints the measured share as its own number beside the score. 719 out of 1000 at an evidence quotient of 81 per cent means 19 per cent of the possible points went unchecked, and the cover page says which. The section on what was not checked stops being a footnote and becomes arithmetic.

For us two missing sources alone cost 33 points of evidence quotient: Google field data and a source for the link profile. That is in the report because it is our problem, not the measured site's.

points earned
measurable points
evidence quotient 81 %
measured
A tool produced the value.
derived
Inferred from something measured, the question itself not asked directly.
human
A named person decided, with the reasoning printed alongside.
not measured
Out of the denominator. Neither point nor deduction.
not applicable
Does not exist here. No forms means no criteria about forms.
Ten categories

The weights are fixed beforehand .

Within each category the criteria total exactly 100, and the ten weights total exactly 1000. A script checks both on every run. The profile is chosen before the measurement and printed in the report, so the distribution cannot be fitted to the result afterwards.

Profile

Trades, law firm, practice, agency. Being found decides everything.

CategoryWeightOur valueEvidence
01Indexing and findability
140
6795 %
02Speed
120
6372 %
03Accessibility
90
9766 %
04Visual quality
90
7184 %
05Content and message
120
8092 %
06Interaction
90
9278 %
07Security and privacy
100
6662 %
08Code and architecture
70
7392 %
09Delivery and operations
80
6290 %
10Trust and legitimacy
100
5273 %
Σ1000719 / 100081 %

“Our value” is our own site on 29 August 2026, profile “local service provider”. 52 out of 100 for trust and legitimacy is the weakest row in the table. It is here for the same reason it is in the report.

Critical failures

Some things cannot be averaged away.

A beautiful site nobody can find is not a good site with one problem. It is a site that does not do its job. Seven states therefore put a ceiling on the result, applied after the calculation: 780 points but not indexable comes out at 400. If several apply, the lowest wins.

  • Site not indexable400
  • Primary language version not indexed550
  • Site does not open on a phone400
  • No HTTPS or broken certificate450
  • Mandatory legal information missing600
  • Secrets exposed in client or repository300
  • Forms submit unencrypted350

All seven checked, none triggered — the reason our cover page says “none”.

The scale

  • 900–1000ReferenceNo faults. Only growth left.
  • 800–899StrongWorks the way it should.
  • 700–799GoodNoticeable gaps, none of them critical.
  • 600–699ServiceableFunctions, but loses customers in specific places.
  • 450–599WeakSystemic problems. A work programme is needed.
  • 300–449FailingDoes not serve its purpose.
  • 0–299BrokenRebuilding costs less than repairing.
Test bench

The report checks itself.

A tool never tested against known truths is Lighthouse, only home made. While building ours, the contrast script reported 89 violations; all 89 were the script's own fault, because Tailwind serves colours in oklab and the maths had been written for RGB. After the repair: one real violation across 104 elements. Since then no report ships without passing three gates.

01

Data

Totals add up, every evidence class is known, no value exceeds its maximum. Above all: no point without an artefact. Entries like “good” or “fine” are rejected as judgement rather than accepted as evidence.

02

Arithmetic

The score is recomputed by a second, independent implementation of the same formula that borrows no line from the report generator. If the two disagree, nothing is printed. Ceilings apply here, and only after the calculation.

03

Appearance

A fingerprint over the stylesheet and page order, plus a real measurement in a browser: no page may overrun its type area. If the measurement returns nothing, the report counts as unchecked and does not go out.

A self test then breaks the data ten different ways and demands that every single break is caught. A gate that never fires is decoration, not a check.

How the report is made

Measured by tools, answered for by a person.

A person sets the brief, obtains the access and chooses the profile. Tools and AI collect the values and do the arithmetic. A named person checks the result, signs it and carries responsibility for it.

This is written here because a report that looks like handwork where scripts did the measuring gives a false impression. We would rather you knew exactly what you are holding.

Our own report

We start with ourselves .

All 21 pages of the publication version from 29 August 2026, uncut. 719 out of 1000, evidence quotient 81 per cent, €437 of lost monthly revenue and a 52 for trust. Anyone selling a rating without showing their own is selling an opinion.

Page 17 shows real measurements of real competitors. In this version the names are replaced by “Anbieter A” to “Anbieter F”: without recognisability, § 6 UWG on comparative advertising does not apply, and the figures are unchanged. The named comparison exists only in the client's report.

We assembled the field ourselves: six agencies in and around Ulm, measured across eleven publicly retrievable signals, one sided, on a single date. It is not a market share and says nothing about the work these providers do for their clients. Anyone named can request the raw data for their row.

What you may do with it

The report, its texts, tables and design are ours. Reprinting, republishing and use in your own offering require our written permission; the right to quote is untouched. “White Whale Score” is our mark and appears only above reports we produced. The method itself we published on purpose: measuring your own site by it is expressly allowed.

Your site

Tell us what to measure.

Three answers are enough for a quote. The button opens WhatsApp with a message already written, which you can read before sending — nothing reaches us until you send it yourself.

What is the site?

Decides the weights of the ten categories.

Which access can you give us?

Every missing access costs evidence quotient, not points. It works without any of it; the report then says how much it could not see.

The form works only in your browser. Your entries are assembled into a message that you send yourself.

FAQ

+What does the White Whale Score cost?

As a standalone audit you ask for a quote; the price depends on the size of the site. In our plans it is included: in the Basic plan, and through the Basic scope in Premium and Scale as well.

+Why 1000 points and not 100?

Because ten categories with their own weights otherwise collapse into decimal places. Across 1000 points it stays visible that findability weighs 140 for a local service provider while code weighs 70.

+Is this not just Lighthouse under another name?

No, and one number shows it: on 29 August 2026 Lighthouse gave our homepage 100, 100 and 100. The White Whale Score gave the same page 719 out of 1000 on the same day. Lighthouse checks whether a site is built correctly. The Score checks whether it does its job.

+How long does an audit take?

The collection itself is a working day. The measurement window is printed in the report, because values from different windows cannot be compared.

+Are competitors named?

In the report to you, yes, with measured values and without a single evaluative word. In any version that becomes public they appear as “Anbieter A” to “Anbieter F”. Anyone named can request the raw data for their row and have an error corrected within 14 days.

+Can I measure my own site with the method?

Yes. The method is published so that it can be checked, and you may apply it to your own site. What you may not take is our report as a template and our name above your result.

+What happens after the report?

The report contains a measures section where every line carries its gain in points: the script provisionally sets a criterion to its maximum and recalculates the whole score. So you see what a repair is worth before you commission it, from us or from anyone else.