An audit, end to end.
An audit goes from a URL plus a short business brief to a ranked backlog of A/B-test hypotheses — each grounded in a named behavioural principle and scored for priority, with focused single-test audits adding an adversarial second model pass that argues against the recommendation. The output is a JSON-shaped contract: 23 fields per hypothesis. The whole pipeline finishes in a few minutes.
A few more recent additions: Hypothesisly drafts your brand voice from your live site copy, so the brief is half-written before you start (you refine it before re-auditing). It puts per-vertical benchmarks next to your own numbers, flags tactics to avoid in regulated niches, and reports opportunity sizing as a range, not a single number. Every brief is versioned, and each audit pins to the exact brief that was live when it ran.
What happens after you enter a URL.
Three streams of evidence converge into one prompt, then narrow to a ranked backlog. The model never sees a bare URL; it sees your page, how it performs, and what your analytics show.
The model sees what a careful auditor would see.
A multimodal prompt assembled in build-payload.ts — three image blocks plus a structured text block, with real evidence attached.
Three image blocks — desktop, mobile, and a throttled-3G mobile render — plus a structured text block carrying the business context, DOM signals, Core Web Vitals, the live GA4 snapshot, the prior audits we’ve already run on this domain, and the outcomes you’ve logged.
It’s framed as evidence of substance, not a list of claims — the model reasons over what’s actually there.
- · Business context
- · DOM signals
- · Core Web Vitals
- · Live GA4 snapshot
- · Prior audits on this domain
- · Logged test outcomes
Five stages, in order.
The same discipline runs on every audit. Each stage can stop the process or lower the confidence tier — by design.
Does the site have enough traffic for an 8-week test to detect a real lift? Below threshold, we return BORDERLINE and refuse to ground hypotheses in numbers we don’t trust.
Structural signals: page type, heading hierarchy, form fields, schema markup, trust badges, Core Web Vitals.
A DOM observation carries full weight only when GA4 or the flow simulation confirms it. On one signal it still appears, ranked lower.
Each surviving hypothesis is grounded in a named principle — from Cialdini, Fogg M/A/T, prospect theory, or cognitive load theory. Mechanism over hunch.
10 questions, up to 14 priority points: six standard questions plus four double-weighted evidence questions. The backlog ranks objectively.
We argue against your next test before you run it.
On a focused, single-test audit, a second pass tries to break the recommendation before it reaches you. Here’s a generic example: the hypothesis, then the self-critique that travels with it.
If we make the primary add-to-cart button a full-width, thumb-reachable tap target on mobile product pages, mobile add-to-cart rate will rise — because the control button is below a comfortable touch-target size and sits outside the thumb zone.
The tier tells you what we actually had.
Every hypothesis is labelled with how much evidence backed it. Less data in means a lower tier out — never a confident answer dressed over a guess.
Each audit starts where the last one left off.
The payload for every run already carries the prior audits on your domain and the outcomes you’ve logged. What won becomes a prior; what lost gets down-weighted. The more you test, the more each recommendation is shaped by your store’s own evidence rather than a generic playbook. That’s per-customer today; cross-customer pattern learning, abstracted to pillar and mechanism (never raw data), rolls out through 2026.
The honest boundaries.
Hypothesisly doesn’t run the test, doesn’t change your site, and doesn’t predict the future. Pair it with whichever testing tool you already use. We make the recommendation defensible; you run the experiment. Two more lines we hold by design: the funnel walk stops at the cart and never enters checkout or payment, and if a site blocks automated access we stop rather than evade it.
Want to see it on your site?
We’ll run a real audit and walk you through what it found — and what it couldn’t be sure about.
15 minutes · a real audit on your site · no slide deck.