The Hypothesisly guide.
Everything you need to run your first audit and read the results — written for someone who has never used the product, and never run an A/B test. New to the vocabulary? Jump to the glossary first; every term is defined in plain English.
What Hypothesisly is
Hypothesisly is a CRO agent for ecommerce. CRO stands for conversion-rate optimization — the practice of getting more of your existing visitors to buy, rather than paying for more traffic. You give it a store URL and a little context; it inspects the page, reads your analytics, and hands back a prioritised list of A/B-test ideas worth running, each backed by evidence from your own site and argued against before it reaches you.
The honest analogy: think of a senior CRO consultant who shows up, audits your store, reads your analytics, and writes a ranked test backlog with the reasoning and research behind each idea — then red-teams their own recommendations before handing them over. Hypothesisly does that work on demand, in a few minutes. You still decide what to test, and you still run it.
Most conversion programmes don’t stall on running tests — they stall on deciding what to test next. Coming up with specific, evidence-grounded ideas is slow and usually needs a seasoned CRO in the room. This does the analysis so you start from a backlog, not a blank page.
Before you start
You’ll get the most out of an audit if you have:
- A live ecommerce URL — usually a product page, a collection/category page, the cart, or the homepage.
- Around 1,000+ monthly sessions on the page you want to improve. Below that, an A/B test usually can’t gather enough data to reach a trustworthy result — and Hypothesisly will tell you so rather than pretend otherwise (see the
BORDERLINEverdict under Reading your report). - Google Analytics 4 (GA4), ideally — connecting it lets the agent see real behaviour, which raises how much weight its recommendations can carry. It’s not required; you can run on the page’s structure alone.
It’s built for ecommerce operators, CRO and growth teams, and agencies running optimisation across several clients. If your site is below the traffic floor or you have no analytics, it still works — the report just says, clearly, how confident it can be.
Getting access
Hypothesisly is in invite-only early access while we’re working with a small group of agencies and operators. To get in:
- Book a demo or request early access. On a demo we run a real audit on your URL and walk you through what it found — no slide deck.
- Once your account is set up, sign in at
app.hypothesisly.com. - Inside, your work is organised as Account → Site → Campaign. Add the site you’re optimising, then create a campaign to group the audits you run against it.
Creating an audit
An audit starts with a short three-step intake wizard. It turns a raw URL into a brief the agent can actually reason about, so spend a minute here — the more context you give, the more specific the output.
- Business context. Your niche, average order value, your primary goal, and your traffic mix and device split. This tells the agent what “good” looks like for a store like yours.
- The problem. Your key metrics and the specific thing you’re trying to fix — for example, “mobile add-to-cart is low” or “people abandon at the cart.” The agent aims its hypotheses at your stated problem.
- The page & data. The exact URL to audit, plus how you want to connect analytics (the next section).
If your URL looks like a homepage but your problem mentions checkout or cart, the wizard asks which page you really mean — so the audit lands on the page that matches the problem.
Connecting your data
How you connect analytics sets the audit’s confidence tier — a label on every recommendation saying how much evidence stands behind it. More real behavioural data in means a higher tier out. You have five options:
- GA4 — OAuth. One-click connect to Google Analytics 4. The simplest way to share live behaviour.
- GA4 — Service account. Add a viewer email to your GA4 property instead of using OAuth — handy for agencies managing client properties.
- CSV upload. Export your GA4 figures and upload them. Static rather than live, but still real numbers.
- Manual entry. Type in a few key metrics (conversion rate, sessions, device split) if you can’t connect GA4.
- DOM-only. Skip analytics entirely; the audit runs on the page’s structure alone.
Connecting Microsoft Clarity on top of live GA4 adds heatmap and session-recording signals and reaches the highest tier. Here’s what each path unlocks:
| Tier | How you connect | What the report can claim |
|---|---|---|
| Brief | Live GA4 + Microsoft Clarity + a qualified brand brief | The most complete tier — behavioural + qualitative signal. |
| Full | Live GA4 (OAuth or service account) | Behavioural validation and lift estimates we stand behind. |
| High | CSV upload of GA4 data | Static behavioural context — directional, not live. |
| Estimated | Manual entry of key metrics | Directional only; no confident claims. |
| Structural | DOM + screenshots only (no analytics) | Surface issues, flagged honestly as un-validated by behaviour. |
Running the audit
Submit the wizard and the agent does the rest — you don’t have to do anything while it works. Behind the scenes it:
- Inspects your page. A headless browser loads it and reads around two dozen on-page signals (page type, forms, calls-to-action, reviews, trust badges, structured data, heading structure), measures Core Web Vitals on both desktop and a throttled-3G mobile connection, and takes screenshots in three viewports.
- Walks the funnel. On product pages it navigates product → add-to-cart → cart → the start of checkout, noting timing, errors, and pop-ups along the way.
- Reads your behaviour data. If you connected GA4, it pulls the funnel that matters: revenue by channel and device, where people drop off, conversion by device, new vs. returning, and your top pages.
- Reasons, then attacks its own work, and ranks what survives (covered next).
A typical audit finishes in a few minutes. Larger backlog runs are processed in the background, so you can close the tab and come back to the result.
Want the full pipeline, stage by stage? See How it works.
Reading your report
You get a prioritised backlog of hypotheses. A hypothesis is a single, testable claim: “if we change X, metric Y will move, because Z.” Each one is a structured record of 23 fields — not a paragraph of prose — so you can act on it without guesswork. The fields group into:
- The idea — the observation, the exact change to make, and the behavioural mechanism that explains why it should work.
- The evidence — the data source behind it, a research citation, and a statistical note.
- The maths — expected lift, an opportunity-size range (a point estimate plus a low/high and the assumptions), and the effort to build it.
- The plan — a recommended run-time in weeks, a tracking spec (what to instrument), and an interaction spec.
PXL score (0–14)
Every hypothesis carries a PXL priority score from 0 to 14, and the backlog is ordered by it — highest impact-versus-effort first. The score comes from ten questions: six standard ones, plus four evidence questions that count double. You can see the breakdown, so the ranking is something you can inspect, not a black box.
The five CRO pillars
Each hypothesis is tagged to one of five pillars, and a healthy backlog spreads across them rather than clustering: Trust & Credibility, Offer Clarity, Friction, Urgency & Motivation, and Page Flow & Hierarchy.
Confidence tier & qualification
Each hypothesis shows the confidence tier from the table above, so you know how much weight it can bear. The audit as a whole also carries a qualification verdict — READY or BORDERLINE, with the reason — so you know up front whether your traffic can support a trustworthy test.
How findings are ranked
The two-sources rule
A finding from the page’s structure alone is treated as a lead, not a conclusion. It carries full weight only once a second, independent signal — a behavioural pattern in your GA4 — agrees with it. Single-signal findings still appear, but they’re labelled and ranked lower. When the evidence is thin, the report doesn’t hide it: it shows the idea and tells you the single cheapest action that would confirm it (a promotion path).
Adversarial falsification
On focused, single-test audits at the higher-confidence tiers, the agent argues against its own top recommendation before you see it — so a weak idea doesn’t arrive dressed up as a strong one. For each, it names:
- the strongest failure mode (how it could flop);
- the likeliest confounder (what else could move the metric);
- an alternative explanation for the same evidence;
- and the discriminating test that would tell them apart.
This pass runs on focused audits at the higher tiers — not on every hypothesis in a large backlog.
Acting on it & re-auditing
- Export the brief. Each hypothesis can be exported as a Markdown brief — the idea, the evidence, and the tracking spec in one document you can hand to whoever builds the test.
- Run the test in your own tool. Hypothesisly produces the plan; you run the experiment in whatever A/B platform you already use. It never touches your live site.
- Log the outcome. Record each test as won, lost, or inconclusive, with the actual lift. This is the important part: those results feed back into the inputs for your next audit on that domain.
- Re-audit over time. Run the same domain again and the agent treats the history as a series — it confirms what looks resolved, revisits what didn’t land, and avoids re-proposing the same thing.
Two more inputs shape later audits: vertical benchmarks (per-industry conversion, cart-abandonment, and returns context) and portfolio learnings (your own logged outcomes, abstracted to pillar and mechanism, used as priors on new audits).
What it won’t do
A few limits are deliberate — they’re features, not gaps:
- It doesn’t run your tests. It produces the hypotheses and the test plan; you stay in control of what ships to your live site.
- The funnel walk stops at the cart. It goes as far as the begin-checkout step — never into checkout, and never submitting payment on a site.
- It plays fair with other sites. If a site blocks automated access, the agent detects it and stops, rather than trying to evade the protection.
- It doesn’t claim to be “better.” It shows its evidence, its reasoning, and its confidence, and lets the results speak.
Data & privacy
Your data — and your clients’ data — is scoped per account; no account reads another’s. Analytics tokens are encrypted at rest, the database is EU-region, and the audit data passes through Anthropic’s commercial API, which does not train models on what you submit. The full detail — hosting, encryption, what leaves our systems, and your export and deletion rights — is on the Security & data page.
FAQ
- Do I need Google Analytics?
- No — you can run a DOM-only audit on the page’s structure. But connecting GA4 lets the agent see real behaviour, which raises the confidence tier and unlocks lift estimates.
- How much traffic do I need?
- Roughly 1,000+ monthly sessions on the page you’re testing, so an A/B test can reach a trustworthy result. Below that, the audit returns
BORDERLINEand won’t fake confidence it doesn’t have. - Does it run the A/B test for me?
- No. It writes the hypotheses and the test plan; you run the experiment in your own testing tool and stay in control of what ships.
- Which testing tools does it work with?
- Any of them. Because it hands you a plan rather than executing the test, the output works alongside whatever platform you already use.
- How long does an audit take?
- A typical audit completes in a few minutes.
- Is my clients’ data isolated?
- Yes — data is scoped per account and encrypted at rest. See Security & data for the specifics.
- What does it cost?
- Pricing is being finalised with the early-access cohort, and it’s free during the private beta. See Pricing.
Glossary
- CRO (conversion-rate optimization)
- Improving the share of visitors who take the action you want (usually a purchase), instead of buying more traffic.
- Conversion rate
- The percentage of visitors who complete the goal — for ecommerce, typically orders divided by sessions.
- A/B test
- An experiment that shows some visitors version A and others version B, then measures which performs better.
- Hypothesis
- A single testable claim of the form “if we change X, then Y will happen, because Z.”
- PXL score
- A 0–14 priority score ranking each hypothesis by impact versus effort, so the backlog is ordered objectively.
- CRO pillar
- The category a hypothesis belongs to — Trust & Credibility, Offer Clarity, Friction, Urgency & Motivation, or Page Flow & Hierarchy.
- Confidence tier
- A label (Brief, Full, High, Estimated, Structural) showing how much evidence backs a recommendation, set by the data you connected.
- Two-sources rule
- A structural finding is promoted only when an independent behavioural signal confirms it; single-signal findings appear lower and labelled.
- Falsification
- A second pass that argues against a recommendation — failure mode, confounder, alternative explanation, and the test that settles it — on focused audits at the higher tiers.
- DOM
- The structure of a web page (its elements and content) that a browser builds when it loads the page.
- Core Web Vitals
- Google’s page-experience measurements for loading speed, responsiveness, and visual stability.
- Funnel walk
- The agent stepping through product → add-to-cart → cart → begin-checkout the way a shopper would, to capture what happens at each step.
See it on your own site.
We’ll run a real audit on your URL and walk you through what it found — and what it couldn’t be sure about.