Method
How a study works, and what it cannot tell you
A study is an observation of simulated participants using your live site. This page describes what they see, what they do not, how we count, and where the numbers stop meaning anything. Every report carries the same information about its own study in an appendix.
Who takes part
Participants are simulated personas drawn from a pool of 10 million that we generated ourselves. The pool is calibrated on age, region, urbanicity and gender. It is not calibrated on the other 1,286 attributes each persona carries, including the traits that shape how someone browses — patience, attention span, decision speed. It is not a representative sample of any population, and no report claims one.
You describe your customers in plain English. We resolve that once into explicit filters on the pool's attributes, print them in the report, and freeze them, so a re-run months later draws the same way. If fewer than 400 people in the pool match, we refuse rather than draw the same people repeatedly. Everyone drawn can read your site's language.
They are simulated participants. They are not real people, not a panel, and not a proxy for your actual visitors. We say so in every report.
No instructions
Each participant is told only that they have access to a product and nobody has told them what to do with it. They decide whether any of it is for them. We call this a roam session.
We started with the other way — give each participant a goal and watch them pursue it. On the same product with the same people, the two framings found different problems. With a goal, participants reported the onboarding overlays. Without one, they found that the price was nowhere to be seen, and none of them chose the flow the goal-driven run had been measuring. A participant given a goal pushes at the door you assigned, and you learn only whether it opens.
A goal, when a study sets one, states what the person wants and never what they will find. Telling a simulated participant how things will go changes what they do: an instruction saying most people leave early produced a study in which everyone left.
What a participant sees, and what they never see
Each participant gets a fresh, real browser on your live site — 1280×800, or a 390×844 phone with touch in a device study. At every step they receive:
- the page's controls as the accessibility tree names them, each checked to be clickable at its centre;
- where things sit on the screen and what is below the fold, and only the text currently on their screen;
- a screenshot of the screen, with no instruction to judge how it looks;
- a measured audit of contrast, text size and tap targets.
They can click, type, choose from a list, scroll, or stop. Between steps they carry forward what they now think the product is, whether it is for them, what has irritated them, what they would use, and a one-word mood.
Each layer answers only what it is good at. Facts on the page — prices, whether a toggle is on, what a button is called — come from the page itself, never from the screenshot, because a model once read six toggles as off when all six were on. Colour and contrast come from measurement, not impression.
What they never see
- Load time. They receive a settled page. Slowness, layout shift and animation are invisible to them.
- Their own history with you. Every participant is a first-time visitor. Only a Second visit study brings the same people back, and it is in beta.
- Controls we cannot drive — a canvas, a drag handle, some custom widgets. When that happens it is recorded as our blindness, never as a finding about your product.
A control group, always
Every study includes a group of people with no plausible use for your product, on the same site, with the same seed and the same browser. It exists because simulated participants have a strong bias to act. In an early test, a third of people who had no reason to want the product still clicked its main button. A click rate measured against zero means nothing; the finding is the difference from the control.
The control also sorts findings. A complaint the control raises as often as your audience is a defect for anyone. One only your audience raises is about them.
Two runs, on different people
Every study except the Quick read runs twice, each run drawn separately. A finding that appears in both runs within ten points holds. One that appears in both but moves more is shaky. One missing from a run is marked as such. We learned this by running the same study twice and getting 13% and 53% on the same measure.
How we count
- People, not mentions. Someone who says the same thing six times counts once.
- Rates only above 60 per run. Themes repeat at 30 people; percentages do not settle until about 60. Below that, a report prints counts and themes, and no percentages. This is why a Quick read has none.
- Judgements as ranges below 120. Measures like "decided it was not for them" moved up to 27 points between identical runs at 30 people, and 2 points at 120.
- Themes are grouped by arithmetic, not by a model, so the same sessions give the same groups every time.
- Findings are ranked by cost. A finding raised by people who then left ranks above one that was only mentioned. Severity — blocker, major, minor, cosmetic — is decided from the sessions by a fixed rule printed in the report.
Simulated participants click far more than real visitors — roughly ten times more, in our measurements — and act rather than read. So a rate in a report is for comparing: phone against desktop, version A against version B, before against after, on the same people. It is not a forecast of your conversion rate. To use a made-up example: "B beat A, 62 to 38 of the same 100 people, and it held on a second run" is the kind of claim we make. "B will lift your conversion by 24%" is one we never will.
Your three guesses
Before a study runs, you write down the three places you believe visitors struggle. The report scores them. Without them, a report is agreeable and cannot be wrong. With them, it is a test you set yourself.
You also name a milestone, the text a visitor sees when they have finished, or say there is none. We will not invent one.
What a study cannot tell you
- How many of your real visitors will do anything. Rates are simulated.
- How much a change will lift conversion. The direction of a difference carries over; its size does not.
- How fast your site is, or whether it shifts while loading.
- What returning or logged-in customers do, unless you give us a test account or buy a Second visit.
- What screen-reader users experience. Our participants do not use one.
- What a population thinks. We tested simulated survey answers against a real survey and closed that product.
- Anything about the traits the pool is not calibrated on, taken as a fact about real people.
We publish these because we measured them. Our own instrument has at times produced most of what it recorded; the report's instrument section says how clean each study was — errors, forced clicks, sessions cut off by the step budget — and a study graded suspect is not delivered.
When we refuse
Before anyone arrives on your site, a study runs pre-flight checks: the audience resolves to real people in the pool, everyone can read the site's language, a test account is clean, the domain is verified. A check that cannot decide refuses. A refused study is refunded in full.
Each refusal costs a study. Each approximation would cost you a wrong headline.