See what breaks
before you ship.
CanaryFlock sends a population of simulated people through your real product and tells you who got through, who gave up, and exactly where. They differ in age, digital fluency, patience, and how much they trust a stranger with their card details — because the person who abandons is rarely the person who built it. Below: the same product, three people, three endings. Pick one and watch what it did to them.
Real sessions from one run. Nobody told the flock the password rule or what ISO-8601 means — neither was stated anywhere on the page, only enforced.
A single agent tells you whether a flow can be completed. A population tells you who completes it.
Three runs free, thirty people each. Work email, no card, no call.
The mistake nobody inside the company can see.
Not the bug that throws an error. The one where every control works, every page loads, and people leave anyway — because the thing that is obvious to you is invisible to somebody who arrived thirty seconds ago. Analytics tells you they left. It does not tell you why, or which of them.
Seven of eight people gave up on this page.
Eight simulated people, one job: work out what this costs and whether you can try it without talking to a salesperson. The main button on this site opened a mail client, which in a browser looks like nothing happening at all. There was no price anywhere.
“I’ve clicked SEND IN THE FLOCK three times now and nothing is happening — no new page, no form, no information. I came here to find out the cost and whether I can try it without talking to sales, but after six attempts I’m still on the homepage with no progress.” canary 0003 · 45–54 · medium digital skill · phone
That run cost $0.24 and took four minutes. Both problems are fixed above: the button goes to the product, and the price is on this page. We are telling you because a demo on somebody else’s product proves nothing about ours.
Five shapes, over and over.
It responds to a click and does nothing a person can see. Everyone who built it knows what it does, so nobody ever watched a stranger press it.
Price, delivery time, whether it works in their country. They look in three places, find none of them, and leave without ever contacting you.
The control is there and reachable. It is called something the company says and the customer does not — so they walk past it, twice.
The password needs a symbol. The date wants one format. The form knows; the person does not, until the third rejection.
They press it. Nothing confirms, nothing moves. So they press it again, and then they leave, and you never learn they were there.
None of these are hard to fix. All of them are hard to see, and every one of them is cheap the week before launch and expensive the week after.
A dress rehearsal with an audience you can afford to lose.
You give it a URL and a job to be done — open an account and subscribe, or find a charger near you and book it. It generates a population, runs each person through your product in their own browser, and records every click, every hesitation and every reason someone quit.
Not random attributes. Age drives digital fluency; income drives price sensitivity; fluency drives patience.
One job to be done, with a plain-language definition of what counts as finished.
Each person drives a real browser, in character, one decision at a time.
Every action stored as an event: page, click, reasoning, and how they felt about it.
Places where many people independently struggled — and who was hit hardest.
Most of them never got to the end.
Each mark is a simulated person moving through the flow. They stop where they stopped. This is a single fifty-person run against a sign-up we had deliberately broken in six different ways.
The eleven who stopped at sign-up mostly could not work out a date format given only as “ISO-8601”. The twenty-two who stopped at payment had just been shown a fee nobody mentioned, and asked for a national ID number.
An average hides the person you are losing.
Real figures from one fifty-person run against a deliberately flawed sign-up. Every row is a group that behaved nothing like the average:
| Who they are | What they did | Rate |
|---|---|---|
| Cautious about sharing data | Finished sign-up | 0% vs 56% |
| Compares prices before buying | Found the cheaper plan | 100% vs 5% |
| Reads the small print | Opened the terms | 92% vs 0% |
| Patient | Got through to the end | 59% vs 6% |
The paid add-on was pre-ticked and disclosed only in the terms. The people who read terms caught it. Nobody else did.
Anything a person does in a browser.
A scenario is one sentence: what this person came to do. The flock works out the rest — where to click, what to type, when it has had enough.
-
subscription software
“Open an account and end up on the plan that suits you.”
Surfaces: password rules nobody stated, pre-ticked add-ons, the step where trial-to-paid dies.
-
online shop
“Buy a pair of running shoes for under £100, delivered this week.”
Surfaces: delivery cost revealed at the last step, address entry, guest checkout dead ends.
-
banking and fintech
“Open an account and pass identity verification.”
Surfaces: document upload on a phone, jargon in error messages, who abandons rather than share a document.
-
marketplace
“Find something available near you and book it.”
Surfaces: empty-result dead ends, map-only discovery, the approval wait nobody explained.
-
support and AI agents
“Get a refund for an order that arrived broken.”
Surfaces: what your assistant does with an angry customer, a confused one, or one who is not a native speaker.
-
before and after
“The same job, on the version you are about to ship.”
Surfaces: whether the change you made helped — and which group it helped, or quietly hurt.
The sign-up button was off the edge of the phone.
a charging slot
Twenty simulated drivers were sent through a peer-to-peer EV charging marketplace to book a charging slot. None of them managed it. On a 390-pixel phone screen the header overflowed and the Sign up button sat between 385 and 506 pixels — past the edge of the display. On a laptop it was fine, which is why nobody had noticed.
People look for a charger on their phone. The product’s entire conversion path was unreachable for them.
Half of every person is measured. We will tell you which half.
Who arrives — their age, and how confidently they use a computer — is drawn from Eurostat, not invented. A cohort for Türkiye and a cohort for Spain are different populations, because they are.
| Market | People who struggle with everyday online tasks | Share | Median |
|---|---|---|---|
| Türkiye | 16.8% | 38 | |
| Italy | 10.7% | 47 | |
| Poland | 10.6% | 43 | |
| European Union | 7.3% | 45 | |
| Spain | 6.1% | 46 | |
| Germany | 5.8% | 45 |
Age. Digital skill given age. Which of them are on a phone. These come from a published survey you can look up, and a cohort cites it.
Patience, trust, how much frustration someone takes before leaving. These are reasoned guesses with no survey behind them — which is why a rate from a run describes the simulation and not your customers.
Everyone in this field claims realistic people. The question worth asking any of them — including us — is which parts are measured and which are asserted, and whether they will show you the line.
Simulated behaviour is not a forecast.
Synthetic people are more patient, more literate and more single-minded than real ones. We report what happened in the simulation and never dress it up as a prediction about your customers.
“31% of the simulated population abandoned at address entry, and low-fluency users were three times more likely to be among them.”
“You will lose 31% of your customers. Fixing this will raise conversion by 14%.”
This layer goes before real-user research and production analytics — not instead of them. It is the cheap pass that finds the obvious failures, so the expensive pass can look for the subtle ones.
Send the flock in first.
Real users shouldn’t be your first warning.
What it costs.
Two numbers on every plan, both fixed, and no meter. What you are buying is the resolution of the answer: thirty people tells you whether a flow works, a hundred tells you which kind of person it fails, two hundred and fifty shows you the segments too small to see any other way.
The whole product, for three runs. Work email, no card, no call. A run that fails does not count against them.
Start a free runOne product, shipped often. Enough to find out whether a flow works before anybody walks into it.
Cohorts big enough that segments become readable — which kind of person gave up, not just how many. Comparison runs: same people, two versions.
Populations large enough for the rare segments, and priority on anything that stops working.
Paid plans are invoiced rather than bought from a form — write to hello@canaryflock.com and we will set you up. Larger populations, a private deployment, SSO or a run triggered from your CI: those are worth a conversation rather than a price on a page. Every plan is read-only on a domain you have not proved you own.
Point it at something you are about to ship.
Three runs, up to thirty people each, at no charge. Sign in with a work email, give it a URL and the job somebody should be able to do there, and you get the findings and the session recordings back. A run that fails does not count against your three.
Start a free runRead-only unless you prove you own the domain: the flock never signs up, buys, or submits anything on somebody else’s product. Questions to hello@canaryflock.com.