Turn cold search traffic into accounts, with no brand to lean on.
IQly has no audience, no product equity and no reason to be trusted. Everything it is worth is concentrated in about fifteen screens.
The single defined KPI
CVR = account creations / page visits
If you read one thing
- The score is given away, not gated. A withheld score makes the account a toll. The account buys precision — where the number places you against people like you.
- The account has no password. An address plus a sign-in link. A password field at the moment of maximum anticipation is the most expensive field in the funnel.
- Question 1 is the landing page. No “Start” button. The cheapest way to lose a click is to ask for one before giving anything.
- Four files, one generator. Every variant is a value change in a single
CONFIGobject, so nothing can differ by accident. - The funnel reports on itself. Every screen, answer and drop-off fires a
dataLayerevent, so the KPI is computable from the file.
Deliverables
Or just click through it here
All four builds run inside this page — no backend, no build step. Switch the build, switch the device.
Open the browser console while you click — the funnel logs every event it would send.
The funnel
Someone searching for an IQ test is not curious about psychometrics. They already hold a private theory about themselves and want a number that settles it. The product is recognition, not information — which is why they will sit through twelve questions and hand over an email.
The central decision: what is free, and what the email buys
The obvious move is to withhold the score. I rejected it. A withheld score makes the gate a toll, and a user who feels tolled at the moment of maximum anticipation abandons — or types a junk address, which corrupts the very metric we are optimising.
So the funnel gives the score away and sells precision instead.
Three quick questions will place it against people like you.
Age, education and region normalise the result; the email delivers the matched report. The exchange is legible: the raw number is earned, the reference group is bought. It is also honest — a score is meaningless without a comparison group, so the demographic questions are not a tax on the user, they are the mechanism.
Supporting decisions
Question 1 is the entry screen
No landing page, no “Start”. The cheapest way to lose a click is to ask for a click before giving anything.
Zero typing until the email
Every question is one tap. The email field is the first keyboard event in the funnel — which is exactly why it is one field.
Difficulty is a curve
Easy items build competence, the hardest sit at 7–9, then it eases back so the test ends on capability rather than frustration.
No right/wrong feedback
A user who knows how they did no longer needs the result. Uncertainty is the engine, and it has to survive to the gate.
The account has no password
It is a real account — an address, opened by a sign-in link, and the result screen confirms it exists. A password buys nothing this product needs and costs the conversion outright.
The ad’s words, said back
The first screen carries IQly · Free IQ test and the free promise in the searcher’s own words. Message match is a CVR argument and a Quality Score one.
Desktop is a different layout, not a wider one
The reading measure stays narrow on both devices — a question set 900px wide is unreadable, and holding one measure means the device is not a confounder in the A/B tests. What changes is what the wider screen is genuinely better at: answers land in a 2×2 grid instead of a stack, halving the eye path per question, and the result becomes one view instead of a scroll — score on the left, comparison group on the right.
The three nudges are true: “Only 19% get this one right” is the item’s actual pass rate, and they attach automatically to the three hardest items in each test. They flatter without revealing.
What the test is actually made of
Every visitor gets a different test, sampled from a 41-item bank against a fixed blueprint, so the difficulty curve is identical while the items are not — answer-sharing stops working and per-item pass rates stay measurable.
The mix is deliberate. A test that claims to measure intelligence and then asks what year a product launched is measuring recall — and recall varies by age, country and schooling, which is the same axis the matched report claims to normalise away. So a twelve-item test is seven reasoning items, three generated visual matrices and two knowledge items, and the two sit at slots 1 and 12, where their job is not measurement: one is a warm-up the user will get right, the other ends the test on capability rather than frustration.
What I killed
- A cognitive-decline angle. It converts, and Lumosity paid the FTC $2M for it. It also contradicts the keyword’s intent.
- A learning-plan offer at the gate. Answers a need the searcher doesn’t have, and signals an upsell behind “free”.
- Results delivered by email. Deferring gratification on mobile is an abandonment event. The email is a copy, not the channel.
- A per-question timer. Adds credibility, adds abandonment. Better tested than assumed.
Measurement — built in, not described
A funnel that cannot report on itself cannot be optimised, so the events are in the file. Open the console and it narrates its own KPI:
page_visit → the denominator
question_answered → per-question drop-off
test_completed → completion, separated from conversion
gate_rejected → friction at the ask
account_created → the numerator
Each event carries the variant name and the inbound gclid, keyword and UTMs, so CVR is readable per variant and per keyword rather than in aggregate. The variant name is read from the filename, so no file carries a config value the others do not.
Primary is the defined KPI. But CVR alone is inflatable — a hard gate buys accounts worth nothing — so I would pair it with deliverable-email rate and per-question drop-off. Those two separate a funnel that converts from one that merely extracts.
Three A/B tests
Each variant changes one thing. All three attack the same question from different sides: how much does a user need to have already received before the account stops feeling like a toll?
1 · Remove the preliminary score
Hypothesis. The control gives the score away to establish good faith, which may leave conversion on the table — a user who already has their number has had the curiosity that carried them here resolved. Withholding it keeps the tension intact through the ask.
Why it might lose. It converts the gate from a trade into a toll. Expect more accounts and worse ones.
KPIs. Primary: CVR. Guardrails: deliverable-email rate, gate abandonment. If CVR rises while deliverability falls, the variant is buying inventory, not users.
2 · Ask demographics before scoring
Hypothesis. In the control the profile questions sit after a score the user already holds, so they are three screens of friction. Moving them earlier reframes them as part of being scored rather than an obstacle to a result.
Why it might lose. It front-loads friction at the point of lowest accumulated investment, and it delays the score.
KPIs. Primary: CVR. Diagnostics: demographic-block completion, step-level drop-off. This is the cleanest read on whether profile questions are friction or scaffolding.
3 · Six questions instead of twelve
Hypothesis. Length trades two forces. Longer means more sunk investment and more credibility at the gate; shorter means fewer chances to quit and a lower promised cost on entry — “90 seconds” instead of “3 minutes”. Twelve was a guess; this measures which force dominates.
Why it might lose. A six-question result is easier to dismiss as a toy, which undercuts why the matched report is worth anything.
KPIs. Primary: CVR. Diagnostics: test completion rate and CVR among completers, measured separately — those moving in opposite directions is the expected result and the actual finding.
Keeping it one variable. The six items are a marked, representative sample of the same blueprint — same curve, same mix of reasoning and visual. Taking the first six slots instead would have quietly stripped every visual item, and the test would have measured length and item type at once.
How they were built
All four files come from one source and one CONFIG object. Variant 2 is a single value change:
{ demographics: { position: 'before-scoring' } }
Screen order is derived from config rather than hardcoded. Each variant carries one independent variable plus the copy that variable forces — Variant 1 also changes the button label because there is no report to show; Variant 3 also changes “3 minutes” to “90 seconds” because the promise would otherwise be false. Both are declared in config rather than hand-edited into the page, which is the difference between a dependent change and a second variable. A diff of any variant against the control is two to four lines, all of them inside CONFIG. Adding a fourth is a minute of work.
Three creatives
Three formats, three different jobs. Not one design resized — a 320×50 slot and a full screen ask for different things, and treating them as one layout wastes the larger surface and overloads the smaller one.
The constant: message match
Someone arriving from “free iq test” should land on a page that looks like the ad they clicked. That is a conversion argument — and an auction one. Google scores landing-page experience, so a creative that oversells relative to its destination raises CPC on the other side of the equation.
Same gradient family, same typeface, same corner radii, and the charcoal CTA as the only dark element in all three. The experience is continuous from impression to account.
Where AI helped
Not by writing the funnel. By making a system where I would otherwise have hand-built a one-off — and then by catching things review would not.
Structure from the model, facts from me
A single wrong fact in a test that claims to measure intelligence destroys the credibility of every other item. So the bank is separated by factual risk:
| Layer | Produced how | Risk |
|---|---|---|
| Visual items | Generated from a transformation rule | None — there is no fact to get wrong |
| Reasoning traps | Canonical puzzles, settled answers | Low |
| Knowledge items | Authored, then verified individually | Managed — and only here |
41 items in the bank, and the blueprint draws seven reasoning items, three generated visual matrices and two knowledge items per test — the composition argued for in Part 1. Every visitor gets a different test with an identical difficulty curve.
What the automated checks caught
Checks run over the whole bank and every build, so adding 200 more questions costs writing, not reviewing. They found:
- A correct answer that explained its own trick — “No — he would have to be dead”.
- Three items where the right answer was visibly the longest option, markable without reading the question.
- A percentile that read as praise for a below-average score — 82 shown as “Top 88%”.
- An infinite redirect loop in Variant 1, and demographics silently skipped in Variant 2 — both invalidated-experiment bugs, not cosmetic ones.
- The privacy line vanishing from all four files after a config refactor, failing silently.
- Transparent corners on all three creatives, which read as black notches on a dark page.
The point
I did not build a funnel. I built a generator, and the attached files are four configurations of it. That is the difference between a deliverable and something a team can keep testing on Monday.