Arpeely · Conversion Funnel Challenge

How do you
actually compare?

A complete funnel for IQly — from a Google Search click on “free iq test” to a created account. Three parts: the funnel, three A/B tests, three creatives.

The brief in one line

Turn cold search traffic into accounts, with no brand to lean on.

IQly has no audience, no product equity and no reason to be trusted. Everything it is worth is concentrated in about fifteen screens.

The user typed the word free. Then we ask them for an account. Every decision here was measured against whether it breaks that promise.

The single defined KPI

CVR = account creations / page visits

If you read one thing

  1. The score is given away, not gated. A withheld score makes the account a toll. The account buys precision — where the number places you against people like you.
  2. The account has no password. An address plus a sign-in link. A password field at the moment of maximum anticipation is the most expensive field in the funnel.
  3. Question 1 is the landing page. No “Start” button. The cheapest way to lose a click is to ask for one before giving anything.
  4. Four files, one generator. Every variant is a value change in a single CONFIG object, so nothing can differ by accident.
  5. The funnel reports on itself. Every screen, answer and drop-off fires a dataLayer event, so the KPI is computable from the file.

Deliverables

Or just click through it here

All four builds run inside this page — no backend, no build step. Switch the build, switch the device.

Open full screen →

Open the browser console while you click — the funnel logs every event it would send.

Part 01

The funnel

Someone searching for an IQ test is not curious about psychometrics. They already hold a private theory about themselves and want a number that settles it. The product is recognition, not information — which is why they will sit through twelve questions and hand over an email.

Entry — the ad’s words, said back
Entry — the ad’s words, said back
One tap per screen, no typing
One tap per screen, no typing
A true pass rate, not a compliment
A true pass rate, not a compliment
The score, given away
The score, given away
Normalisation, not a marketing tax
Normalisation, not a marketing tax
A real account. No password.
A real account. No password.
Account created, and the matched report
Account created, and the matched report

The central decision: what is free, and what the email buys

The obvious move is to withhold the score. I rejected it. A withheld score makes the gate a toll, and a user who feels tolled at the moment of maximum anticipation abandons — or types a junk address, which corrupts the very metric we are optimising.

So the funnel gives the score away and sells precision instead.

Preliminary score: 108. Higher than 70% of everyone who has taken this test.
Three quick questions will place it against people like you.

Age, education and region normalise the result; the email delivers the matched report. The exchange is legible: the raw number is earned, the reference group is bought. It is also honest — a score is meaningless without a comparison group, so the demographic questions are not a tax on the user, they are the mechanism.

Supporting decisions

Question 1 is the entry screen

No landing page, no “Start”. The cheapest way to lose a click is to ask for a click before giving anything.

Zero typing until the email

Every question is one tap. The email field is the first keyboard event in the funnel — which is exactly why it is one field.

Difficulty is a curve

Easy items build competence, the hardest sit at 7–9, then it eases back so the test ends on capability rather than frustration.

No right/wrong feedback

A user who knows how they did no longer needs the result. Uncertainty is the engine, and it has to survive to the gate.

The account has no password

It is a real account — an address, opened by a sign-in link, and the result screen confirms it exists. A password buys nothing this product needs and costs the conversion outright.

The ad’s words, said back

The first screen carries IQly · Free IQ test and the free promise in the searcher’s own words. Message match is a CVR argument and a Quality Score one.

Desktop is a different layout, not a wider one

The reading measure stays narrow on both devices — a question set 900px wide is unreadable, and holding one measure means the device is not a confounder in the A/B tests. What changes is what the wider screen is genuinely better at: answers land in a 2×2 grid instead of a stack, halving the eye path per question, and the result becomes one view instead of a scroll — score on the left, comparison group on the right.

The same funnel at desktop width
The same question at desktop width — answers in a grid, not a stack

The three nudges are true: “Only 19% get this one right” is the item’s actual pass rate, and they attach automatically to the three hardest items in each test. They flatter without revealing.

What the test is actually made of

Every visitor gets a different test, sampled from a 41-item bank against a fixed blueprint, so the difficulty curve is identical while the items are not — answer-sharing stops working and per-item pass rates stay measurable.

The mix is deliberate. A test that claims to measure intelligence and then asks what year a product launched is measuring recall — and recall varies by age, country and schooling, which is the same axis the matched report claims to normalise away. So a twelve-item test is seven reasoning items, three generated visual matrices and two knowledge items, and the two sit at slots 1 and 12, where their job is not measurement: one is a warm-up the user will get right, the other ends the test on capability rather than frustration.

What I killed

  • A cognitive-decline angle. It converts, and Lumosity paid the FTC $2M for it. It also contradicts the keyword’s intent.
  • A learning-plan offer at the gate. Answers a need the searcher doesn’t have, and signals an upsell behind “free”.
  • Results delivered by email. Deferring gratification on mobile is an abandonment event. The email is a copy, not the channel.
  • A per-question timer. Adds credibility, adds abandonment. Better tested than assumed.

Measurement — built in, not described

A funnel that cannot report on itself cannot be optimised, so the events are in the file. Open the console and it narrates its own KPI:

page_visit        → the denominator
question_answered → per-question drop-off
test_completed    → completion, separated from conversion
gate_rejected     → friction at the ask
account_created   → the numerator

Each event carries the variant name and the inbound gclid, keyword and UTMs, so CVR is readable per variant and per keyword rather than in aggregate. The variant name is read from the filename, so no file carries a config value the others do not.

Primary is the defined KPI. But CVR alone is inflatable — a hard gate buys accounts worth nothing — so I would pair it with deliverable-email rate and per-question drop-off. Those two separate a funnel that converts from one that merely extracts.

Part 02

Three A/B tests

Each variant changes one thing. All three attack the same question from different sides: how much does a user need to have already received before the account stops feeling like a toll?

1 · Remove the preliminary score

Hypothesis. The control gives the score away to establish good faith, which may leave conversion on the table — a user who already has their number has had the curiosity that carried them here resolved. Withholding it keeps the tension intact through the ask.

Why it might lose. It converts the gate from a trade into a toll. Expect more accounts and worse ones.

KPIs. Primary: CVR. Guardrails: deliverable-email rate, gate abandonment. If CVR rises while deliverability falls, the variant is buying inventory, not users.

2 · Ask demographics before scoring

Hypothesis. In the control the profile questions sit after a score the user already holds, so they are three screens of friction. Moving them earlier reframes them as part of being scored rather than an obstacle to a result.

Why it might lose. It front-loads friction at the point of lowest accumulated investment, and it delays the score.

KPIs. Primary: CVR. Diagnostics: demographic-block completion, step-level drop-off. This is the cleanest read on whether profile questions are friction or scaffolding.

3 · Six questions instead of twelve

Hypothesis. Length trades two forces. Longer means more sunk investment and more credibility at the gate; shorter means fewer chances to quit and a lower promised cost on entry — “90 seconds” instead of “3 minutes”. Twelve was a guess; this measures which force dominates.

Why it might lose. A six-question result is easier to dismiss as a toy, which undercuts why the matched report is worth anything.

KPIs. Primary: CVR. Diagnostics: test completion rate and CVR among completers, measured separately — those moving in opposite directions is the expected result and the actual finding.

Keeping it one variable. The six items are a marked, representative sample of the same blueprint — same curve, same mix of reasoning and visual. Taking the first six slots instead would have quietly stripped every visual item, and the test would have measured length and item type at once.

How they were built

All four files come from one source and one CONFIG object. Variant 2 is a single value change:

{ demographics: { position: 'before-scoring' } }

Screen order is derived from config rather than hardcoded. Each variant carries one independent variable plus the copy that variable forces — Variant 1 also changes the button label because there is no report to show; Variant 3 also changes “3 minutes” to “90 seconds” because the promise would otherwise be false. Both are declared in config rather than hand-edited into the page, which is the difference between a dependent change and a second variable. A diff of any variant against the control is two to four lines, all of them inside CONFIG. Adding a fourth is a minute of work.

Part 03

Three creatives

Three formats, three different jobs. Not one design resized — a 320×50 slot and a full screen ask for different things, and treating them as one layout wastes the larger surface and overloads the smaller one.

The constant: message match

Someone arriving from “free iq test” should land on a page that looks like the ad they clicked. That is a conversion argument — and an auction one. Google scores landing-page experience, so a creative that oversells relative to its destination raises CPC on the other side of the equation.

320x50 banner
320 × 50 · mobile banner
Room for one idea. The keyword as an eyebrow, the differentiating question below, CTA on the right. Nothing competing for 50 pixels.
250x250 square
250 × 250 · square
Shows a real item from the test. The viewer starts solving before they click — the ad is the product.
1080x1920 vertical
1080 × 1920 · full screen
Shows the payoff, not the pitch — the actual result screen, so the demographic questions later read as the thing that was advertised.

Same gradient family, same typeface, same corner radii, and the charcoal CTA as the only dark element in all three. The experience is continuous from impression to account.

Appendix

Where AI helped

Not by writing the funnel. By making a system where I would otherwise have hand-built a one-off — and then by catching things review would not.

Structure from the model, facts from me

A single wrong fact in a test that claims to measure intelligence destroys the credibility of every other item. So the bank is separated by factual risk:

LayerProduced howRisk
Visual itemsGenerated from a transformation ruleNone — there is no fact to get wrong
Reasoning trapsCanonical puzzles, settled answersLow
Knowledge itemsAuthored, then verified individuallyManaged — and only here

41 items in the bank, and the blueprint draws seven reasoning items, three generated visual matrices and two knowledge items per test — the composition argued for in Part 1. Every visitor gets a different test with an identical difficulty curve.

What the automated checks caught

Checks run over the whole bank and every build, so adding 200 more questions costs writing, not reviewing. They found:

Every one of these would have shipped unnoticed. None of them was visible in the code.

The point

I did not build a funnel. I built a generator, and the attached files are four configurations of it. That is the difference between a deliverable and something a team can keep testing on Monday.

Try the flow Write-up (PDF)