A/B tests for empty states
An empty state is the first product a new user meets. Five shippable hypotheses, the psychology, and the metrics.
An empty state is not a missing illustration. It is the first product a new user uses.
Signup created an account. The next screen is a list, a board, a canvas, or a dashboard with nothing in it. Activation is decided here: between the account row and the first object in the database. If this screen apologises, the user learns the product is empty. If it offers one action, the user learns the product is a tool.
Most teams never test it. They ship a mascot, a grey "Nothing here yet", and three equal buttons. Then they blame onboarding for the leak.
Context
An empty state is any in-product container that would normally hold the user's work, and currently does not. First session, first project, first search with zero results.
The job of that screen is not to look designed. The job is to produce one real object: a project, a doc, a campaign, a spec. That object is the activation event. Time-to-first-object is the clock.
Blank canvases fail for a specific reason. They ask a novice to do expert work: invent the first object from nothing, with no example of what "good" looks like. Hick's law says decision time rises with the number of options. A blank editor with infinite possibility is the largest option set you can put on a screen.
Magic Hour tested the opposite. Instead of dropping people into an empty editor, they showed ready-made templates. First-week retention moved about 18 percent. The learning was not that templates are pretty. It was that cutting cognitive load at the first step beats showing the product's full range.
Do not start by testing the illustration. Test the action, the starter object, and the number of choices.
What a shippable hypothesis looks like
Write it as a bet a sceptic can kill.
If we change X on this empty state, Y (first object created, or first-week retention) moves, because Z (a named CRO mechanism).
X is a UI change on one screen. Y is an event you already log, or can log this week. Z is a mechanism, not a vibe. "Because it feels friendlier" is not Z. "Because Hick's law predicts three equal CTAs raise decision time" is Z.
Run the tests below one at a time, assigned at user grain, on users who have zero objects. Guardrail: do not celebrate a rise in "clicked CTA" if first-object creation is flat. Clicks without objects are a prettier dead end.
Hypothesis 1. One first action, not an apology
Hypothesis. If we replace "Nothing here yet" plus an illustration with a headline that names the object and a single CTA that creates it, more new users will create a first object in the same session, because the screen stops asking them to interpret emptiness and starts asking them to take one implementation intention.
Why it should win. Cognitive load is the working memory spent figuring out what the screen wants. An apology ("Your list is empty") spends that memory on a status the user can already see. An implementation intention is a when-then plan: when I am on this screen, I create a project. The CTA is that plan, visible.
The goal-gradient effect (Hull, 1932) says people work harder as a goal gets closer. Naming the object ("Create your first project") makes the goal specific. "Get started" does not.
Implementation. Kill the illustration on the treatment. Headline: "Create your first {object}." One sentence on what that object is for. One primary button: "Create {object}." Import, invite, and sample sit under a text link. Flag the empty-state component. Do not change the create-object modal in the same test.
Metrics. Primary: first object created within the session, among new users who saw the empty state. Secondary: time-to-first-object. Guardrail: next-session return.
Expected impact. Medium-to-high on activation. The usual miss is keeping "Get started" as the label, which is not an action.
Hypothesis 2. Templates instead of a blank canvas
Hypothesis. If we show three starter templates on the empty state, and make "start from scratch" a secondary path, more new users will create a first object and more of them will still be active at day 7, because templates convert an unbounded invention task into a bounded edit task.
Why it should win. A blank canvas is not zero options. It is every option. Templates collapse that set to three named starts. People also simulate the filled product faster when they can see a filled product. Editing a template produces a first object with less expertise than composing one. Watch the template count. Three beats twelve. Twelve is another blank canvas, just with thumbnails.
Implementation. Three templates, each named for an outcome ("Weekly standup notes", not "Template A"). Clicking a template creates the object prefilled. "Start blank" is a text link. Do not auto-create all three. Do not mix this test with a change to the editor chrome.
Metrics. Primary: first object created. Secondary: day-7 retention, and the share of first objects that are still non-empty 24 hours later. Guardrail: objects created from templates that never get edited. If that share is high, you moved a vanity create event.
Expected impact. High on products where the first object has structure (docs, boards, campaigns, specs). Lower where the first object is a single field.
Hypothesis 3. Sample data that looks like a filled account
Hypothesis. If we seed a new account with one obvious sample {object}, labelled as sample and deletable in one click, more new users will reach a second real object within seven days, because they learn the product by using it instead of imagining it.
Why it should win. A filled screen teaches the information architecture for free: where a title lives, where the status lives, where the primary action lives. That is cheaper than a tooltip tour.
The trap: if sample data looks like their data, they will not trust the product. Label it. Make delete obvious. Measure real objects, not interaction with the sample.
Implementation. One sample, not a demo universe. Name it "Sample: {outcome}." Persistent "Remove sample" control. The empty state still shows if they delete it. Assign at signup, before the first home render, so you do not flash empty then fill.
Metrics. Primary: first user-authored object (exclude the sample id). Secondary: time-to-first-real-object, and whether they opened the sample. Guardrail: support tickets about "who created this?", and deletion of the sample without a follow-up create.
Expected impact. High for dense UI (tables, boards, analytics). Medium for simple lists. Skip it if the sample cannot be sandboxed from billing, permissions, or customer data.
Hypothesis 4. Endowed progress on the empty screen
Hypothesis. If we put a three-to-five step first-win checklist on the empty state, with the steps they already completed (account created, workspace named) marked done, more new users will finish the remaining step that creates the first object, because the goal-gradient effect is stronger when some of the path is already behind them.
Why it should win. Hull's goal-gradient: approach motivation rises as the remaining distance falls. Endowed progress (Nunes and Dreze) is the design version. Give people a head start on a punch card and they finish more often than an empty card of the same remaining length.
A checklist also chunks the work. "Activate" is a programme. "Create a project" is a step. Keep it short. Five steps is the ceiling. If step one is "Watch a 4-minute video", you have designed a course, not a first win.
Implementation. Three steps. Step 1 and 2 pre-checked if they are true. Step 3 is the object-create CTA, visually primary. Persist the checklist until the first object exists, then kill it. Do not add a product tour in the same experiment.
Metrics. Primary: first object created. Secondary: checklist completion and time-to-first-object. Guardrail: session length (a checklist can inflate time without creating objects) and day-7 retention.
Expected impact. Medium-to-high. The failure mode is a checklist of product features instead of the one activation event.
Hypothesis 5. One primary CTA, not three equal doors
Hypothesis. If we keep a single filled CTA for the first-win action and demote import, invite, and "start from scratch" to secondary text, more new users will complete any first action, because Hick's law says three equal choices raise decision time and drop completion.
Why it should win. Hick's law (Hick and Hyman): reaction time grows roughly with the log of the number of options. On an empty state the options are rarely equal in value. Creating the object that is this product is the activation event. Invite is a later loop. Import is a power-user path. Presenting them as peers makes the important one look optional.
Import and invite are useful. They are not first.
Implementation. Treatment: one filled button, two text links underneath, or an overflow "More ways to start." Control: the current three-button row. Same labels, different hierarchy. Do not rewrite copy in the same test if you can avoid it. Hierarchy is the treatment.
Metrics. Primary: any first successful action within the session. Secondary: mix of which action won, and first-object creation. Guardrail: invite rate. A stronger create CTA can suppress invites in week one.
Expected impact. Medium. Cheap to ship. Often the right first test: it needs no new objects or template content.
Which test to run first
| Test | Psychology | Effort | First metric | When to pick it |
|---|---|---|---|---|
| One first action | Cognitive load, implementation intention | Low | First object in-session | Copy is currently an apology |
| One primary CTA | Hick's law | Low | Any first action | Three equal buttons already exist |
| Templates | Hick's law, concreteness | Medium | First object + day-7 retention | First object has structure |
| Endowed checklist | Goal-gradient, endowed progress | Medium | First object | You already have a first-win definition |
| Sample data | Concreteness, learning by using | High | First real object | Dense UI, sample can be sandboxed |
Run the low-effort tests first if you have not instrumented the empty state. A hierarchy test that ships this week beats a sample-data platform that ships next quarter.
Metrics, assignment, and how long to wait
Log four events before you start:
empty_state_viewed(which surface, which cohort: first session vs later)empty_state_cta_clicked(which action)first_object_created(object type, from_template / from_scratch / from_sample)day_7_retained
Assign the flag at user grain, at first empty-state view, sticky. Do not re-assign on refresh. Exclude users who already have objects. If the empty state also appears for returning users who deleted everything, split that cohort. Their psychology is not first-session.
Wait at least one weekly cycle. Peeking at day two will lie to you about week-one retention. If traffic is thin, first-object-in-session on the same page is a valid learning metric. Couple the change to a response on the same visit. Do not test the empty state against paid conversion three steps later and wait six months.
Write it as a ticket, then ship
Each of the five is already a ticket if you paste the hypothesis, the treatment, and the metrics into Linear and attach a flag. The missing piece on most teams is the UI. Someone has to draw the treatment.
If you have a screenshot of the current empty state, that is the whole brief. MAGE takes the screenshot plus one sentence for the aim ("More new users create a first project from this empty state") and returns a spec: the hypothesis with the CRO mechanism named, and the UI to test. Lite gives ASCII. Starter and up give high-fidelity mockups. Paste that into the ticket.
An empty state you do not test is a first impression you left to the illustration library. Pick the lowest-effort row in the table, instrument the four events, and ship one treatment this week.
MAGE
Ready-to-run experiment tickets.
Upload a screen. Get a spec written like a growth engineer would write it.
Try MAGE
