How to write hypotheses that actually ship

Most experiment hypotheses die in docs or decks. The form that forces a kill criterion, names a mechanism, and lands in a ticket this week.

Kameron Tanseli

Kameron Tanseli

Head of growth engineering

Most hypotheses never leave the document they were written in.

Someone opens a Notion page, types "Improve the pricing page", adds a few bullets, and the work is considered done. The flag never gets created. The UI never gets drawn. The metric never gets instrumented. Six weeks later the same page is still leaking, and the team has a folder of untested ideas.

The failure is not laziness. It is the form. A hypothesis that cannot be killed, that does not name a mechanism, and that stays at the level of a desirable outcome will not become a ticket. It will become a discussion. Discussions do not ship.

This post is the form that does ship: a bet a sceptic can kill, written so the team can attach a flag this week.

A hypothesis is a bet a sceptic can kill

Karl Popper's criterion of falsifiability is usually treated as philosophy. In a growth team it is an engineering constraint. A claim is useful if a result can kill it. "We should improve conversion" cannot be killed. Any movement in the numbers can be narrated as progress, and any flat week can be blamed on seasonality. The claim survives every outcome, so it teaches nothing.

A shippable hypothesis dies cleanly.

If we remove the optional company field on the pay step, checkout completion rises, because each extra input raises cognitive load at the moment of commitment.

If completion is flat after one weekly cycle, the field was not the load. You learned something about this screen and you still have the pixels. That death is the point. You did not waste a design cycle on an unkillable story.

The same criterion kills the other common failure: the prediction that is too vague to measure. "Users will feel more confident" is not an event your analytics can log. "First-object creation in session rises" is. The kill criterion has to be an event you already log, or can log before the flag goes live.

Near construal makes the how visible

Trope and Liberman's construal level theory (Psychological Review, 2010) says psychological distance changes the grain of thought. Far events are represented by why and desirability. Near events are represented by how and feasibility. A quarterly OKR is far. A Linear ticket assigned this week is near.

"Simplify onboarding" is high-level construal. It names a desirable outcome and leaves the how unspecified. Everyone can agree because the agreement costs nothing. Engineering, design, and data then invent three different hows, and the flag that ships is one of those inventions rather than a deliberate test of a mechanism.

"Replace the three equal CTAs on the empty state with one primary Create project button" is low-level construal. The how is already in the sentence. The sceptic can point at the current screen and ask whether the hierarchy is the real problem. The designer can draw the treatment without inventing a new problem. The engineer can attach the flag without a clarifying meeting.

The tickets post covers why the artefact is a ticket; this post is about the quality of the bet inside it. Write the hypothesis at the near grain or it will never leave the far document.

Implementation intentions for the team

Gollwitzer (American Psychologist, 1999) showed that when-then plans raise follow-through versus mere goals. "I will exercise more" is a goal. "When it is 7am on a weekday, I will run" is an implementation intention. The situation cues the action, so working memory is not spent reconstructing the plan.

A growth hypothesis is the team's when-then. When this flag is assigned, ship this UI change on this screen, measure this event, kill it if the metric is flat. Without the when-then, the when is "after the next planning cycle" and the then is reconstructed from whoever still remembers the meeting. Cognitive load in the team is the same budget Sweller described for the user. Standup spent on "what did we mean by simplify pricing" is extraneous load. The written hypothesis is the plan, visible.

Pre-commit the kill in the same sentence. Loss aversion runs on the team too: once a treatment has been designed, shipping it feels like protecting an endowment. Write the fail condition before the mockup exists. "If Y is flat after one weekly cycle, we revert" belongs in the hypothesis, not in a later conversation.

The form that forces the mechanism

The sentence that ships is always the same shape.

If we change X on this screen, Y moves, because Z.

X is a single UI change on one viewport. It has to be visible in a screenshot. "Improve hierarchy" is not X. "Make the primary CTA the only filled button and demote the other two to text links" is X.

Y is an event you already log, or can log this week, measured on the same visit or within one weekly cycle. "Users feel less friction" is not Y. "Checkout completed" or "first_object_created" is Y. Guardrail metrics (refunds, support tickets, time-to-first-action) sit beside it so a lift that destroys quality is not celebrated.

Z is a named mechanism, not a vibe. "Because it feels cleaner" is not Z. "Because Hick's law predicts that three equal CTAs raise decision time on a screen where the alternatives are not equal" is Z. "Because each extra input raises cognitive load at the moment of commitment" is Z. The mechanism has to be falsifiable in the same way the metric is. If the result is flat, you learn that this mechanism did not apply here, or that the change did not instantiate the mechanism. Both are useful.

Leave Z out and you have a UI change with no theory. If it wins you do not know what to copy to the next surface. If it loses you do not know what to try next. Leave Y out and you will celebrate clicks. Leave X out and you have a research note.

Conditions under which the form is enough

The form produces a shippable ticket when three conditions hold.

  1. X is a change on one screen that can be captured in a screenshot today.
  2. Y is an event that already exists in the event taxonomy, or can be added before the flag goes live.
  3. Z is a mechanism that has been observed in prior work (your own past tests, published studies, or the CRO literature) and can be pointed at on this screen.

If the change spans three surfaces, the hypothesis is a programme. Split it until X fits on one viewport. If Y requires a new data pipeline that takes two sprints, write a smaller Y for the same X. If Z is "because users will like it", rewrite Z until it names a mechanism a sceptic can argue with.

A hypothesis that fails any of the three will still generate discussion. It will not generate a flag this week.

Failure modes that look like good writing

Four patterns survive review and still never ship.

The unkillable claim. "This will improve the user experience." No result can kill it. Rewrite until a flat metric ends the bet.

The mechanism-less change. "We will add social proof above the fold." X and Y may be present, but without Z the team cannot decide what to copy if it wins or what to try if it loses. Name the mechanism (trust reduction of perceived risk, specific social proof vs generic, placement relative to the decision).

The far construal. "We should reduce friction in the funnel." The how is missing. Drop the grain until X is a concrete pixel change.

The multi-surface bet. "Improve activation." Activation is a sequence. Write one hypothesis per surface, each with its own flag. Bundling them puts you back at high-level construal.

Each of these can be fixed by rewriting the same sentence until X, Y, and Z are all present and near.

Metrics for the hypothesis itself

Before the flag runs, the hypothesis already has a metric: does it become a ticket with a flag attached this week? Track that. A team that writes twenty elegant hypotheses and ships two is still losing experiments between the document and production. The volume of shippable hypotheses is the leading indicator of experiment velocity.

Once the flag is live, the usual rules apply. Assign at user grain, sticky. Wait at least one weekly cycle. Do not celebrate a rise in an intermediate click if the primary Y is flat. Log the kill decision in the same ticket so the next person can see why the treatment was reverted.

Write it as a ticket, then ship

Paste the hypothesis into Linear. Attach a screenshot of the current screen (the control) and a description or mock of the treatment. Instrument Y and the guardrails. Pre-write the kill criterion. Assign the flag this week.

The missing piece on most teams is the UI of the treatment. Someone has to draw what the user will see. If you have the screenshot and the one-sentence aim, that is the whole brief. MAGE takes the screenshot plus the aim and returns a spec: the hypothesis with the CRO mechanism named, and the UI to test. Lite gives ASCII. Starter and up give high-fidelity mockups. Paste that into the ticket.

A hypothesis that cannot be killed is a story. A hypothesis that names the mechanism, the metric, and the pixel change is a ticket. Write the second kind, attach a flag, and ship one treatment this week.

MAGE

Ready-to-run experiment tickets.

Upload a screen. Get a spec written like a growth engineer would write it.

Try MAGE