Growth

ICE score calculator and growth experiment planner

By Charles Summers · Updated · Free, no signup

Short answer

This scores a growth experiment using the ICE framework (Impact, Confidence, Ease) and returns more than a number: a properly formed hypothesis, the minimum sample size question you need to answer before starting, a kill criterion, and where the test belongs in your sprint order. Confidence is weighted more heavily than the raw ICE average because low-confidence high-impact ideas are the most common way growth teams waste a quarter.

Use the growth experiment planner

What does this tool actually do?

This scores a growth experiment using the ICE framework (Impact, Confidence, Ease) and returns more than a number: a properly formed hypothesis, the minimum sample size question you need to answer before starting, a kill criterion, and where the test belongs in your sprint order. Confidence is weighted more heavily than the raw ICE average because low-confidence high-impact ideas are the most common way growth teams waste a quarter..

It runs entirely in your browser. Nothing you type is sent to a server, no account is required, and there is no usage limit, because there is no cost per run to control.

What ICE is good for, and what it quietly hides

ICE scores an idea on Impact, Confidence and Ease, each on a ten point scale, and ranks the backlog by the result. Sean Ellis popularised it precisely because it is crude. Its value is not that it predicts outcomes, because it does not. Its value is that it forces a vague argument into three named dimensions, so when two people disagree about an idea they can locate the disagreement. Most backlog arguments are actually confidence arguments dressed up as impact arguments, and the framework surfaces that in about a minute.

The hidden failure is that scores are not comparable across people. Everyone rates their own idea eight, eight, eight. Confidence is where the inflation concentrates, because impact can be sanity-checked against the metric and ease can be sanity-checked against an engineer, while confidence is pure self-report. A backlog where the top ten items all score above 7 is not a prioritised backlog, it is a list in the order people submitted things, with numbers attached for decoration.

The fix is not a better formula. It is anchoring each dimension to something external. Impact should be expressed as an estimated change in the actual metric, not a feeling. Ease should be expressed in person-days by whoever will build it, not by whoever wants it built. Confidence should be tied to an evidence tier, described below, so that a 9 means something specific and cannot be awarded for enthusiasm.

Anchoring confidence to evidence, not enthusiasm

Confidence is a claim about how much you already know, which means it can be graded. Use a fixed ladder and make people name the tier out loud rather than picking a number. The effect on a scoring session is immediate, because someone has to say the words this is tier one, we made it up in the meeting, and that is much harder than typing 8 into a spreadsheet cell.

  • 9 to 10. You have run essentially this test before, on a comparable surface, and it won. Re-running a proven pattern on a new page is the cheapest reliable growth work available and it is systematically underrated.
  • 7 to 8. Your own data points at this specific cause. A funnel step with an unexplained drop, session recordings showing the same stall, a support ticket theme repeating fifty times a month.
  • 5 to 6. A credible external result plus a plausible mechanism in your context. Competitors doing it consistently counts here, but only weakly, because you cannot see whether it worked for them.
  • 3 to 4. Informed opinion from someone with real domain knowledge and no supporting data. Worth testing, worth being honest that it is a guess by an expert.
  • 1 to 2. It sounded good in the meeting. These should still get built occasionally, because cheap wild ideas are where the outsized wins hide, but they should never outrank a tier four with the same impact estimate.

The arithmetic that decides whether the test can run at all

Before scoring anything, check whether the surface has enough traffic to produce an answer. Sample size scales with the inverse square of the effect you want to detect, which is brutal at low conversion rates. At a 3% baseline, detecting a 20% relative lift (3% to 3.6%) takes roughly 13,000 visitors per variant at the conventional 95% confidence and 80% power. Detecting a 10% relative lift on the same baseline takes closer to 50,000 per variant. A page seeing 2,000 visitors a month cannot run either test. Not slowly, at all.

When the maths says no, you have four honest options and one dishonest one. Move the test to a higher-traffic surface. Test a bigger swing, since a redesigned page detects faster than a button colour. Measure a proxy closer to the change, such as click-through on the step rather than final conversion, accepting that a proxy win is weaker evidence. Or run it as a judgement call, ship it, and label it a decision rather than a result. The dishonest option is running it anyway and reading the noise.

Duration has a separate floor. Run for at least two complete weeks so the sample covers both weekday and weekend behaviour and any weekly buying cycle. Stopping a test the first time it crosses significance, then declaring a win, inflates your false positive rate well above the 5% the test assumes, because you have effectively taken many bites at the same cherry. Fix sample size and end date in advance and write them into the test plan.

Running a programme rather than a series of tests

Calibrate expectations early. Across large-scale experimentation programmes, the widely reported pattern is that roughly a third of tests improve the target metric, a third do nothing, and a third make things worse. Teams new to experimentation typically report higher win rates, which usually means they are calling noise and shipping things that will not hold. A mature win rate in the 10% to 30% range is a sign of honest measurement, not of a weak team.

The compounding asset is the log, not the wins. Record every test with its hypothesis, its ICE inputs, the result, and one sentence on what it taught you, including the losers. Losers are more informative than winners because they usually falsify an assumption several other backlog items were quietly resting on. A year of that log is what stops a team from re-running the same failed pricing page test every time a new PMM joins.

On cadence, two to six shipped tests a month is realistic for a small team, and the constraint is almost never ideas. It is decision latency. Give every test one named owner, one primary metric agreed before launch, and a decision date in the calendar. Tests that stay live because nobody wants to call them are the main reason experimentation programmes stall in month four.

Numbers worth knowing

MetricTypicalWhat it means
Experiment win rate10% to 30%Counting only tests that improve the primary metric and hold up afterwards. Reported win rates above 50% almost always mean tests are being stopped early.
Minimum test duration2 full weeksCovers the weekly cycle. Shorter runs oversample whichever days happened to fall inside the window, which is why Monday-to-Friday tests flatter promotions.
Sample for a 20% lift on a 3% baselineabout 13,000 per variantAt 95% confidence and 80% power. Halve the effect you want to detect and this roughly quadruples, which is why small sites should test big changes.
Tests shipped per month, small team2 to 6The bottleneck is deciding and building, not generating ideas. If your backlog has 80 items and you ship two a month, the backlog is a wish list.

Mistakes that quietly cost you results

Stopping the test the moment it hits significance
Checking repeatedly and stopping on the first green result means your real false positive rate is far above the 5% you think you are running at. Set the sample size and end date before launch and hold to both.
Running tests the traffic cannot support
A test that needs 13,000 per variant on a page getting 2,000 a month never produces an answer, it produces a coin flip dressed as data. Check the sample requirement before scoring the idea, and route low-traffic pages to judgement calls instead.
Letting people score their own ideas on confidence
Confidence is the only unverifiable input, so it absorbs all the wishful thinking and the ranking becomes a popularity contest. Force an evidence tier to be named aloud, and have someone other than the author confirm it.
Changing several things at once and calling it an experiment
A new headline, new hero image and new form in one variant tells you the page moved but not which change moved it, so the learning does not transfer anywhere else. Bundle only when you are optimising for speed and say so explicitly.
Logging only the tests that won
Losers falsify assumptions that other backlog items depend on, so discarding them guarantees the same idea returns in six months with a new sponsor. One line per test on what it disproved is the highest-return documentation in growth.

What does the output look like?

This is the exact output the tool produces from the example inputs. It is generated by the same code that runs when you click the button, so what you see here is what you get.

EXPERIMENT Add social proof above the pricing table HYPOTHESIS We believe that add social proof above the pricing table will improve trial-to-paid conversion rate from a baseline of 4.2%. We will know we are right when trial-to-paid conversion rate moves and holds for two consecutive weeks. SCORES Impact 7/10 Confidence 5/10 Ease 9/10 ----------------------- ICE (mean) 7.0 Confidence-weighted 15.8 VERDICT: DE-RISK FIRST The score is inflated by optimism. Run a cheap validation (5 customer calls, a fake-door test, a survey) before committing build time. EVIDENCE CHECK Confidence at 5/10 is workable but thin. One customer conversation or one lookalike case study would firm this up cheaply. KILL CRITERION If trial-to-paid conversion rate has not moved measurably within 7 days, stop and write down why. Do not extend the test to find a result. BEFORE YOU START 1. How many users or sessions do you need for the result to mean anything? 2. What is the one number you will look at, and who checks it? 3. What will you do if it works? If there is no next step, the test is not worth running.

Frequently asked questions

How is the ICE score calculated?

The classic version is the mean of Impact, Confidence and Ease. This tool reports that, and also a confidence-weighted score that squares the confidence term. The reason is practical: a 10/2/10 idea and a 7/7/7 idea have nearly the same plain ICE score, but the first is a guess and the second is a plan. Weighting confidence separates them.

What is a good ICE score?

Scores are only meaningful relative to your other ideas, not against an absolute bar. A backlog where everything scores above 8 means the scoring is not honest yet. Expect a spread, and expect most ideas to land between 4 and 7.

Why does the output include a kill criterion?

Because the most expensive experiment is the one nobody ends. Deciding in advance what result would make you stop is the difference between a test and a preference. The tool proposes one based on your ease score, since cheap tests deserve shorter leashes.

Should low-ease experiments ever go first?

Only when the impact is high and confidence is genuinely high, which is rare. The sequencing note in the output tells you which bucket your idea falls into and what should be true before you commit engineering time to it.

Related free tools

Some links on this site are affiliate links, which means Hacking Demand may earn a commission if you buy through them at no extra cost to you. This does not influence which tools are listed. The tools on this page are free and have no affiliate relationship of any kind.