Working with the agent

A/B tests

Split your traffic between the current experience and the new one, and let your customers decide. Run it with HI Engine, or hand it to Intelligems, Shoplift or Shopify Rollouts.

Shipping a change to everyone assumes it works. A test finds out.

Testing is the default path for an approved build. Publish straight to everyone when the change fixes something plainly broken, or when your traffic is too thin to measure anything. Otherwise, test.

Set up a test

Write the hypothesis

The test setup starts with one sentence: what you expect to change, and why.

HI Engine drafts it from the idea's evidence. Edit it if you disagree — the hypothesis is what the result gets read against, so it is worth a minute.

Pick the tool

Choose which platform runs the split. The options are below.

Launch it

With HI Engine, the test starts when you confirm. With the other tools you launch it yourself, using the instructions HI Engine prepares.

Which tool to use

HI Engine

Recommended. We run the split, measure the revenue difference, and keep the change only if it wins.

Intelligems

We write the launch instructions and tell you when they are ready. You run the test and control when it starts and stops.

Shoplift

A Shopify A/B testing platform. You run the test; describe the variant so HI Engine can read the result against it.

Shopify Rollouts

Shopify's native A/B testing. Same shape as Shoplift — you run it, HI Engine tracks the outcome.

If you use a platform not on that list, choose Other and name it. HI Engine still records the hypothesis, the variant and the result against the task.

HI Engine

The default. Half your visitors see the current experience, half see the approved change. Each visitor stays in the same variant across visits.

You do not describe the variant — HI Engine built it, so it already knows what is being tested.

Intelligems

HI Engine writes the launch instructions for the test and notifies you when they are ready. You start and stop the test inside Intelligems, so the timing stays in your hands.

Use this when you already run price or offer tests in Intelligems and you want every experiment in one place.

Shoplift and Shopify Rollouts

You run the test in your own tool. HI Engine records the hypothesis and the result against the task, so the outcome still feeds the ranking.

For these, describe the variant when you set the test up. That description is what HI Engine reads the result against later.

Reading the result

HI Engine reports the conversion rate of each arm, the difference between them, and how likely that difference is to be noise.

ResultWhat it meansWhat to do
Variant B winsThe difference is real at the stated confidencePublish B
Variant A winsThe change hurtRoll back, and read why
No measurable differenceThe test could not separate themRoll back, and spend the traffic on a bigger change

A result is a winner only when the difference is statistically distinguishable from noise. Otherwise the honest answer is "no measurable difference", and HI Engine says exactly that.

A flat result is a real result

"No difference" saves you from shipping a change that costs effort and returns nothing. It also tells HI Engine to rank that category of idea lower for your store in future.

How long a test needs

A test ends when the result is statistically significant: when the gap between the two arms is large enough, over enough visitors, that chance is an unlikely explanation for it.

Two arms almost never convert at exactly the same rate, even when the change did nothing. Split one unchanged page down the middle and one half still comes out ahead. Significance is the test that separates a real effect from that everyday variation.

Significance is reported as a confidence level. At 95% confidence — the common standard — a difference this large would turn up by chance in fewer than 5 of every 100 tests where the change did nothing. It is not the probability that your variant is better. It is a statement about how easily chance alone could have produced what you are looking at.

Three things set the duration:

FactorEffect on duration
Your trafficMore visitors per day, sooner the test resolves
Your conversion rateA low rate needs more visitors to produce the same number of orders, and orders are what the test counts
The size of the effectA large improvement shows up quickly. A small one needs a large sample to separate from noise

A store with heavy traffic can read a result in days. A quieter store needs weeks, and for a small effect it may never reach significance at all. HI Engine states the expected duration when the test starts, and it does not declare a winner before the test reaches significance.

Significant is not the same as worth doing

A large sample can make a very small difference significant. That proves the effect is real; it does not prove it is worth the maintenance. Read the size of the lift, not only whether it cleared the bar.

Rules that keep tests trustworthy

  • Do not stop a test early because it looks good. Early leads reverse constantly, and stopping the moment a test looks like a winner is how a coin flip becomes a finding. Wait for significance.
  • Do not change the variant mid-test. That restarts the measurement.
  • Do not run a test through a large campaign. A traffic mix that changes shape mid-test contaminates both arms.
  • Run one test per surface at a time. Two tests on the same page cannot be read apart.

These apply whichever tool runs the split.

What a test feeds back

Every completed test calibrates the ranking for your store, no matter which platform ran it. A category that wins on your store gets more confidence in future scoring; a category that does not gets less. See How impact is estimated.

On this page