Testing & Measurement

A/B Test

A controlled experiment comparing two variants to measure which performs better.

A/B testing is the discipline of comparing one variable across two otherwise identical cells. On ad platforms, true A/B tests use the platform's built-in split (Meta A/B, Google Experiments) so audience splits are clean.

Definition

An A/B test randomly splits traffic between two variants (A and B) and measures the difference in a target metric.

Why it matters

It's the cleanest way to isolate the effect of a single change — creative, headline, price, page — from all the noise around it.

A/B Test in practice

An A/B test randomly splits traffic between two variants (A and B) and measures the difference in a target metric. Measurement concepts exist because ad platforms report the results they can see, and that is not the same as the results you caused. The gap shows up whenever a channel reports growth that the P&L never receives. Knowing which question a method answers keeps you from over-reading a dashboard. Read it next to Multivariate Test, Statistical Significance, Control Group.

How to get it right

Decide the question, the metric and the stopping rule before launch. Change one variable per test, hold budget and audience constant, and let the test run through at least one full purchase cycle. Write the result down — including the null results, which are the ones teams repeat most often.

What to measure and watch

Check whether you have the sample size to detect the effect you care about before declaring a winner; small differences need far more data than most accounts generate in a week. Cross-check platform-reported results against a holdout or an aggregate model when the stakes are large. Why this matters commercially: It's the cleanest way to isolate the effect of a single change — creative, headline, price, page — from all the noise around it.

Where A/B Test sits in an agentic creative workflow

Testing at any useful rate needs supply. Xeli produces clean, single-variable variants across your catalog — same layout, one changed element — so tests are properly controlled and the production queue is never the bottleneck. In the context of testing & measurement, that means the concept stops being something a person re-applies by hand every campaign and becomes a rule the system enforces on every asset it produces.

Failure modes worth naming

The recurring problems are predictable: calling the test on day 1; testing two things at once; ignoring statistical significance. Each of these is a process gap rather than a knowledge gap — which is why the fix is usually a checklist, a template or an automated rule instead of more training.

Common mistakes

Frequently asked questions

An A/B test randomly splits traffic between two variants (A and B) and measures the difference in a target metric.

It's the cleanest way to isolate the effect of a single change — creative, headline, price, page — from all the noise around it.

Platform reporting is correlational and biased toward the channel that served the ad. Structured methods like Incrementality and Holdout Test isolate the lift you actually caused, which is the number that belongs in a budget decision.

Calling the test on day 1. Testing two things at once. Ignoring statistical significance.

Closely connected concepts include Multivariate Test, Statistical Significance, Control Group.