Performance
Rank creative before it costs you anything
Testing is how you learn, but testing everything is how you burn budget. Most accounts spend a meaningful share of monthly media discovering that a third of their creative was never going to work.
Xeli scores candidates before launch across the dimensions that actually correlate with early performance: how fast the hook lands, whether the value proposition survives a muted autoplay, contrast and legibility at feed size, and how different the concept is from what you already have live.
Why pre-launch judgement fails
- 01Creative reviews happen at full resolution on a large monitor, not at thumbnail size on a phone with sound off.
- 02Teams unconsciously pick the concept that looks nicest rather than the one that stops a scroll.
- 03Every new variant competes with your own existing winners for delivery, so weak entrants quietly cost twice.
- 04Post-hoc reporting tells you what failed after the money is gone.
How it works
- 1Simulate the surface
Each asset is rendered at real feed dimensions, muted, and cropped as the placement would crop it.
- 2Score the dimensions
Hook clarity, message legibility, brand fit, placement safety and novelty against your live set each get a sub-score.
- 3Rank and cut
The set is ordered, with reasons attached. Low scorers are regenerated rather than shipped.
- 4Launch a slate
The top slice goes live as a structured test with enough budget per variant to reach a readable result.
- 5Feed results back
Live CTR, hook rate, thumb-stop and ROAS update the scoring weights for your account specifically.
Reports in revenue. Not impressions.
| SKU | ROAS | CTR | Score | Status |
|---|---|---|---|---|
| Glow Serum | 5.2× | 3.4% | 92 | Scaling |
| Sleep Balm | 4.6× | 2.9% | 88 | Scaling |
| Reset Mask | 2.1× | 1.2% | 54 | Fatigued |
| Body Oil | 4.1× | 2.6% | 81 | Live |
Your whole catalog. On-brand ads on the other side.
Watch raw catalog photos glide through Xeli. Every SKU comes out with a different brand-native ad layout, background, typography, offer and CTA. Drag to explore, hover to pause.








































- ✕ No branding
- ✕ No messaging
- ✕ No differentiation
- ✓ On-brand DPA creatives
- ✓ Conversion-focused messaging
- ✓ Targeted ICP testing
What you get
Video is scored the way most of it is watched: no sound, first two seconds, phone-sized.
A concept too close to a creative already live is flagged, because near-duplicates cannibalise rather than expand.
Every score carries the reason, so the team learns the pattern instead of trusting a black box number.
Weights shift toward what has historically worked in your vertical and your account, not a generic benchmark.
Where it stops
The honest limits. If any of these are dealbreakers, better to know now.
- —A score is a prior, not a verdict. The auction remains the only real test, and occasional low scorers win.
- —Prediction quality improves with account history. A brand-new account starts on vertical priors and calibrates over the first few weeks.
- —Scores measure creative, not offer. A weak discount on a weak product will not be rescued by a strong hook.
Questions people ask
How accurate is it?
Treat it as triage rather than prophecy. The reliable gain is removing clear losers before spend, which raises the average quality of what you test rather than guaranteeing any single winner.
Does it replace A/B testing?
No. It decides what enters the test. You still need enough spend per variant to read a result, which our creative testing framework covers in detail.
What signals feed the model?
Rendered-asset features plus your own historical performance: CTR, hook rate, thumb-stop rate, frequency decay and ROAS by concept.
Want this running on your catalog?
We onboard a small number of brands at a time so setup gets real attention.