How to Find the Best-Performing Ad Through A/B Ad Testing

How to Find the Best-Performing Ad Through A/B Ad Testing

Table of Contents
    Add a header to begin generating the table of contents

    Let me start with the question I get asked the most by marketing teams: “How many versions of my ad should I actually test?”

    And the second question, which always follows: “How many people do I need to see each one before I can trust the result?”

    Too many variants can make respondents jaded. This blog contains the framework: how many variants to test, how many respondents you need per variant, and how to know when a result is real versus random noise. If you want the broader picture of the testing process first, our team’s guide on what ad testing is and how to do it covers the setup. This blog is about numbers.

    A Common Mistake of Ad Testing

    When teams get excited about testing, they go big. Five headlines. Three thumbnails. Two calls to action. Suddenly they’re running 8, 10, 15 variants in a single study.

    It feels rigorous. It’s actually the fastest way to fool yourself.

    Here’s why. Every variant you add splits your audience into a smaller slice. If you survey 600 people across 3 variants, each version gets a solid 200 responses. Split those same 600 people across 12 variants and each one gets just 50. Far too few to tell a real winner from a lucky one.

    The more variants you test, the higher the odds that one of them looks great purely by chance. Test enough versions and randomness alone will hand you a “winner.” Statisticians call this the multiple comparisons problem. The reason so many “data-backed” ads underperform the moment real money hits them.

    Test fewer variants, deeper. Two to four strong variants beats a dozen half-baked ones every single time.

    Ad Testing Variants: How Many Should You Test?

    The honest answer depends on your stage:

    Early concept stage — test 2 to 3 directions. You’re choosing between fundamentally different ideas, not polishing. Two or three distinct concepts is plenty. You want a clear directional signal.

    Copy or headline stage — test 3 to 4 variants. Once the concept is locked, you can afford a few more. The differences are smaller and you’re optimizing rather than choosing.

    Final creative refinement — test 2 variants, head to head. At this point a clean A/B is the cleanest, most trustworthy read you can get.

    Ad Testing on A Sample Size of Participants

    The number of respondents you need per variant depends on three things:

    1. The effect size you care about. Are you trying to detect a huge difference (one ad is clearly better) or a subtle one (a 3% lift)? Subtle differences need far bigger samples.
    2. Your confidence level. The industry standard is 95% confidence meaning if you ran the test 100 times, you’d get the same answer 95 of them.
    3. Your baseline response rate. How people are already reacting affects how many you need to detect a shift.

    As a practical starting point:

    • For directional reads (which concept feels stronger), 50 to 100 matched respondents per variant is often enough to act on.
    • For decisions you’re putting real budget behind, aim for 150 to 250+ per variant to reach statistical confidence on meaningful differences.
    • For small lifts you need to prove (think large-scale media spend), you may need 300 to 500+ per variant.

    The more people you need per variant, which is exactly why testing fewer variants matters. Your respondents are on a fixed budget.

    Before you launch anything, run your specific numbers through our sample size calculator. It takes thirty seconds and saves you from the single most common ad-testing mistake: declaring a winner that isn’t one.

    What “statistical significance” actually means for your ad

    Statistical significance just answers one question: “Is the difference between my variants real, or could it have happened by random chance?”

    Say Ad A scores 68% purchase intent and Ad B scores 64%. Is A genuinely better, or is that 4-point gap just noise from a small sample? Significance testing tells you. If the result is significant at 95% confidence, you can trust that A really is winning. If it’s not, that gap is meaningless and shipping A over B is a coin flip dressed up as a decision.

    Two practical takeaways (the goal is quick decisions):

    • A bigger gap needs fewer people to prove. If one ad is crushing the other, even a modest sample will show it clearly.
    • A small gap needs a big sample — or it doesn’t matter. If two ads are nearly tied even at proper sample size, here’s a liberating truth: it probably doesn’t matter which you pick. Ship either and move on.

    The hidden tax on this whole process: speed

    More respondents per variant, proper confidence levels points toward bigger, slower studies. You’ve got a launch date, and recruiting hundreds of matched respondents and manually reading their feedback can eat the very week you need for the test.

    This is the real reason teams skip testing or cut corners on sample size. Not because they don’t understand the calculation. Because the traditional way of getting enough quality responses is too slow to fit a real campaign calendar.

    There’s a deeper problem hiding underneath the sample-size question — one that no amount of “more respondents” fixes on its own.

    The problem bigger samples can’t solve

    Even with a perfectly sized sample, traditional ad testing has a flaw baked into it. You’re measuring what people claim, not what they actually felt.

    Ask someone what they thought of your ad and you get a socially filtered answer. People want to be agreeable. They tell you it was “good, very good”.  You walk away thinking your ad is a winner, when that warm response was just good manners. This gets worse across cultures and regions, where politeness norms differ so much that even your benchmarks stop being comparable.

    A survey can only tell you about the ad as a whole. It can’t tell you that viewers checked out at second 17, right before your brand appeared. That moment-by-moment truth the exact spot where your ad bleeds attention stays invisible, no matter how many people you survey.

    So you have two compounding challenges: getting enough of the right people, fast, and getting past what they merely claim to what they actually experienced. This is exactly where modern ad testing has leapt forward.

    Merren’s MIRA Changes The Variant-Testing Math

    MIRA by Merren stands for Multi Input Response Agent for Ads . It replaces “what people claim” with what their brains and faces actually do. Instead of relying on a survey verdict, it builds a second-by-second picture of how your ad performs across three fused layers:

    • AI-predicted brain response. MIRA uses a brain-mapping model trained on roughly a thousand hours of fMRI data to predict, second by second, how a human brain reacts to your ad. Based on likability, comprehension, memorability, attention, and brand connection before a single live respondent is needed for this layer.
    • Gaze and facial expression. Real participants watch via webcam, and MIRA tracks exactly where their eyes go and what emotion plays across their face, frame by frame, synced to the moment that caused it.
    • AI emotion analysis. Computer vision reads seven emotion families: delight, engagement, disengagement, friction, aesthetic admiration, negative reaction, and surprise. Each timestamped, so you see the precise emotional state of viewers at the exact instant your brand appears on screen.

    If you want the full breakdown of how a MIRA test runs read the details here: The Future of Ad Testing: MIRA by Merren.

    The Final Framework

    1. Test 2 to 4 variants, not more. Fewer variants, tested deeper, beats a sprawling leaderboard every time.
    2. Match your sample size to your stakes. 50–100 per variant for directional reads; 150–500+ when the real budget rides on it. Run the sample size calculator before you launch.
    3. Check for significance, but don’t worship it. If two ads tie at proper sample size, it doesn’t matter which you pick. Ship and move on.
    4. Go beyond claims. The biggest gains come from measuring what people actually felt, second by second not what they politely reported.

    The marketers who win aren’t the ones who test the most variants. They’re the ones who test the right number, against the right number of people, and measure what actually happened inside the ad.

    When you’re ready to test that way, fast enough to fit your campaign and deep enough to trust the result: see what Merren and MIRA can do. Your next campaign shouldn’t launch on a hypothesis. Test it, prove it, then spend.

    Table of Contents
      Add a header to begin generating the table of contents

      SHARE THIS ARTICLE

      SHARE THIS ARTICLE