Ad Testing Glossary: 62 Terms That Are Claims vs Measurements

Ad Testing Glossary: 62 Terms That Are Claims vs Measurements

Table of Contents
    Add a header to begin generating the table of contents

    This glossary defines the terms used in ad testing by market researchers, advertising teams, brand managers, insights professionals, and creative agencies.

    It is organised around a distinction most glossaries leave out, and which determines whether a term can be trusted in a given study: some of these are claims, and some are measurements.

    Confusing the two is the most common analysis error in ad testing. The sections below are split accordingly.

    New to the subject? Start with Ad Testing: What It Is and the 6 Types.

    What is Ad Testing?

    Ad testing is the process of evaluating advertisements before, during, or after launch to understand how audiences respond to creative content.

    The purpose of ad testing is to identify which advertisement is most likely to achieve business objectives such as:

    • Increasing brand awareness
    • Driving consideration
    • Improving purchase intent
    • Generating conversions
    • Maximizing advertising ROI
    • Strengthening brand perception

    Ad Testing Fundamentals

    Pre-testing

    Evaluation of an advertisement before launch, while changes to the creative are still cheap. The stage with the highest return on research spend. It is the only point at which a finding can still alter the ad.

    In-market Testing

    Evaluation of an advertisement during a live campaign is to catch an underperforming asset before spending the remaining budget. All platform-native testing is in-market testing by definition.

    Post-testing

    Evaluation is done after a campaign is to measure what shifted and inform the next brief. The only stage at which brand lift can be established.

    Ad Concept Testing

    Testing conducted before production begins. Participants evaluate storyboards, scripts, static concepts, animatics, or early creative ideas. The objective is to identify winning concepts before spending on production.

    Copy Testing

    Testing focused on the language in an ad: headlines, body copy, taglines, and calls to action. Copy is the cheapest element of an ad to change. Copy is the one where small differences in wording produce massive differences in response.

    Creative Testing

    Testing focused on the visual and executional layer: images, video, layout, colour, casting, music, and pacing. Usually run on near-final or final assets.

    Animatic

    A rough animated version of a storyboard used as a test stimulus before full production. Fidelity matters here: respondents are more willing to criticise something that looks unfinished. So rough stimuli produce more honest concept feedback than polished ones.

    Ad Variation

    An alternative version of an advertisement containing modifications to elements such as the headline, visuals, music, call to action, product positioning, or brand cues.

    A/B Testing

    A method comparing two versions of an advertisement to determine which performs better. For example, Ad A uses an emotional message and Ad, B uses a product-focused message. Note that A/B testing is a method inside ad testing, not a substitute for it: it learns after spend has begun.

    Multivariate Testing

    An advanced form of testing where multiple creative elements are evaluated together, across combinations of headlines, visuals, offers, messaging, and branding.
    Identifies the highest-performing creative mix, at the cost of requiring a substantially larger sample.

    Incrementality Testing

    Isolating the true effect of a campaign by comparing exposed and unexposed groups. This will establish how much additional value the advertising created rather than how much activity coincided with it.

    Screener

    The set of qualifying questions used to determine whether a respondent enters a study. A good screener disqualifies competitors’ employees, market research professionals. In most categories anyone working in advertising evaluates creative as professionals rather than as consumers.

    2. Test Design

    This section is where most ad tests are decided, before a single response arrives.

    Monadic Testing

    A monadic design splits the sample into matched cells. Each cell sees one execution, and the test compares results across groups. This gives the cleanest read because no respondent sees another version. It requires the largest sample because each cell must independently reach a robust base. You must also match cells on any attribute that could independently drive the response.

    Sequential Monadic Testing

    In this design, each respondent sees one execution and answers the full question set. They then move to the next execution. This approach uses the sample more efficiently and therefore costs less. It has two costs. Results flatten when the executions are similar, and later executions receive less attention than earlier ones. Randomise the order, or the first execution can win on novelty alone.

    Comparative Testing

    This design shows executions side by side and forces a choice. It is fast and decisive, which makes it useful late in the process. However, it measures preference rather than performance. Those are not the same thing. That makes comparative testing a weak instrument for anything other than a final two-way call.

    Cell

    A cell is a matched subgroup of the sample in a monadic design. Each cell sees one execution. Cell size, not total sample size, determines whether a difference between executions is real.

    Control Group

    A control group sees either the original advertisement or no advertisement at all. It provides the baseline against which you measure advertising impact.

    Test Group

    The test group sees the advertisement being evaluated.

    Randomisation

    Randomisation assigns participants to advertisements at random. This ensures that differences in results reflect the creative rather than differences between the groups.

    Clutter Reel

    A clutter reel contains a sequence of unrelated advertisements with the test asset placed somewhere in the middle. This lets you evaluate the ad under conditions that resemble a real ad break or feed.

    Testing an ad in isolation is the design error most likely to produce a confident wrong answer. Nobody watches one ad. People watch five or six in a row, or scroll past yours between two others. Isolation measures performance under conditions that never occur in the market. It also flatters every ad equally.

    An ad that performs well alone but disappears in clutter is a common and expensive finding. You only see that problem if you design the test to expose it.

    Multiple Comparisons Problem

    The more variants you test, the higher the probability that one will appear to win by chance alone. Testing twelve variants across a sample of 600 leaves only 50 responses per variant. It also makes it almost certain that something will look like a winner even when it is not. The practical limit is two to four variants.

    Central Location Testing (CLT)

    Central Location Testing brings respondents to a physical facility to view and evaluate advertising. It remains valuable for creative with visual complexity. However, remote and AI-moderated alternatives increasingly replace it because they offer better cost and reach.

     

    3. Ad Testing Metrics: Claims

    These are things respondents tell you. Asking is the correct instrument. No measurement technology replaces them. Any tool that produces one of these numbers without asking anyone is modelling it rather than measuring it.

    Purchase Intent

    Purchase intent measures how likely consumers say they are to purchase after viewing an advertisement. Scores systematically overstate actual purchase behaviour. Use the metric comparatively, such as Version A against B or pre-exposure against post-exposure, rather than treating it as an absolute predictor of sales.

    Relevance

    Relevance measures how personally meaningful consumers find an advertisement and whether they feel it is meant for someone like them. It reflects a judgement about self-concept, which means you can only measure it by asking.

    The follow-up probe carries more value than the score. It reveals the specific casting, scenario, or language that creates or breaks the connection.

    Brand Lift

    Brand lift measures the increase in awareness, consideration, favourability, or purchase intent generated by advertising exposure. You need a pre-and-post design or an exposed versus unexposed control. That means brand lift spans the campaign. You cannot pre-test it.

    Brand Awareness

    Brand awareness measures whether consumers recognise or remember a brand. It often serves as a primary success metric. Always interpret it against starting awareness. A five-point lift from a base of 20% is not comparable to five points from a base of 80%.

    Brand Favourability

    Brand favourability measures positive sentiment toward a brand following exposure to advertising.

    Consideration

    Consideration measures whether consumers would consider the brand in a future purchase decision.

    Message Takeout

    Message takeout captures the key message consumers remember after watching an advertisement. An open-ended question captures it better than a rating scale. A respondent can rate clarity highly while taking away a completely different message.

    Persuasion

    Persuasion measures an advertisement’s ability to change attitudes or influence behaviour.

    Differentiation

    Differentiation measures how effectively an advertisement distinguishes a brand from its competitors. To measure this against competitor creative, put both ads in front of the same audience under the same conditions.

    Competitor ad volume in a public ad library indicates spend, not effectiveness.

    Top-two-box

    Top-two-box combines the percentage of respondents who select the two most positive options on a scale. It provides the standard way to report intent and agreement metrics.

    State this explicitly because a “30% intent score” means something different depending on whether it represents top-one-box or top-two-box.

    4. Ad Testing Metrics: Measurements

    These measurements come from what you record while a viewer watches. Asking viewers about them produces a whole-ad average filtered through social politeness. That hides where the problem occurs. This is why several of these terms do not appear in older glossaries.

    Attention Grab

    Attention grab measures whether an advertisement earns attention in its opening seconds. It differs from attention hold. Most ads that fail in the market fail here, before any other metric gets a chance to apply.

    You cannot ask viewers to report this reliably. A respondent recruited into a test watches attentively because that is the job they accepted. The test conditions therefore destroy the behaviour you want to measure.

    In the market, a viewer whose attention the ad never captured has no memory of it to report.

    Attention Hold

    Attention hold measures whether attention survives the middle of the advertisement. You measure it as a curve across the runtime rather than as a single score.

    You also cannot ask viewers to report it reliably. Nobody can report their own engagement at second-level granularity. As a result, the question goes unasked and the exact point where the ad loses attention stays invisible.

    Attention grab and attention hold can fail independently. They also require different fixes. Strong grab with weak hold points to a story problem in the middle. Weak grab means nobody evaluates anything that happens later in the ad.

    Attention Curve

    An attention curve plots audience attention across an advertisement, second by second. Ideally, it combines predicted neural responses with real gaze data and brand moments flagged on the same timeline.

    The shape matters more than the average. A curve that dips at second twelve and then recovers indicates a pacing problem in one scene. A curve that declines steadily indicates a premise problem. No amount of re-editing will fix that.

    Attention

    Attention describes an advertisement’s ability to capture and maintain audience focus. It often serves as the strongest single predictor of effectiveness. Split it into grab and hold rather than reporting it as one number.

    Engagement

    Engagement describes sustained attention to an advertisement. In behavioural measurement, it refers to a per-second state.

    This differs from the platform meaning of engagement, which includes clicks, shares, and comments. Those measures capture interaction with the ad unit rather than absorption in the content.

    Disengagement

    Disengagement marks the measured point at which a viewer’s attention begins to leave. It is more actionable than a completion metric because it carries a timestamp. You can therefore identify a specific edit at a specific second.

    Emotional Response

    Emotional response describes the feelings an advertisement triggers. Facial expression analysis can measure these reactions while the viewer watches and score them second by second.

    Asking viewers about their feelings creates three compounding problems.

    First, respondents filter their answers socially because they want to be agreeable. Second, that filter varies by culture. A score that means “strong” in one market can mean “lukewarm but polite” in another. You can even see this gap between a metro and a small town in the same state, not just between countries.

    Third, respondents report one summary feeling for a stimulus that may have produced many different reactions. An ad might make someone laugh at second four and lose them at second nineteen. The respondent may still describe it simply as “quite enjoyable.”

    Friction

    Friction indicates a measured emotional state in which the message is not landing. Read it across an ad’s narrative structure.

    If friction rises during the story build and falls at the payoff, the narrative works as designed. If friction spikes at the payoff, the ad explains itself and then undoes that explanation.

    Delight

    Delight represents a measured state of visible enjoyment. It draws from micro-expressions such as amusement, joy, and satisfaction.

    Its diagnostic value comes from its position. A delight peak matters most when you compare it with where the brand appears.

    Surprise

    Surprise measures a reaction to an unexpected moment in an advertisement. It helps establish whether a creative twist registered at all.

    Ad Recall

    Ad recall measures whether respondents remember seeing an advertisement after exposure.

    Unaided recall asks respondents to name ads without prompting. Aided recall shows them the brand and asks whether they remember it.

    Aided scores are always higher because recognition is easier than retrieval. That means aided recall can flatter weak creative.

    You can measure recall and ask about it, and the strongest read combines both. The survey tells you whether the ad registered. Predicted memory-associated response, read against the seconds when the brand actually appears on screen, tells you whether viewers remembered the brand or simply remembered the joke in the opening shot.

    Memorability

    Memorability provides a diagnostic score for whether an advertisement is likely to be remembered. More importantly, it shows whether memorable moments coincide with brand presence.

    A high memorability score with peaks that miss every brand moment means you created a well-remembered ad for nobody in particular.

    Message Comprehension

    Message comprehension measures whether viewers understood what the advertisement communicated. Read it in two ways.

    Use open-ended verbatims to understand the substance. Use a friction signal across the narrative arc to identify the location of the problem.

    An ad can lose the thread at second nine and recover by second twenty. The verbatim may then show mild confusion without telling you where it happened.

    Brand Attribution

    Brand attribution measures whether viewers correctly identify the brand behind the advertisement. Even highly engaging ads fail when attribution is weak.

    Brand Fit

    Brand fit measures whether an advertisement’s tone, imagery, and values align with how consumers perceive the brand. It is also called brand congruence or brand linkage.

    You can ask about brand fit, and researchers usually ask about it in the abstract after the ad finishes. That misses the more consequential version of the question. See Brand Connection.

    Brand Connection

    Brand connection measures whether the brand landed where the attention actually was. It combines automatic detection of brand moments with the viewer’s emotional and attention state at each of those timestamps.

    Survey-measured brand fit can return 74% agreement on an ad where the logo appears in a disengagement trough. Those findings do not conflict. The ad can fit the brand while nobody watches when it communicates the brand.

    Brand connection catches this problem and points to a specific fix: move the brand to where the attention already is.

    Likability

    Likability provides a diagnostic score for whether viewers enjoy an advertisement. Include it, but keep it in proportion. Likability correlates weakly with effectiveness, and a well-liked ad that sells nothing remains a familiar outcome.

    5. Video and Platform Terms

    Hook

    The hook is the opening stretch of an advertisement. In diagnostic use, it acts as a structural unit that you score separately from the story build and brand payoff because attention grab and attention hold can fail independently.

    Many advertisements win or lose in the first five seconds. This matters particularly in skippable environments, where the opening determines whether viewers watch the remaining runtime at all.

    Scoring the hook separately tells you whether a weak result comes from a bad opening or a bad middle. These are different problems and require different fixes.

    Video Completion Rate (VCR)

    Video Completion Rate measures the percentage of viewers who watch an advertisement to the end.

    Average View Duration

    Average View Duration measures the average time viewers spend watching an advertisement.

    Skip Rate

    Skip Rate measures the percentage of viewers who skip a video advertisement before completion.

    Drop-Off Rate

    Drop-Off Rate shows where audiences stop watching across an advertisement.

    Platform-reported drop-off gives you an aggregate view. It tells you that viewers left around a certain point but not what they experienced when they left.

    When you measure drop-off against a per-second engagement and emotion timeline, the same event becomes diagnostic. You can distinguish a dip that recovers, which indicates a pacing problem in one scene, from a steady decline, which indicates a premise problem.

    Creative Fatigue

    Creative fatigue describes a decline in advertising performance caused by repeated exposure to the same creative.

    Effective Frequency

    Effective frequency measures the number of exposures needed before an advertisement reliably registers.

    A long-standing rule of thumb puts this at roughly three exposures. Treat that figure as directional rather than precise.

    This is one reason media placement typically costs five to ten times more than producing the ad. It is also why pre-testing pays for itself.

    6. Biology-Based Measurement

    Neuromarketing

    Neuromarketing applies neuroscience to understand consumer response below the level of conscious report.

    Historically, this meant facility-based work using EEG caps, fMRI scanners, and per-study costs that put it out of reach for most brands.

    Prediction changed the economics. Models trained on fMRI data can now forecast neural response to a stimulus without scanning anyone. This moves neuromarketing from a specialist commission into a routine step in creative development.

    Predicted Neural Response

    Predicted neural response forecasts how a human brain would respond to each second of an advertisement. A model trained on fMRI data produces the forecast rather than a scanner measuring it.

    The distinction matters, and anyone using the measure should state it plainly. Predicted is not measured.

    For creative decisions, the trade is reasonable, particularly when you validate the prediction against real gaze and facial data from actual viewers. It is not academic neuroscience, and you should not describe it as such.

    Facial Coding

    Facial coding analyses facial expressions to identify emotional reactions.

    Originally, trained coders reviewed footage frame by frame. Today, automated systems can use a viewer’s own webcam to read emotion families continuously. They score each emotion from 0 to 100 and timestamp the peaks.

    The output does not tell you a feeling in isolation. It tells you where the reaction occurs. For example, an ad may produce delight eleven seconds after most viewers have stopped watching.

    Eye Tracking

    Eye tracking measures where viewers look during exposure. It reveals whether they notice branding, products, and key messages.

    One constraint matters before you budget for it. Camera-based eye tracking needs a screen with a camera pointed at a seated viewer. That means it works on laptops and connected TVs.

    It does not work on mobile, where viewers consume a large share of advertising. Facial emotion analysis still works on mobile. Gaze does not. Question any vendor that claims otherwise.

    Biometrics

    Biometrics measure physiological responses to advertising. Historically, these measures included heart rate, skin conductance, eye movement, and facial reaction.

    Current practical use is narrower than the textbook list. Camera-based capture of gaze and facial expression requires nothing worn or attached and works remotely at scale.

    Sensor-based measures such as skin conductance require a facility and a fitted device. That is why they remain rare outside academic work.

    Brand Moment

    A brand moment is a point in an advertisement where the brand appears, either visibly on screen or through spoken mention.

    Per-second diagnostics can detect these moments automatically. They anchor the analysis of brand connection.

    The important question is not how often the brand appears. It is whether the brand appears when anyone is watching.

    Area of Interest (AOI)

    An Area of Interest is a defined region of a frame, such as a logo, pack shot, or headline. Researchers use it to measure gaze data and establish how long viewers actually look at that element.

    7. Research and Statistics Terms

    Sample Size

    Sample size is the number of respondents in a study. Larger samples improve confidence, but matched respondents matter more than volume.

    One hundred people who fit the target audience are more useful than 1,000 people from a generic panel.

    Representative Sample

    A representative sample accurately reflects the target audience for the campaign. That audience frequently includes people who have no current relationship with the brand.

    A sample made entirely of existing customers inflates every positive score.

    Statistical Significance

    Statistical significance indicates whether a difference between advertisements is likely to reflect real performance rather than random variation. Most studies target 90% to 95% confidence.

    Margin of Error

    Margin of error gives the range within which the true value of a measured opinion is likely to sit.

    Weighting

    Weighting applies a statistical adjustment so that results reflect the characteristics of the target population rather than the composition of the sample that happened to respond.

    Norms Database

    A norms database stores historical scores and lets you benchmark a new result against category or market averages.

    Handle cross-market comparisons carefully. Self-report carries a culturally variable social filter, so a norm built in one market can misclassify a result in another.

    Build market-specific norms or lean on measured signals, which do not carry the same filter.

    Audience Segmentation

    Audience segmentation divides audiences into groups based on demographics, behaviours, interests, or attitudes.

    Thematic Saturation

    Thematic saturation occurs when additional interviews stop surfacing new themes in qualitative research. Researchers typically reach it after 8 to 12 in-depth interviews, or two to three focus groups, per audience segment.

    8. Bias and Limitation Terms

    Terms in this section describe the ways ad testing goes wrong. They belong in a glossary because most bad research is not badly executed. Researchers often execute it correctly on an instrument that could not answer the question.

    Social Desirability Bias

    Social desirability bias occurs when respondents give answers they believe are acceptable or helpful rather than accurate.

    In ad testing, it usually presents as warmth. An ad receives a positive response because respondents are being polite, and the researcher reads that politeness as promise.

    Demand Effect

    The demand effect occurs when respondents behave as they believe the study wants them to behave.

    In ad testing, this explains why you cannot reliably ask about attention grab. A respondent recruited to watch an ad watches it attentively. That is precisely the behaviour a real viewer may not exhibit.

    Recency Bias

    Recency bias occurs when respondents give disproportionate weight to the most recent part of a stimulus.

    It creates a significant problem in sequential monadic designs and in any post-exposure question about a long video. The ending can colour the respondent’s verdict on the entire advertisement.

    Order Effect

    Order effect describes the influence of the sequence in which you show executions.

    Randomisation manages this effect. Without randomisation, the first execution shown can win on novelty rather than merit.

    Leading Question

    A leading question signals the answer the researcher wants.

    “This ad makes you want to buy, right?” and “how great was this ad?” produce data that feels like validation but predicts nothing.

    Neutral framing fixes the problem. Ask, “after seeing this ad, how likely are you to find out more?”

    Over-optimisation

    Over-optimisation occurs when you refine an advertisement against test scores until it performs well in testing but poorly in the market.

    This usually happens because you tune the ad to conditions that do not exist in the wild, such as full attention, isolated viewing, and an engaged respondent.

    Merren’s Products

    MIRA (Multi Input Response Agent for Ads)

    MIRA is Merren’s biology-based ad diagnostics product. It surveys nobody. It measures what viewers’ brains and faces do while they watch, second by second, and answers the four questions that cannot be asked.

    Three signals fuse into a single timeline:

    • Predicted neural response. An open-source brain-mapping model built on roughly 1,000 hours of fMRI data, producing on the order of 20,000 data points per second, forecasts how a human brain would respond across every second of the ad. No respondents are required for this layer, and the read arrives in about twenty minutes.
    • Gaze. Real target viewers watch the ad inside a clutter reel through their own webcams, with gaze captured frame by frame, so every reaction is matched to the moment that caused it.
    • Facial emotion. Seven families, each scored 0 to 100 with timestamped peaks: delight, engagement, disengagement, friction, aesthetic appreciation, negative reaction, and surprise.

    Brand moments are detected automatically, both on-screen appearances and spoken mentions, so the viewer’s exact state at the instant the brand lands can be read directly.

    What MIRA outputs. An engagement curve with peaks, troughs, and brand moments flagged on one timeline, plus diagnostics scored 0 to 100 for attention grab, attention hold, memorability, message comprehension, likability, and brand connection. Each score arrives with the evidence behind it and a specific action, for example: brand connection 27, weak, brand visible in only 7% of frames and not on peak moments, action is to move the brand to where the attention already is.

    What MIRA does not do. It does not measure purchase intent, relevance, or brand lift. Those are claims, and claims require asking. It cannot read assets under about five seconds reliably, because a per-second model needs a signal to work with. Its gaze layer needs a laptop or connected TV. It does not benchmark creative against a competitor’s unless both run through the same test under the same conditions.

    Maya AI

    Maya AI is Merren’s AI-moderated interview product, covering the claims half of the list above: relevance, credibility, comprehension in the respondent’s own words, and intent.

    Maya runs interviews from a researcher-approved discussion guide, executed consistently across every respondent, with adaptive probing when someone expresses confusion or strong emotion. Respondents participate via WhatsApp in their own environment, which produces more natural responses than a facility. Analysis is automatic: themes clustered across the sample, verbatims surfaced by theme, and a report structured around the research objective, available within hours of the last interview.

    MIRA measures biology. Maya uncovers the meaning.

    Explore MIRA | Explore Maya AI | Request a pilot

    Related reading

    Table of Contents
      Add a header to begin generating the table of contents

      SHARE THIS ARTICLE

      SHARE THIS ARTICLE