Insights guide

How to compare pre-launch emotion baselines across multiple product categories

2026-10-06 · 8 min read · Rob / EmotionTrac

A practical workflow for insights professionals testing video content with second-by-second facial coding across different product categories.

You're testing video content for a new product launch. Your team needs to know if the emotional response matches expectations, but you're also testing ads for 3 different product categories at once.

The problem: emotional baselines vary by category. A 2-second joy spike in a snack ad means something different than the same spike in a financial services spot. If you compare raw emotion scores without accounting for category context, you'll misread the data.

Here's a practical workflow for setting and comparing pre-launch emotion baselines when you're working across multiple product categories.

Step 1: Define your category groups before testing

Start by grouping your products into clear categories. Don't overthink this. Use the categories your audience already understands.

Examples:

If you're testing 6 ads and 3 fall into FMCG, 2 into financial services, and 1 into entertainment, you've got 3 category groups. Tag each video with its category before you start collecting data.

Step 2: Collect second-by-second emotion data from opt-in panelists

Run your video tests with a panel that matches your target audience. With EmotionTrac, panelists opt in and their webcam captures facial expressions while they watch. The system analyzes micro-expressions frame by frame using facial action coding (FACS).

You'll get emotion scores for each second of video. Common outputs include joy, surprise, confusion, and negative affect. The data comes in as time-series curves, one per panelist per video.

Make sure your sample size is adequate for each category. If you're testing 3 FMCG ads with 30 panelists each but only 1 financial services ad with 10 panelists, your financial baseline won't be reliable. Aim for at least 25-30 panelists per video to start building stable baselines.

Step 3: Calculate category-level emotion averages

Aggregate your data by category. For each emotion metric, calculate the mean score across all videos in that category.

Example: if your 3 FMCG ads show average joy scores of 0.42, 0.38, and 0.45 (on a 0-1 scale), your FMCG category baseline for joy is 0.42. Do this for every emotion you're tracking.

Don't skip the variance check. Calculate standard deviation within each category. If your FMCG joy scores range from 0.38 to 0.45, that's a tight cluster. If they range from 0.20 to 0.65, you've got high variance and your baseline is less useful.

High variance usually means one of two things: your category definition is too broad, or one video is an outlier. Review the outliers and decide if they belong in the category or need their own group.

Step 4: Normalize scores within each category

Once you have category baselines, normalize your individual video scores against them. This lets you compare performance across categories without getting distorted by baseline differences.

A simple approach: calculate z-scores. For each video, subtract the category mean and divide by the category standard deviation. A z-score of +1.0 means the video performed 1 standard deviation above the category average. A z-score of -0.5 means it's half a standard deviation below.

Now you can compare a snack ad's joy response to a banking ad's joy response on the same scale. The raw scores might be 0.50 vs 0.30, but if the normalized scores are both +0.8, they're both performing well relative to their category norms.

Step 5: Track peak moments and duration separately

Category baselines help with overall comparisons, but they don't tell you everything. You also need to track emotional peaks and how long they last.

For each video, identify the top 3 emotion peaks. Note the timestamp, the emotion type, and the intensity. A 3-second joy spike at 0.75 intensity is different from a 0.5-second spike at 0.60 intensity, even if the average scores are similar.

Duration matters too. Some categories (like entertainment) often show short, intense emotion bursts. Others (like healthcare) might show longer, steadier engagement. Don't flatten these patterns into a single average or you'll lose critical context.

Step 6: Build a reference library over time

After you've tested 10-15 videos per category, you'll have enough data to build a reference library. This becomes your benchmark for future tests.

Store your baselines in a simple spreadsheet or database. Include:

Update your baselines quarterly or after every 5-10 new tests. Emotional norms change. What worked in Q1 might not match Q4 audience responses.

Step 7: Flag cross-category anomalies

Sometimes a video performs unusually well or poorly compared to both its category baseline and all other categories. These are your anomalies, and they're worth investigating.

Example: a financial services ad shows a joy peak that's +2.5 standard deviations above the financial category baseline and also exceeds the FMCG baseline. That's rare. Dig into what's happening in that moment. Is it humor? A surprising visual? A relatable scenario?

Anomalies can reveal creative approaches that break category conventions. They're not always good (sometimes they're confusing or off-brand), but they're always informative.

Common mistakes to avoid

Don't compare raw scores across categories without normalization. A 0.40 joy score in one category isn't the same as 0.40 in another.

Don't ignore sample size. A baseline built on 3 videos and 50 panelists isn't as stable as one built on 10 videos and 300 panelists.

Don't assume baselines are static. Audience expectations and emotional norms shift. Update your reference data regularly.

Don't forget to check for technical issues. If one video shows unusually flat emotion curves across all panelists, it might be a lighting problem, a technical glitch, or a video that's genuinely boring. Review the raw footage and panelist feedback before you trust the data.

What this workflow gets you

You'll know if your new product video performs above or below category norms before launch. You'll spot emotional patterns that work across categories and patterns that only work in specific contexts. And you'll build a reference library that makes every future test faster and more accurate.

This isn't about chasing perfect emotion scores. It's about understanding what's normal for your category, what's exceptional, and what's worth changing before you spend media budget on a video that doesn't connect.

Try it: Schedule an EmotionTrac demo and see second-by-second emotion tracking in action. Or visit Insights for more information.

Sources and further reading