Most performance marketing teams know their attribution models are imperfect but do not have a reliable alternative for understanding which channels are actually driving incremental revenue versus which ones are claiming credit for purchases that would have happened anyway. Incrementality testing is the methodological answer to that question. It is also genuinely achievable for DTC brands operating without a data science team, if the test is designed with appropriate constraints from the start.
What Incrementality Testing Actually Measures
The core concept is a controlled experiment. You identify a group of potential customers, split them into an exposed group (who see your ads) and a holdout group (who do not, or who see a placeholder ad that suppresses your normal creative), then measure the difference in conversion rates between the two groups. The measured difference is your incremental lift: the purchases that happened specifically because of the ad exposure, not because of organic intent or other factors.
This is different from what attribution models measure. Attribution models observe conversion paths and distribute credit based on rules. Incrementality testing observes outcomes and measures causation. The gap between what attribution models say about a channel's contribution and what incrementality tests show is often large, and the gap is not random. It reliably shows that conversion-stage channels like branded search and retargeting are overcredited by attribution models, while upper-funnel channels like prospecting campaigns are undercredited.
The Basic Design: Single-Channel Holdout
The simplest incrementality test is a holdout on a single channel. You take a representative sample of your audience (say, 20 percent) and suppress your ads from reaching them on that channel for the test period, typically 2 to 4 weeks. You then compare the conversion rate of the holdout group to the conversion rate of the exposed group, controlling for differences in the groups.
The mechanics for this vary by channel. On Meta, you can set up a holdout within a campaign using the campaign-level holdout feature in Ads Manager. On Google, you use Campaign Experiments to run a control cell with modified campaign settings. TikTok's split testing feature can serve a similar function, though the holdout methodology is less sophisticated than Meta's implementation.
The key design constraint is that the holdout group must be a random sample of your actual target audience, not a convenient split like "users who live in a different city" or "users on a different device type." Geographic and demographic splits introduce confounders that contaminate the result. A truly random user-level holdout, maintained consistently throughout the test period, is what produces interpretable lift estimates.
The Power Calculation Problem
Before running any incrementality test, you need to estimate whether your account has enough volume to produce a statistically meaningful result in a practical timeframe. This is where many DTC brands hit a wall: they do not have enough conversion events per channel to reach statistical significance within 2 to 4 weeks.
The rough minimum: you need enough conversions in both the control and exposed groups that the measured conversion rate difference, if it is real, would be detectable above statistical noise. For a channel where the baseline conversion rate is 2 percent and you expect an incremental lift of 15 percent above baseline (meaning the true conversion rate for the exposed group is 2.3 percent), you need around 3,000 to 4,000 conversions per cell to reach 80 percent statistical power at a standard 95 percent confidence interval. That is 6,000 to 8,000 total conversions for the test, spread across 2 to 4 weeks.
For brands spending $5,000 to $15,000 per week on a given channel, this threshold is often out of reach. The practical alternatives are to test over longer periods (accepting more external variation), to accept lower statistical confidence (80 percent CI instead of 95 percent), or to test at higher holdout fractions (30 to 40 percent holdout instead of 20 percent, which reduces the spend during the test period but increases cell sizes). None of these is ideal. They are the honest tradeoffs of working with limited volume.
Multi-Channel Holdout Tests: What Gets More Complex
Running incrementality tests across multiple channels simultaneously creates a genuine design challenge: interaction effects. If a customer is held out of your Meta campaigns but still sees your Google and TikTok ads, the measured "incrementality" of Meta includes the cannibalized revenue that your other channels absorbed. Conversely, if you hold someone out of all channels simultaneously, you lose the ability to attribute the measured lift to any specific channel.
The standard approach for multi-channel incrementality testing is to run single-channel holdouts sequentially, not simultaneously. Test Meta for four weeks, then test Google for four weeks, then test TikTok. This avoids interaction effects but takes significantly longer and means you are measuring each channel's incrementality in a different context (different seasons, different creative, different competitive conditions).
A more sophisticated option, available if you have the volume, is a full factorial test design where you create user segments that are held out of different combinations of channels and compare all combinations. This lets you measure channel interaction effects directly. It requires roughly 4x the volume of a single-channel holdout and is only feasible for accounts with substantial conversion volumes per channel.
We are not saying the sequential approach is perfect. Its limitation is real: it cannot capture channel interactions. But for most DTC brands, it is the practical starting point.
Interpreting the Results: What the Numbers Mean for Allocation
A well-designed holdout test gives you an estimated incremental ROAS for each channel you test: how much incremental revenue did each dollar of spend on that channel actually generate, net of what would have happened organically. This number is almost always different from your platform-reported ROAS, and almost always different from what your attribution model tells you.
The typical finding for DTC brands running their first incrementality test is something like this: Google branded search shows platform-reported ROAS of 10 to 15x, but incremental ROAS of 1.5 to 3x (it is mostly capturing organic intent). Meta prospecting shows platform-reported ROAS of 2 to 3x, but incremental ROAS of 3 to 5x (it is actually generating demand that would not otherwise have existed). TikTok shows high variance, with incremental ROAS anywhere from below 1x to above 5x depending on creative quality and audience saturation at the time of the test.
The implication for allocation is clear: you were probably underinvesting in channels with high incremental ROAS and overinvesting in channels with low incremental ROAS. The branded search overspend is the most common finding. Teams discover they have been scaling branded search budgets based on flattering last-click ROAS numbers, while the incremental test shows that most of those conversions would have happened via organic search if the paid campaign had not been running.
Incrementality and Forward Forecasting
Incrementality testing gives you the causal truth about channel contribution over the test period. But test results are backward-looking by definition. They tell you what a channel's incremental value was during the test. They do not tell you what it will be next week or next month, which is what matters for allocation decisions.
The way we think about this in the context of Flyweel's forecasting is that incrementality test results are high-quality calibration inputs for the channel-level ROAS models. A channel that shows consistently high incremental ROAS across multiple test periods has a fundamentally different expected contribution than a channel where incremental lift tests close to zero. That historical pattern feeds into the forecast signal.
The practical cadence that works for most DTC brands: run incrementality tests once per quarter per channel, treat the results as ground truth for calibrating your understanding of each channel's actual contribution, and use that calibrated understanding as the baseline for forward allocation decisions. The two methods are complementary, not competing.