Academy

/

Geo-Lift and Holdout Testing for Influencer Marketing: A Practical How-To

Guide

Tracking & ROI

Brands

Geo-Lift and Holdout Testing for Influencer Marketing: A Practical How-To

A geo-lift test measures a campaign's real effect by comparing sales in a group of geographic areas that see the campaign (the test group) to a matched holdout of comparable areas that don't (the control group). The difference in sales between the two groups — adjusted for the natural baseline movement — is the campaign's incremental lift. The method is typically used for paid ads, where exposure can be geographically confined; on organic influencer content it's harder to run cleanly, because a post can't be kept out of a specific area.

What a geo-lift test is, and why it exists

A geo-lift test is a practical way to run an incrementality test: instead of randomly splitting individual users into an exposed group and a holdout group, you split geographic areas — postal-code clusters, regions or whole markets. Some areas see the campaign (the test group); others deliberately don't (the control or holdout group). The difference in sales between the two groups — adjusted for whatever change both groups would have seen anyway — is the campaign's real, incremental lift. Influencer marketing attribution explained introduces incrementality testing and holdout groups as a concept; this article is the practical how-to for designing, running and reading a geographic version of that test.

User-level holdout vs geographic holdout

Both variants answer the same question, but need very different conditions to work:

 User-level holdoutGeographic holdout (geo-lift)
How the groups are formedRandom selection of individuals within a shared audienceWhole geographic areas (postal codes, regions, countries) are assigned to test or control
What it needs technicallyThe ability to exclude specific individuals from exposure (usually only possible in paid media)The ability to confine exposure to a geography — again, usually only possible with paid geo-targeting
Does it work on organic content?Barely — you can't stop a specific person from seeing a postOnly partly — see the organic-content section below
Typically used forSingle-channel, paid campaigns with user-level data accessCampaigns spanning multiple channels, or where user-level tracking is weakened or impossible
StrengthCan be a more precise, less noisy measurement where the data existsNeeds no cookies or device IDs — robust against tracking limitations
WeaknessSensitive to cookie lifetime and cross-device gapsNeeds enough geographic units to find statistically comparable pairs

How to build a geo-lift test in practice: 5 steps

Step 1 — Pick a metric and a baseline period. Decide which number the test should measure (orders, revenue, app installs), and collect at least 4-8 weeks of historical data for each candidate area before you design anything else. Without a solid baseline you can't tell a real campaign effect from ordinary weekly noise.

Step 2 — Select and pair geographic areas. The single most important design choice is that the test and control areas resemble each other historically — same seasonal pattern, same price sensitivity, same category behaviour. Meta's own open-source tool for this, GeoLift, uses what it calls a synthetic control: instead of matching the test area to one single control area, the tool builds a weighted blend of several control areas, constructed to mimic the test area's actual sales pattern in the period before the campaign. The better the synthetic control tracks the test area's historical curve, the more trustworthy the lift measured afterward.

Step 3 — Run a power analysis. Before running anything, you should calculate how large a lift the test can even reliably detect, given the number of areas and the test length — the minimum detectable effect (MDE). According to GeoLift's own documentation, its worked examples test combinations of 2-4 test markets against the rest as control, with a holdout ranging from 50-100% of control-market conversions withheld from the campaign, and a test length that at minimum covers one full purchase cycle for the category — the tool's own walkthrough trials both 10- and 15-day windows for a fast-moving retail category. Those specific numbers are illustrations from Meta's own documentation, not universal rules — your category's purchase cycle and data volume decide what makes sense for you.

Step 4 — Run the test. On paid media, you can use the platforms' own tools: Google Ads' Conversion Lift based on geography divides a country into algorithmically generated regions (Google calls them "Google Marketing Areas"), assigns them to test or control, and requires the campaign to target a single country to be eligible. On organic influencer content, that kind of geo-targeting isn't technically possible — see the dedicated section below.

Step 5 — Read the result. Compare the test area's actual performance during the campaign against the synthetic or matched control's expected performance over the same period. The difference is the incremental lift. Per Google Ads' own definition, the resulting iROAS (incremental ROAS) is then calculated as incremental revenue divided by total ad spend — the same logic as in how to calculate influencer marketing ROI, just with a lift-based revenue figure instead of a last-click one.

Why geo-lift is harder on organic influencer content than on paid ads

Paid ads can be geo-targeted technically — the platform simply doesn't serve the ad to users in the control area. An organic influencer post can't do the same: a creator can't stop a specific follower from seeing a post, and content spreads via reposts, screenshots and DMs across whatever geographic line you try to draw. That produces contamination — the control group ends up partly exposed anyway, which pulls the measured lift downward and understates the campaign's real effect.

Two practical workarounds are typically used instead of a clean geo-fence on organic content:

  • Staggered rollout: the campaign launches in the test area earlier than in the control area, so you can compare the two areas' trajectories during the overlap window, before the control area is exposed at all.
  • Convert it to a geo-targeted paid boost: if you whitelist the creator's content and run it as a paid ad via Partnership Ads or Spark Ads, exposure becomes technically geo-targetable again, because it's now a paid ad with the platform's own geo-targeting — the cost is that you're measuring the whitelisted ad, not the pure organic effect.

The Denmark problem: why a clean national geo-split rarely works

Google Ads' own documentation for Conversion Lift based on geography requires a campaign to target a single country, and needs that country to contain enough statistically distinct regions to run a reliable test — the tool itself shows a feasibility rating (High/Medium/Low), and low population or low geographic variation pulls that rating down. Denmark, with a total population under six million and relatively homogeneous digital behaviour across regions, typically runs into exactly that constraint: there are too few, too similar regions to find statistically reliable test/control pairs inside the country alone.

Two realistic alternatives for a Denmark-only brand:

  • A time-based on/off holdout instead of a geographic split: run the campaign in selected weeks and pause it in others, then compare sales in the "on" and "off" periods — adjusted for seasonality, day-of-week and any other campaigns running at the same time. It's methodologically weaker than a geo split (you can't isolate the effect as cleanly, because time rather than geography is the varying factor), but it's technically achievable for any Denmark-only brand.
  • Use a comparable market as control, if the brand also sells in at least one other Nordic market with similar consumer behaviour — for example, holding Denmark as the test market and another market as control, or the reverse. That requires the two markets to actually have comparable baseline sales patterns, which should be checked against historical data before the test is designed, not assumed.

Worked example: a difference-in-differences calculation

The numbers below are hypothetical and illustrate the calculation method only. This is not a real Make Influence customer case, and none of the figures are benchmarks for what a campaign typically delivers.

Assume a brand identifies two comparable, hypothetical areas — Area A and Area B — selling at almost the same level over a 4-week baseline period before the campaign: Area A averages 82 orders per week per 10,000 residents, Area B 79 — a gap under 4%, which makes them a reasonably matched pair.

Area A becomes the test region and is exposed to an influencer campaign for 4 weeks. Area B stays as control. Sales rise in both areas during the campaign window — partly from the campaign, partly from ordinary seasonal movement that both areas experience regardless:

  • Area A rises to 101 orders/week per 10,000 residents — an increase of 101 − 82 = 19.
  • Area B rises to 84 orders/week per 10,000 residents — an increase of 84 − 79 = 5, attributable to the seasonal effect both areas share.
  • The difference-in-differences lift = 19 − 5 = 14 orders per 10,000 residents per week attributable to the campaign specifically — once the shared seasonal effect is subtracted out.

Assume Area A has 400,000 residents, i.e. 40 units of 10,000. Incremental sales work out to 14 × 40 = 560 orders per week, and over the 4-week campaign: 560 × 4 = 2,240 incremental orders in total. At an average order value of DKK 350, that's incremental revenue of 2,240 × 350 = DKK 784,000. If the campaign cost DKK 180,000 in creator fees and any paid boost, iROAS = 784,000 ÷ 180,000 ≈ 4.4× — DKK 4.4 of incremental revenue for every krone spent, measured via lift rather than last-click-tracked codes or links.

Which approach fits your situation

 Platform-run lift study (Google/Meta)Open-source tool (e.g. GeoLift)DIY difference-in-differences
Who it's forBrands with access to a Google or Meta account representative and a budget that qualifies for the studyBrands with internal data resources able to set the tool up themselvesAny brand with historical sales data per area, regardless of technical maturity
What it needsAccess through the platform; the campaign must target a single countryR skills or an analyst who can set it up; enough historical sales data per areaA spreadsheet, two comparable areas, and the discipline to keep the campaign out of the control area
StrengthHandles market selection, power analysis and statistics automaticallyFree, transparent methodology, used by large advertisersSimple to understand and explain to a decision-maker
WeaknessNot available to every account; typically needs a certain budget levelRequires statistical competence to set up and interpret correctlyLess statistically robust — sensitive to how well the two areas actually match

Common mistakes

  • Picking two areas that aren't actually comparable. A different seasonal profile, different concurrent campaigns, or different category maturity ruins the test even when the population sizes look similar.
  • Ignoring contamination on organic content. If the control area gets exposed anyway via sharing or reposting, the test understates the real effect — without you necessarily noticing.
  • Running the test too short to cover one purchase cycle. A test that ends before the typical customer has time to act systematically undercounts lift.
  • Overlooking external events during the test window. A sale, a competitor campaign, or a weather event in one area can mimic or mask a campaign effect.
  • Attempting a pure geographic split inside Denmark without knowing Google's own feasibility constraint. See the Denmark section above — it's rarely the right starting method for a Denmark-only brand.

Decision framework: when a geo-lift test makes sense

  • IF you run paid advertising with geo-targeting, and have budget for a multi-week test → a platform-run lift study (Google or Meta) or GeoLift is the most robust route.
  • IF you sell only in Denmark and don't have access to a lift study → a time-based on/off holdout is the more realistic alternative, not an internal geo split.
  • IF the campaign is organic influencer content with no paid boost → expect contamination, and consider a staggered rollout or whitelisting the content instead, if you need to measure it geographically cleanly.
  • IF you don't have enough volume to detect a reliable lift (too few orders, too small a campaign) → use last-click and your operational numbers instead, and treat them as a floor on the real effect, as described in influencer marketing attribution explained.

Make Influence's operational perspective

In Make Influence's experience, most Danish DTC brands have neither the volume nor the geographic spread needed to get a platform-run geo-lift test into "High" feasibility. Our own tracking model is built at order level — tracking links and discount codes with a 30-day cookie window, see how influencer tracking actually works — which answers a different question than a geo-lift test: which orders can be traced to which creator, not how many orders the campaign as a whole created above baseline. Our recommendation for brands without the resources for a full geo-lift programme is rarely to build one from scratch — it's to supplement day-to-day last-click reporting with a periodic, simpler on/off test, and to use discount codes or tracking links consistently enough that the daily tracking is at least comparable over time.

FAQ

What's the difference between a geo-lift test and a standard holdout test?

A standard holdout test randomly splits individuals into an exposed group and a holdout group. A geo-lift test splits whole geographic areas the same way. Geo-lift is typically used when user-level tracking is weak or impossible, while a user-level split needs access to individual user data.

How long should a geo-lift test run?

At minimum, the test should cover one full purchase cycle for your category — shorter tests systematically undercount lift. Meta's own GeoLift tool uses examples of 10-15 days in its documentation for a fast-moving retail category; a category with a longer consideration period needs a correspondingly longer test.

Can I run a geo-lift test on organic influencer content, or only on paid ads?

It's technically possible, but markedly harder, because an organic post can't be confined to a specific geographic area. See the section above on staggered rollouts and whitelisting as practical workarounds.

Do I need Google's or Meta's own tool, or can I work it out myself?

Both are possible. The platforms' own tools handle market selection and statistics automatically, but typically need access through an account representative. A simple difference-in-differences calculation, like the worked example above, can be done in a spreadsheet with no platform access at all — it's just less statistically robust.

Is a geo-lift test realistic for a Denmark-only brand?

Rarely in its pure form, because Denmark is typically too small and too homogeneous to meet the platforms' own requirements for geographic variation. A time-based on/off holdout is the more realistic alternative for most Denmark-only brands.

How many incremental sales do I need to see before the result is reliable?

It depends on your baseline volume and how large a lift you're trying to detect — that's exactly what a power analysis calculates before the test starts. As a rule of thumb: the smaller the campaign's expected effect relative to the natural noise in sales, the more orders and the longer a test you need for a reliable result.

Make Influence

Want influencer marketing to be easier?

Find creators with real audience data, run collaborations in one place, and see clicks and sales per creator while the campaign is live.

Book a demoCreate account

Make Influence

Get paid for the audience you built

Apply to campaigns from brands that are actively looking, follow your own clicks and sales, and get paid without chasing invoices.

Create creator profileMore creator guides

Make Influence

One place for the whole collaboration

Briefs, agreed terms, tracking links and results sit together — so brands and creators see the same numbers.

See how it worksBrowse the Academy