Guide
Tracking & ROI
Brands
A geo-lift test measures a campaign's real effect by comparing sales in a group of geographic areas that see the campaign (the test group) to a matched holdout of comparable areas that don't (the control group). The difference in sales between the two groups — adjusted for the natural baseline movement — is the campaign's incremental lift. The method is typically used for paid ads, where exposure can be geographically confined; on organic influencer content it's harder to run cleanly, because a post can't be kept out of a specific area.
A geo-lift test is a practical way to run an incrementality test: instead of randomly splitting individual users into an exposed group and a holdout group, you split geographic areas — postal-code clusters, regions or whole markets. Some areas see the campaign (the test group); others deliberately don't (the control or holdout group). The difference in sales between the two groups — adjusted for whatever change both groups would have seen anyway — is the campaign's real, incremental lift. Influencer marketing attribution explained introduces incrementality testing and holdout groups as a concept; this article is the practical how-to for designing, running and reading a geographic version of that test.
Both variants answer the same question, but need very different conditions to work:
| User-level holdout | Geographic holdout (geo-lift) | |
|---|---|---|
| How the groups are formed | Random selection of individuals within a shared audience | Whole geographic areas (postal codes, regions, countries) are assigned to test or control |
| What it needs technically | The ability to exclude specific individuals from exposure (usually only possible in paid media) | The ability to confine exposure to a geography — again, usually only possible with paid geo-targeting |
| Does it work on organic content? | Barely — you can't stop a specific person from seeing a post | Only partly — see the organic-content section below |
| Typically used for | Single-channel, paid campaigns with user-level data access | Campaigns spanning multiple channels, or where user-level tracking is weakened or impossible |
| Strength | Can be a more precise, less noisy measurement where the data exists | Needs no cookies or device IDs — robust against tracking limitations |
| Weakness | Sensitive to cookie lifetime and cross-device gaps | Needs enough geographic units to find statistically comparable pairs |
Step 1 — Pick a metric and a baseline period. Decide which number the test should measure (orders, revenue, app installs), and collect at least 4-8 weeks of historical data for each candidate area before you design anything else. Without a solid baseline you can't tell a real campaign effect from ordinary weekly noise.
Step 2 — Select and pair geographic areas. The single most important design choice is that the test and control areas resemble each other historically — same seasonal pattern, same price sensitivity, same category behaviour. Meta's own open-source tool for this, GeoLift, uses what it calls a synthetic control: instead of matching the test area to one single control area, the tool builds a weighted blend of several control areas, constructed to mimic the test area's actual sales pattern in the period before the campaign. The better the synthetic control tracks the test area's historical curve, the more trustworthy the lift measured afterward.
Step 3 — Run a power analysis. Before running anything, you should calculate how large a lift the test can even reliably detect, given the number of areas and the test length — the minimum detectable effect (MDE). According to GeoLift's own documentation, its worked examples test combinations of 2-4 test markets against the rest as control, with a holdout ranging from 50-100% of control-market conversions withheld from the campaign, and a test length that at minimum covers one full purchase cycle for the category — the tool's own walkthrough trials both 10- and 15-day windows for a fast-moving retail category. Those specific numbers are illustrations from Meta's own documentation, not universal rules — your category's purchase cycle and data volume decide what makes sense for you.
Step 4 — Run the test. On paid media, you can use the platforms' own tools: Google Ads' Conversion Lift based on geography divides a country into algorithmically generated regions (Google calls them "Google Marketing Areas"), assigns them to test or control, and requires the campaign to target a single country to be eligible. On organic influencer content, that kind of geo-targeting isn't technically possible — see the dedicated section below.
Step 5 — Read the result. Compare the test area's actual performance during the campaign against the synthetic or matched control's expected performance over the same period. The difference is the incremental lift. Per Google Ads' own definition, the resulting iROAS (incremental ROAS) is then calculated as incremental revenue divided by total ad spend — the same logic as in how to calculate influencer marketing ROI, just with a lift-based revenue figure instead of a last-click one.
Paid ads can be geo-targeted technically — the platform simply doesn't serve the ad to users in the control area. An organic influencer post can't do the same: a creator can't stop a specific follower from seeing a post, and content spreads via reposts, screenshots and DMs across whatever geographic line you try to draw. That produces contamination — the control group ends up partly exposed anyway, which pulls the measured lift downward and understates the campaign's real effect.
Two practical workarounds are typically used instead of a clean geo-fence on organic content:
Google Ads' own documentation for Conversion Lift based on geography requires a campaign to target a single country, and needs that country to contain enough statistically distinct regions to run a reliable test — the tool itself shows a feasibility rating (High/Medium/Low), and low population or low geographic variation pulls that rating down. Denmark, with a total population under six million and relatively homogeneous digital behaviour across regions, typically runs into exactly that constraint: there are too few, too similar regions to find statistically reliable test/control pairs inside the country alone.
Two realistic alternatives for a Denmark-only brand:
The numbers below are hypothetical and illustrate the calculation method only. This is not a real Make Influence customer case, and none of the figures are benchmarks for what a campaign typically delivers.
Assume a brand identifies two comparable, hypothetical areas — Area A and Area B — selling at almost the same level over a 4-week baseline period before the campaign: Area A averages 82 orders per week per 10,000 residents, Area B 79 — a gap under 4%, which makes them a reasonably matched pair.
Area A becomes the test region and is exposed to an influencer campaign for 4 weeks. Area B stays as control. Sales rise in both areas during the campaign window — partly from the campaign, partly from ordinary seasonal movement that both areas experience regardless:
Assume Area A has 400,000 residents, i.e. 40 units of 10,000. Incremental sales work out to 14 × 40 = 560 orders per week, and over the 4-week campaign: 560 × 4 = 2,240 incremental orders in total. At an average order value of DKK 350, that's incremental revenue of 2,240 × 350 = DKK 784,000. If the campaign cost DKK 180,000 in creator fees and any paid boost, iROAS = 784,000 ÷ 180,000 ≈ 4.4× — DKK 4.4 of incremental revenue for every krone spent, measured via lift rather than last-click-tracked codes or links.
| Platform-run lift study (Google/Meta) | Open-source tool (e.g. GeoLift) | DIY difference-in-differences | |
|---|---|---|---|
| Who it's for | Brands with access to a Google or Meta account representative and a budget that qualifies for the study | Brands with internal data resources able to set the tool up themselves | Any brand with historical sales data per area, regardless of technical maturity |
| What it needs | Access through the platform; the campaign must target a single country | R skills or an analyst who can set it up; enough historical sales data per area | A spreadsheet, two comparable areas, and the discipline to keep the campaign out of the control area |
| Strength | Handles market selection, power analysis and statistics automatically | Free, transparent methodology, used by large advertisers | Simple to understand and explain to a decision-maker |
| Weakness | Not available to every account; typically needs a certain budget level | Requires statistical competence to set up and interpret correctly | Less statistically robust — sensitive to how well the two areas actually match |
In Make Influence's experience, most Danish DTC brands have neither the volume nor the geographic spread needed to get a platform-run geo-lift test into "High" feasibility. Our own tracking model is built at order level — tracking links and discount codes with a 30-day cookie window, see how influencer tracking actually works — which answers a different question than a geo-lift test: which orders can be traced to which creator, not how many orders the campaign as a whole created above baseline. Our recommendation for brands without the resources for a full geo-lift programme is rarely to build one from scratch — it's to supplement day-to-day last-click reporting with a periodic, simpler on/off test, and to use discount codes or tracking links consistently enough that the daily tracking is at least comparable over time.
A standard holdout test randomly splits individuals into an exposed group and a holdout group. A geo-lift test splits whole geographic areas the same way. Geo-lift is typically used when user-level tracking is weak or impossible, while a user-level split needs access to individual user data.
At minimum, the test should cover one full purchase cycle for your category — shorter tests systematically undercount lift. Meta's own GeoLift tool uses examples of 10-15 days in its documentation for a fast-moving retail category; a category with a longer consideration period needs a correspondingly longer test.
It's technically possible, but markedly harder, because an organic post can't be confined to a specific geographic area. See the section above on staggered rollouts and whitelisting as practical workarounds.
Both are possible. The platforms' own tools handle market selection and statistics automatically, but typically need access through an account representative. A simple difference-in-differences calculation, like the worked example above, can be done in a spreadsheet with no platform access at all — it's just less statistically robust.
Rarely in its pure form, because Denmark is typically too small and too homogeneous to meet the platforms' own requirements for geographic variation. A time-based on/off holdout is the more realistic alternative for most Denmark-only brands.
It depends on your baseline volume and how large a lift you're trying to detect — that's exactly what a power analysis calculates before the test starts. As a rule of thumb: the smaller the campaign's expected effect relative to the natural noise in sales, the more orders and the longer a test you need for a reliable result.
Make Influence
Find creators with real audience data, run collaborations in one place, and see clicks and sales per creator while the campaign is live.
Book a demoCreate accountMake Influence
Apply to campaigns from brands that are actively looking, follow your own clicks and sales, and get paid without chasing invoices.
Create creator profileMore creator guidesMake Influence
Briefs, agreed terms, tracking links and results sit together — so brands and creators see the same numbers.
See how it worksBrowse the Academy