What you'll learn in this article
- What the headline uplift number actually means and why it is an average, not a promise
- What the feature really does under the hood before you ever measure it
- The exact clean-test design I use to isolate the feature from every other variable
- How I read incremental conversions and CPA together instead of conversions alone
- How to tell real account-level lift from traffic quietly reallocated inside the account
- Where I have seen genuine uplift survive, and where it collapsed once I netted it out
Every time a rep or a deck waves the ai max for search performance uplift google ads number at me, I have the same reaction: fine, but on which account, measured how, and net of what? An average uplift across thousands of advertisers tells me almost nothing about the specific account in front of me. What I care about is whether the feature produced conversions that would not otherwise have happened, or simply moved existing conversions from one campaign to another and relabelled them as a win. Those two outcomes look identical on a single campaign card and completely different at the account level, and the whole job of a clean test is to force that distinction into the open.
I have been running this feature on live accounts long enough to have both outcomes in my notes: accounts where the lift was unambiguously real and survived every sanity check, and accounts where the same feature produced a beautiful-looking number that dissolved the instant I looked at the account total. The difference was never the feature itself. It was the shape of the account I switched it on in, and whether I had the discipline to measure it against a proper control instead of reading the campaign card and celebrating. This article is how I actually do it: what the feature does, the test design, the numbers I read, and the reasoning that decides whether I call the uplift real or fake.
The uplift claim, and what it really says
Google's published figure is specific and worth quoting precisely: advertisers who activate the feature in their search campaigns will typically see roughly 14% more conversions or conversion value at a similar CPA or ROAS, per Google's own official documentation on the AI Max feature, drawn from internal 2025 data for non-Retail advertisers. Read the sentence carefully, because the caveats are the whole story. It is an average, so half the advertisers in it did better and half did worse. It is conditioned on a similar CPA or ROAS, which means the comparison is only fair if efficiency held constant. And it is a Google-measured aggregate, not an independently audited incrementality study on your traffic.
None of that makes the number dishonest. It makes it a hypothesis. The claim I am willing to act on is not "the feature gives 14%" but "it might give my account somewhere around that, and the only way to know is to measure it against a proper control." Half the trouble people have with this metric is that they treat a portfolio average as a per-account guarantee, then feel cheated when their own account lands somewhere else on the distribution. If you want the feature itself explained before you test it, I keep a fuller overview of what AI Max is that lays out the machinery this uplift is supposed to come from. The test is what turns the marketing average into a number I can defend to a client without flinching.
What the feature actually does before you measure it
You cannot measure an uplift honestly if you do not know where it is supposed to come from, so it is worth being precise about the mechanics. Inside Google Ads this is an ai powered matching and creative layer that sits inside a Search campaign, not a separate campaign type, and it changes three things at once. First, it loosens matching so the campaign becomes eligible on search queries your existing keyword structure never captured, which is the part that either finds genuine search opportunities or quietly starts eating traffic you already owned. Second, it can rewrite and generate headlines and descriptions to fit the query and the user, rather than serving only the fixed assets you wrote. Third, it leans on final url expansion to route people to the page most likely to convert them.
That routing behaviour is the one most people misread, so I spell it out. When expansion is on, the system can override your set landing page and instead pick from relevant pages across your site, choosing based on your landing page content and the intent behind the query. The stated goal is to send users to whatever destination the model thinks will convert best, which in practice means it tries to route users to the most relevant URL rather than the one you hard-coded. That is powerful when your site has deep, well-structured content and dangerous when it does not, because a model that can pick the destination can also pick a worse one. Knowing this up front tells me exactly which levers could be producing any lift I later measure, and which ones I need to watch for damage.
The reason I care about the plumbing is attribution. If I see a lift, I want to be able to say whether it came from broader matching finding new demand, from better-fitting creative lifting CTR and conversion rate, or from expansion sending people to a stronger page. Those are different stories with different implications for rollout, and a test that just reports "conversions went up" collapses all three into a number I cannot reason about. So before I touch the experiment controls, I write down which of these mechanisms I expect to matter on this specific account.
The clean test I actually run
A clean test has exactly one moving part. I set up a campaign experiment that splits traffic between two arms: a control with the feature off and a treatment where I enable ai max and nothing else. Everything else is held identical across both, same budget logic, same bidding strategy, same target CPA or ROAS, same creative, same negatives. If any of those differ, I am no longer measuring the feature; I am measuring a bundle of changes and guessing which one moved the needle. The single-variable discipline is the entire point, and it is where most in-house tests quietly fall apart because someone "also refreshed the ads while they were in there."
Bidding stays automated on both arms, and that is non-negotiable for a reason that comes straight from how the feature works. The search-term matching that drives most of the ai max google ads performance uplift conversions does not function under manual CPC, because the system relies on the signals that Maximize Conversions or Maximize Conversion Value feed it. Test it on manual bids and you measure a deliberately hobbled version, then wrongly conclude it does nothing. This is the same reason the broader set of controls I keep switched on for the AI Max layer assumes conversion-based bidding as a baseline rather than an option.
There is one preparation step I never skip, and it is the one that saves branded-heavy accounts from a false positive. Before the experiment starts, I add negative keyword coverage for the brand terms I do not want the loosened matching to absorb, because otherwise the treatment arm can look like a hero simply by swallowing branded demand that would have converted anyway. Getting that hygiene right up front is what lets me later trust that any query the feature newly reached was genuinely incremental rather than borrowed. Then I let it run: past the learning period, and until each arm has enough conversions that the gap between them is signal, not the random noise of a short window.
Reading the numbers without fooling myself
When the test matures I never look at conversions in isolation, because conversions alone are the easiest metric to be seduced by. I read conversions and CPA together, at the account level, over the same window. The logic is simple: if the treatment arm produced more conversions while CPA stayed at roughly the same efficiency, that is a real uplift, exactly the "more conversions at a similar CPA" shape the headline promises. If conversions rose but CPA rose with them, the campaign did not get more efficient; it just spent more to buy more, which any campaign can do without a fancy feature.
My single most useful diagnostic here is the search terms report. It tells me exactly which new queries the loosened matching started serving on, and that is where I separate a story I like from a story I can trust. If the incremental conversions are coming from queries that are genuinely new to the account, non-brand, and aligned to intent, I believe the lift. If they are coming from terms that are brand, near-brand, or obvious variants of what another campaign was already winning, I know I am looking at cannibalization dressed as growth. I also skim which destinations expansion chose, because a cluster of high performing conversions all landing on one aggressively-swapped page usually means the model found a genuinely better route rather than manufacturing demand.
The second discipline is refusing to read the campaign in isolation from the rest of the account. A feature that expands match keywords can lift its own campaign's conversions by quietly absorbing queries that another campaign used to serve. So I always pull the account total, not the campaign card, and I check whether a gain in the treatment arm is mirrored by a loss somewhere else. That cross-check is the difference between a number I trust and a number I have merely admired. It is the same account-level reasoning I apply when I compare the feature head to head in my AI Max versus Performance Max breakdown, because both features can look like heroes locally while doing nothing globally.
Where I saw real lift and where it was reallocation
On a couple of lead-gen accounts with genuine unmet non-brand demand, the uplift was real. Turning the feature on found queries the tight keyword structure had never reached, the treatment arm added conversions, and the account total went up at a CPA that held. When I netted out every other campaign, the gain was still there. That is incremental: new conversions that did not exist before, not borrowed from a neighbour. The pages the model routed to were a good sign too, because expansion consistently pushed new visitors onto deep service pages that converted better than the generic landing page the old keyword ads had used. On those accounts I rolled the feature out and the client number improved for real.
On branded-heavy accounts the story flipped. The loosened matching started scooping up branded and near-branded queries that would have converted anyway through the brand campaign. The campaign card showed a satisfying jump, but the brand and exact-match campaigns lost almost exactly what the treatment arm gained, and the account total barely moved. That is reallocation wearing an uplift costume: the feature was not creating demand, it was re-routing demand I already owned and charging me for the privilege. The tell was always the same, a local win paired with an almost equal loss elsewhere, which is why the account-level read is the only one I trust and why the brand-term negatives I set up before the test matter so much once the feature goes live.
There is a subtler failure mode in the middle that is worth naming, because it fools people who do check the account total. On one mid-market retailer the account total genuinely rose, but the CPA rose with it, and when I traced the extra conversions they were low-value, low-intent queries the model had reached by casting wide. Volume up, quality down, efficiency flat-to-worse. That is not a win I would present as an uplift, and it is a reminder that "more conversions" and "more good conversions" are different claims. The number that survives all three checks, account-level, efficiency-held, quality-intact, is the only one I put in a deck.
My verdict on the uplift number
So is the uplift real? My honest answer after running these tests is: it depends entirely on account shape, and the average number cannot tell you which side of it you land on. Accounts with real unmet non-brand demand and a restrictive existing keyword structure are where the feature has the most room to find genuinely incremental queries, and that is where I have seen the headline shape hold up under a clean, account-level read. Accounts dominated by branded search are where the feature most often produces impressive-looking numbers that dissolve the moment you net out cannibalized brand traffic. The feature is the same in both cases; only the account differs, and that is the whole point.
It is also worth being clear-eyed that this is where the platform is heading regardless of any single test. Google's whole automation stack is built on the same google ai premise: hand the system signals and destinations and let it find conversions you would not have reached by hand. This feature is that philosophy applied inside Search, and the strategic question is not whether to resist it but whether, on your account, it produces incremental value or just reshuffles what you already had. A clean test is how you answer that for real rather than on faith.
That is why I never present the 14% to a client as an expectation. I present it as the reason to run a test, then I present my own measured number as the answer. The reasoning matters more than the result: a disciplined single-variable experiment, read on incremental conversions and CPA at the account level, cross-checked against the queries the model actually served, is what separates a feature that earned its rollout from one that just moved money around the account and hoped I would not check the total. Run the test properly and ai max for search stops being a marketing promise and becomes a decision you can actually defend.