What you'll learn in this article
- Why expanding a working campaign fails for reasons that have nothing to do with the keywords you chose
- The order I pull expansion candidates in, and why the account's own data always comes before any tool
- The three questions a candidate has to survive before it earns a place in a live campaign
- How I size a batch against conversion volume so a bad addition can't damage the campaign that funds it
- How to tell a keyword that failed from a keyword that was judged too early, and what each one costs you
The hardest keyword work I do isn't on new campaigns. It's on campaigns that already perform. A launch has nothing to lose, so a bad list costs you a fortnight and some budget. A campaign returning reliably for eight months has everything to lose, and the question of how to find keywords for Google Ads campaign expansion becomes a question about risk management rather than about research.
I've broken campaigns doing this. Not by picking obviously bad terms, but by adding twelve reasonable ones on a Tuesday and watching cost per acquisition drift upward for three weeks while I tried to work out which of the twelve was responsible. The answer, most of the time, was that no single one was. The batch was the problem.
What follows is the process I settled on after enough of those episodes: how I source candidates, how I qualify them, how I contain them, and how I read what happens next.
Why expansion breaks a working campaign, and it isn't the keywords
The instinctive explanation for a post-expansion dip is that the new terms were bad. Sometimes they are. More often the mechanism is indirect, and understanding it changes what you do far more than any list of good keywords would.
The first mechanism is budget displacement. A campaign at or near its daily cap is already spending everything it has on the traffic it knows. New keywords don't get given extra money, they compete for the existing money. Even genuinely good additions cannibalise impression share from proven terms during their unproven phase, so the campaign gets worse before it can get better, and it gets worse in a place you weren't looking.
The second is the bidding model. Smart Bidding is predicting conversion probability from patterns it has learned on your current traffic mix. Change that mix abruptly and its predictions degrade, not because the new terms are unprofitable but because it has no basis for pricing them yet. The reporting shows this as instability that looks exactly like a strategy problem, and I've watched people respond by changing the strategy, which compounds it.
The third is diagnostic, and it's the one that actually costs the most. Add twelve keywords simultaneously and you've created a single experiment with twelve variables. If performance drops you know something is wrong and nothing about what. Every subsequent decision is guesswork, and the usual resolution is to pause everything new, which discards the good additions along with the bad ones and leaves you exactly where you started minus the spend.
Once I understood that the third mechanism does more long-term damage than the first two combined, the whole approach reorganised itself around one constraint: never add more at once than you can attribute afterwards.
Where expansion candidates come from, in strict order
For a live campaign, keyword research in Google Ads is mostly a data problem rather than a discovery problem. The account has been buying information for months. Most people expand by opening a research tool, which means ignoring evidence they already paid for in favour of estimates.
First: the search terms report
This is the only source where a candidate arrives with performance history attached. A query in that report already triggered your ads, already cost money, and already did or didn't convert. Google's documentation is direct about this dual role, describing how to use search terms data to refine your keyword list, which is the mechanism I lean on hardest. What I'm hunting for is a converting query that isn't matched by any exact keyword I own. That term is already working while being priced as an approximation of something else, and promoting it to its own keyword is the lowest-risk addition available.
The pattern I look for second is a cluster of related queries sharing a modifier my list doesn't contain. Ten variations of the same intent with a phrasing I never bid on is a theme the account has discovered on my behalf. The cluster is the candidate, not any individual query in it.
Second: the shape of your own converting terms
Before any tool, I look at what converts and ask why. If the winners share a structure, that structure is a hypothesis I can test in vocabulary the account hasn't seen. This is inference rather than data, but it's inference grounded in the account's own outcomes, which puts it well ahead of any external estimate. Match type matters here more than people expect, because a phrase variant of a proven exact term can capture an adjacent demand pool at almost no risk, and choosing between those options is a decision I break down in my guide to how keyword match types work.
Third, and only then: research tools
Planner and third-party tools earn their place when the account has genuinely exhausted its own data, when you're entering new geography or a new product line, or when you need volume estimates to size an opportunity. Their output is estimated, unvalidated by your account, and priced by a model rather than by your auction. That doesn't make it useless, it makes it third. The full logic of what those tools do and don't tell you is something I cover in my walkthrough of keyword research for Google Ads from scratch.
The three questions a candidate has to survive
Having candidates isn't having decisions. Each one goes through the same three questions, in order, and failing any of them ends it.
Is this new demand or a redistribution of demand I'm already buying? This is the question nearly everybody skips and it's the one that determines whether expansion is growth or theatre. If a new term overlaps heavily with an existing keyword's matching, adding it doesn't reach anyone new. It splits the same traffic across two keywords, halves the data behind each, slows learning on both, and produces a report that looks like activity. I check by searching the term in the existing keyword's matched queries. Heavy overlap means the answer is a match type adjustment, not a new keyword.
Does the intent match an ad and a landing page I already have? A keyword doesn't perform alone. If the closest ad group's ads aren't what I'd write for this term, then adding it means testing a keyword and a mismatched ad simultaneously, and when it underperforms I won't know which one failed. Either it goes somewhere its ads fit or it gets its own group with its own ads. There's no third option that isn't just deferring the problem.
Can I afford to be wrong about it? Estimated cost per click against my target cost per acquisition and a conservative conversion rate gives the honest downside. If a term needs a conversion rate materially above the campaign average to break even, I'm not expanding, I'm gambling on outperformance from traffic I've never bought. Sometimes that's a bet worth taking. It should be taken knowingly, and it should be small.
Roughly a third of my candidates survive all three, and the ones that fail on the first question outnumber the ones that fail on the third.
Batching: sizing additions so a mistake can't hurt
The batch, not the keyword, is the real unit of risk. I size it against conversion volume rather than picking a fixed number, because the same five keywords are trivial on a campaign with two hundred monthly conversions and catastrophic on one with eight.
The constraint I use: a batch shouldn't be able to consume more than about a tenth of campaign budget in its first two weeks. On low-volume campaigns that means three or four terms. On high-volume ones it can mean twenty. Either way, the worst case is a tenth of the budget spending badly for a fortnight, which is an annoyance rather than a crisis.
Containment does the rest. New terms enter at a match type at least as narrow as their proven neighbours, usually phrase or exact, because broad on an unproven term means paying to discover things about intent I could have inferred for free. Every batch gets labelled with its date, so that in six weeks I can pull performance by label rather than reconstructing history from memory. And where the batch is large enough or its intent distinct enough, it goes into its own ad group, which isolates both the experiment and the ad copy testing it.
The negative keyword pass happens before launch, never after. Any obvious informational or unqualified variant that the new terms will match gets excluded on day one. Waiting to see what comes in means paying for the education, and the education is predictable. This is the same discipline I apply account-wide when building exclusion lists during research rather than after it, which I've written up in detail in my approach to finding negative keywords in Google Ads.
One thing I hold constant through all of it: I don't change bids, budgets or bid strategy in the same week I expand. If two things change and performance moves, the attribution is gone. That discipline has saved me more analysis time than any tool I've ever bought.
Reading the result without fooling yourself
The judgement window is the hardest part, because the two available errors are both expensive and they pull in opposite directions.
Killing early feels disciplined. It usually isn't. A keyword with forty clicks and no conversions on a campaign converting at three percent has produced no information whatsoever, and pausing it is a coin flip dressed as a decision. My threshold is clicks worth several times my target cost per acquisition, and never less than two weeks in calendar time regardless of how fast the clicks accumulate, because weekday and weekend intent differ enough to mislead over shorter windows.
Waiting too long is the other failure and it's less discussed because it doesn't feel like a mistake. A term burning multiples of target CPA with no conversions and nothing promising in its search terms doesn't need a longer window. That evidence has already arrived. The window exists to protect terms with ambiguous signals, not to postpone obvious verdicts.
What I actually check at the end of a window, in this order: did the new terms convert on their own; did the existing terms lose impression share or conversion volume; and what search terms did the new keywords bring in. The second is the one people forget, and it's where the real cost of a bad expansion shows up, because a batch can look acceptable in isolation while quietly starving the terms that were paying for everything.
The third often matters more than the first. A keyword that didn't convert but pulled in three excellent queries isn't a failure, it's a signpost. I've kept underperforming keywords for a cycle purely because their search terms were pointing somewhere worth going, and promoted those queries instead. The keyword was wrong; the direction was right.
The cadence, and why slower compounds faster
On a stable account I run one expansion batch every two to three weeks. It feels slow. Over a year it's fifteen to twenty batches, each one attributable, each one either kept or removed on evidence, and the account ends up with a keyword list where I can explain the presence of every term.
Compare that with quarterly expansion: four large batches a year, each unattributable, each producing a fortnight of uncertainty, and a list nobody can justify by the end of it. Fewer, bigger interventions don't save time. They convert time into ambiguity.
The pause between batches isn't idleness either. It's when the previous batch's search terms report matures, and that report is where the next batch usually comes from. Expansion done this way is a loop rather than a project: each addition generates the evidence for the next one, and the account teaches you what to buy next if you leave enough time between questions to hear the answer.
So, condensed: expansion breaks working campaigns through budget displacement, bidding instability and lost attribution, not through bad keyword choice. Source from the search terms report before any tool. Qualify each candidate against overlap, ad fit and affordable downside. Size batches against conversion volume, contain them with narrow match types and pre-emptive negatives, and change nothing else that week. Then read the result at two weeks minimum, looking at what the batch did to the incumbents as carefully as what it did for itself. Every campaign I've grown without breaking was grown that way, and every one I broke was broken by doing all of it at once.