What you'll learn in this article
- How far the planner estimate has drifted from the billed cost of google adwords keywords in my accounts
- Why a keyword has no price at all only a distribution of auctions
- The four mechanisms that make the estimate wrong in a predictable direction
- What I budget from instead, and how I build the forecast I actually show a client
- What a large estimate-to-reality gap is telling you about the account
Every client conversation about budget starts the same way: someone opens the planner, reads a number next to a keyword, multiplies it by an imagined click volume and arrives at a monthly figure. Then the campaign launches and the invoice doesn't match. Sometimes it's half. More often it's double.
I've stopped treating that as an error in the tool. The estimate answers a question nobody is asking what the average advertiser paid, historically, across a geography and a match behaviour that has nothing to do with your account. What you pay is the outcome of your specific auctions, with your relevance, your bid, your negatives, your schedule. Those are different objects that happen to be denominated in the same currency.
What follows is what I've measured across the Search accounts I manage, and what I do differently now that I've stopped expecting the two numbers to converge.
The gap I measure between estimate and invoice
My habit for the last few years has been to record the planner's suggested range at launch and compare it to the realised average after the first ninety days. The pattern that comes out of that log is more consistent than I expected.
In broad consumer categories with heavy competition, the realised figure has typically landed below the top of the suggested range and near or under its midpoint because the range aggregates advertisers with far worse relevance than a properly built account, and their clearing prices pull the average up.
In narrow, low-volume, high-intent commercial terms, the opposite happens more often than not. The realised cost of keywords in google adwords has run above the top of the range in a meaningful minority of the launches I've tracked, usually because the estimate was built on a thin sample of historical auctions that a couple of new entrants have since repriced.
The variance inside a keyword is bigger than the gap between keywords
The thing that surprised me most when I started segmenting properly is how little a single average describes. Take one keyword in one account for one month and split it by device, hour and geography, and the cheapest slice frequently costs a third of what the most expensive slice costs. The "price" you were quoted sits somewhere in the middle of that spread describing none of it.
This is why I no longer describe a keyword as expensive or cheap. I describe the profitable slice of it as affordable and the rest as something I'm choosing to buy or not buy. That reframing has changed more budgets than any bidding change I've made.
The estimate degrades faster than people expect
An estimate pulled in January and used to justify a June launch has been wrong in almost every account I've inherited. Auction depth in service and B2B verticals moves quarter to quarter, and seasonal categories can double and halve within a single planning cycle. I treat any estimate older than a month as a rough directional read and nothing more.
Why the tool number and the billed number diverge
Four mechanisms explain nearly every gap I've had to explain to a client, and none of them is a fault in the estimate.
1. You are never charged your bid
The most common misreading is treating the suggested bid as a price. It isn't it's the ceiling you authorise. Google's documentation on average cost-per-click is explicit that the reported average is simply total click cost divided by clicks, built on the actual amount charged per click, and that this figure will differ from the maximum you set. Every keyword-level cost you see is a quotient of outcomes, not a rate card.
That single distinction resolves most of the confusion. The estimate describes what a bid might need to be; the invoice describes what a series of auctions actually cleared at. They were never the same quantity.
2. The estimate assumes an average account, and yours isn't
Historical CPC ranges blend advertisers with excellent relevance and advertisers with none. Build a tight ad group with copy that matches the query and the clearing price sits at the bottom of that blend; import a loose structure and it sits at the top. In thin accounts especially I treat keyword quality as the main lever on realised cost, because relevance differences never average out at low volume.
3. Match behaviour changes what you actually bought
An estimate is attached to a keyword; your spend is attached to the query set that keyword bought. Loosen the match and you pull in cheaper adjacent traffic that drags the average down while quietly buying intent you didn't want. Tighten it and the average rises even though the account got better. Any comparison that ignores which keyword match types generated each number is comparing traffic mixes and calling it price.
This is also why I mirror negative lists before comparing anything. In most accounts the exclusions shape the realised cost of google adwords keywords more than the bids do.
4. Automation optimises for a target, not for the click price
Once a value or acquisition target is running, the system will pay well above any estimate for an auction it predicts converts, and well below it for one it doesn't. The resulting average is an artefact of the prediction, not a description of the market. Reading it as a market price and then adjusting the bid manually is how people fight their own keyword bidding strategy for a quarter without noticing.
In accounts where I've moved from manual to value-based bidding, the average click cost has usually risen while cost per acquisition fell. The estimate had nothing useful to say about either.
How I build a budget I'm willing to defend
I still open the planner. I just use it for a different job than most people do.
Use the estimate for ranking, not for arithmetic
The tool is genuinely good at telling me that keyword A costs several times keyword B. It is poor at telling me what either costs. So I use it to sort a keyword list into cost tiers and to spot the terms that will eat a small budget in a week and I never multiply it by anything.
Work backwards from the value of the outcome
The only number that constrains what I can pay is what a conversion is worth and how often clicks become one. Close rate and job value give me an affordable cost per click; the estimate tells me whether that number is anywhere near the market. If it isn't, the answer is a different keyword set or a better landing page, not a bigger bid.
Buy the data before you buy the plan
For any account where the estimate matters commercially, I run a deliberately small two-to-three week test on a tight exact-match subset with mirrored negatives. That produces a realised figure from my own auctions, which is worth more than any historical range. Every forecast I've built on live data has survived the client conversation; the ones built on planner numbers rarely did.
Forecast as a range with the assumptions attached
I present a low, expected and high case, and state what would move the outcome between them: competitive entry, seasonality, relevance improvement, match loosening. A single number invites the client to treat it as a commitment. A range with mechanisms invites them to reason about it.
Re-measure on a fixed subset, quarterly
I keep a small stable set of keywords and check the realised average against the original estimate every quarter, with click counts visible next to the averages. That log is how I've learned which verticals reprice fast, and it's the reason I catch new entrants early. I fold it into the same optimisation routine so it doesn't become a separate task nobody does.
What I infer from the size of the gap
Realised cost far below estimate usually means relevance, not luck. When an account clears well under the suggested range, it's normally because the structure is tight and the copy matches. That advantage is real and it erodes if the account is left alone.
Realised cost far above estimate means the sample was stale. Thin keywords produce estimates from few historical auctions. A single funded entrant reprices them, and the tool won't reflect it for months.
A cost that barely moves month to month means nobody is contesting the keyword. That's either an opportunity or a signal the query isn't worth contesting. Checking which has been more valuable than reacting to the number.
A rising average with falling acquisition cost is the system working. Automation buying more expensive, better-qualified auctions looks like inflation on a click report and like progress on a business report. I read the second one.
Wide intra-keyword variance means the segmentation is the opportunity. When device or hour slices differ by multiples, the win is in bid adjustments and scheduling, not in a better keyword.
Estimates that match reality precisely make me suspicious. In practice that's usually happened in accounts running loose match on generic terms, where the traffic mix is so close to the market average that the account has no edge at all.
What I stopped doing
Multiplying an estimate by an imagined click volume. That calculation has never once matched an invoice in my accounts, and it sets an expectation I then have to walk back.
Quoting a single figure to a client. A single number becomes a promise the moment it's written down. A range with named assumptions doesn't.
Treating a suggested bid as a price. It's a ceiling. Confusing the two makes people bid to the estimate and then wonder why the average crept toward it.
Reusing an estimate more than a month old. Auction depth moves faster than planning cycles in almost every vertical I work in.
Comparing account-level averages across time. Blending verticals, match types and devices into one figure produces a number that describes no auction anyone is bidding in.
Reacting to a monthly move on low click counts. At the volumes most keywords generate, a handful of unusual auctions can shift a monthly average by a third and mean nothing at all.
The practical takeaway
The estimate isn't lying to you, and it isn't going to become accurate. It describes a historical population of advertisers; you are one advertiser with a specific structure, a specific relevance profile and a specific set of exclusions. The gap between the two numbers is not noise to be eliminated it's a measurement of how much your account differs from the average one.
So use the tool to rank keywords and to spot budget-eaters. Derive what you can afford from conversion value rather than from market data. Buy three weeks of real auction data before committing to any meaningful plan. Present a range with its mechanisms rather than a figure. Then re-measure quarterly on a fixed subset so you notice repricing before the invoice does.
Do that and the estimate becomes useful for the first time, precisely because you've stopped asking it to be accurate. The advertisers who get hurt are the ones who read a number next to a keyword and mistake a historical average for a price they're about to be charged.