QuantumAdsLab Logo Quantum Ads Lab
A dashboard of rising green conversion bars stamped with a red self-graded A+ seal, illustrating an automated bidding system reporting on its own performance
The optimizer grades its own homework — and it always passes

The Optimizer Grades Its Own Homework

Summary

What you'll learn in this article

  • Why automated bidding both spends your budget and grades its own results — the referee and the player are the same entity
  • How a soft conversion goal (a form view, a newsletter tick) can fill your dashboard with green while real sales sit flat
  • Why Performance Max can claim credit for brand searches you'd have gotten for free, inflating reported ROAS with near-zero incremental revenue
  • Why a broken account and a brilliant one can show the identical green dashboard, and why nobody audits a number that's going up
  • The checks the optimizer doesn't get to write itself: incrementality tests, geo holdouts, new-customer counts and the P&L

Here is a thing we do in paid media and rarely say out loud. We hand Google's automation the bid, the budget, the audience, the placement, the timing, and then we open Google's report to find out how well Google did. The same system that spends the money writes the report card on how the money was spent. We call the report card “results” and we plan the next quarter around it.

I want to state the problem as plainly as the person I'm writing this with: a system that reports on its own performance will report success it did not produce. Not because it lies. Because it is grading its own homework, against a rubric it wrote, using a red pen it owns.

This isn't a conspiracy and it isn't a bug. Smart Bidding, Performance Max, Demand Gen: these are genuinely good tools, and I run them on most mature accounts. But every one of them collapses two jobs that should never sit with the same party. The optimizer decides where your money goes, and then the optimizer measures whether that was a good idea. When the referee and the player are the same entity, the score stops being evidence.

Let me make it concrete, because “self-reported” is too abstract to act on.

The conversion it counts is the conversion it chose

Automated bidding optimizes toward a conversion action you nominally chose but it interprets. Point it at a soft signal (a form view, an add-to-cart, a “lead” that is really a newsletter box tick) and it will become exquisitely good at manufacturing that signal. The dashboard fills with conversions. The cost per conversion drops. Every arrow is green. Meanwhile the only number that pays your salary, customers who actually bought , sits flat, and nothing on the screen tells you, because the screen is measuring the thing the system was told to maximize, not the thing you needed. It did the homework it assigned itself, and it gave itself an A.

The credit it claims is credit it borrowed

Performance Max is the sharpest example I have. Left unconstrained, it will happily intercept people already searching your brand name: people typing “quantumadslab” into Google, who were going to reach you anyway, for free and it will book each of those clicks as a conversion it produced. The reported ROAS is beautiful. I've seen 8x, 10x, glowing. The incremental revenue, the revenue that exists because the campaign ran and would not exist otherwise, can be close to zero. You are paying a toll on traffic you already owned, and the report thanks itself for the traffic.

You cannot see this on the dashboard, and that is the whole point. The counterfactual, the sale that would have happened without the ad, never appears as a line item. There is no red number labelled “revenue you'd have gotten for nothing.” The system reports what it touched, not what it changed, and those two things are wildly different in exactly the cases where it matters most.

The failure that shows up as a success

This is where my half of the argument meets the other one, the one The Filter Lab makes about market data. In both worlds the danger isn't a failure that screams. It's a failure that smiles.

A broken paid-media account and a brilliant one can produce the same green dashboard. That sentence should frighten anyone who allocates budget off a screen. When the tool that grades the work is the tool that did the work, breakage doesn't render as an error, it renders as achievement. The campaign harvesting your own brand, the campaign optimizing to a junk conversion, the campaign quietly serving your ad to people who'd never buy: all three light up the same shade of green as the campaign that's genuinely printing money. The report has no vocabulary for its own mistakes, because a mistake, from inside the optimizer's frame, is just another conversion it decided to count.

That's why these failures stay invisible for months. Nobody audits a number that's going up.

A real account, anonymized

Last quarter a client came to me delighted. Their Performance Max campaign was reporting an 8-to-1 return, and an agency pitching them had used that number as proof the account was healthy and should simply be scaled. Scale the winner, the logic went.

Before adding a euro, I did the one thing the platform can't do for you: I measured it against something the platform doesn't control. We carved brand terms out of PMax and ran a clean geo holdout (some regions saw the campaign, matched regions didn't) and compared actual revenue in the bank, not conversions in the interface. A large slice of that “8x” was brand harvesting and audiences that converted in the holdout regions too, with no ads at all. The incremental return, the money that existed because of the spend, was a fraction of the headline. Not zero, but nowhere near a figure you'd bet a scaling budget on.

We reallocated toward the campaigns that survived a holdout and cut the toll we were paying on our own name. Spend went down, revenue held. On the dashboard, it looked like we'd broken something, the reported ROAS dropped, while the bank account said the opposite. That gap, between the report getting worse and the business getting better, is the entire subject of this post.

Grade the homework yourself

The fix isn't to distrust automation. It's to refuse to let the automation be the only witness to its own trial. The score only becomes evidence again when it's produced by something the optimizer can't edit.

So I measure against things that live outside Google's report. Incrementality tests and holdouts, where a matched group sees no ads and tells me what would have happened anyway. New-customer count rather than blended conversions, so brand harvesting can't disguise itself as growth. And above all my client's own P&L, revenue that actually landed, margin that actually cleared, because that number is written by reality, not by the system asking to be judged. When the platform's report and the bank disagree, the bank wins, every time.

An optimizer grading its own homework will always pass. That's not a reason to fire the optimizer. It's a reason to be the examiner it doesn't get to be: to keep one measurement, always, that the machine has no hand in writing.

My co-author is about to show you the same trap in market data, where a measurement can be internally consistent and still describe nothing. Different domain, one level down: mine is a system grading its own performance, his is a system grading its own measurement and in both, the moment the thing being checked also holds the pen, a number that looks right and a number that is right quietly stop being the same thing.

Read the market-data half at The Filter Lab (link to follow).

FAQ: Self-reported metrics in Google Ads

What does it mean for a platform to "grade its own homework" in paid media?
It means the same system that decides where your budget goes — the bid, the audience, the placement, the timing — is also the system that reports back on whether that spend worked. Smart Bidding, Performance Max and Demand Gen all collapse these two jobs into one party. That isn't a bug or a conspiracy, but it does mean the report is not independent evidence: it's the optimizer scoring itself against a rubric it wrote.
Why can Performance Max show an inflated ROAS?
Left unconstrained, Performance Max can intercept people already searching your brand name — traffic you would have gotten anyway, for free — and book each of those clicks as a conversion it produced. The reported ROAS can look excellent (8x, 10x) while the incremental revenue, the revenue that exists because the campaign ran, is close to zero. The dashboard has no line item for the sale that would have happened without the ad, so this gap never shows up on its own.
What is an incrementality test or geo holdout, and why does it matter?
An incrementality test compares matched groups where one is exposed to a campaign and the other isn't, then looks at actual revenue rather than platform-reported conversions. A geo holdout does this by region: some areas see the campaign, matched areas don't, and the difference in real revenue is the incremental effect. It matters because it's one of the few measurements the optimizer doesn't get to write itself — it's produced by something the platform doesn't control.
How do I know if my Smart Bidding conversions are real?
Check three things outside the platform's own report: whether the conversion action you're optimizing toward is a soft signal (a form view, an add-to-cart) rather than an actual sale; whether new-customer count is growing, not just blended conversions, since brand harvesting can disguise itself as growth; and whether your P&L — revenue that actually landed — agrees with the dashboard. When the platform's report and the bank disagree, the bank is right.