Note · Practice
The Machine Says It’s Working
Optimization score is a report card the student wrote. Knowing whether an AI-driven account actually works is less about the dashboard and more about deciding, in advance, what would have counted as failure.
On this page
Google Ads stopped being a set of levers a while ago. The bidding is a model, the matching is a model, Performance Max is a black box wearing a campaign type’s clothes, and the interface has quietly reorganized itself around telling an advertiser that everything is fine. Optimization score sits at the top of the page like a report card written by the student. Recommendations auto-apply themselves unless someone goes and turns that off. The system is confident. Confidence is not evidence.
So the question — how does anyone actually know it’s working — turns out to be less about finding the right dashboard and more about deciding, in advance, what would have counted as failure. Most accounts I’ve looked at have never done that. They have conversions. They have a cost per conversion. They have a number that goes down sometimes and someone screenshots it.
Conversion counting is not measurement
The first thing that breaks in an AI-driven account is conversion quality, and it breaks silently. Smart Bidding optimizes toward whatever it’s told to value. Tell it a form fill is worth one, and it will find the cheapest human beings on the internet who will fill out a form — people who wanted a free estimate, people who wanted a price, people who wanted somebody to talk to at 11pm. The model is not wrong. The model is doing exactly what it was told, extremely well, which is the whole problem (this is the part that gets called “the AI got worse” in the account review, when what actually happened is that the AI got better at a bad instruction).
Lead gen accounts are where this is worst, because the thing being bought is not the conversion. The conversion is a proxy. The actual thing is a booked job, a signed contract, revenue that lands in a bank account thirty to ninety days later, and the gap between those two events is where every bad Google Ads account lives.
- Form fill, phone call, chat, click-to-call, WhatsApp tap — all firing as one generic “conversion” with one generic value
- Calls counted at 30 seconds, when the actual qualifying call runs six minutes
- Thank-you page conversions triggering on repeat views, bookmarks, refreshes
- Every lead worth $1, when the roof replacement is worth $14,000 and the gutter cleaning is worth $200
Fixing that is not a settings change, it’s a plumbing project. Offline conversion imports — GCLID or GBRAID captured on the form, stored in the CRM, pushed back to Google when the lead becomes a quoted job and again when it becomes a signed one — are the single highest-leverage thing available in the AI era and almost nobody does them, because they require the sales side and the ads side to share a database. Enhanced Conversions for Leads does a hashed-email version of the same trick for accounts that can’t hold a GCLID. Either way the point is identical: the model gets to see which leads were real. Once it sees that, bidding stops being cost-per-form and starts being cost-per-actual-money, and the campaign quietly stops buying the cheap garbage.
Attribution got blurrier and that’s mostly fine
Data-driven attribution is the default now, last-click is gone as an option in most places, Consent Mode v2 is modeling a chunk of European conversions outright, and iOS keeps eating the rest. A meaningful percentage of the conversions in the interface were not observed — they were estimated. That sounds like a scandal until it’s compared to the alternative, which was pretending a last-click cookie was a fact.
The practical consequence is that the conversion number in Google Ads should be treated as directionally useful and precisely wrong, and any decision that depends on a 4% difference between two weeks is a decision built on sand. Comparisons that survive:
- Rolling 28-day windows against the prior 28, never week-over-week
- Year-over-year for anything seasonal (a roofer in June and a roofer in January are different businesses)
- Conversion lag accounted for — a 30-day window means the last 30 days are always understated, and the account always looks like it’s collapsing on a Tuesday
Incrementality, or the only question that matters
Search advertising has an old, unfashionable disease: brand terms. Somebody searches the company name, clicks the ad instead of the organic listing directly beneath it, converts, and the campaign takes credit for a customer who was already walking through the door. Performance Max has an aggressive version of this, since it will happily serve on brand queries and shopping and YouTube and Gmail and then report all of it as one glorious blended ROAS. An account can look like it’s printing money while the incremental contribution is close to nothing.
The tests that answer this are unglamorous and take weeks:
- Geo holdouts — matched market pairs, ads off in one, on in the other, compare total business (not platform-reported) outcomes
- Drafts and Experiments for anything inside the same campaign type — a real 50/50 split with a statistical readout, which Google will actually calculate
- Brand exclusion lists on PMax, run for a month, watching whether non-brand volume actually existed underneath
- Budget step tests — up 40%, hold four weeks, watch whether total conversions moved or just the cost per one
None of these produce an answer on Thursday. That’s the tax. An account that can’t wait four weeks for an answer is an account that will keep making decisions off noise for years.
What success looks like written down
The version of this that works starts from the business and reads backward, and it fits on an index card. Average revenue per closed job. Gross margin on it. Close rate on qualified leads. Qualification rate on raw leads. Multiply through and there’s a maximum allowable cost per lead, and that number — not a benchmark, not an industry average, not what a competitor claims on a podcast — is the target. Everything above it is a losing account and everything comfortably below it is an account that should probably be spending more.
Then the scoreboard is short. Cost per qualified lead, tracked monthly. Lead-to-sale rate by campaign, which requires the CRM to hold the source. Total booked revenue attributable to paid, blended, no heroics. Impression share lost to budget, as the ceiling indicator — the one number that says “there is more of this available and it isn’t being bought.” Search terms and asset group performance reviewed for relevance drift, because broad match plus a bidding model will wander into adjacent categories that convert on paper and never on a job site.
And a discipline about touching things. The learning period is real, two weeks is roughly the floor, and every mid-flight tCPA adjustment restarts it (the most common cause of a bad AI account is not the AI, it’s an anxious human editing the target every four days and then concluding the machine is broken). Change one variable. Wait. Read.
The uncomfortable part
An advertiser who cannot answer “what would make me shut this off” does not have a measurement system, they have a subscription. The AI-era Google Ads account rewards exactly one skill above all others, and it isn’t keyword research or ad copy or bid strategy selection — it’s the willingness to build the feedback loop that tells the model what a good customer actually looks like, and then to leave it alone long enough to find some. The rest is dashboard decoration. Optimization score can sit at 100% on an account that is quietly setting money on fire, and it frequently does, and Google will send an email about it either way.
The machine is only as honest as the data it’s fed.