Google Ads Account Structure Grading Criteria

Google Ads account structure should now be graded on data flow, not organizational tidiness.

Cover illustration for “Google Ads Account Structure Grading Criteria”

Google Ads account structure has been graded by three different rulebooks since 2014, and each one replaced the one before it. From roughly 2014 to 2019, the single keyword ad group, or SKAG, was the standard: one keyword per ad group, graded on Quality Score alignment, bid precision, and exact-match fidelity. That logic worked as long as match types stayed rigid. It broke in 2018, when Google widened close variant matching and the premise that one keyword equaled one search query stopped being true. From about 2019 to 2022, themed ad groups took over, graded on intent coherence and whether each ad group held somewhere between 3 and 20 semantically related keywords. That grouping size mattered because Responsive Search Ads and early Smart Bidding needed enough volume inside each ad group to function. Then, from roughly 2023 to 2025, the grading frame shifted again, this time to campaign-type segmentation: how cleanly Search, Performance Max, Shopping, and brand campaigns were kept apart so they weren't competing against each other in the same auctions.

Each of these rulebooks made sense for its moment. SKAGs matched a world of rigid match types. Themed ad groups matched a world where Smart Bidding needed volume to learn. Campaign-type segmentation matched a world where PMax and Search were new enough to each other that cannibalization was the main risk to manage. The mistake isn't that any of these approaches was wrong when it was built. The mistake is holding an account to one of these old rulebooks today, as if account structure were a fixed discipline instead of a moving target that tracks how the platform actually works.

Diagram: The Three Eras of Google Ads Account Structure. Visualizes: Show three consecutive eras of Google Ads account structure grading, each defined by its dominant logic and timeframe.

Why structural tidiness no longer predicts performance

An account that looks clean in the interface, with dozens of tightly labeled campaigns and ad groups, can work directly against the AI that now runs most of the bidding. Smart Bidding builds a separate prediction model for each bid strategy, and that model needs a steady flow of conversion events to learn what a good auction looks like. Splitting a budget across many campaigns doesn't create many strong models. It creates many weak ones, each starved of the data it needs to tell a good result from a lucky one.

An underfed model doesn't just perform worse. It drops back into a learning phase every time someone touches it, and it overreacts to single conversions because it has nothing else to weigh them against. One account rebuilt from 43 campaigns and 178 ad groups, running a substantial monthly budget, averaged fewer than a dozen conversions per campaign per month. Target CPA goals swung wildly from one week to the next, and campaign-level CPA numbers told the team almost nothing useful, because there wasn't enough data behind any single number to trust it.

That's the real test a grader needs to apply. An account with many campaigns, careful naming, and deep segmentation can still fail where it counts most: it can't hand the algorithm enough signal to learn from. Account structure used to be treated as the lever that determined performance directly. Now it functions as the framework that decides whether the algorithm can do its job well or poorly, and any grading criteria written for 2026 have to start from that premise, not from the old one.

Conversion data quality as the first grading criterion

Before anyone asks how many campaigns an account should have, or how its keywords are grouped, the account has to pass a more basic test: is the conversion data clean enough for the algorithm to learn anything from it? A conversion in Google Ads is simply the action told to Google to optimize toward, like a demo request, a booked call, or a qualified lead. It is not revenue, and account structure exists to protect that signal, so bad structure doesn't dilute it further.

Grading conversion data starts with checking what's actually being counted as a conversion. A Performance Max campaign set to optimize toward form fills, instead of qualified leads, will do what it's told: it will find the cheapest form fill available anywhere on the internet and buy thousands of them. The campaign will look highly efficient and deliver almost nothing a sales team can use. Grading an account on structure alone, without checking what the conversion action actually measures, misses this kind of problem completely.

Tracking method matters just as much as the conversion action's definition. Enhanced Conversions and server-side tracking now belong in the same structural audit as campaign counts and keyword groupings, because a technically well-organized account with weak, privacy-degraded measurement is still feeding the algorithm corrupted numbers. Conversion volume looks like a budget constraint, not a structural one. It's both. Structure decides how the total pool of conversions gets distributed across campaigns, and a poorly structured account can take a healthy conversion volume and spread it so thin across too many containers that no single campaign ever gets enough to learn from.

Campaign consolidation as a grading criterion, and the two tests that justify a split

There's no fixed number of campaigns that makes an account correctly structured. What there is, is a pair of tests that any proposed split has to pass, and accounts that skip these tests and multiply campaigns anyway are structurally weak by the standards that matter now.

Both tests have to return yes before a new campaign is justified. First: does this segment genuinely need a different budget or a different performance target than everything else? Second: will this segment, on its own, still clear the minimum conversion volume a bid strategy needs to function? A product line that converts at a low cost per acquisition and one that converts at a much higher cost belong in separate campaigns, because a single target can't serve both well. Six locations in the same metro area, selling the same service at the same price, usually fail this test, no matter how much a client wants a campaign named after each city. If the economics are the same, splitting the campaign only divides the conversion data without changing anything the algorithm needs to optimize differently.

A campaign split earns its place when at least one of these genuinely differs: the budget cap or floor, the bid strategy or target, the geography, language, or compliance requirement, the primary conversion action, or a landing page story different enough that it can't reasonably be tested inside one campaign. When none of these differ, the right tool is a label, not a new campaign. Labels let a team report on segments separately without paying the algorithmic cost of splitting the data that feeds Smart Bidding.

Brand campaigns are the one case where separation should hold regardless. Brand traffic converts at a CPA so different from non-brand traffic that blending the two will distort whatever target gets set, usually by making non-brand look worse than it is. Blended reporting on brand and non-brand together can also make a flat account look like it's growing, when the growth is coming entirely from branded search that would have converted anyway.

Ad group construction: intent, landing page, and data consolidation within a campaign

Inside a campaign, the right way to organize ad groups has moved away from keyword taxonomy and toward intent paired with a distinct landing page. That shift carries its own set of checks. Every ad group should be built around one intent that genuinely earns its own ad copy and its own landing page. If two ad groups point to the same URL and say roughly the same thing, those two ad groups are really one ad group split across two containers, and keeping them separate just cuts the conversion data behind both in half.

Exact match keywords still have a job, but only in two situations: when a search term needs a different landing page than the rest of the ad group, or when a near-variant match would be both expensive and semantically wrong. "Emergency plumber" and "plumber jobs" make the point well. Google's matching systems can treat those two phrases as close enough to serve under the same ad group, but a plumbing business knows they represent completely different users with completely different intent, and only exact match protects that distinction.

Grading ad group hygiene comes down to two questions. Are there ad groups sending traffic to the same destination with nearly identical copy? And are there ad groups so narrow that they can't accumulate enough impressions for Responsive Search Ads to actually test different combinations of headlines and descriptions? Match type, in this framework, is a tuning dial. Splitting by match type only makes sense when a strict budget cap or a genuinely different economic target requires it, not as a default way to build out a campaign.

How Performance Max fits into the graded structure

Performance Max doesn't replace the structural thinking that applies to Search campaigns. It adds its own grading criteria, built around feed quality, how clean the conversion signal is, and how much risk it carries of cannibalizing other campaigns, and those criteria need to be checked on their own terms.

For lead generation accounts, the bar is stricter. A PMax campaign should only go live after brand exclusions are set and after someone has confirmed the conversion actions it's optimizing toward actually represent qualified leads. Without that groundwork, PMax will optimize toward whatever is cheapest to generate, which is usually the lowest-quality form fill available, not the lead that's actually worth paying for.

Campaign count discipline applies here just as it does with Search. Start with a single PMax campaign. Add a second one only when a specific product group needs a genuinely different ROAS target, a separate budget, a different conversion goal, or some other strategic difference that can't be served inside the first campaign. Piling on more PMax campaigns without one of those reasons just repeats the fragmentation problem described earlier, with the added complication that PMax's full-inventory reach makes cannibalization harder to spot.

AI Max deserves a separate mention, because it's easy to confuse with PMax and the two carry different grading requirements. AI Max is a setting inside a standard Search campaign. It adds broad-match keyword expansion and automated asset generation, but it operates inside that Search campaign's existing structure. Performance Max is a fully separate campaign type that runs across all of Google's inventory. A structural audit has to treat these as two distinct elements, each with its own data needs and its own cannibalization risks.

Naming conventions, labels, and reporting architecture as grading criteria

Naming conventions and labels aren't decoration. They decide whether anyone can act on performance data quickly, and whether a structural decision made six months ago can still be audited today. A workable naming convention passes a simple test: can someone new to the account, or an outside auditor, tell a campaign's type, geography, bid strategy, and primary conversion action just from its name, without opening it?

Labels are the right tool for any segmentation that exists purely for reporting. If a split doesn't unlock a different budget, a different bid target, a different geography, a different compliance setting, or a different conversion action, it should be a label attached to an existing campaign, not a new campaign of its own.

Shared negative keyword lists, applied at the account level, keep spend from drifting into irrelevant search queries across every campaign at once. An account missing these lists has a structural gap that no amount of tidy campaign naming fixes. Shared budgets and portfolio bid strategies, used where they genuinely apply, also cut down on manual intervention and keep the algorithm's learning more stable, since fewer separate containers means fewer separate models to retrain. The clearest sign of a reporting failure is simple: if understanding how the account is performing requires a complicated spreadsheet pulled together outside Google Ads, the structure itself isn't doing its job. Good structure should make the right numbers visible inside the platform, without needing translation.

A grading rubric that puts all the criteria in order of priority

The criteria above don't carry equal weight, and they don't stand alone. Measurement has to be sound before campaign architecture can be judged fairly, and campaign architecture has to be sound before ad group construction and day-to-day hygiene are worth evaluating.

Tier 1 covers measurement integrity. Are the conversion actions tracking real business outcomes, like qualified leads or booked calls? Is the tracking technically sound, with Enhanced Conversions and server-side tracking in place where they're relevant? Is each campaign getting enough conversions to actually support the bid strategy running on it?

Tier 2 covers campaign consolidation. Does every campaign in the account pass both split-justification tests: a genuinely different economic target, and enough conversion volume to stand on its own? Is brand kept separate from non-brand? Is Performance Max isolated with the right brand exclusions in place, and has it been checked for cannibalization against Search?

Tier 3 covers ad group and keyword construction. Is each ad group built around a single intent with its own landing page? Are keyword lists lean, backed by strong negative lists? Are Responsive Search Ads set up with enough combinations and enough traffic to actually test what works?

Tier 4 covers operational hygiene. Can someone unfamiliar with the account read the naming conventions and understand them? Are labels doing the work that extra campaigns used to do? Are shared negative lists and bid strategy libraries actually in place?

An account that scores well against this rubric can look sparse compared to what passed for good structure a few years ago: fewer campaigns, fewer ad groups, shorter keyword lists. That sparseness is what an account looks like when its structure is built to feed an algorithm enough clean data to do its job.

Diagram: Four-Tier Grading Rubric for Account Structure. Visualizes: Show a prioritized four-tier hierarchy of account grading criteria, where each tier must be sound before the next is meaningful.

Daria Solomennik

Senior Audit Correspondent

Daria spent eight years inside performance marketing agencies before going independent, where she developed a rigorous methodology for dissecting Google Ads accounts that bleed budget without results. Her audit teardowns have been cited by several paid search communities as reference material for account managers.