An AI ad creative review process is the stage most paid-media pipelines never designed, because until recently they never needed one. Generation became cheap and instant; checking did not. Meta’s most recently disclosed figure — more than four million advertisers using its generative ad-creative tools, given on the Q4 2024 earnings call — describes an industry where creative volume is set by compute rather than by headcount.
The consequence is structural, not moral. When a team could produce twelve concepts a week, every concept got a human look by default. When the same team can produce two hundred, the default breaks and nothing replaces it. Creative that would once have died in a review meeting now reaches a live auction because no stage in the pipeline was ever assigned the job of stopping it — and on some surfaces, such as Google’s generative Performance Max asset-group build flow, generation and publication sit inside the same workflow with no separate brand-or-claim approval step in between.
This playbook covers the gate that fills that hole: what it checks, how it differs from performance testing, which parts the ad platforms already ship (and where their own documentation says they stop), how to set a sampling rate by risk tier instead of pretending you will review everything, what approval SLA the sampling implies, and the four metrics that tell you the gate is catching things rather than just adding days.
- 01Volume outran review, not judgment.Meta’s most recently disclosed count is 4M+ advertisers on its generative ad-creative tools, from the Q4 2024 earnings call. Nothing about that number says the creative is worse — it says the ratio of assets to reviewers changed and the review stage never got redesigned.
- 02A review gate is a different artifact from a test.Performance testing asks which creative wins after publish. A review gate asks which creative may publish at all. They use different inputs, different owners and different pass criteria, and one cannot substitute for the other.
- 03Platform AI-label settings are necessary, not sufficient.Google ships an AI content label setting across five of its ad surfaces, adds a visible on-ad overlay for campaigns targeting the EU, India and New York, and embeds SynthID plus C2PA metadata — while stating plainly that the setting does not guarantee regulatory compliance.
- 04Sampling by risk tier is the only workable answer.Reviewing everything is arithmetic that does not close. Reviewing nothing is the current default. Tiering the output and assigning each tier a review rate is what makes the gate survivable at volume.
- 05Without an SLA and an escape metric, a gate becomes a queue.A gate with no published turnaround gets routed around within a month. Publish the SLA per tier, then track escape rate — the share of published creative that later has to be pulled — as the number that says whether the gate is real.
01 — The Volume ProblemCreative output is now set by compute, not headcount.
The clearest public marker of the shift comes from Meta’s Q4 2024 earnings call, reported in February 2025: more than four million advertisers were using its generative AI ad-creative tools, roughly four times the figure disclosed about six months earlier. That is the most recently disclosed number we could verify at the time of writing, and it should be read as a 2024-25 marker rather than a live count — but the direction it establishes has not reversed. Meta’s CFO said on the same call that Image Animation, launched in October 2024, already had hundreds of thousands of advertisers using it monthly within roughly three months.
Adoption at that scale changes the shape of a creative pipeline in a specific way. The bottleneck used to be production: a designer’s week, a studio booking, a shoot. Review was free because it happened inside production. Once generation collapses to minutes, production stops being the constraint and review becomes the only remaining one — except no one budgeted it, staffed it, or gave it a definition of done.
The generation-to-live path has also shortened on the Google side. Google documents a workflow for building a Performance Max asset group using generative AI, available in the United States, in which assets are created inside the campaign-build flow itself. That is genuinely useful, and it means the only thing standing between a generated asset and a live auction is whatever discipline the team supplies. If you want the volume, you have to supply the gate. For the upstream half of the problem — the brief that the gate then checks creative against — see our workflow for generating the brief the gate checks creative against.
02 — Gate Versus TestTesting asks which creative wins. A gate asks which may run.
The most common failure we see is a team that believes it already has a review process because it has a testing process. Those are different stages with different failure modes. A performance test tells you a variant underperformed; it tells you nothing about whether the variant misstated a price, drifted off the brand’s typographic system, implied an endorsement nobody gave, or needed a disclosure it did not carry. Worse, the test only produces its verdict after the creative has already been served to real people.
Keep the two artifacts separate and let each do its own job. If you need the performance half, we have written a performance-testing framework for AI-generated variants and CTR and ROAS benchmark data for AI-generated creative separately; if you are automating the loop end to end, there is also an agentic testing workflow. None of those replace the gate, and the gate does not replace them.
Which creative wins?
Runs after publish. Inputs are impressions, CTR, ROAS and holdout structure. Owned by paid media. Its verdict arrives once real budget and real audiences have already touched the asset — which is exactly what makes it useless as a safety control.
Which creative may publish?
Runs before publish. Inputs are the brief, the brand system, the claim evidence file and the disclosure rules for the target markets. Owned jointly by creative ops and whoever carries claim risk. Binary pass or fail with a named reason, not a score.
How is this creative disclosed?
Runs alongside both, at the platform layer. Handles AI-disclosure metadata and, in some markets, a visible on-ad overlay. It answers a labeling question only — it says nothing about whether the creative is on-brand or whether its claims are substantiated.
“The ‘AI wow factor’ can fool people into thinking they’ve created a good end product.”— Gareth Morgan, Head of Global Brand Design at Revolut, quoted in Superside
That is the specific cognitive failure a gate exists to interrupt. Novelty reads as quality. A generated image that looks technically impressive collects approval that a mediocre stock photo of the same concept would never have received, and the approval happens fast because everyone in the room is reacting to the medium rather than the message. A gate with named checks and a written pass criterion takes the judgement out of the room and puts it against a list.
03 — Gate DesignFour checks, each with an owner and a pass criterion.
A gate that is only a vibe check will be routed around the first time a campaign is late. Four checks is the smallest set we have found that covers the real failure modes without turning into a committee. Each needs a named owner, a written pass criterion, and a place the result is recorded — the record is what lets you compute the metrics in Section 08.
Brand drift
Does the asset sit inside the brand system: palette, type, logo lockup, product rendering, tone of the headline? Generated assets drift most on the things a brand guideline states least precisely. Pass criterion: a named reviewer can point to the master the asset derives from.
Claim and compliance
Every factual, comparative or endorsement-shaped statement in the asset maps to a line in an evidence file. Anything generated that has no evidence line does not get softened — it gets cut. This is the check with legal exposure attached, and the one AI generation makes easiest to fail.
Sampling at volume
You cannot review everything, so decide deliberately what fraction of each tier gets human attention and write it down. An undeclared sampling rate is still a sampling rate — it is just one nobody chose and nobody can defend.
Approval SLA
A published turnaround per tier, measured from the moment an asset enters the queue to the moment it has a decision. Without it the gate becomes a bottleneck, and bottlenecks get bypassed quietly rather than escalated openly.
04 — Platform ControlsWhat the platforms actually ship — and where they stop.
Before building anything, take credit for what already exists. Google documents an AI content label setting that rolls out across Google Ads, Display & Video 360, Campaign Manager 360, Merchant Center and Ads Editor, letting advertisers designate assets as AI-created or AI-edited. Where a campaign targets a market whose rules require it, Google places a visible overlay on the ad itself rather than tucking the disclosure into the “How this ad was made” panel. Google also prompts advertisers to review assets when ads in a campaign may require a label before those assets run unlabeled.
Two details in Google’s own documentation are worth internalising before you treat the setting as a solution. Google can apply labels to an advertiser’s assets on its behalf in some circumstances — when it is legally required to, or when it receives signals from other platforms — and once applied that way, the labels cannot be overwritten by the advertiser. And Google embeds non-visible SynthID watermarks alongside C2PA machine-readable provenance metadata in images and videos generated inside its ads tools, which means the provenance signal travels with the file whether or not your reporting captures it.
Google surfaces
Google Ads, Display & Video 360, Campaign Manager 360, Merchant Center and Ads Editor all carry the AI content label setting, so the designation follows the asset rather than living in one tool.
Named jurisdictions
Google cites AI regulations in the EU, India and New York as requiring disclosures or labels on ads carrying certain AI-generated or AI-edited assets; for campaigns targeting those regions the label appears on the ad itself.
SynthID plus C2PA
Assets generated inside Google Ads tools carry a non-visible SynthID watermark and C2PA machine-readable metadata. Provenance is embedded by default — your gate should read it, not recreate it.
The labeling layer is deep enough to be its own subject, and it varies by platform and market in ways a single paragraph flattens. We have covered it separately in the platform-by-platform labeling reference and, for the European side specifically, in our guide to EU provenance-marking rules. For the Meta side of the tooling picture, see our advertiser guide to Meta’s AI creative ad tools. Feed all three into the gate as inputs to Check 02, not as substitutes for it.
05 — Claim SubstantiationThe claim check has the sharpest edges.
One clarification first, because a lot of writing on this subject invents authority that does not exist. We could not locate an AI-specific advertising rule from the FTC covering generated ad creative at the time of writing. What does exist, and what the claim check should be built against, is the FTC’s Endorsement Guides, revised in June 2023. They require that endorsements and testimonials in advertising — across television, print, radio, online, podcasts and social — be truthful and not misleading, and that an unrepresentative testimonial be paired with what consumers can generally expect.
The Endorsement Guides are not regulations in their own right. The FTC’s own framing is that it can investigate and sue under the FTC Act’s prohibition on unfair or deceptive practices where advertisers ignore them. That is the enforcement mechanism your claim check is ultimately protecting against, and it is why the check has to be binary. A generated testimonial-style line that no customer ever said is not a stylistic choice to be softened in review; it is a fabricated endorsement, and the correct gate outcome is a fail.
In practice, three generated patterns account for most of what our claim check catches. Composite or invented testimonials, where the model produces a plausible customer voice from nothing. Comparative superlatives the brand has never substantiated, because superlatives are what language models reach for when a brief says “make it punchy.” And numeric specificity with no source — a percentage, a saving, a rating — that reads as data because it is formatted as data. Each fails for the same structural reason: no evidence line, no publish.
06 — Sampling MathReview rate should scale with consequence, not volume.
The instinct when volume rises is to demand that everything be reviewed. The arithmetic never closes, so what actually happens is that review quietly becomes optional and nobody says so. The alternative is to decide, explicitly, that different creative carries different consequence and to set a review rate per tier.
One published operational data point is worth borrowing as an anchor. The creative-services firm Superside describes generating roughly 50 to 100 AI creative options per project and narrowing to 5 to 10 refined finalists, inside a three-stage model of strategic direction before generation, expert curation of outputs, and human refinement including compliance verification before anything ships. Those are vendor-published figures from Superside’s own case data rather than an audited industry average, so read the ratio as one firm’s operating shape — between roughly a twentieth and a fifth of raw generation surviving to reviewed output, centring near a tenth — and not as a benchmark to hit.
The table below is our own synthesis. The tiers combine three inputs: the generate-wide, review-narrow ratio above as the low-consequence baseline; the FTC substantiation expectation as the trigger for full review of claim-bearing creative; and the EU / India / New York disclosure requirements as the trigger for the top tier. The minute costs and SLAs are Digital Applied planning figures, not measured industry values — the point is the shape of the curve, and you should recalibrate the numbers against your own escape rate within a quarter of running it.
| Risk tier | Required checks | Review rate | Min / reviewed asset | Min / 100 generated | SLA and owner |
|---|---|---|---|---|---|
| Low-consequence output — sampled | |||||
| Tier 0 — reformat of an approved master | Brand drift only: automated diff against the master plus a human spot-check | 10% | 3 | 30 | Same day · creative ops |
| Tier 1 — net-new variant, no product claim | Brand drift plus a full copy read against the brief | 25% | 6 | 150 | Same day · creative ops |
| Consequence-bearing output — reviewed in full | |||||
| Tier 2 — claim-bearing or endorsement-shaped | The above plus claim substantiation against the evidence file | 100% | 12 | 1,200 | 24 hours · brand and legal |
| Tier 3 — regulated category, or EU / India / New York targeting | The above plus platform AI-label verification and legal sign-off | 100% | 25 | 2,500 | 48 hours · legal-led |
The fifth column is derived, not asserted: review rate × 100 generated assets × minutes per reviewed asset. Tier 0 costs 0.10 × 100 × 3 = 30 reviewer-minutes per hundred generated; Tier 3 costs 1.00 × 100 × 25 = 2,500. That is a spread of roughly 83× across the same nominal unit of work, and it is the reason a flat review policy is always wrong in one direction or the other. Applied uniformly at the Tier 3 rate, a team drowns; applied uniformly at the Tier 0 rate, the claim-bearing work never gets a proper read.
07 — SLA And CapacityThe small tier eats the largest share of capacity.
Run the tier rates against a realistic monthly mix and the staffing question answers itself. Take 1,000 generated assets in a month split 60 / 25 / 10 / 5 across the four tiers — 600 reformats, 250 net-new variants, 100 claim-bearing assets and 50 in the regulated tier. Tier 0 costs 600 × 0.10 × 3 = 180 reviewer-minutes. Tier 1 costs 250 × 0.25 × 6 = 375. Tier 2 costs 100 × 1.00 × 12 = 1,200. Tier 3 costs 50 × 1.00 × 25 = 1,250. The total is 3,005 minutes, just over 50 hours — roughly 31% of one reviewer’s month at 160 working hours.
Share of monthly review capacity by risk tier · 1,000 generated assets
Source: Digital Applied model — tier rates from the table above applied to an illustrative 60/25/10/5 monthly volume mixThe inversion in that chart is the planning insight. Tier 3 is 5% of the assets and 41.6% of the review capacity; Tier 0 is 60% of the assets and 6.0% of the capacity. Together, the two consequence- bearing tiers are 15% of volume and 81.5% of the reviewer’s month. Every hour you spend automating the brand-drift diff on reformats buys back a rounding error. Every hour you spend making the claim-substantiation step faster — a better-structured evidence file, a shorter legal loop — buys back real capacity.
That also sets the SLA. Same-day is achievable for the sampled tiers because their total load is 555 minutes a month across 1,000 assets. Twenty-four hours is honest for Tier 2 because it needs a second person. Forty-eight hours for Tier 3 is not slowness, it is the legal loop being told the truth about its own turnaround. Publish those numbers, then hold to them — an SLA the gate misses twice teaches the organisation to route around the gate, which is a worse outcome than never having built it.
Looking forward, we expect the pressure on this design to come from the generation-into-campaign flows rather than from volume alone. As more surfaces let a marketer produce and publish inside one workflow, a gate that sits as a separate downstream step becomes structurally easy to skip. The version that survives is the one that moves upstream into the brief and the evidence file — constraining what can be generated in the first place — and keeps only the sampled human read at the end. Teams that treat the review rate as a budget line now will find that transition much cheaper than teams still treating review as goodwill.
08 — Gate MetricsFour numbers that say the gate is working.
A gate with no instrumentation is indistinguishable from a delay. These four metrics are cheap to compute if the four checks record their outcomes, and between them they catch both failure directions: a gate that has stopped catching anything, and a gate that has started blocking everything.
Catch rate
The share of reviewed assets that fail a check. A healthy gate has a non-zero catch rate. When it trends toward zero, either the upstream brief genuinely improved or reviewers have stopped reading — check which by re-reviewing a small blind sample.
Escape rate
The share of published creative later pulled for a brand, claim or disclosure reason. This is the only metric that measures the outcome the gate exists for. Every escape should be traced back to the tier it came from and the tier rate adjusted accordingly.
Time in gate
Median and 90th-percentile hours from queue entry to decision, per tier, against the published SLA. Watch the 90th percentile rather than the median: the tail is where people learn that the gate is unreliable and start going around it.
Coverage
Actual review rate per tier versus the rate you published. Drift here is almost always silent and almost always downward, and it is the earliest available signal that the gate is being quietly abandoned under deadline pressure.
Read them as a pair of pairs. Catch rate and escape rate tell you whether the gate is finding real problems; time in gate and coverage tell you whether the organisation still believes in it. A gate can fail on either axis, and the second failure mode is far more common — the checks are fine, but the queue got long during a launch, an exception was granted, and the exception became the process. If you want help designing the gate into an existing paid-media operation, that is the kind of work our paid media team builds alongside the campaign structure itself, and it usually sits next to the broader operating-model work in our AI transformation engagements.
09 — ConclusionThe gate is the part nobody budgeted.
Generation got a budget line. Review still has not.
Every element of this playbook exists because the economics of producing an ad changed and the economics of checking one did not. The platforms have built real controls — Google’s label setting spans five surfaces, adds a visible overlay in the markets it names, and embeds SynthID and C2PA provenance at generation — and Google says in its own documentation that none of that establishes regulatory compliance. That boundary is honest, and it is where your gate begins.
The design that survives contact with volume is unglamorous: four named checks, a review rate that scales with consequence rather than with output, a published SLA per tier, and four metrics recorded well enough to argue with. In the illustrative model above, 15% of the assets consume 81.5% of the review capacity — which means the strategic question is not “how do we review more,” it is “which 15% actually needs a person.”
The forward risk is that generation and publication keep collapsing into the same workflow, and a review stage that sits outside that workflow becomes easy to skip on a deadline. The answer is to push the constraints upstream into the brief and the evidence file, so that most of what gets generated is already inside the lines, and reserve human attention for the small tier where being wrong is expensive. That is not a tooling decision. It is a decision about what your brand is willing to say when nobody is reading every word.