Between April and July 2026 we pulled 13,572 real, running ads from Meta's Ad Library — the public archive of every ad on Facebook and Instagram — and put 3,260 of them through a human design review. Just over half, 50.8%, were rated bad. Not mediocre. Bad. Only 19.2% earned a good.
These are not ads from beginners. A large share came from 311 hand-picked big-brand advertisers — the companies whose creative teams are supposed to know what they're doing.
If you want the short version of what actually separates the winners: they bake the headline into the image, they put a number in that headline, they keep the body copy short, and they sign their work at the bottom like a painter. The data behind each of those claims is below.
- Of 3,260 human-reviewed ads, 50.8% rated "bad" on a three-point design rubric; only 19.2% earned "good".
- 96.6% of ads whose image our vision model rated "good" have text baked into the image, against 67% of "fair" ones. Among ads that survived 90+ days, 80.2% carry baked-in text.
- The median Facebook ad runs 37 days. One in four runs past 90 days; one in eight passes 180.
- Numbers work: 36.4% of good-rated headlines contain a digit, against 17.9% of bad-rated ones. Ads that survived 90+ days used digits nearly twice as often as ads that died inside two weeks.
- Big brands put their logos at the bottom: 57% of labeled ads place the logo in the bottom row of a nine-point grid, and only 24% at the top.
The dataset: 13,572 ads, 3,260 human reviews
Everything in this report comes from our own ad corpus — the teacher set we built to train our creative engine. The pipeline:
- Collection. 13,572 active ads pulled from Meta's Ad Library API between April 14 and July 7, 2026, across 2,312 distinct advertiser pages. 311 of those advertisers were hand-curated big brands whose full active libraries we pulled; the rest came in through keyword search.
- Coverage. Eleven verticals: finance (1,844 ads), ecommerce (1,785), handmade goods (1,475), health and wellness (1,429), SaaS and productivity (1,234), marketing (1,172), developer tools (1,119), education (1,118), legal (998), HR and hiring (814), and CRM and sales (584).
- Labeling. 3,260 ads got a human review on a three-point rubric — bad / fair / good — across three dimensions: headline, image, and product-message fit. A frontier vision model graded another 1,643 on the same rubric, and 707 winners got a full creative deconstruction (composition, typography, color, the transferable idea). Ad embeddings live in pgvector so we can retrieve by similarity.
- Longevity. For nearly every ad we estimate days active — how long the advertiser has kept paying for it. Ad libraries don't publish performance, so lifespan is our proxy: nobody keeps spending on a loser for months. It's revealed preference, in budget form.
One bias worth stating up front: the human-review queue prioritized designed ads — compositions with intent — and auto-skipped bare product photos. The 50.8% failure rate is measured on the ads that were trying.
Half of these ads would fail a design review
The full distribution across the 3,184 human image ratings: 1,619 bad (50.8%), 953 fair (29.9%), 612 good (19.2%).
The failures are rarely exotic. The most common negative labels in our review data are the boring ones: low image quality, no clear focal point, too busy or cluttered, generic stock-photo feel. The most common positive labels on winners are equally unglamorous: professional feel, clean composition, product clearly shown. Nobody fails at advertising creatively; everybody fails the same four ways.
We asked an AI to grade the same ads. It disagreed with us 74% of the time.
For 955 ads, both a human and a frontier vision model rated the image. They agreed on only 26% of them. The model was systematically generous: of the 689 ads it called "good", the human confirmed just 149 — and rated 269 of them outright bad.
Vision models score polish. They see clean gradients, sharp product renders, and balanced layouts and call it good. What they miss is the question a media buyer asks in the first second: would this stop anyone, and would they know what's being sold? That gap — 74% disagreement on a three-point scale — is why we still put a human in the loop on creative quality, and why we treat "an AI rated our ads highly" as close to meaningless.
The one trait 96.6% of top-rated ads share
For every AI-graded ad we also detect whether the image carries baked-in text — a headline, an offer, any deliberate typography rendered into the creative itself.
The split is the widest we measured anywhere in the corpus:
| Image rating | Ads with text baked into the image |
|---|---|
| good | 96.6% (of 914 ads) |
| fair | 67.0% (of 576 ads) |
| bad | 26.1% (of 23 ads — small sample) |
Almost every ad the model rated "good" is a designed composition with words on it, not a photo with a caption underneath. And this isn't circular reasoning from our own rubric: among ads that survived 90+ days in the wild — a quality signal our rubric has no say in — 80.2% carry baked-in text.
Meta's old "20% text rule" trained a generation of advertisers to keep images clean and let the caption do the talking. The brands spending the most money today do the opposite: the headline lives on the image, where it gets read even when the caption gets scrolled past.
The headline belongs on the image. Feed placement shows your creative before your copy — an image that doesn't state the offer is a billboard with no words.
Survivors look different: numbers in, "free" out
The median ad in our corpus runs 37 days. A quarter of them (25.6%) pass 90 days, and 12% pass 180 — evergreen workhorses that quietly absorb most of the spend.
We compared the headlines of 90-day survivors (3,347 ads) against the small cohort that vanished within two weeks of appearing (66 ads):
| Headline trait | Survived 90+ days | Died inside 14 days |
|---|---|---|
| Contains a digit | 19.9% | 10.6% |
| Contains "free" | 6.1% | 13.6% |
| Contains "!" | 6.4% | 3.0% |
| Median body copy length | 137 characters | 233 characters |
Two patterns stand out.
Numbers age well. Survivors use digits nearly twice as often. The same signal shows up in quality ratings: 36.4% of good-rated headlines contain a digit against 17.9% of bad-rated ones, and good-rated headlines use percentage figures at three times the rate. A specific claim — "25% lower TCO", "14-day rollout" — keeps earning clicks month after month.
"Free" burns fast. Ads that died young said "free" more than twice as often as survivors. A giveaway headline spikes, exhausts its audience, and gets killed. It isn't that "free" never works — 9.2% of good-rated headlines use it — it's that it works the way a promo works: briefly, on purpose. If your evergreen campaign leans on "free", the data says you're running a sprint in a marathon lane.
Survivors also carry half the body copy of the ads that died (137 vs. 233 median characters). Long persuasion essays don't survive contact with the feed.
Live today.
The full campaign — copy, images, targeting — generated for your site and deployed paused for your approval.
The boring-headline paradox
Here's the finding that surprised us most. We expected long-running ads to have the best headlines. The opposite is true: ads whose headline our model rated bad had a median lifespan of 101 days. The "good" headlines? 46 days.
| AI headline rating | Median days active |
|---|---|
| bad | 101 |
| fair | 76 |
| good | 46 |
Before you fire your copywriter, look at which headlines score bad and live forever: they're labels, not claims. Brand names. And — our favorite discovery in the whole corpus — city names. Whole families of ads titled "Billings, Montana", "North Platte, Nebraska", "Ottawa, Kansas": programmatic templates stamped out per city, running indefinitely because they're cheap, evergreen, and nobody reviews them individually.
So the paradox resolves into something more useful: lifespan measures commitment, not craft. The ads that run longest are the generic, always-on plumbing of big-brand media plans. The sharp, specific headlines belong to campaigns — they're born, they win, and they're retired on schedule. Judge an ad by what it's for. And note what "good" looks like when it does appear: good-rated headlines average 7.3 words — a full clause with an actual claim — while bad-rated ones cluster at a 4-word label.
Brands sign at the bottom, and nobody asks questions
We labeled logo placement on a nine-point grid for 183 reverse-engineered ads from the corpus. Everything you'd assume about logos being top-left is wrong:
| Position | Share of ads |
|---|---|
| Bottom row (left + center + right) | 57.4% |
| Top row | 23.5% |
| No logo at all | 18.0% |
| Dead center | 1.1% |
Among ads that show a logo anywhere, 70% put it in the bottom row — bottom-left the single most popular cell. Big brands treat the logo like a signature on a painting: the idea gets the canvas, the name gets the corner. (Nearly one in five top-tier ads skips the logo entirely and lets the product carry the branding.)
Punctuation follows the same quiet confidence. Question-mark headlines — the "Struggling with X?" formula every ads course teaches — are nearly extinct: 2.7% of good-rated headlines, 0.8% of fair, and precisely zero among bad. Exclamation marks appear in fewer than 7% of headlines in any cohort. One thing the winners do more: talk to the reader. "You" or "your" appears in roughly 16% of good- and fair-rated headlines, but only 7.6% of bad ones.
What we changed after seeing the data
This corpus isn't an academic exercise — it's the teacher set behind our creative engine. The findings above are encoded into how it works: every ad it generates bakes the headline into the image, prefers specific numeric claims where the business has them, keeps body copy short, and treats the logo as a signature rather than a header. When we tested four image models on the same brief, and when we compared LLM ad copy against human copy, the same lesson kept repeating: the craft is in the constraints, not the generator.
If you run your own ads, steal the constraints directly. Put the offer on the image in real typography. Put a number in the headline. Cut the body copy in half. Move the logo to the bottom-left. And if the headline says "free", set a calendar reminder to kill the ad in three weeks — the data says you'll be about on schedule.
Download the dataset
We released a 500-ad slice of this corpus as an open dataset: each ad rated by our human reviewer AND by the vision model, with the model's full written reasoning per ad. Free under CC BY 4.0:
- Hugging Face — ad-creative-quality-human-vs-llm
- Kaggle — Human Expert vs LLM Judge: Facebook Ad Quality
If you use it in research or a write-up, cite this report as the source.
FAQ
How many Facebook ads were analyzed in this study?
13,572 active ads collected from Meta's Ad Library between April 14 and July 7, 2026, across 2,312 advertiser pages in 11 verticals. Of those, 3,260 received a human design review, 1,643 were graded by a vision model, and 707 top performers received a full creative deconstruction.
What makes a Facebook ad "good" by this rubric?
Each ad was rated bad, fair, or good on three dimensions: headline quality, image quality, and product-message fit. The most common traits of good-rated ads: a clean composition with one focal point, the product or service clearly shown, a professional finish, and a headline making a specific claim — usually with the headline text baked into the image itself.
Should you put text on Facebook ad images?
Yes. 96.6% of ads whose image rated "good" carry baked-in text such as a headline or offer, against 67% of "fair" ads. Among ads that survived 90 or more days, 80.2% have text on the image. Meta no longer penalizes image text the way its retired "20% text rule" did.
How long does a typical Facebook ad run?
The median ad in this dataset ran 37 days. 62.6% ran at least 30 days, 25.6% passed 90 days, and 12% passed 180 days. Long-running ads skew toward evergreen brand campaigns; short-lived ads skew toward promotions and giveaways.
What is the best headline length for a Facebook ad?
Good-rated headlines in this dataset average 7.3 words (about 42 characters) — long enough to state an actual claim. Bad-rated headlines cluster around 4-word labels, most often just a brand or place name.
Where should the logo go in an ad?
The bottom row. 57.4% of labeled ads place the logo in the bottom third of the image (70% among ads that show a logo at all), with bottom-left the most common cell. 18% of top-tier ads show no logo and let the product carry the branding.
Can AI judge ad creative quality?
Not yet on its own. Across 955 ads rated by both a human reviewer and a frontier vision model, the two agreed only 26% of the time — the model over-rewards polish and misses whether an ad would actually stop a scroller. AI grading is useful as a first pass; the final call still needs human taste.
Live today.
The full campaign — copy, images, targeting — generated for your site and deployed paused for your approval.

We build AdControlCenter — AI-powered ad management for small businesses, online stores, SaaS companies and service providers. We write what we'd want to read: real numbers, no fluff, the things we wish we'd known when we started.
More from the team →



