Field data · B2B HR services · 4,646 prompt checks
ChatGPT ads went from 0 to 90% of our prompts in six weeks
The same 101 B2B HR services prompts, run every day from June 1 through July 16, 2026. The switch didn’t flip once. It flipped, flipped back, and then held.
On June 1 we started running the same 101 prompts through ChatGPT every day. All of them B2B HR services queries — the questions our client’s buyers actually ask. Not a rotating sample. Not a fresh set each week. The identical 101, 46 days straight, logging one thing: did the response come back with a sponsored placement attached to it?
For the first eleven days, the answer was no. Every time. That’s 1,111 checks and zero ads.
On July 16, 85 of those same 101 prompts returned an ad.
The endpoints aren’t the interesting part. The shape in between is, because it isn’t a ramp. Inventory got switched on, yanked back, switched on harder, yanked back again, and then — starting July 1 — it stayed on.
Here is every check we logged. One column per day, one mark per prompt. Filled means an ad came back.
Prompts returning a sponsored placement, by day
← Scroll to see all 46 days →
Three phases, not one ramp
Split the 46 days at the two inflection points and the rollout reads as three separate systems, not one curve.
Phase 1 · Jun 1–18
0.4%
avg of 101 prompts/day
7 ads in 1,818 checks
ads on 4 of 18 days
Dark, with a flicker
Statistically indistinguishable from nothing.
Phase 2 · Jun 19–30
22.9%
avg of 101 prompts/day
range 5.9% – 92.1%
std dev 25.0 points
The tuning window
Ads every single day, but the daily number is all over the map.
Phase 3 · Jul 1–16
88.1%
avg of 101 prompts/day
range 79.2% – 95.0%
std dev 4.3 points
Plateau
Sixteen days, never below 80 prompts. The variance collapsed.
Phase 1: eleven days of nothing, then a flicker
June 1 through June 11: zero. Then June 12 returned two sponsored prompts, 2.0%. Back to zero for three days. Two on the 16th. One on the 17th. Two on the 18th.
Seven ads in 1,818 checks. Watching a dashboard, you’d have called it noise and gone back to work. That’s the shape of a test, not a rollout — a handful of prompts switched on and off inside a corpus where nothing else was moving.
Phase 2: the tuning window
June 19 is the first real break: 12 prompts, 11.9%. From that day the count never returns to zero. It doesn’t climb either. It thrashes — 12, 7, 9, 7, 8 — then 26 on the 24th and 20 on the 25th.
Then June 26: 93 of 101. That’s 92.1% in a single day, up from 19.8% the day before.
And then it came apart. 47 on the 27th. 12 on the 28th. Six on the 29th. Three days after touching near-full inventory, we were back under 6% — lower than June 19.
A rollout curve doesn’t go 93 to 6 in three days. A throttle does.
That’s the most useful thing in the dataset. Someone turned inventory all the way up on a Friday, watched the weekend, and turned it back down. June 30 recovered to 31. Then July 1 hit 95 and never let go.
June 26 wasn’t just us
We’re one prompt set, so the fair question is whether we measured ChatGPT or measured our own corpus. Independent rollout tracking published by Cloro describes the same June shape from a completely different angle: a mid-June collapse to well under 1% of responses, a re-ramp starting June 26, and a plateau around half of US responses in the week ending July 3.
Different corpus, different method, same inflection date. That’s a rollout event, not a quirk of our prompts.
The gap between their number and ours is the part worth sitting with. Roughly 51% of US responses against 88% of our prompts isn’t a contradiction. It’s a selection effect. Their corpus is everything anyone asks ChatGPT — recipes, code, homework, therapy. Ours is 101 prompts in a category where every question is a purchase in progress.
Ad coverage isn’t spread evenly across ChatGPT. It concentrates where buying decisions get made, which is why a general-purpose corpus reads at half our rate. If your customers ask commercial questions, you’re on our side of that gap, not theirs.
Phase 3: the plateau
July 1: 95 of 101. July 2: 96, the high-water mark of the study at 95.0%. Sixteen consecutive days, never once below 80. Average 88.9 prompts a day.
The average isn’t the number that matters. The standard deviation is. Phase 2 ran at 25.0 points. Phase 3 runs at 4.3. When a system stops swinging and starts holding a narrow band, it has left testing and entered production.
So don’t read the July 6 dip to 79.2% as a pullback. Inside a plateau this tight, an 80–96 band is fill rate: which advertisers had budget that day, which prompts fell outside category eligibility, which responses didn’t clear the bar. Inventory noise, not a trend.
Sixteen days is a short window, and the obvious caveat applies: this plateau has held for two weeks, not two quarters. June proved the throttle moves.
The same story, weekly
Daily counts are noisy. Weeks aren’t. Share of the 4,646 checks that returned an ad, by calendar week:
Hatched bar = partial week (Jul 13–16 only)
What this changes for you
-
Your answer has a neighbor now
AEO work earns you a mention inside the response. As of July, something sits underneath that response nine times out of ten — and it can be a competitor with a budget and no organic presence at all. Getting cited is no longer the end of the page.
-
Any plan built on “ads are rare in ChatGPT” expired on July 1
If your planning treated sponsored placements as a curiosity to deal with later, that assumption now has a date stamp on it. On commercial prompts, near-total coverage is the current baseline.
-
88% is our corpus, not a universal number
Coverage tracks commercial intent, so your category has its own number. This is what B2B HR services looked like in July 2026. Independent tracking put the US-wide figure near half of all responses over the same window — well below our set, and still enormous.
-
Baseline now, because the throttle can move again
June 26 to June 29 went 93 to 6 in three days. Start measuring after the next swing and you won’t know whether what you’re seeing is normal. A daily log is cheap. Reconstructing one after the fact is impossible.
Every day, June 1 – July 16
Every daily figure, denominator fixed at 101. Marked rows are the two inflection days.
| Date | Ads | % of 101 |
|---|
How we measured this
- The set
- 101 B2B HR services prompts, unchanged for the entire window, checked once per day from June 1 to July 16, 2026. 46 days × 101 prompts = 4,646 individual checks.
- The source
- Daily prompt runs via Otterly.ai. We’re an Otterly agency partner and use the platform for client AI visibility monitoring — disclosing that rather than burying it. The data is what the tool returned. The reading of it is ours.
- The metric
- Binary, per prompt, per day: did the response carry at least one sponsored placement? We did not record ad count, ad position, or advertiser. A prompt with three ads counts the same as a prompt with one.
- The denominator
- Fixed at 101 every day, on purpose. Holding it constant keeps every day comparable to every other day and never lets a thin day inflate a percentage.
- Limits worth stating
- One prompt set, one category, one checker. These are coverage rates for our corpus — the share of our B2B HR services prompts that returned an ad — not ad frequency for any given ChatGPT user, and not a claim about ChatGPT overall. Ad eligibility varies by account tier, geography, and topic category, and OpenAI excludes sensitive categories from ads entirely. Read the 88% as a category signal, not a platform statistic.
What to do about it
Find out what ChatGPT returns for your prompts
We built this dataset the same way we build them for clients: pick the prompts your buyers actually type, run them daily, and watch what happens around your answer. If you’d rather know your number than guess at it, that’s the work.
Book an AEO CallWe’ll run your prompts, not ours, and show you the coverage rate for your category.