ChatGPT looks like the advertising environment marketers have spent years asking for.
Instead of inferring intent from a three-word search, its ad system can use the current conversation as a relevance signal without giving advertisers access to private chats.

A buyer can compare software, explain constraints, reject alternatives, and reveal what matters before a sponsored placement appears. On paper, that should create unusually precise commercial intent.
The early evidence is less comfortable.
Independent studies are finding loose ad relevance, some advertisers are getting real clicks with weak post-click behavior, and agencies still report basic measurement gaps.
The useful question is no longer whether ChatGPT can become an ad channel. It already is one.
The question is whether advertisers should trust the channel enough to scale it.
The early evidence is uncomfortable
The most important early results do not all measure the same thing. That is useful because the weaknesses appear at different layers of the advertising system.
| Evidence | What happened | What it suggests | Limitation |
|---|---|---|---|
| Seer Interactive | ~55,000 ChatGPT responses and ads analyzed; 38% fully on-topic, 58% category-adjacent, 4% unrelated | Ad matching is often broader than search-style relevance | Observational relevance study, not conversion data |
| Barn2 | $477 spend, 19,154 impressions, 370 clicks, 0 conversions | Real paid traffic can still be commercially weak | One advertiser, four products, six weeks |
| Barn2 organic comparison | Organic ChatGPT traffic showed 443 sec avg. time on site and 5 sales; paid traffic showed 37 sec and 0 sales | The issue may lie in how paid ads are matched to intent | One business and niche products |
| Digiday / Accuracast | Lead forms were submitted while Ads Manager still showed zero conversions two weeks later | Measurement can fail even when campaign outcomes occur | Agency example, not industry benchmark |
None of these examples proves that ChatGPT Ads are universally ineffective.
They do show three separate failure modes that are appearing often enough to investigate: relevance, traffic quality, and measurement.
| Early evidence is mixed. The correct response is investigation, not a universal verdict. |
|---|
Unlock insights that drive growth
Why ChatGPT Ads looked inevitable
Search advertising became valuable because a query exposes intent. ChatGPT potentially has something richer: the question, the follow-up, the constraints, the objections, and the decision criteria in one session.

A conversation about choosing accounting software can reveal company size, industry, workflow requirements, integration needs, and budget sensitivity before the user ever clicks a commercial result.
OpenAI built the ad system around that premise.
Its current advertiser documentation says delivery can consider conversation context and intent, landing pages, ad copy, advertiser context hints, and permitted personalization signals.
The platform also moved quickly beyond early reach buying. Ads Manager now supports Views, Clicks, and Conversions objectives, with CPM, CPC, and conversion-optimized buying depending on campaign eligibility.
| ChatGPT does not suffer from a shortage of intent signals. The question is whether its ad system knows how to use them yet. |
|---|
Problem 1: Relevant chat, irrelevant ad?
The first mistake is assuming that because ChatGPT understands a conversation, its ad system must automatically understand which advertiser belongs beside that conversation.
Those are different problems. The model generates an answer.
The advertising system still has to identify eligible advertisers, interpret their context hints, score relevance, and run an auction without degrading the user experience.
Seer Interactive analyzed roughly 55,000 ChatGPT responses and ads and found that only 38% were fully on-topic, while 58% were category-adjacent and 4% had no meaningful connection.
Category-adjacent does not automatically mean useless. It can work for broad awareness. It becomes more concerning when advertisers expect high-intent performance comparable with mature search advertising.
Seer also reported improvement from earlier observations, which matters. The relevance problem appears to be changing rather than fixed, but precise alignment is still inconsistent.
| The problem may not be whether ChatGPT understands the conversation. It may be whether the ad auction has the right advertiser for that conversation. |
|---|
Context hints are not keywords
Part of the confusion comes from language that sounds familiar to performance marketers.
OpenAI lets advertisers provide context hints describing the conversations, topics, or situations where an offer may be useful.
But OpenAI explicitly says these hints are not exact-match keywords and do not guarantee delivery in specific conversations.
They are inputs to relevance, not a way to buy a specific query.
| Advertiser expectation | What context hints actually do |
|---|---|
| Target this exact query | Describe situations where the offer may be relevant |
| Only show beside this topic | Guide matching without guaranteeing placement |
| Behave like search keywords | Act as one relevance input among several |
| Define the full audience | Add context about product, customer need, or use case |
Google trained advertisers to think keyword, intent, auction. ChatGPT Ads asks them to think conversation, inferred context, eligible advertisers, relevance, then auction. That is a different operating model.
Intent is not always purchase intent
There is another assumption worth challenging. A detailed conversation can reveal strong intent without revealing that the user is ready to buy.
Someone can spend ten minutes comparing project-management tools because they are writing a report, helping a colleague, or learning the category.
The conversation is rich, but the commercial timing is weak.
Search advertising often benefits from explicit transactional language such as pricing, demo, alternative, discount, or buy. ChatGPT conversations can contain far more context while still remaining exploratory.
That creates a new matching challenge.
The ad system has to distinguish between topic relevance and purchase readiness, then decide whether a sponsored placement adds value at that point in the conversation.
This may help explain why a category-adjacent ad can feel reasonable to the system yet produce poor economics for an advertiser.
The topic can be right while the buying moment is wrong.
| Rich conversational context is not the same as high commercial readiness. |
|---|
See what's working. Fix what's not. Grow faster.
Problem 2: Real clicks, weak traffic
Relevance only matters if it changes behavior. OpenAI can deliver impressions and clicks. The harder question is whether those clicks represent people who are genuinely likely to buy.
Barn2 provides the clearest public advertiser test so far. Over six weeks, the company spent $477 and recorded 19,154 impressions, 370 clicks, and zero conversions across four products.
The important detail is that the clicks were not imaginary. Barn2 checked Google Analytics and found 401 sessions from 374 people, very close to the volume OpenAI reported.
The traffic existed, but neither Barn2 site recorded a sale from the paid ChatGPT cohort during the campaign.
| The traffic existed. The value did not. |
|---|
Organic ChatGPT makes the gap stranger
Barn2 had a useful comparison hiding in plain sight: the company was already receiving organic referral traffic from links inside normal ChatGPT answers during the same six-week period.
| Metric | Organic ChatGPT | Paid ChatGPT Ads |
|---|---|---|
| Sessions | 552 | 210 |
| Average time on site | 443 sec | 37 sec |
| Pages per session | 2.18 | 1.09 |
| Engaged visits | 60.9% | 20.5% |
| Sales | 5 | 0 |
Organic ChatGPT visitors stayed almost 12 times longer, viewed more pages, engaged at roughly three times the rate, and generated five sales. Paid visitors generated none.
That makes the simplistic conclusion, “ChatGPT users do not buy,” difficult to defend. The same business was already converting visitors sent organically by ChatGPT.
The more specific question is whether the paid system was placing Barn2 beside the right conversations.
Barn2 sells niche WordPress and WooCommerce products, where broad category matching can become expensive quickly.
| The early problem may not be ChatGPT traffic. It may be paid ChatGPT traffic. |
|---|
One campaign is not a benchmark
Barn2 is unusually transparent about its limitation: one advertiser, four products, six weeks, and 370 clicks are not an industry benchmark.
A broad consumer product could behave very differently from a niche B2B or software offer. Campaign structure, creative, landing pages, geographies, and conversion windows can also change the outcome.
That is why the Barn2 result should not be used to declare the channel dead.
It should be used to identify a hypothesis worth testing: targeting breadth may be especially risky for narrow customer profiles.
The Seer relevance data makes that hypothesis harder to dismiss because it identifies a similar issue at a much larger observational scale, even though it does not measure conversions.
Problem 3: Measurement is still catching up
The third weakness is more subtle because it can make a functioning campaign look broken.
Advertisers need to know whether poor performance is real or whether the platform failed to observe the conversion.
Digiday reported that Accuracast ran a UK lead-generation campaign on ChatGPT alone and watched form submissions arrive, yet Ads Manager still showed zero conversions two weeks later.
That is a different failure from weak targeting. The campaign may have created value, but the platform could not reliably prove that value inside its own reporting.
| Poor advertising performance and poor advertising measurement can look identical inside a dashboard. |
|---|
The problem compounds. Weak attribution confidence encourages smaller budgets. Smaller budgets produce less conversion data. Less data slows optimization, which keeps performance uncertain and budgets cautious.
The billion-dollar paradox
This is where the ChatGPT Ads story becomes more interesting than a simple performance review.
The platform can succeed commercially for OpenAI before its ad performance becomes consistently predictable for advertisers.

OpenAI reported that ChatGPT Ads reached a $1 billion annualized revenue run rate in under 200 days and attracted tens of thousands of advertisers.
By late September, the product had expanded to more than 60 countries.
That proves advertiser demand and distribution. It does not prove advertiser return on investment.
Novelty, first-mover experimentation, brand budgets, and fear of missing a new channel can all drive spend before benchmarks stabilize.
| Platform revenue measures advertiser demand. It does not measure advertiser ROI. |
|---|
Why this does not mean ChatGPT Ads are doomed
The strongest argument against writing off the channel is how quickly the product is changing. Several limitations visible in early campaigns are exactly the areas OpenAI is now expanding.
Conversions campaigns can optimize toward supported downstream events.
Eligible advertisers can use conversion-optimized CPC or CPM buying.
OpenAI Pixel and Conversions API can be used together for more resilient measurement.
Ads Manager can use attribution windows, click references, advanced matching, and modeled measurement where available.
Static tracking parameters such as UTMs persist on ad clicks for independent analytics.
OpenAI has added Sponsored Agents plus HubSpot and Shopify integrations as the ad product expands.
The correct conclusion is not that ChatGPT Ads failed. It is that the platform is evolving much faster than reliable cross-industry performance benchmarks can form.
OpenAI also reports positive advertiser outcomes.
Its August milestone post cited an ecommerce advertiser achieving 3x ROAS over 28 days.
It also cited a technology partner reporting that more than 80% of ad-driven traffic came from new customers.
Those examples are useful counterweights to the negative tests, but they are first-party examples rather than independent benchmarks. The fairest reading is that performance appears uneven, not uniformly poor.
What ChatGPT Ads offer today
The current product is already more sophisticated than the earliest beta. A concise view helps separate outdated criticism from the problems that still matter.
| Area | Current capability |
|---|---|
| Placement | Sponsored placements shown separately below ChatGPT responses |
| Matching | Conversation intent plus ad, landing-page, context, and permitted personalization signals |
| Context hints | Natural-language guidance, not exact-match keywords |
| Objectives | Views, Clicks, and Conversions |
| Buying | CPM, CPC, and eligible conversion-optimized buying |
| Measurement | Pixel, Conversions API, attribution windows, click reference, advanced matching |
| Reporting | Impressions, clicks, spend, CTR, average CPC, average CPM, conversions |
| URL tracking | Static UTM-style parameters supported on landing pages |
| Availability | More than 60 countries as of late September 2026 |
This matters because early campaign results should be interpreted against the product version that actually ran them.
A campaign from June may not reflect the same measurement stack available in October.
What ChatGPT Ads cost
There is no universal ChatGPT Ads price. Cost depends on objective, bid, auction competition, relevance, geography, and delivery conditions.
For Clicks campaigns, OpenAI currently recommends a starting maximum CPC bid of $3 to $5.
That is a bid recommendation, not a promise that every click will cost $3 to $5.
The more useful economic question is whether a paid visit becomes a customer. A cheap click can still be expensive if the audience is poorly matched.
| Effective CAC = ChatGPT Ads spend ÷ customers acquired |
|---|
What advertisers should test instead
The early evidence suggests advertisers should treat ChatGPT Ads as a structured learning program, not as a mature budget line that deserves immediate scale.

Measure relevance before conversion. Check whether the traffic looks like the audience the campaign was designed to reach.
Compare paid vs. organic ChatGPT traffic. Large quality gaps can reveal a paid matching problem rather than weak ChatGPT users.
Judge post-click behavior before scaling. Engagement, landing-page progression, and meaningful events can expose low-quality traffic early.
Validate platform conversions independently. Ads Manager, website analytics, CRM data, and commerce data should be reconciled rather than assumed identical.
Keep early budgets experimental. Scale only after the channel produces enough downstream evidence to support the economics.
| Test the channel. Do not yet trust the channel. |
|---|
A credible test needs downstream metrics
CTR can look healthy while the campaign fails commercially. A credible test follows the journey far enough to see whether the click turns into value.
| Layer | Metrics |
|---|---|
| Delivery | Impressions, spend, CPM |
| Interest | Clicks, CTR, CPC |
| Visit quality | Engagement, pages per session, meaningful events |
| Conversion | Signup, lead, demo, purchase |
| Qualification | MQL, SQL, opportunity |
| Customer | Paid customer, order value, activation |
| Economics | CAC, revenue, payback, revenue per visitor |
| Attribution | First touch, assists, final touch, model-assigned revenue |
A campaign that wins on CTR but loses on customer acquisition cost is not a performance success. The platform metric is only one layer of the decision.
Use OpenAI and independent analytics together
OpenAI Ads Manager should remain the source for platform delivery: impressions, clicks, spend, bids, and OpenAI-attributed conversions.
Independent analytics should show landing-page behavior, return visits, assisting channels, and whether leads become customers.
The CRM or commerce system should then confirm the commercial outcome: qualified opportunity, Closed Won, purchase, refund, renewal, or another value event.
When several channels appear before that outcome, multi touch attribution can help distinguish an early ChatGPT Ads interaction from the final converting touch.
| No single dashboard should be allowed to grade its own homework. |
|---|
UTMs matter more on an immature platform
OpenAI supports static tracking parameters on ChatGPT Ads landing pages and says those parameters persist on clicks. That makes independent validation possible without waiting for every native integration.
Consistent UTM parameters identify the source, paid medium, campaign, and creative variation. Their naming must remain consistent across the ad platform, analytics, and CRM.
For example, source can identify ChatGPT, medium can distinguish paid traffic, campaign can identify the initiative, and content can separate creative variants when the workflow supports it.
Do not use UTMs as proof of causality. They preserve campaign context. They do not solve cross-device identity, zero-click influence, or attribution-model disagreements.
Where Usermaven fits today
Usermaven does not currently provide a native ChatGPT Ads integration that automatically imports OpenAI impressions, spend, or clicks.
Its documented paid-ad connections currently cover Google, Meta, LinkedIn, and Microsoft Ads.

Usermaven becomes relevant after an identifiable ChatGPT Ads visitor reaches the website. Tagged campaign traffic can be analyzed alongside behavior, conversion events, journeys, and downstream commercial outcomes.
A paid ads attribution workflow compares platform claims with independent conversions and revenue, rather than treating the ad dashboard as the final source of truth.
Website interactions can be evaluated through website analytics, while Funnels can show whether ChatGPT Ads visitors progress from landing page to signup, demo, activation, or purchase.
Customer journeys become useful when the ad starts a journey that later continues through organic search, direct return, email, or sales activity.
Configured Events help connect those journeys with actions such as qualified leads, purchases, and other conversion milestones.
If the business outcome happens later, revenue attribution can preserve the acquisition context while the customer moves through multiple sessions and touchpoints.
The Measurement Trust Center can flag collection, identity, and integration gaps before teams rely on the downstream numbers.
The distinction is simple: Ads Manager explains what happened to the ad. Independent analytics should explain what happened to the customer.
Should marketers test ChatGPT Ads now?
Yes, but the answer depends on what the business is willing to learn and how much uncertainty it can tolerate.
Strong candidates for testing
Broad enough addressable markets that can tolerate imperfect matching.
Products where customers research, compare, and explain needs before buying.
Clear landing pages and conversion events that make traffic quality easy to judge.
Independent analytics and CRM or commerce data for validation.
Enough budget and conversion volume to learn without requiring immediate efficiency.
Use more caution when
The ideal customer profile is very narrow or technically specialized.
Clicks are expensive relative to customer value.
Conversion volume is too low to distinguish signal from noise.
The business lacks independent post-click and revenue measurement.
Performance targets leave little room for experimental spend.
Barn2 is most relevant to the second group. Its experience should not be generalized to every category, but niche advertisers should pay close attention to the mismatch risk.
Final verdict
ChatGPT Ads may eventually become one of the most important performance channels on the web.
The ingredients are unusually strong: massive usage, rich conversational intent, and rapidly improving advertising infrastructure.
But those ingredients are not the same as proven performance.
Early evidence still points to a gap between understanding a conversation and reliably turning that understanding into precise ad matching, valuable traffic, and trusted measurement.
The sensible position today is neither hype nor dismissal. ChatGPT Ads deserve structured testing, independent validation, and patience. They have not yet earned blind scaling.
The first budget should therefore buy knowledge as much as traffic: which conversations produce qualified visits, which conversions can be independently verified, and which customer outcomes justify more spend.
Book a demo to see how Usermaven can connect identifiable paid acquisition visits with customer journeys, conversions, pipeline, and revenue.
FAQs
1. What are ChatGPT Ads?
ChatGPT Ads are sponsored placements shown separately from ChatGPT responses. OpenAI uses conversation context and other relevance signals to determine when an eligible advertisement may be useful.
2. Do ChatGPT Ads work?
Early evidence is mixed. Some advertisers have reported weak conversion quality, while OpenAI says other advertisers are achieving positive returns. There is not yet enough independent cross-industry data for a universal verdict.
3. Why are some ChatGPT Ads underperforming?
The strongest current concerns are loose relevance, low-quality post-click traffic for some advertisers, and measurement gaps that can make campaign performance difficult to verify.
4. How much do ChatGPT Ads cost?
There is no fixed universal price. OpenAI currently recommends a starting maximum CPC bid of $3 to $5 for Clicks campaigns, but actual costs depend on auction conditions, relevance, geography, and competition.
5. How does ChatGPT Ads targeting work?
Delivery can consider conversation context and intent, the landing page, ad copy, advertiser context hints, and permitted personalization signals. Context hints guide relevance but do not guarantee a specific placement.
6. Are ChatGPT Ads based on keywords?
Not in the same way as traditional search ads. Advertisers can provide context hints that may mention topics or keywords, but OpenAI says those hints are not exact-match keywords.
7. How are ChatGPT Ads conversions measured?
Advertisers can use the OpenAI Pixel, Conversions API, or both. OpenAI evaluates conversion events against the campaign configuration, attribution window, click reference, and other eligible measurement signals.

Written by
Junaid Ahmed
Content Writer & Digital Marketer
Junaid Ahmed is a content and copywriter with 3+ years of experience creating research-driven content across SaaS, B2B, ecommerce, and digital marketing. He specializes in turning complex topics into clear, practical content that helps marketers better understand their challenges, evaluate solutions, and make informed decisions. His work spans educational content, industry insights, and actionable marketing guides.
All articles by Junaid →
