Red flags when hiring an AI search agency
The category is new enough that almost nobody buying it has bought it before. Twelve ways these agencies fail, and the question that exposes each one.
Agencies in this category fail in a small number of repeatable ways, and most of them are audible on a first call. The sales-process tells are guarantees, an agency that will not name the people who will do the work, and speed promised on competitive commercial prompts; Google and OpenAI both state in their own public documentation that placement is not guaranteed. The proof tells are a portfolio with no losses and no context, Domain Rating used as the evidence, and renewal or cancellation terms you have to go hunting for. The delivery tells are a content plan with nothing happening on your own domain, and setup treated as a finished deliverable. The measurement tells are mentions, visibility, citations and share of voice used as one number, and a monthly count with no baseline beside it. For every claim, ask for a number, a name and a date.
On this page
You are about to pay someone to get your company named inside ChatGPT, Perplexity, Gemini and Google AI Overviews. The category is new enough that almost nobody buying it has bought it before, so the usual defense (you have hired three of these, you know the tells) does not apply. The decks are good. The evidence behind them is often thin.
Agencies in this category fail in a small number of repeatable ways, and most of those ways are audible on a first call. Twelve of them are below, each with the question that exposes it.
The 12 red flags, and the question that exposes each one
| # | Red flag | Why it predicts failure in AI search | Verification question to ask |
|---|---|---|---|
| 1 | Guaranteed citations, guaranteed rankings, guaranteed ROI | Google's documentation says indexing and serving are not guaranteed. OpenAI's says placement is not guaranteed. No vendor sits upstream of the platforms. | "Google and OpenAI both say placement isn't guaranteed. So what exactly are you guaranteeing, and what happens to my money if it doesn't happen?" |
| 2 | Fast wins promised on competitive, buyer-intent prompts | AI Overviews cluster on low-volume informational queries, and a top-10 organic ranking does not secure inclusion. Speed on commercial prompts is the hardest version of the job. | "Name the first five prompts you targeted for your most recent SaaS client, and the volume and intent behind each." |
| 3 | No named humans attached to your account | Prompt selection, entity language and answer-block structure are judgment calls that do not survive being templated or passed down a production chain. | "Send me the LinkedIn profiles of the three people who will do the work, and tell me how many other accounts each is on right now." |
| 4 | A portfolio of wins with no losses and no context | The FTC's endorsement guides call for disclosing what results are generally expected when a featured result is exceptional. A results wall with no misses is the curation pattern the FTC flags on reviews. | "Show me an engagement that underperformed, what you thought went wrong, and what you changed after it." |
| 5 | Domain Rating or a logo wall used as the proof | Check the DR of the URLs actually cited on these prompts and the range is wide, and Google rank does not carry over to AI citation. Authority is a proxy, not the mechanism. | "Forget DR. What's the citation rate on the pages you built for your last three clients, and can you show me the export?" |
| 6 | Renewal and cancellation terms you have to go looking for | With cited domains reshuffling 40% to 60% month over month, a long notice period closes your exit before the first honest read on the work arrives. | "What's the notice period, how do I cancel, and can you point me at the clause that says so?" |
| 7 | A content plan with nothing happening on your own domain | Roughly 57% of AI citations resolve to brand-owned sites. A plan that never touches your domain skips the surface most citations come from. | "How many pages on my site will you have edited by day 30, and which ones?" |
| 8 | Off-site placements sold as the engine of citation share | For SaaS and software, the median share of citations coming from earned media is 11.4%, against 59% for pharma and biotech. Placements amplify. They are not the base. | "What split do you expect between citations from my own site and citations from placements, and why that split for us?" |
| 9 | Setup treated as a finished deliverable | Half of the content answer engines cite is under 13 weeks old, and 40% to 60% of cited domains change month over month. A one-time build decays on a schedule. | "Show me a page you refreshed for another client: the before, the after, and what triggered it." |
| 10 | Mentions, visibility, citations and share of voice used interchangeably | The four measure different things. The blur lets a flat month be reported as a good one. | "What do you mean by visibility, share of voice and citation rate, and which one of them goes in my monthly report?" |
| 11 | One monthly snapshot, no baseline, no competitor set | Cited domains reshuffle 40% to 60% month over month, so a single month's count cannot be separated from ordinary platform churn. | "Show me a real client report with the same metric across six months, with the competitor set next to it, redacted if you need to." |
| 12 | Reporting that never reaches revenue | Visibility and share of voice are upstream signals. A program can win both and produce no pipeline. | "Which tool of mine would that number come from, and have you wired that up for a client before?" |
Sales-process flags: what the pitch gives away
1. Anything guaranteed
Google's Search Central documentation is blunt about it: "Indexing and serving isn't guaranteed," and there are "no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary." OpenAI says the same thing about ChatGPT search: "Placement is not guaranteed." Both companies publish this in their own help documentation, for free, in plain English.
An agency guaranteeing a citation is therefore claiming a lever that the platform itself says does not exist. That claim also has a legal problem. The FTC's Operation AI Comply sweep, announced in September 2024, brought five actions against companies using AI branding to make claims they could not back. Lina Khan, the FTC chair at the time, put it: "there is no AI exemption from the laws on the books." DoNotPay settled for $193,000 over a complaint alleging it had marketed a "robot lawyer" it never tested against a human lawyer's work. The FTC's 2024 complaint alleged that FBA Machine cost consumers more than $15.9 million selling AI-powered tools with "guaranteed income." The FTC's advertising substantiation policy requires an objective claim to be substantiated before it is made. The evidence has to exist at the moment of the claim. A folder assembled after your complaint arrives is too late.
This is the one I would walk out on. Everything else on the list is a fixable process problem. A guarantee is a statement about what the vendor believes it can get away with.
A good answer: what they commit to instead of a guarantee, and what happens to your money if they miss it.
2. Speed promised on the hardest prompts
A first AI citation genuinely can land inside 24 hours. The conditions are narrow: a brand with at least modest existing authority, a prompt with no entrenched citation winners, a page structured so the answer is easy to lift, and fast indexing. Google's general guidance for a recrawl is wider, anywhere from several days to several months. Both statements are true at the same time. The problem is an agency that quotes the fast case as the normal case and prices the rare outcome as the default. What a real first 90 days looks like is a useful sanity check against whatever they hand you.
Watch what they promise speed on. Semrush's study of 200,000 AI Overview keywords found that 82% of desktop AI Overviews and 76% of mobile ones appeared on keywords with under 1,000 monthly searches, that 80% of desktop ones targeted informational intent, and that transactional keywords made up under 3% of AI Overview triggers. The study also found that top-10 organic rankings did not ensure inclusion. Fast wins on low-competition informational prompts are real and worth having. Fast wins on the prompts your buyers actually type are a different job, and anyone promising them in month one is guessing.
A good answer: which prompts are winnable for you now, which are a year out, and why.
3. No names attached to your account
The people in the pitch may not be the people who do the work. In this category that costs more than it does elsewhere. The work that decides citation outcomes is judgment work: which prompts to target, what entity language to use, how to structure an answer block a model will lift, what schema goes under it. Judgment does not survive being templated or handed down a production chain.
So ask early: who will touch my account, what does each person own, and is any of it subcontracted? Then ask the one a reseller cannot answer without a discovery call it has never had: which prompts would you target for us, and who writes the sentences the model is meant to lift? "Our team" is not an answer, and neither is a promise to introduce you after signature. If the question is treated as rude, that is the reply.
A good answer: names, roles, a straight yes or no on subcontracting, and how many other accounts each person carries right now.
Proof and contract flags: what the evidence gives away
4. Wins only
Every agency's case-study page is a highlight reel. That is fine until it is offered as the expected outcome. The FTC's endorsement guides say an endorsement "can't be used to make a claim the marketer of the product couldn't legally make," and that where a featured result is exceptional, the advertiser must disclose what results are generally expected unless it has proof the featured one is typical. The same guides treat curated reviews as a deception pattern, and they name two versions of it: delaying only negative reviews "could create a biased picture," while displaying five-star reviews first regardless of date leaves a consumer with "a misleading picture of what users think." Verifying an agency's results before you hire goes through what to check line by line.
So ask for the miss. Not as a trap. An agency that can describe an engagement that underperformed, name what it misread, and say what changed in the method afterwards is an agency with a method. One that has never had a bad quarter in a category this young is telling you it does not keep records.
A good answer: a miss they can describe in detail, and the method change that came out of it.
5. Domain Rating as the proof
DR is a handicap. It describes a starting line. Run one of the prompts you care about, then check the DR of every URL the engines cite back: the range is wide, and the biggest publishers in it do not dominate. The Semrush data points the same way: a top-10 Google position did not ensure AI Overview inclusion across 200,000 keywords. Citation share tracks structure, freshness and where the model looks for sources rather than accumulated authority.
This matters commercially. If an agency's pitch rests on its own DR or its client logos, it is selling you a proxy it cannot transfer to you anyway. Ask what happened to the pages it built: were they retrieved, were they cited, on which engines, over what window. An agency doing real generative engine optimization has those numbers per page. An agency doing classic link building with a new name on it has a DR chart. Handed a deck of logos, that is the question I would spend first.
A good answer: per-page citation evidence with the engine and the window named, and the export behind it.
6. Terms you have to go hunting for
Auto-renewal, a notice period measured in quarters and a cancellation that runs through a phone call decide whether you can act on everything else on this list. In this category that has a specific cost. With cited domains reshuffling 40% to 60% month over month, it takes months of trend before anyone can honestly say whether the work moved anything, and a twelve-month term with a long notice period closes your exit before that read arrives.
The regulatory backdrop here is unsettled. The FTC's click-to-cancel rule, announced in October 2024, is not something you can quote at anyone today. The Commission's own Negative Option Rule page shows it has since revised the rule to conform to federal court decisions (February 2026) and opened a new advance notice of proposed rulemaking (March 2026). The principle it set out, that cancelling should be as easy as signing up, is still a reasonable thing to ask for, and the Commission cited nearly 70 consumer complaints a day about recurring-billing practices in 2024, up from 42 a day in 2021. The Section 5 prohibition on deceptive and unfair practices has not moved.
What you want is boring and specific: the notice period in writing, the renewal mechanics in writing, and an exit that does not depend on catching someone on the phone. Then one more, because the asset that compounds here is your own citation history: who owns the tracked prompt set, the monitoring seat and the baseline when the engagement ends? An agency holding that seat keeps your measurement history when you walk. Published pricing is a reasonable first signal, and the contract is its own separate check. How AEO agency pricing actually works covers what the money buys at each band.
A good answer: the notice period in writing, and your prompt set, your monitoring seat and your baseline leaving with you.
Delivery flags: what the work plan gives away
7. Nothing happening on your own domain
Profound's analysis of 11.84 billion citations found that roughly 57% of AI citations resolve to brand-owned sites, with the mix varying sharply by engine: ChatGPT leans on brand sites least at 47%, Gemini most at 69%. Your own domain is the single largest citation surface you have, and it is the only one you control.
So a plan that is all outreach, all social, all "thought leadership placements" and no work on your site is a plan that skips the main thing. Look for the unglamorous items: schema, canonical tags, answer blocks that can be lifted without rewriting, entity language that is consistent across your pages, product and solution pages that actually state what you do. This is the part of SEO and AEO work that produces no screenshot for a monthly report, which is exactly why it gets dropped.
A good answer: a named list of the pages on your domain they will have changed, with dates.
8. Placements sold as the engine
Off-site placements are a minority of your citation volume with outsized effect per placement. A single strong placement on the right property can move share of answer within days. That is real, and it is worth paying for.
It is not the base of the program, and for your vertical the data is unusually clear. In the same Profound study, SaaS and software companies see a median of 11.4% of citations coming from earned media. Pharma and biotech sit near 59%, because compliance friction pushes that industry's content off its own sites. Most SaaS does not have that problem. If an agency's plan is 80% placements, ask what it thinks your citation mix will look like at the end and why it differs from the median for your category.
A good answer: a number for the split and a reason tied to your company. A bad one repeats the word authority.
9. Setup treated as done
Profound measured citation drift by running roughly 80,000 prompts per platform twice, a month apart, and comparing which domains got cited: 59.3% of cited domains changed on Google AI Overviews, 54.1% on ChatGPT, 53.4% on Microsoft Copilot, 40.5% on Perplexity. The same comparison run from January to July shows drift climbing to 70% to 90%. For reference, Google ships four to five core algorithm updates a year. Classic search is the stable one.
Profound's citation-decay work separately found that half of the content answer engines cite is less than 13 weeks old, and tracks each cited URL through a lifecycle: first cited, rise to peak, then a half-life measured in days from peak.
A build that ends with a handover is a product that starts losing value the week it ships. The refresh discipline is the deliverable, and it is also the easiest line on an invoice to fill without doing. So make it checkable. A real refresh names the page, the trigger and the change: this page lost its cited slot on Perplexity, so the direct-answer block was rewritten and the numbers updated the same week. Ask what gets re-examined monthly and what triggers a rewrite. Then ask who owns that work after month three. Month six is when a client finds out either way, which is why I would want all of that in writing before signing. An organic growth program with no answer for month four is a project wearing a retainer's clothes.
A good answer: a page they refreshed for another client, the trigger that started it, and the before and after.
Measurement flags: what the reporting gives away
10. Four words used as one
Peec's own documentation separates them cleanly. Visibility is the percentage of AI responses where your brand appears. Share of voice is your share of all tracked-brand mentions in those responses. Position is your average rank when you do appear. Then there is a second family entirely, at the source level: retrieved (did a URL from your domain show up as a source), retrieval rate, and citation rate (was that source actually referenced in the visible answer, or just read in the background).
The split that catches people: brand visibility and source visibility are different. You can be cited as a source without your brand being named, and named as a brand without your site being used. An agency that cannot draw that distinction on a call will not draw it in your report either, and a report that quietly swaps the generous metric in when the strict one goes flat is a sales document.
A good answer: one definition per metric, said the same way twice, and the one that goes in your monthly report.
11. A snapshot with nothing to compare it to
No trade body sets a format for these reports, so judge one against the churn instead. Given 40% to 60% domain churn month over month, "we got you 14 citations this month" is a fact about the platform. It cannot be separated from the reshuffling that happens on its own. The honest version shows direction across several periods, against your own prior baseline, against a named competitor set, and broken out per engine instead of blended into one flattering average.
Ask to see a real report, redacted if it has to be. You are looking for six months of the same metric defined the same way. If the metric definition changed halfway through, ask why.
A good answer: six months of one metric, per engine, with your baseline and a named competitor set beside it.
12. Nothing that reaches revenue
Visibility, share of voice, sentiment and position are upstream signals, and they are worth tracking. You are buying something further downstream, and a program can win all four while your pipeline does nothing. In AI search that gap has a specific cause: the answer satisfies the question inside the interface, and the reader never lands on your site at all.
The question is whether anyone on the agency side has wired the signal to a commercial number: signups, booked demos, assisted pipeline, branded search lift. Measuring AEO ROI sets out what can and cannot be attributed honestly. Ask which number in the monthly report connects to revenue and how the attribution works. If the answer is a shrug, you are buying a dashboard. And if visibility climbs while nothing converts, you have a conversion problem sitting underneath a search problem, which is worth knowing early.
A good answer: the tool of yours the number comes from, and a client they have already wired it up for.
How to run the call
The twelve questions in column four are ammunition. Spend them in an order that gets you the most information before anyone gets defensive. If you are still deciding whether to hire at all, compare the agency route honestly against building the capability in-house. For some teams that is the right answer, and a good agency will say so. Across a shortlist, run the structured version of the selection process rather than a vibe check across three calls.
- Before the call. Read their methodology page and their most recent case study. Write down every number on both. Check whether each one names a time window, an engine and a prompt set. A number missing all three is decoration. A published methodology that names its engines, its measurement window, its price and a number that went down is what a full answer looks like. If there is no published methodology at all, that is your first question.
- First ten minutes. Ask flag 3 (who touches the account) and flag 1 (what is guaranteed). These are the cheapest questions and the most diagnostic. You will know a lot by minute eight.
- The middle. Ask flags 7, 8 and 9: what happens on my domain, what split do you expect between owned and earned, what happens after launch. This is where a real operator gets more specific and a reseller gets more abstract.
- Last ten minutes. Ask flags 10, 11 and 12: define the metrics, show me six months of one, tell me which one touches revenue. Ask for a redacted report before you decide.
- After the call. Get the contract and read the renewal and notice clauses before the second conversation rather than the week you sign.
One flag is a conversation. Two in the same group is the operating model.
Frequently asked questions
Answers to the questions readers ask most about this topic.
Can an AI search agency guarantee that ChatGPT will cite my site?
No. OpenAI's own help documentation says placement in ChatGPT search "is not guaranteed," and Google's Search Central documentation says "indexing and serving isn't guaranteed" with "no additional requirements to appear in AI Overviews." No agency sits between you and those systems. An agency can make citation more likely by structuring content the engines can lift, keeping it fresh and making the site crawlable. A guarantee is a claim the platforms themselves decline to make, and the FTC's position is that a performance claim must be substantiated before it is made.
How long does a first AI citation actually take?
A first citation can land within 24 hours when the conditions line up: a brand with modest existing authority, a prompt with no entrenched citation winners, a page whose answer is easy to lift, and fast indexing. Google's general guidance on recrawling is wider, from several days to several months. Holding a slot in the cited-source set takes weeks of surviving re-evaluation. Reaching a dominant share of answer on a competitive prompt cluster takes months. An agency that sells the first speed and bills for the third is the pattern to watch for.
What should a monthly AI search report contain?
No trade body publishes a required format, so judge it against the volatility instead. Because 40% to 60% of cited domains change month over month, a single-month citation count proves nothing on its own. A useful report shows the same metrics across multiple periods against your own baseline, breaks results out per engine rather than blending them, names the competitor set you are measured against, and connects at least one line to a commercial outcome. If a metric definition moves between reports, ask why before you ask anything else.
Does a high Domain Rating mean an agency will get me cited?
It is a weak signal at best. Semrush's study of 200,000 AI Overview keywords found that top-10 organic rankings did not ensure AI Overview inclusion. Run one of the prompts you care about and check the Domain Rating of every URL the engines cite back: the range is wide, and the biggest publishers in it do not dominate. Citation selection tracks content structure, freshness and where each engine looks for sources. Treat DR as a handicap that describes the starting line rather than as the lever, and ask for per-page citation evidence instead.
What is the difference between visibility, share of voice and citation rate?
Visibility is the percentage of AI responses where your brand appears at all. Share of voice is your slice of all tracked-brand mentions inside those responses, so it moves with your competitors as well as with you. Citation rate belongs to a different family: it measures how often your domain, once pulled in as a source, is explicitly referenced in the visible answer rather than read as background. You can be used as a source without being named as a brand, and named as a brand without your site being used. Any agency reporting to you should be able to say which one it is reporting and why.
Is it a red flag if an agency will not publish its methodology or real numbers?
Yes, and it is one of the easier ones to check before you ever book a call. A methodology worth the name states which engines it tracks, how often, what counts as a win, and what it does when a prompt is already owned by an entrenched competitor. Published results should carry a time window, a named engine and a prompt set. The FTC's endorsement guides also call for disclosing what results are generally expected when a featured result is exceptional, so a page of outliers with no context is a problem on two fronts at once.
How much of my AI citation volume should come from off-site placements?
Less than most pitches imply, if you sell software. Profound's analysis of 11.84 billion citations found SaaS and software companies see a median of 11.4% of citations from earned media against 59% for pharma and biotech, while roughly 57% of AI citations overall resolve to brand-owned sites. Off-site placements are a minority of citation volume with outsized effect per placement, and one strong placement can move share of answer quickly. Any plan that treats them as the primary driver for a SaaS brand is working against its own category's numbers.


