How to verify an AEO agency's results before you hire one
Hire the AEO agency that can trace a claimed result from the exact prompt to the answer, source, comparison record, and verified business outcome. Require the denominator and keep the original records. Reject screenshots without context, changing prompt sets hidden inside percentages, and revenue claims without a matching CRM record.
Hire the AEO agency that can trace a claimed result from the exact prompt to the answer, source, comparison record, and verified business outcome. Require the denominator and keep the original records. Reject screenshots without context, changing prompt sets hidden inside percentages, and revenue claims without a matching CRM record.
On this page
- 1. Ask for the exact prompt, engine, and date
- 2. Make the agency define its denominator
- 3. Keep the same set of questions
- 4. Check the source evidence before you accept the logo
- 5. Separate visibility, leads, and revenue
- 6. Ask what would disprove the agency's case
- Build one result file another person can verify
- Resolve conflicting records before you score the result
- Choose the agency that fits the decision you need to make
- Put the reporting standard into the engagement
- Frequently Asked Questions
An AEO agency can show you a chart or a prompt screenshot. Neither proves much on its own.
AEO means work that aims to improve how a company appears in AI answers. The practical question is whether you can inspect a claimed result and connect it to a business outcome.
Short answer: Hire the AEO agency that can trace a claimed result from the exact prompt to the answer, source, comparison record, and verified business outcome. Require the denominator and keep the original records. Reject screenshots without context, changing prompt sets hidden inside percentages, and revenue claims without a matching CRM record.
Use this checklist before you sign.
| What the agency claims | Ask to see | What it proves | What it does not prove |
|---|---|---|---|
| "We got you cited" | The full answer, source link, prompt, engine, and date | The answer links to the page as a source | That the engine relied on it accurately, recommended the company, or led to buyer action |
| "We improved visibility" | The prompt list, answer count, period, and denominator | The brand appeared in a defined share of sampled answers | That the sample stayed consistent or that revenue rose |
| "We drove pipeline" | The attribution rule, CRM record, and outcome date | A recorded lead or deal has an auditable source path | That every buyer who used AI was captured |
1. Ask for the exact prompt, engine, and date
"We appear in ChatGPT" is not a result. It is a headline without the evidence underneath.
Ask the agency to open the answer in front of you. Record the exact question. Record the engine and product mode. Record the date. If the answer has sources, open the cited page and check that it supports the statement beside it.
A saved answer records one observation on one date. I checked the linked OpenAI and Google guidance on 8 September 2026. OpenAI says search citations can be incomplete, outdated, or wrong, and advises users to inspect the cited source. Google describes AI Overviews as answers with links to relevant web results. A link is useful evidence, but it has a narrow meaning.
The difference matters in a buying decision. Use "page citation" and "brand recommendation" as operational labels in the agency's report. They are not vendor-defined terms. Under that rule, a page citation records a source link in an answer. A brand recommendation records that the answer names the company as a suitable choice. Those observations can occur together.
Good answer: "On 3 September, this exact prompt in this engine cited this page. Here is the answer, the source link, and the saved record."
Red flag: a cropped screenshot with no prompt, date, source link, or way to repeat the check.
2. Make the agency define its denominator
Percentages can sound precise while hiding the sample.
If an agency says your visibility rose from 10% to 25%, ask: 25% of what? The answer should name the number of prompts, the engines, the number of responses, and the date range.
A useful report might say: "Across 40 fixed buyer questions and 120 sampled answers from three named engines, the brand appeared in 30 answers this month." You can then check the arithmetic: 30 divided by 120 is 25%.
That is a hypothetical example. It is not a performance benchmark. Its value is the structure. You need the numerator, the denominator, and the rules that created both.
Good answer: the report includes the complete prompt list and a count of every sampled answer.
Red flag: a percentage appears alone, or undisclosed prompt changes are used to claim an improvement.
3. Keep the same set of questions
A before-and-after result only works when the conditions match.
The agency should keep the prompt set fixed for the main comparison. It should name the engines, retain the market setting where relevant, and use equal-length periods. These controls improve comparability. They do not prove that the agency caused the change. If the conditions change, the report can still help, but it should describe the result as directional.
For example, an agency might add a new product category to the tracking list. That may be sensible. It also means the new total cannot cleanly compare with the old total. Ask for a separate view of the original set of questions.
Good answer: "The core set of questions has not changed; any added prompts appear in a separate expansion report."
Red flag: the agency cannot give you the prior prompt list or explain why the denominator changed.
4. Check the source evidence before you accept the logo
An answer can name a company and cite a different source. It can cite your page but describe another company as the better fit. It can also mention you in a list without linking to you.
Ask the agency to classify each observation with its written operational definition. The definitions should cover source links, brand mentions, recommendations, and negative or mixed framing. Then read a small sample yourself.
Google's guidance for publishers asks whether a page shows who created it, gives readers a reason to trust it, and provides useful original information. Those are practical checks for agency proof too. Can you see where the claim came from? Does the agency explain the method? Can you inspect the work?
Good answer: a report saves the answer text, the linked sources, and the agency's classification rule.
Red flag: a dashboard reduces every answer to one green score.
5. Separate visibility, leads, and revenue
Visibility is an early signal. It is not a sale.
An answer-level report can show whether AI tools mention or cite your company for defined buyer questions. Under LoudFace's public method, web analytics can report visits with a recorded AI referrer. Your CRM can show a booked call, an opportunity, or closed revenue when the tracking path and sales record support it.
Each layer answers a different question. Do not let an agency use an answer mention to imply pipeline. Do not let it use a tagged visit to imply revenue. Ask for the label that matches the evidence.
Under that method, AI-referred visits and captured leads are reported as floors when attribution is thin. A careful agency states that limit. It does not fill the gap with a confident number.
LoudFace's public methodology gives one useful model: record signals separately by engine, then read them against search demand, captured leads, and CRM outcomes. Its Toku case study shows the kind of public record a buyer can inspect. The point is not the framework's branding. It is the visible chain from an answer observation to a commercial record.
Good answer: "This is answer visibility. These are attributed visits. These are confirmed opportunities. We do not claim the rest."
Red flag: "AI search generated revenue" with no attribution rule or commercial record.
6. Ask what would disprove the agency's case
Good measurement leaves room for a disappointing answer.
Ask for the agency's failure conditions before it starts. It should state the report it will produce if citations rise but qualified calls do not. It should state the rule that ends a tactic after a failed prompt test.
You are looking for a team that can separate an observation from a conclusion. AEO covers content, technical access, source quality, and how a company is represented online. OpenAI says placement in ChatGPT search results is not guaranteed. Set failure conditions around the records the agency can produce.
The agency you hire should make its work legible. You should be able to see the prompt, the answer, the source, the comparison rules, and the business record. If you cannot inspect those five things, you are buying a story about results.
Build one result file another person can verify
A result should remain understandable after the salesperson, analyst, or client contact leaves. Put the supporting records for one claim in a file that another person can open without oral context.
| File part | Keep | Return for correction when |
|---|---|---|
| Reported claim | The exact wording and its evidence label | The headline implies a recommendation or sale that the records do not show |
| Prompt entry | The literal question, engine, mode, relevant market, and date | Only a topic label or rewritten prompt appears |
| Answer record | The relevant text, source links, and surrounding context | A cropped image hides the prompt or surrounding answer |
| Source inspection | Whether the linked page supports, partly supports, or does not support the statement | A visible link is treated as proof without opening the page |
| Classification | The written rule for a mention, citation, recommendation, or commercial label | One score combines different observation types |
| Comparison record | The numerator, denominator, cohort, engines, interval, periods, and changed conditions | The percentage cannot be recalculated from the file |
| Attribution record | The attribution rule, available analytics event, lead record, CRM outcome, and end of the join | The file jumps from visibility to pipeline |
| Ownership | The client can retain an export | The record disappears with access to the agency's dashboard |
A visible answer with no recorded visit stays an answer observation. A recorded visit with no matched lead stays an observed visit. Use only the label that the record supports.
Give each result file a stable identifier. Keep the original capture beside later corrections, and record who changed a classification and why. A corrected dashboard should not erase the record used for the original claim.
Choose the review sample before the sales call. Request one result the agency considers successful and one result that disappointed the client. Then ask for a disputed classification if one exists. Clean evidence tests record quality. Messy evidence tests whether the method survives disagreement and correction.
The file format should also suit the people who must inspect it. A legal reviewer may need a fixed document with dates and source copies. An analytics lead may need a structured export that can be recalculated. Ask each reviewer to confirm that the record answers their decision before you accept the format.
Redaction does not make this impossible. A sample can hide a client's name, contract amount, or sensitive prompt wording while preserving the record type and method. Replace the private value with a consistent label. Keep the dates, definitions, arithmetic, and connection between records visible. If redaction removes the information needed to test the claim, describe it as an unverified private example.
The result file also shows weak methods early. An agency may have a strong strategy but poor record keeping. Another may have a polished dashboard built on classifications nobody checks. Neither problem appears in a logo slide. Both appear when you ask someone outside the account team to reconstruct one result.
A useful sales demonstration therefore has one job: open a reported claim and reconstruct it. The agency should move from the prompt entry to the answer, then to the source inspection and comparison record. If it claims a business outcome, it should continue to the matching commercial record. The demonstration succeeds when the records explain the claim without the presenter rescuing it.
Resolve conflicting records before you score the result
The records often conflict. The prompt list changes. A human reviewer disagrees with an automated label. Analytics records a visit but the CRM has no source. The agency's quality shows in how it resolves those conflicts.
A hypothetical report says visibility rose from 10% to 25%. The agency tracked 40 buyer questions across three engines, which produced 120 sampled answers in each period. The earlier period contains 12 appearances. The current period contains 30. The arithmetic is correct.
The prompt register reveals a comparability problem. Eight of the 40 questions changed between periods. The unchanged cohort contains 32 questions and 96 answers per period. It records 10 appearances in the earlier period and 19 in the current period. That comparable cohort rose from about 10.4% to about 19.8%. The 25% headline describes the full current sample, while the 19.8% figure supports the cleaner before-and-after comparison.
The report can keep both numbers. They answer different questions. Use the comparable cohort as the primary result when the report claims improvement. Show the full current sample as an expanded-scope snapshot. Calling the whole change a single 10%-to-25% gain would fail to disclose the changed questions.
Now inspect the answer labels. One saved answer links to the company's page but never names the company. The automated report calls it a brand mention. Under the written definitions used here, that record is a page citation without a brand mention. Correct the label before recalculating the totals. Do not change the definition to protect the headline.
A second answer names the company and presents it as suitable for a specific buyer. It links to an independent page rather than the company's site. That can be a brand recommendation without a citation to the company's own page. The source domain does not erase the recommendation. It changes what the citation proves.
A third answer names the company but includes a negative qualification. Count the mention if the rule counts all mentions. Preserve the negative framing in a separate field. A positive mention rate that drops the qualification gives a buyer the wrong picture of how the brand appears.
The commercial records create another conflict. In this hypothetical review, analytics shows four visits with an AI referrer. The CRM contains two opportunities whose contacts recall using an AI tool during research. Only one opportunity has a recorded path that connects it to one of those visits. The defensible report contains four observed AI-referred visits and one attributed opportunity. The second opportunity can appear as buyer-reported influence, but it should not share the stronger attribution label.
This adjudication uses a simple order. Preserve the raw record first. Apply the written definition next. Recalculate the metric after correcting classifications. Then state the highest evidence layer supported by the joined records. When two records conflict, keep both and explain the conflict. Do not select the one that produces the better result.
| Conflict | Defensible treatment | Inflated treatment |
|---|---|---|
| Prompt set changed | Compare the stable cohort and show the expansion separately | Blend every prompt into one before-and-after percentage |
| Citation without a brand name | Record a page citation only | Count it as a recommendation or positive mention |
| Brand named with a qualification | Preserve the mention and the framing | Count only the favorable part |
| Engine results disagree | Report each engine separately | An average conceals the disagreement |
| A referrer exists without a lead join | Report an observed visit | Call the visit pipeline |
| Buyer recalls AI research without a tracked path | Label it as buyer-reported influence | Claim attributed revenue |
The choice of comparison method is my editorial recommendation. It is not an industry standard issued by OpenAI, Google, or another vendor. I recommend the stable-cohort view because it makes the before-and-after claim easier to inspect. An agency can use another method if it defines that method before the result and keeps changed conditions visible.
The same principle applies to missing data. A result can be marked "unavailable." That label tells the buyer which evidence is available. An estimate can be useful for planning, but it should not replace an observed value in a performance claim. The report should mark the estimate, explain its basis, and keep it outside the verified total.
Choose the agency that fits the decision you need to make
Check evidence quality first. It is not the only selection criterion. Once each finalist can support its claims, compare the operating model against the decision your team needs to make. The same vendor will not be the best fit for every buyer.
Imagine a company with no consistent prompt register and no reliable connection between analytics and its CRM. Agency A offers a sophisticated visibility dashboard, but it expects clean inputs from the client. Agency B starts with a measurement baseline and defines the records before it reports progress. Agency A may have better software. Agency B is the safer choice for this buyer because the first problem is measurement design.
Now imagine a company with a mature analytics team and a stable set of buyer questions. Its main need is content and source improvement. Agency A can work inside the existing measurement rules and lets the client retain every answer record. Agency B insists on replacing the prompt set with its own template. Here, Agency A is a better fit because it can improve the work without breaking the comparison history.
A third buyer works in a category where every public claim receives legal review. Its agency must preserve the text, source, date, reviewer decision, and revision history behind each result. A vendor that provides downloadable records has a clear advantage. A vendor that offers only a dashboard screenshot creates extra review work, even if its strategy is strong.
Consider a smaller marketing team that needs help making decisions each month. One finalist sends a large export and leaves interpretation to the client. Another separates observations from conclusions, shows what changed, and names the next decision. The second vendor provides more value because the team needs an accountable operator who interprets the data.
Evaluate price after the evidence and fit review. A cheaper proposal can cost more if the client must rebuild the records, reconcile unclear labels, or recover data at the end. An expensive proposal is not safer by default. The buyer should compare the work that remains after each agency delivers its report.
Use a side-by-side decision note instead of a single total score. A total can hide a deal-breaking weakness. Record the evidence result, operating fit, access terms, and unresolved risk separately.
| Buyer need | Stronger vendor response | Weak response |
|---|---|---|
| Establish a baseline | Defines the prompt cohort and evidence labels before work starts | Starts with a headline visibility score |
| Preserve an existing comparison | Works with the stable cohort and isolates additions | Replaces the prompt set without a bridge |
| Support legal or executive review | Keeps source-level records and reviewer notes | Supplies screenshots and verbal explanations |
| Connect work to pipeline | States the join rule and its limits | Treats every AI-referred visit as revenue influence |
| Reduce the client's reporting load | Explains the decision that follows from each finding | Delivers an export with no conclusion |
| Protect continuity | Gives the client usable exports and definitions | Keeps the method inside a proprietary account |
Two agencies can pass the evidence test and still suit different teams. One may be stronger at technical access. Another may produce better editorial work. A third may integrate more cleanly with the client's analytics. The buyer should make those tradeoffs after removing vendors whose claims lack supporting records.
Do not reward a vendor for claiming certainty where none exists. A good agency can hold a firm position and still state the limit of the record. "We improved citation frequency in the stable prompt cohort" is a useful claim. "We caused every influenced sale" is not credible when the source path is incomplete.
The best finalist should also handle disagreement without defensiveness. Give each agency the same ambiguous answer record and ask how it would classify it. The exact label matters less than a consistent rule, a preserved answer, and a willingness to correct the total. This exercise reveals more than another success story.
End the selection note with a concrete decision. Examples include: run a limited engagement with the vendor that can establish the baseline; keep the current measurement system and hire the stronger content operator; or pause the purchase because neither finalist can provide retained source records. Base the choice on the evidence.
Put the reporting standard into the engagement
A good sales demonstration can still become a weak monthly report. Convert the accepted evidence standard into the reporting brief or agreement. The document should state what the agency reports, how it proves each label, who can inspect the records, and what happens when the method changes.
The fixed comparison starts with a prompt register. Name the engines, market conditions, observation interval, and comparison period. Give additions a separate expansion label until enough comparable history exists. This keeps sensible experimentation from rewriting the baseline.
Plain classification rules let a reviewer repeat the decision. A mention means the answer names the company. A page citation means the answer links to the page. A recommendation means the answer presents the company as suitable. Give negative or mixed framing its own field, and keep the answer text behind every classification.
Commercial labels need different records. An observed visit needs the recorded referrer. A captured lead needs the defined conversion event. An attributed opportunity needs the accepted CRM join. Closed revenue needs the matching outcome and date. Buyer-reported influence remains useful, but it stays separate from tracked attribution.
Before access ends, the client should retain the prompt register, monthly reports, saved answers, definitions, and supporting-record exports. The agency can retain its software and internal annotations. Agree on the export format and delivery point before signing.
A correction rule protects the audit trail. If an error changes a reported metric, preserve the prior report, issue the corrected value, and explain the cause. Quietly replacing a dashboard value destroys that record.
Set failure criteria for the work. Do not set guarantees about the answer engines. A useful criterion names the observed condition and the decision it triggers. If citations rise but qualified calls do not, the agency may review prompt intent and the buyer path. If a page earns links but the answer describes the company inaccurately, the agency may correct the page's factual support. If a tactic produces no meaningful change under the agreed comparison, the team may stop it.
These criteria should allow a negative report. A null result is information. A decline is information. A changed engine response is information. The agency earns confidence by showing what happened and making the next decision explicit.
A redacted sample report can test the standard before signing. It should contain the same fields planned for the client account. Look for a prompt register entry, saved answer, classification, denominator, comparison note, and downstream record where applicable. Also request an example of a bad month. The explanation should show the evidence that led to a changed decision.
The monthly review can then focus on exceptions instead of replaying the entire method. Inspect changed prompts, disputed classifications, large movements, and new commercial joins. Sample a few unchanged records to confirm the system still works. Record one decision for each material finding.
| Reporting question | Required record | Decision it supports |
|---|---|---|
| Did visibility change? | Stable cohort, numerator, denominator, and periods | Continue, revise, or stop the tested work |
| Did the type of appearance change? | Saved answers and classification rules | Improve sources, positioning, or page facts |
| Did buyers reach the site? | Referrer-backed visit records | Inspect the path from answer to page |
| Did a commercial event occur? | Lead or CRM record with the accepted join | Credit only the outcome the record supports |
| Did the method change? | Change log and separate expansion view | Protect the old comparison or establish a new baseline |
| Can the client audit the result later? | The client keeps exports and definitions | Require better access if the records are missing |
When an engine disagrees with another engine, keep the names visible. One may cite the page while another omits it. That split can guide source work or show that the result is limited to one product. A blended score may be convenient, but it should never be the only view.
When the agency and client disagree about a classification, return to the saved answer and the written rule. The reviewer records the decision and reason. If the rule itself needs to change, apply the new rule prospectively or recalculate the earlier periods. Do not change it only for the favorable answer.
The standard should remain usable after the engagement ends. A new employee should be able to read an old report and locate the records behind it. A new agency should be able to continue the stable cohort or explain why it needs a new baseline. The client then retains the measurement records instead of losing access to them.
The final hiring decision becomes simpler after this work. Choose the agency whose claims have supporting records. Confirm that its operating model fits your present problem. Confirm that its records remain usable after the contract ends. If no finalist clears those conditions, keep looking.
Author disclosure: Arnel Bukva runs LoudFace, an organic-growth agency. LoudFace publishes its measurement approach at loudface.co/methodology.
Frequently asked questions
Answers to the questions readers ask most about this topic.
Can an AEO agency guarantee a citation?
No. OpenAI says placement in ChatGPT search results is not guaranteed. An agency can commit to a documented baseline, a fixed prompt set, and a repeatable reporting method. It cannot honestly promise that a named page will appear in a specific ChatGPT search answer on a specific date.
What should an AEO agency include in a monthly report?
Ask for the complete prompt set, the engines sampled, the response count, and the period. The report should show the answer text or saved record behind each classification. It should also separate a fixed comparison set from any new prompts. That gives you a report you can audit instead of a percentage that changes meaning each month.
How do I check whether an AI answer cited my page?
Open the saved answer. Confirm the exact prompt, engine, date, and source link. Then open the linked page and check that it supports the statement beside it. A screenshot without those records is weak evidence. OpenAI also advises readers to inspect cited sources because results and citations can be incomplete, outdated, or wrong.
What is the difference between an AI mention and a citation?
Use clear operational definitions in the report. A mention records that an answer names the company. A page citation records a source link to the page. A recommendation records that the answer presents the company as a suitable choice. These observations can overlap, but they answer different questions. Do not let one label stand in for another.
Can an AEO agency claim that AI search created revenue?
Only when it can show the attribution rule and the commercial record. An answer mention is visibility. A referrer-backed visit is an observed visit. A booked call or closed deal needs the matching CRM record. Where referrer data is missing, report the observed visits or leads as floors. Do not convert incomplete attribution into a revenue claim.


