Guide

How to verify an AEO agency's results before you hire one

Hire the AEO agency that can trace a claimed result from the exact prompt to the answer, source, comparison record, and verified business outcome. Require the denominator and keep the original records. Reject screenshots without context, changing prompt sets hidden inside percentages, and revenue claims without a matching CRM record.

The short answer

Hire the AEO agency that can trace a claimed result from the exact prompt to the answer, source, comparison record, and verified business outcome. Require the denominator and keep the original records. Reject screenshots without context, changing prompt sets hidden inside percentages, and revenue claims without a matching CRM record.

On this page
  1. 1. Ask for the exact prompt, engine, and date
  2. 2. Make the agency define its denominator
  3. 3. Keep the same set of questions
  4. 4. Check the source evidence before you accept the logo
  5. 5. Separate visibility, leads, and revenue
  6. 6. Ask what would disprove the agency's case
  7. Build one result file another person can verify
  8. Resolve conflicting records before you score the result
  9. Choose the agency that fits the decision you need to make
  10. Put the reporting standard into the engagement
  11. Frequently Asked Questions

An AEO agency can show you a chart or a prompt screenshot. Neither proves much on its own.

AEO means work that aims to improve how a company appears in AI answers. The practical question is whether you can inspect a claimed result and connect it to a business outcome.

Short answer: Hire the AEO agency that can trace a claimed result from the exact prompt to the answer, source, comparison record, and verified business outcome. Require the denominator and keep the original records. Reject screenshots without context, changing prompt sets hidden inside percentages, and revenue claims without a matching CRM record.

Use this checklist before you sign.

What the agency claimsAsk to seeWhat it provesWhat it does not prove
"We got you cited"The full answer, source link, prompt, engine, and dateThe answer links to the page as a sourceThat the engine relied on it accurately, recommended the company, or led to buyer action
"We improved visibility"The prompt list, answer count, period, and denominatorThe brand appeared in a defined share of sampled answersThat the sample stayed consistent or that revenue rose
"We drove pipeline"The attribution rule, CRM record, and outcome dateA recorded lead or deal has an auditable source pathThat every buyer who used AI was captured
Scroll for the full table

1. Ask for the exact prompt, engine, and date

"We appear in ChatGPT" is not a result. It is a headline without the evidence underneath.

Ask the agency to open the answer in front of you. Record the exact question. Record the engine and product mode. Record the date. If the answer has sources, open the cited page and check that it supports the statement beside it.

A saved answer records one observation on one date. I checked the linked OpenAI and Google guidance on 8 September 2026. OpenAI says search citations can be incomplete, outdated, or wrong, and advises users to inspect the cited source. Google describes AI Overviews as answers with links to relevant web results. A link is useful evidence, but it has a narrow meaning.

The difference matters in a buying decision. Use "page citation" and "brand recommendation" as operational labels in the agency's report. They are not vendor-defined terms. Under that rule, a page citation records a source link in an answer. A brand recommendation records that the answer names the company as a suitable choice. Those observations can occur together.

Good answer: "On 3 September, this exact prompt in this engine cited this page. Here is the answer, the source link, and the saved record."

Red flag: a cropped screenshot with no prompt, date, source link, or way to repeat the check.

2. Make the agency define its denominator

Percentages can sound precise while hiding the sample.

If an agency says your visibility rose from 10% to 25%, ask: 25% of what? The answer should name the number of prompts, the engines, the number of responses, and the date range.

A useful report might say: "Across 40 fixed buyer questions and 120 sampled answers from three named engines, the brand appeared in 30 answers this month." You can then check the arithmetic: 30 divided by 120 is 25%.

That is a hypothetical example. It is not a performance benchmark. Its value is the structure. You need the numerator, the denominator, and the rules that created both.

Good answer: the report includes the complete prompt list and a count of every sampled answer.

Red flag: a percentage appears alone, or undisclosed prompt changes are used to claim an improvement.

3. Keep the same set of questions

A before-and-after result only works when the conditions match.

The agency should keep the prompt set fixed for the main comparison. It should name the engines, retain the market setting where relevant, and use equal-length periods. These controls improve comparability. They do not prove that the agency caused the change. If the conditions change, the report can still help, but it should describe the result as directional.

For example, an agency might add a new product category to the tracking list. That may be sensible. It also means the new total cannot cleanly compare with the old total. Ask for a separate view of the original set of questions.

Good answer: "The core set of questions has not changed; any added prompts appear in a separate expansion report."

Red flag: the agency cannot give you the prior prompt list or explain why the denominator changed.

An answer can name a company and cite a different source. It can cite your page but describe another company as the better fit. It can also mention you in a list without linking to you.

Ask the agency to classify each observation with its written operational definition. The definitions should cover source links, brand mentions, recommendations, and negative or mixed framing. Then read a small sample yourself.

Google's guidance for publishers asks whether a page shows who created it, gives readers a reason to trust it, and provides useful original information. Those are practical checks for agency proof too. Can you see where the claim came from? Does the agency explain the method? Can you inspect the work?

Good answer: a report saves the answer text, the linked sources, and the agency's classification rule.

Red flag: a dashboard reduces every answer to one green score.

5. Separate visibility, leads, and revenue

Visibility is an early signal. It is not a sale.

An answer-level report can show whether AI tools mention or cite your company for defined buyer questions. Under LoudFace's public method, web analytics can report visits with a recorded AI referrer. Your CRM can show a booked call, an opportunity, or closed revenue when the tracking path and sales record support it.

Each layer answers a different question. Do not let an agency use an answer mention to imply pipeline. Do not let it use a tagged visit to imply revenue. Ask for the label that matches the evidence.

Under that method, AI-referred visits and captured leads are reported as floors when attribution is thin. A careful agency states that limit. It does not fill the gap with a confident number.

LoudFace's public methodology gives one useful model: record signals separately by engine, then read them against search demand, captured leads, and CRM outcomes. Its Toku case study shows the kind of public record a buyer can inspect. The point is not the framework's branding. It is the visible chain from an answer observation to a commercial record.

Good answer: "This is answer visibility. These are attributed visits. These are confirmed opportunities. We do not claim the rest."

Red flag: "AI search generated revenue" with no attribution rule or commercial record.

6. Ask what would disprove the agency's case

Good measurement leaves room for a disappointing answer.

Ask for the agency's failure conditions before it starts. It should state the report it will produce if citations rise but qualified calls do not. It should state the rule that ends a tactic after a failed prompt test.

You are looking for a team that can separate an observation from a conclusion. AEO covers content, technical access, source quality, and how a company is represented online. OpenAI says placement in ChatGPT search results is not guaranteed. Set failure conditions around the records the agency can produce.

The agency you hire should make its work legible. You should be able to see the prompt, the answer, the source, the comparison rules, and the business record. If you cannot inspect those five things, you are buying a story about results.

Build one result file another person can verify

A result should remain understandable after the salesperson, analyst, or client contact leaves. Put the supporting records for one claim in a file that another person can open without oral context.

File partKeepReturn for correction when
Reported claimThe exact wording and its evidence labelThe headline implies a recommendation or sale that the records do not show
Prompt entryThe literal question, engine, mode, relevant market, and dateOnly a topic label or rewritten prompt appears
Answer recordThe relevant text, source links, and surrounding contextA cropped image hides the prompt or surrounding answer
Source inspectionWhether the linked page supports, partly supports, or does not support the statementA visible link is treated as proof without opening the page
ClassificationThe written rule for a mention, citation, recommendation, or commercial labelOne score combines different observation types
Comparison recordThe numerator, denominator, cohort, engines, interval, periods, and changed conditionsThe percentage cannot be recalculated from the file
Attribution recordThe attribution rule, available analytics event, lead record, CRM outcome, and end of the joinThe file jumps from visibility to pipeline
OwnershipThe client can retain an exportThe record disappears with access to the agency's dashboard
Scroll for the full table

A visible answer with no recorded visit stays an answer observation. A recorded visit with no matched lead stays an observed visit. Use only the label that the record supports.

Give each result file a stable identifier. Keep the original capture beside later corrections, and record who changed a classification and why. A corrected dashboard should not erase the record used for the original claim.

Choose the review sample before the sales call. Request one result the agency considers successful and one result that disappointed the client. Then ask for a disputed classification if one exists. Clean evidence tests record quality. Messy evidence tests whether the method survives disagreement and correction.

The file format should also suit the people who must inspect it. A legal reviewer may need a fixed document with dates and source copies. An analytics lead may need a structured export that can be recalculated. Ask each reviewer to confirm that the record answers their decision before you accept the format.

Redaction does not make this impossible. A sample can hide a client's name, contract amount, or sensitive prompt wording while preserving the record type and method. Replace the private value with a consistent label. Keep the dates, definitions, arithmetic, and connection between records visible. If redaction removes the information needed to test the claim, describe it as an unverified private example.

The result file also shows weak methods early. An agency may have a strong strategy but poor record keeping. Another may have a polished dashboard built on classifications nobody checks. Neither problem appears in a logo slide. Both appear when you ask someone outside the account team to reconstruct one result.

A useful sales demonstration therefore has one job: open a reported claim and reconstruct it. The agency should move from the prompt entry to the answer, then to the source inspection and comparison record. If it claims a business outcome, it should continue to the matching commercial record. The demonstration succeeds when the records explain the claim without the presenter rescuing it.

Resolve conflicting records before you score the result

The records often conflict. The prompt list changes. A human reviewer disagrees with an automated label. Analytics records a visit but the CRM has no source. The agency's quality shows in how it resolves those conflicts.

A hypothetical report says visibility rose from 10% to 25%. The agency tracked 40 buyer questions across three engines, which produced 120 sampled answers in each period. The earlier period contains 12 appearances. The current period contains 30. The arithmetic is correct.

The prompt register reveals a comparability problem. Eight of the 40 questions changed between periods. The unchanged cohort contains 32 questions and 96 answers per period. It records 10 appearances in the earlier period and 19 in the current period. That comparable cohort rose from about 10.4% to about 19.8%. The 25% headline describes the full current sample, while the 19.8% figure supports the cleaner before-and-after comparison.

The report can keep both numbers. They answer different questions. Use the comparable cohort as the primary result when the report claims improvement. Show the full current sample as an expanded-scope snapshot. Calling the whole change a single 10%-to-25% gain would fail to disclose the changed questions.

Now inspect the answer labels. One saved answer links to the company's page but never names the company. The automated report calls it a brand mention. Under the written definitions used here, that record is a page citation without a brand mention. Correct the label before recalculating the totals. Do not change the definition to protect the headline.

A second answer names the company and presents it as suitable for a specific buyer. It links to an independent page rather than the company's site. That can be a brand recommendation without a citation to the company's own page. The source domain does not erase the recommendation. It changes what the citation proves.

A third answer names the company but includes a negative qualification. Count the mention if the rule counts all mentions. Preserve the negative framing in a separate field. A positive mention rate that drops the qualification gives a buyer the wrong picture of how the brand appears.

The commercial records create another conflict. In this hypothetical review, analytics shows four visits with an AI referrer. The CRM contains two opportunities whose contacts recall using an AI tool during research. Only one opportunity has a recorded path that connects it to one of those visits. The defensible report contains four observed AI-referred visits and one attributed opportunity. The second opportunity can appear as buyer-reported influence, but it should not share the stronger attribution label.

This adjudication uses a simple order. Preserve the raw record first. Apply the written definition next. Recalculate the metric after correcting classifications. Then state the highest evidence layer supported by the joined records. When two records conflict, keep both and explain the conflict. Do not select the one that produces the better result.

ConflictDefensible treatmentInflated treatment
Prompt set changedCompare the stable cohort and show the expansion separatelyBlend every prompt into one before-and-after percentage
Citation without a brand nameRecord a page citation onlyCount it as a recommendation or positive mention
Brand named with a qualificationPreserve the mention and the framingCount only the favorable part
Engine results disagreeReport each engine separatelyAn average conceals the disagreement
A referrer exists without a lead joinReport an observed visitCall the visit pipeline
Buyer recalls AI research without a tracked pathLabel it as buyer-reported influenceClaim attributed revenue
Scroll for the full table

The choice of comparison method is my editorial recommendation. It is not an industry standard issued by OpenAI, Google, or another vendor. I recommend the stable-cohort view because it makes the before-and-after claim easier to inspect. An agency can use another method if it defines that method before the result and keeps changed conditions visible.

The same principle applies to missing data. A result can be marked "unavailable." That label tells the buyer which evidence is available. An estimate can be useful for planning, but it should not replace an observed value in a performance claim. The report should mark the estimate, explain its basis, and keep it outside the verified total.

Choose the agency that fits the decision you need to make

Check evidence quality first. It is not the only selection criterion. Once each finalist can support its claims, compare the operating model against the decision your team needs to make. The same vendor will not be the best fit for every buyer.

Imagine a company with no consistent prompt register and no reliable connection between analytics and its CRM. Agency A offers a sophisticated visibility dashboard, but it expects clean inputs from the client. Agency B starts with a measurement baseline and defines the records before it reports progress. Agency A may have better software. Agency B is the safer choice for this buyer because the first problem is measurement design.

Now imagine a company with a mature analytics team and a stable set of buyer questions. Its main need is content and source improvement. Agency A can work inside the existing measurement rules and lets the client retain every answer record. Agency B insists on replacing the prompt set with its own template. Here, Agency A is a better fit because it can improve the work without breaking the comparison history.

A third buyer works in a category where every public claim receives legal review. Its agency must preserve the text, source, date, reviewer decision, and revision history behind each result. A vendor that provides downloadable records has a clear advantage. A vendor that offers only a dashboard screenshot creates extra review work, even if its strategy is strong.

Consider a smaller marketing team that needs help making decisions each month. One finalist sends a large export and leaves interpretation to the client. Another separates observations from conclusions, shows what changed, and names the next decision. The second vendor provides more value because the team needs an accountable operator who interprets the data.

Evaluate price after the evidence and fit review. A cheaper proposal can cost more if the client must rebuild the records, reconcile unclear labels, or recover data at the end. An expensive proposal is not safer by default. The buyer should compare the work that remains after each agency delivers its report.

Use a side-by-side decision note instead of a single total score. A total can hide a deal-breaking weakness. Record the evidence result, operating fit, access terms, and unresolved risk separately.

Buyer needStronger vendor responseWeak response
Establish a baselineDefines the prompt cohort and evidence labels before work startsStarts with a headline visibility score
Preserve an existing comparisonWorks with the stable cohort and isolates additionsReplaces the prompt set without a bridge
Support legal or executive reviewKeeps source-level records and reviewer notesSupplies screenshots and verbal explanations
Connect work to pipelineStates the join rule and its limitsTreats every AI-referred visit as revenue influence
Reduce the client's reporting loadExplains the decision that follows from each findingDelivers an export with no conclusion
Protect continuityGives the client usable exports and definitionsKeeps the method inside a proprietary account
Scroll for the full table

Two agencies can pass the evidence test and still suit different teams. One may be stronger at technical access. Another may produce better editorial work. A third may integrate more cleanly with the client's analytics. The buyer should make those tradeoffs after removing vendors whose claims lack supporting records.

Do not reward a vendor for claiming certainty where none exists. A good agency can hold a firm position and still state the limit of the record. "We improved citation frequency in the stable prompt cohort" is a useful claim. "We caused every influenced sale" is not credible when the source path is incomplete.

The best finalist should also handle disagreement without defensiveness. Give each agency the same ambiguous answer record and ask how it would classify it. The exact label matters less than a consistent rule, a preserved answer, and a willingness to correct the total. This exercise reveals more than another success story.

End the selection note with a concrete decision. Examples include: run a limited engagement with the vendor that can establish the baseline; keep the current measurement system and hire the stronger content operator; or pause the purchase because neither finalist can provide retained source records. Base the choice on the evidence.

Put the reporting standard into the engagement

A good sales demonstration can still become a weak monthly report. Convert the accepted evidence standard into the reporting brief or agreement. The document should state what the agency reports, how it proves each label, who can inspect the records, and what happens when the method changes.

The fixed comparison starts with a prompt register. Name the engines, market conditions, observation interval, and comparison period. Give additions a separate expansion label until enough comparable history exists. This keeps sensible experimentation from rewriting the baseline.

Plain classification rules let a reviewer repeat the decision. A mention means the answer names the company. A page citation means the answer links to the page. A recommendation means the answer presents the company as suitable. Give negative or mixed framing its own field, and keep the answer text behind every classification.

Commercial labels need different records. An observed visit needs the recorded referrer. A captured lead needs the defined conversion event. An attributed opportunity needs the accepted CRM join. Closed revenue needs the matching outcome and date. Buyer-reported influence remains useful, but it stays separate from tracked attribution.

Before access ends, the client should retain the prompt register, monthly reports, saved answers, definitions, and supporting-record exports. The agency can retain its software and internal annotations. Agree on the export format and delivery point before signing.

A correction rule protects the audit trail. If an error changes a reported metric, preserve the prior report, issue the corrected value, and explain the cause. Quietly replacing a dashboard value destroys that record.

Set failure criteria for the work. Do not set guarantees about the answer engines. A useful criterion names the observed condition and the decision it triggers. If citations rise but qualified calls do not, the agency may review prompt intent and the buyer path. If a page earns links but the answer describes the company inaccurately, the agency may correct the page's factual support. If a tactic produces no meaningful change under the agreed comparison, the team may stop it.

These criteria should allow a negative report. A null result is information. A decline is information. A changed engine response is information. The agency earns confidence by showing what happened and making the next decision explicit.

A redacted sample report can test the standard before signing. It should contain the same fields planned for the client account. Look for a prompt register entry, saved answer, classification, denominator, comparison note, and downstream record where applicable. Also request an example of a bad month. The explanation should show the evidence that led to a changed decision.

The monthly review can then focus on exceptions instead of replaying the entire method. Inspect changed prompts, disputed classifications, large movements, and new commercial joins. Sample a few unchanged records to confirm the system still works. Record one decision for each material finding.

Reporting questionRequired recordDecision it supports
Did visibility change?Stable cohort, numerator, denominator, and periodsContinue, revise, or stop the tested work
Did the type of appearance change?Saved answers and classification rulesImprove sources, positioning, or page facts
Did buyers reach the site?Referrer-backed visit recordsInspect the path from answer to page
Did a commercial event occur?Lead or CRM record with the accepted joinCredit only the outcome the record supports
Did the method change?Change log and separate expansion viewProtect the old comparison or establish a new baseline
Can the client audit the result later?The client keeps exports and definitionsRequire better access if the records are missing
Scroll for the full table

When an engine disagrees with another engine, keep the names visible. One may cite the page while another omits it. That split can guide source work or show that the result is limited to one product. A blended score may be convenient, but it should never be the only view.

When the agency and client disagree about a classification, return to the saved answer and the written rule. The reviewer records the decision and reason. If the rule itself needs to change, apply the new rule prospectively or recalculate the earlier periods. Do not change it only for the favorable answer.

The standard should remain usable after the engagement ends. A new employee should be able to read an old report and locate the records behind it. A new agency should be able to continue the stable cohort or explain why it needs a new baseline. The client then retains the measurement records instead of losing access to them.

The final hiring decision becomes simpler after this work. Choose the agency whose claims have supporting records. Confirm that its operating model fits your present problem. Confirm that its records remain usable after the contract ends. If no finalist clears those conditions, keep looking.

Author disclosure: Arnel Bukva runs LoudFace, an organic-growth agency. LoudFace publishes its measurement approach at loudface.co/methodology.

FAQ

Frequently asked questions

Answers to the questions readers ask most about this topic.

Can an AEO agency guarantee a citation?

No. OpenAI says placement in ChatGPT search results is not guaranteed. An agency can commit to a documented baseline, a fixed prompt set, and a repeatable reporting method. It cannot honestly promise that a named page will appear in a specific ChatGPT search answer on a specific date.

What should an AEO agency include in a monthly report?

Ask for the complete prompt set, the engines sampled, the response count, and the period. The report should show the answer text or saved record behind each classification. It should also separate a fixed comparison set from any new prompts. That gives you a report you can audit instead of a percentage that changes meaning each month.

How do I check whether an AI answer cited my page?

Open the saved answer. Confirm the exact prompt, engine, date, and source link. Then open the linked page and check that it supports the statement beside it. A screenshot without those records is weak evidence. OpenAI also advises readers to inspect cited sources because results and citations can be incomplete, outdated, or wrong.

What is the difference between an AI mention and a citation?

Use clear operational definitions in the report. A mention records that an answer names the company. A page citation records a source link to the page. A recommendation records that the answer presents the company as a suitable choice. These observations can overlap, but they answer different questions. Do not let one label stand in for another.

Can an AEO agency claim that AI search created revenue?

Only when it can show the attribution rule and the commercial record. An answer mention is visibility. A referrer-backed visit is an observed visit. A booked call or closed deal needs the matching CRM record. Where referrer data is missing, report the observed visits or leads as floors. Do not convert incomplete attribution into a revenue claim.

Written by
Arnel Bukva
Arnel Bukva
Founder & Head of Growth

Arnel Bukva is the founder of LoudFace, a B2B SaaS organic growth agency that ships AEO (Answer Engine Optimization), SEO, and Webflow programmes for Series A to C companies. His work focuses on AI-cited content systems that move pipeline rather than vanity traffic, with named client outcomes including Toku (consistently the top-cited vendor on stablecoin payroll prompts in AI search) and TradeMomentum (a major climb in organic impressions). One of the earliest Webflow users (2017), he has spent the past several years at the intersection of technical SEO and AI search, building the prompt-graph methodology LoudFace uses across every client engagement.

On the record
Published
Sep 8, 2026
Category
Guide
Reading time
23 min read
LoudFace — strategy callB2B SaaS only

Ready to grow your business?

Let’s discuss how we can help you achieve your goals. 30 minutes, no pitch deck. We’ll look at your site together and name what should move first: build, growth, or both.

Book a callBuild and growth, one team
Cover — LIQID, built by LoudFaceloudface.co