| Grader | You enter | Signup to see the score | What it actually checks | Engines in the free result | What it said about LoudFace | Paid next step |
|---|---|---|---|---|---|---|
| HubSpot AI Search Grader | Company, location, product, industry | No (full report gated) | What models say about you "based on their training data" | GPT-5.4 mini, Perplexity, Gemini | 37, 56 and 45 across three cards | HubSpot AEO, $50/month |
| Mangools AI Search Grader | Brand name, niche description | No (three models behind login) | How often a generated prompt set names you | GPT-5 Nano, Mistral Small, Llama 3.3 | 0, and "consumer electronics" | Mangools plans |
| Foglift AEO checker | One URL | No | Page signals: schema, headings, FAQ, crawler access | No AI engine (AI answers are a locked preview) | AI Readiness 76, homepage only | Paid monitoring |
| Ahrefs Free AI Visibility Checker | Brand name | No, but a human check | Mentions across six AI surfaces | ChatGPT, Gemini, Perplexity, Copilot, AI Overviews, AI Mode | No result (blocked by a human check) | Brand Radar from $199/mo |
| LLM Pulse AEO Grader | Website, email, country | Email and a CAPTCHA | An AEO score out of 100, which engines lift your content as the answer, content gaps, competitors side by side | AI Overviews, ChatGPT, Perplexity, AI Mode, Gemini | Not run (email gate) | From EUR 49/month |
| aeograder.org | Your website URL and brand name | No, for the estimate | A Gemini estimate with Google Search grounding | Gemini (free); four engines for $10 | Not run | $10 full report |
What an AEO grader actually measures
The phrase "AEO grader" covers two different tools, and which one you use changes what the score means.
The first kind reads your page. Foglift says its AEO checker "measures how easily an AI answer engine can understand, extract, and attribute a page." It scores eight things: structured data (20%), heading clarity (15%), FAQ quality (15%), content depth (15%), entity identity (10%), citation formatting (10%), AI crawler access (10%) and topical authority (5%). That is a technical audit with an AI label. It is useful. It tells you nothing about whether ChatGPT recommends you.
The second kind asks models about you. HubSpot, Mangools, Ahrefs, LLM Pulse and aeograder.org all generate a set of prompts, send them to one or more models, and count how often your brand appears and how it is described. That sounds like visibility. Whether it is depends on which models they ask, whether those models can search, how many prompts they send and how many times.
HubSpot is the clearest about its own scope. Its grader is "a free, one-time check that reveals what ChatGPT, Perplexity, and Gemini say about you based on their training data." Training data is what a model learned before it was released. It is not what ChatGPT reads when a buyer asks it a question today.
We ran LoudFace through the free graders on 24 September 2026
We used our own brand, because we know the answer. LoudFace is an AI-native organic growth agency for B2B SaaS, and we track how AI engines answer our buyers' questions every day in Peec. So we can check each grader's verdict against a measured number instead of a feeling.
Same inputs everywhere: "LoudFace", "AEO agency for B2B SaaS", United States, and loudface.co where a URL was asked for. No logins, no paid tiers.
HubSpot AI Search Grader: 37, 56 and 45
HubSpot returned three cards, one per model, each out of 100:
- The card HubSpot labels "Powers ChatGPT": 37
- The card labelled "Real-Time AI-Powered Answers": 56
- The card labelled "Supports Google Results": 45
HubSpot weights five dimensions: sentiment (up to 40 points), presence quality (20), brand recognition (20), share of voice (10) and market competition (10). Sentiment carries 40 of the 100 points. So a brand that models describe warmly can score well while rarely being named. Our share-of-voice line read 1, 2 and 1 out of 10.
Every card said "You're on the right track." The full breakdown sits behind an "Unlock Insights" step. Our first submit cleared the form and returned nothing; the second worked.
Mangools AI Search Grader: 0, and we apparently make headphones
Mangools said LoudFace is "present in 0% of the prompts we checked" and gave an AI Search Score of 0.
The more useful part was the brand description. Mangools' free model described LoudFace as a brand that "specializes in bold, expressive consumer electronics and accessories." It then suggested we try niches such as 'loudface distinctive headphones'.
The prompt list had a second problem. Next to the agency prompts it generated from our niche, it listed 15 locked prompts, behind an unlock button, about search engines in general: top search engines for mobile devices, the best search engines for academic journals and the like. The brands it ranked for our category included G2, Capterra, Deloitte Digital and Bain & Company. Those are real companies. They are not who a B2B SaaS founder hires for AI search.
Mangools does say, on the results page, that each run generates a new prompt set, "leading to slightly different results." Take that at face value. You are looking at one draw.
Foglift AEO checker: 76, from the homepage alone
Foglift gave loudface.co a Site Technical Health score of 89, SEO 100 and AI Readiness 76. It flagged a missing Content Security Policy header and "No FAQ section found" on the homepage, and it noted that the AI Readiness score was "based on partial data."
None of that involved an AI engine. The "How does AI see Loudface?" panel was a locked preview, and the free monitoring on offer is a weekly Perplexity check. As a quick technical scan of one page, it is fine. As a read on AI visibility, it is the wrong instrument.
Ahrefs and LLM Pulse: no result
Ahrefs' checker promises "No signup required." On submit it showed a "Verify you are human" check. LLM Pulse asks for your email and a CAPTCHA, then sends the report to your inbox and adds you to its newsletter. We record both as not run. If you run them yourself, expect the same gates.
What we measured over the same 30 days
We track LoudFace's buyer prompts in Peec across ChatGPT, Perplexity and Google AI Overviews. The figures below leave out the four tracked prompts that name LoudFace, so every answer counted replies to a question that does not mention us. For 25 August to 23 September 2026, reading visibility (the share of answers that name the brand) per engine:
| Engine | Answers that named LoudFace | Visibility |
|---|---|---|
| ChatGPT | 1,405 of 5,157 | 27.2% |
| Perplexity | 520 of 5,156 | 10.1% |
| Google AI Overviews | 431 of 4,999 | 8.6% |
On ChatGPT that makes LoudFace second of the tracked agencies, behind Omniscient at 33.8%.
Put that next to the graders and three things stand out.
- HubSpot put our strongest engine last. Its highest card was the one labelled "Real-Time AI-Powered Answers" (56); its lowest was the ChatGPT card (37). Measured over 30 days, ChatGPT is our strongest engine by more than double.
- Mangools' zero is far from the measured number. Mangools found us in 0% of its prompts. ChatGPT named us in more than a quarter of its answers to our non-branded buyer questions.
- Mangools also put us in a category we are not in. Those 1,405 ChatGPT answers named LoudFace in reply to tracked agency-hiring questions that do not mention us by name. A small model, apparently answering from memory, filed us under headphones.
A reasonable objection: our prompt set is ours, and a grader's is not. That is true. A grader picks the prompts for you, and you only learn whether they match your buyers by reading them. Some of Mangools' were about search engines for academic journals.
Why the free graders get it wrong
HubSpot asks a model's memory. Buyers get live search.
ChatGPT, Perplexity and Google's AI features look things up when a question needs current information. OpenAI says "ChatGPT can search the web to answer questions with current information and links to relevant sources." Perplexity says it "uses advanced AI to search the internet in real-time." Google says AI Overviews and AI Mode may issue "multiple related searches across subtopics and data sources" (it calls this query fan-out) and that a supporting page must be indexed and eligible to appear in Search.
Now look at what the graders ask. HubSpot says its check reads training data. Mangools' free tier runs GPT-5 Nano, Mistral Small and Llama 3.3; Claude, Gemini and DeepSeek sit behind a login. Three graders do use live answers. Ahrefs says it queries ChatGPT, Gemini, Perplexity, Copilot and Google AI Overviews with search-backed prompts. LLM Pulse says it runs "live prompts across five AI engines", and it sends the report only after an email and a CAPTCHA. aeograder.org's free estimate is one engine, Gemini with Google Search grounding; its four-engine read is the $10 report.
Our own answer data shows why that matters. In a read of 120 answers from 26 August to 1 September 2026 (40 per engine), LoudFace was named 17 times, and every one of those 17 answers had retrieved a loudface.co page. None came from an answer that had not read our site. For a brand like ours, visibility is produced at answer time, by retrieval. A grader that asks a small model from memory cannot see that. We think that explains both the 0 and the headphones. We cannot prove it from outside Mangools, but the pattern fits.
If your brand is old, famous and written about everywhere, memory and search agree, and a grader will look accurate. If you are a growing B2B SaaS company, what a model remembers about you can lag well behind what live search finds, and a grader that reads memory will report the older picture.
One run is one sample
AI answers change between runs of the same question. SparkToro and Gumshoe had 600 volunteers run 12 prompts through ChatGPT, Claude and Google AI a combined 2,961 times. Rand Fishkin's summary: "If you ask an AI tool for brand/product recommendations a hundred times nearly every response will be unique in three ways: the list presented, the order of the recommendations, the number of items on that list." The study recommends 60 to 100 runs per prompt for a stable read.
A free grader gives you one run of a handful of prompts. That is enough to see whether you exist. It is not enough to tell a 20% brand from a 30% brand, and nowhere near enough to measure month-on-month change.
The score blends things you need apart
HubSpot's total mixes sentiment, recognition, presence and share of voice into one number. Mangools weights models by market share. A single blended score hides the engine you are losing. Our 30-day data shows a spread from 27.2% on ChatGPT to 8.6% on AI Overviews. Each engine also trusts a different set of sources. A blended grade cannot tell you which gap to work on.
Which free grader is good for what
Each tool is fine at the thing it was built for. Use them for that.
- HubSpot AI Search Grader. Best for: a quick read on how models describe your brand's tone and category from memory. Weak on: how often you are named when a buyer asks today. HubSpot's paid AEO tool tracks 25 prompts across three engines for $50/month; read the grader as the free trial of that.
- Mangools AI Search Grader. Best for: a zero-signup sanity check of whether small models know your brand at all. If it describes you as something you are not, that is a real entity problem worth fixing. Weak on: prompt relevance and anything B2B.
- Foglift AEO checker. Best for: a free technical scan of one page (schema, headings, FAQ, crawler access). Weak on: anything to do with what AI engines actually answer.
- Ahrefs Free AI Visibility Checker. Best for: seeing which topics and cited domains show up next to your brand across six AI surfaces, if you get past the human check. It is a preview of Brand Radar, which starts at $199/mo.
- LLM Pulse AEO Grader. Best for: a five-engine snapshot that includes live prompts, if you are happy to trade your email for it.
- aeograder.org. Best for: a grounded Gemini estimate for free, or a four-engine read for $10.
If you need a paid tracker rather than a one-off check, we compared them in our AEO tools ranking.
How to use a free AEO grader without fooling yourself
- Find out what it asks. Model memory, live search or your page's code. The tool's own FAQ usually says. If it says "training data", read the result as a brand-perception check. It does not measure visibility.
- Read the prompts it generated. If they are not questions your buyers ask, the score is about a different market. Rewrite the niche description and run it again.
- Run it three times. If the number moves a lot, you have learned how much one run is worth.
- Read engines separately. A middling average can hide one engine where you are strong and one where you are missing. Find the engine you lose on.
- Check how it describes you. A wrong category or product is the most useful thing a free grader can tell you. It points to an entity problem, and that is fixable.
- Never compare two tools' scores. HubSpot's 45 and Mangools' 0 are on different scales, from different models, on different prompts.
- Never judge an agency on a grader. Ask for tracked visibility per engine, on your buyer prompts, over at least 30 days. We explain what to ask in how to verify an AEO agency's results.
When a free grader is no longer enough
You have outgrown free graders when a decision depends on the number. Signing an agency, cutting a content line, reporting to a board: each needs a fixed prompt set, per-engine readings, many runs per prompt and a comparison against named competitors.
That is what continuous tracking does, and it is how we run our own program and every client's GEO program. It is also how our own case study can show that, on the 23 non-branded prompts we have tracked since April 2026, LoudFace went from 0.06% of AI answers in April to 16.9% over 18 to 28 September 2026: the same prompts, read the same way, over months. Per engine, in that window, ChatGPT named us in 24.9% of its answers on those prompts, Perplexity in 13.8% and Google AI Overviews in 11.8%. One free run tells you far less.
Then take the last step the graders never take. Visibility is a means. What matters is whether AI answers send you pipeline, so we measure AEO against revenue, not against a score out of 100.
If you want that read for your own brand, per engine and against your competitors, start with an AI visibility audit.

