What We Actually Learned Running AI-Search Programs for B2B SaaS Clients (2026)
Six lessons from a year and a half running AI-search programs for B2B SaaS clients, every number pulled live: wedge strategy, the three clocks, the invisible quarter, and why you re-measure your own best case study.
On this page
- Lesson 1: Dominate a corner before you chase the category
- Lesson 2: There isn't one speed. There are three.
- Lesson 3: The "6-12 months" claim is half right, and the half that's wrong is the half that matters
- Lesson 4: The first quarter produces nothing you can put on a slide
- Lesson 5: We once mistook a login spike for AEO working
- Lesson 6: The numbers move, including the ones already published
- Frequently Asked Questions
Two blog posts on our own site, one on AEO agency pricing and one ranking AEO agencies, get pulled into 5.45% of the AI chats across the B2B-SaaS-agency questions we track. When they get pulled in, they get quoted 1.24 times per pull on average. They don't just get retrieved and sit there ignored. That's 557 citations in a single 30-day window, from two pages we built to sell the agency. Farming citations was never the point of writing them. I pulled those numbers this morning, August 2026, filtering our tracked agency prompts to those two URLs, so anyone can reproduce them.
I'm leading with it because it's the only kind of proof that means anything to me anymore. I've spent the better part of a year and a half running AI-search programs for B2B SaaS clients, and I've watched the industry's talking points, and my own past claims, drift out of date at almost the same speed the AI engines themselves update. Six lessons below, what actually held up across those engagements, what I got wrong, and one number I quoted with confidence a few months ago that I'd phrase differently today. Every figure is either pulled live from our tracking as of this morning or checked against a source I can point to.
Most of what gets published about AEO reads like it was written by someone who ran the numbers once, wrote the case study, and never opened the dashboard again. I understand the temptation. A clean win makes a better slide than a moving target. But a program that actually works keeps moving after the case study ships, and the competitive field keeps reacting to whatever worked. I'd rather tell a client what changed than protect a number I already published.
Lesson 1: Dominate a corner before you chase the category
Toku sells stablecoin payroll to crypto and Web3 companies. Pulled live today, Toku shows up in 91.21% of sampled AI chats answering "what are the best stablecoin payroll solutions for crypto and Web3 companies," with an average cited position of 2.7. Ask a broader question instead, roll in the generic global-payroll and EOR prompts Toku is also tracked on, and that number collapses to 21.79% visibility across the generic prompt set.
That gap is the whole strategy working exactly as designed. Toku doesn't compete on "best EOR provider." It competes on "payroll that runs in stablecoins," and outside that lane it mostly doesn't show up at all, because Deel and Remote already own the generic prompts and aren't going anywhere. Fighting them there would mean spending a year of budget to move a number that was never going to move much. Picking the corner they don't care about moved a different number to 91% inside the same year.
| Metric (30-day window, live) | Wedge prompt only | Generic prompt set |
|---|---|---|
| Toku visibility | 91.21% | 21.79% |
| Toku share of voice | 23.38% | 15.54% |
| Avg. position when cited | 2.7 | not applicable, blended |
The mistake I see other agencies make, and made myself early on, is reporting the blended number to a client because it feels like the more complete answer. It isn't. It's the less useful one. A client who owns a narrow wedge at 91% and hears "21.79% visibility across your category" will conclude the program isn't working, when the program is doing exactly what a wedge strategy is supposed to do. Reporting the blended figure without the wedge figure next to it is technically honest and still the wrong number to lead with, and leading with the wrong number is how a program that's working gets cancelled by someone who never saw the number that mattered.
I wrote a longer breakdown of why picking a sub-category nobody owns beats chasing the whole market, and the Toku numbers are the cleanest live proof of it I have. The full arc of that engagement, from zero to the wedge number above, is in the Toku case study.
Trying to contest them there would have meant spending a year of program budget to nudge a number that was structurally never going to move much, in exchange for a slide that says the brand is now barely visible on a hundred generic prompts instead of invisible on all of them. Nobody signs a renewal for a slide like that.
The same logic holds outside Toku, even where I can't hand you a live number as clean as this one. Every B2B SaaS client I've run this program for has a category with an entrenched leader nobody's dislodging in the next eighteen months, and a narrower slice of that category where the leader has never bothered to show up in force. The instinct most marketing teams have is to go pick a fight with the leader on the leader's own turf, because that's where the search volume lives on paper. The move that actually produces a number worth reporting is finding the slice the leader ignored and owning that slice completely, then letting the case study argue for the wedge on its own instead of arguing for it in a sales deck.
By engine, the split is lopsided
The wedge prompt breaks down unevenly by AI engine. Perplexity puts Toku's visibility at 96.77%, position 1.9. ChatGPT sits at 93.55%, position 3.8. Google AI Overview comes in at 82.76%, position 2.4. Most people I talk to optimize for ChatGPT first, because it's the one they personally have open in a browser tab all day. Across the generic prompt set, though, Google AI Overview actually carries the largest single share of Toku's total AI mentions at 53.7%, ahead of ChatGPT's 26.8% and Perplexity's 19.4%. If you're only watching ChatGPT, you're watching the smaller half of the picture, and you're watching it because it's the one you personally use. It isn't the one that moves the most volume.
Lesson 2: There isn't one speed. There are three.
Clients ask "when will this work" as if AI citation moves at a single pace. It doesn't. I've come to think about it as three separate clocks running at once inside the same program, and conflating them is the single fastest way to burn a client's trust in the first quarter.
The first clock is hours to days. A clean, well-formatted answer block on a page that already ranks can get pulled into a live AI answer almost immediately, because engines like Google AI Overview build responses at query time off the current index rather than answering purely from a fixed, pre-computed snapshot. The second clock is weeks. Early, narrow citations on long-tail questions start appearing while the broader entity graph is still forming underneath everything else. The third clock is months. Stable, dominant share on your head terms takes quarters to compound, because it depends on an accumulated trust signal an engine builds about your brand over time, across many pages and many mentions. A single well-written page shipped last Tuesday doesn't build that on its own.
The trap is that the first clock is real and it's fast, and watching it move teaches everyone involved, client and agency both, to expect the third clock to move at the same speed. It won't, and pretending it might is how a program earns an unfair reputation for being slow in month four, right when the second and third clocks are actually doing the work that was always going to take that long.
| Phase | Typical window | What's actually happening |
|---|---|---|
| Foundation | Days 1-30 | entity setup, structured content, no visible output |
| First citations | Weeks 2-8 | narrow, long-tail AI mentions start appearing |
| Referral growth | Days 60-180 | measurable traffic starts showing in analytics |
| Compounding | Months 6-12+ | share on head terms stabilizes and becomes dominant |
I wrote the full mechanics of this out separately in how long AI citations actually take, because it's the single most common source of a client conversation going sideways somewhere around month two, right at the point where the first clock has already delivered something small and the third clock hasn't delivered anything yet.
The practical fix is naming which clock a number belongs to before you show it to anyone. A same-day AI Overview pickup on a fresh answer block is a real result, and I'll take credit for it, but I won't let a client generalize from it to "so the whole program should move this fast." A long-tail citation appearing in week three is a real result too, and it tells you the second clock started on schedule. It says nothing about whether the third clock is anywhere close to finished. Treating all three as one undifferentiated "AEO progress" number is how a program that's on track gets read as behind schedule by someone comparing the wrong clock to the wrong expectation.
Lesson 3: The "6-12 months" claim is half right, and the half that's wrong is the half that matters
This is the industry line I hear most often, usually stated as a flat fact with no qualifier attached: AEO takes six to twelve months to work. I went looking for where that number actually comes from before writing this, rather than assume it's a strawman other agencies invented to sound authoritative, and it turns out it's real. It's just describing something narrower than it sounds like it's describing when someone repeats it in a sales call.
Two agency sources published in recent months make close to the same claim, independently of each other. One states that first AI mentions typically appear within one to two weeks, with consistent citation patterns taking eight to twelve weeks, then frames six-plus months as the point where ROI becomes stable, arguing that timeframe is "the time necessary to build the foundation, earn early citations, stabilize them, and accumulate enough attribution data to prove revenue impact." The other lays out a near-identical phase structure: foundation in the first 30 days, first citations in weeks two through eight, measurable referral traffic by day 90, and compounding pipeline impact across months six through twelve, adding that "the companies that quit at day 60 usually stop right before the compounding starts."
What both sources actually agree on, buried under their own headline
Read past the six-to-twelve-month headline and both sources say first pickup is fast. One to two weeks in one case, two to six weeks in the other. The six-to-twelve-month figure they're both actually describing is the climb to stable, dominant share, not the wait for any AI mention at all to show up. That's the correction I'd make to the industry talking point. It's a precise fix. It isn't a rejection of the whole idea. Dominance taking months is accurate. The claim breaks when it gets flattened, in a pitch or a client kickoff, into "you'll see nothing for six months," because that's a different claim, a much scarier one, and it's the one that makes a client quit at the exact point where the foundation is about to start paying off.
There's a mechanical reason the fast clock and the slow clock coexist on the same program. Google's own developer documentation describes AI Overviews and AI Mode using a technique it calls "query fan-out," issuing multiple related searches at the moment a question is asked, rather than answering purely from a fixed, pre-computed snapshot. That's live, query-time retrieval, and it's a big part of why a single strong page can get pulled into an answer almost immediately after it's published or updated. Compare that to a large language model's own base knowledge, the thing people usually mean when they picture "AI." OpenAI's published model documentation lists its current flagship family's knowledge cutoff as a fixed date, locked until the next model release actually ships. One surface reads the live web every single time it answers a question. The other is frozen between releases, blind to anything published after its cutoff unless a browsing tool gets invoked mid-conversation. Google AI Overview carrying 53.7% of Toku's total mentions is the fast clock running on the one surface built from the ground up to run fast, and everyone chasing the slow-moving model instead is fighting the wrong clock.
I want to be precise about one thing here, because it's easy to overstate. Google's documentation doesn't commit to a specific "within hours" freshness figure, and I'm not going to put an exact service-level number in Google's mouth that Google itself hasn't published. What I can point to is the mechanism, live retrieval versus a fixed cutoff, and let Toku's own recomputed number carry the concrete proof instead of a made-up SLA. That's the difference between describing how a system works and inventing a stat to make the description sound more precise than it actually is, and it's a distinction I think most AEO content skips past because the vague version sounds more confident.
The practical upshot for anyone signing a contract on the strength of the "6-12 months" line: ask which half of it the agency is promising. If they mean you'll see your first AI mention somewhere in that window, walk away, because both of the sources making this claim publicly, sources with every incentive to make AEO sound slower and harder than it is, put first pickup at weeks not months. If they mean stable, dominant share on your head terms will take that long, that's a defensible claim and roughly matches what I've seen across our own engagements. The words "six to twelve months" can describe either one of those outcomes, and a client who doesn't ask which one is buying a much vaguer promise than they think they're buying.
Lesson 4: The first quarter produces nothing you can put on a slide
Every agency, including us on our worse days, skips the foundation quarter in the pitch deck, because there's nothing in it worth putting on a slide. No citation count worth showing a prospect. No dashboard line trending up and to the right. Just structured content going in, an entity graph slowly forming underneath the surface, and a client staring at a flat line while the invoice for that quarter arrives right on schedule.
I wrote about this at length in the invisible quarter, because it's the exact stretch where a program that's genuinely working and a program that's quietly failing look identical from the outside. Both of the industry sources behind the "6-12 months" claim above independently describe the same invisible early phase from their own side of the table, with no reason to agree with each other on it: foundation work in the first 30 days that produces no visible citation and no measurable traffic yet. That convergence is worth sitting with. Two agencies with different clients, different methodologies, and no incentive to validate our framing landed on the same shape of the first month, purely because the shape is a structural fact of how these systems work. Neither side is telling a story to sound smart.
The tension isn't a mystery once you sit on both sides of the table. Finance sees a real invoice going out and a dashboard reading zero, and "we spent real money and we're in zero AI answers" is a hard sentence to defend in a budget review, regardless of whether the foundation is two weeks from compounding. The person who approved the program is the one spending down their credibility defending a flat line, and cutting the program looks decisive in a way that waiting doesn't. On our side of the table, an agency staring down a renewal conversation it isn't confident about feels the same pull toward manufacturing visible motion, chasing a same-day AI Overview pickup on an easy page instead of finishing the harder foundation work that doesn't show up for another month. Both sides end up optimizing for the next meeting instead of the actual result, and neither side is being dishonest exactly, just human.
What to watch instead of the citation count during the quiet stretch
The leading indicators that actually tell you whether the quiet stretch is working, before a single citation shows up anywhere, are branded search creeping up on terms that didn't exist before the program started, impressions accumulating in Search Console even at bad average positions, and AI crawler frequency rising in server logs as the engines start reading the new structure. None of those show up in a citation counter. All three tend to show up weeks before the citation counter does, if anyone bothers to watch them instead of refreshing the dashboard that isn't moving yet.
Lesson 5: We once mistook a login spike for AEO working
Here's the mistake, and I'm naming it as one rather than dressing it up as a hypothetical cautionary tale about some unnamed agency. Early in an engagement, we watched a client's branded search volume climb and treated the climb as a sign the AEO program was landing. It wasn't, or at least not entirely. A meaningful chunk of that spike turned out to be existing users searching the brand name to find the login page and get into their own dashboard. It had nothing to do with new buyers discovering the brand for the first time through an AI answer somewhere upstream. We caught it, corrected the read before it reached the client, and it changed how I look at every branded-search chart handed to me since.
The fix is a simple rule I now apply before reporting any branded-search number to any client: new, non-branded-intent queries that didn't exist before the program started are signal. Branded queries carrying login, sign in, dashboard, or account intent are noise. That's an existing customer base going about a normal Tuesday, and it says nothing about a program result. Mixing the two inflates the number in the short term and, worse, teaches a client to expect a repeat of something that was never really about the AEO work in the first place, which is a much harder conversation to have two quarters later when that number quietly flattens out.
| Query pattern | What it actually tells you |
|---|---|
| New, non-branded questions the brand wasn't ranking for before the program | Real signal, the AEO program surfacing the brand to people who didn't already know it |
| Branded queries carrying login, sign in, dashboard, or account intent | Noise, existing users doing routine account access, unrelated to new-buyer discovery |
I keep coming back to how easy that mistake was to make, and how easy it would have been to never catch. Branded search going up looks like a win from across the room. Nobody's instinct, mine included, is to open the query list and check whether "login" or "sign in" is doing the heavy lifting. The instinct is to screenshot the chart. The discipline that actually protects a client is checking the ugly, unglamorous query-level detail before the good-looking chart goes anywhere near a client. That discipline has to survive on the days when the pressure is to show a win, and it can't only show up on the easy days.
Why I'm telling you this instead of the version where we got it right the first time
It would be easier to write a piece full of clean wins. It would also be a worse piece, and less useful to anyone actually running one of these programs. I can't re-pull the exact underlying query-level numbers behind that specific catch from where I'm sitting today, and I'm not going to invent a figure to make the story sound more dramatic than the correction actually was. What I can tell you is that the rule survived the mistake. I apply it to every branded-search chart that lands on my desk now, on every client we run, including the ones that never made a mistake like this one.
Lesson 6: The numbers move, including the ones already published
Toku's stablecoin-payroll wedge sits at 91.21% visibility today, share of voice at 23.38%, average position 2.7, all pulled live this morning before I started writing. I've quoted a lower visibility figure for this same engagement in earlier pieces of mine, and the honest thing to say is that the number moved up since then, which is good, and something else moved with it at the same time that isn't as clean a headline.
On this same wedge prompt, Toku is no longer the only name that shows up. Since I first published a figure for this engagement, other providers have moved into the same lane, and the gap at the top has tightened on a live pull. That isn't a pricing comparison, and it isn't a recommendation to switch providers, and I'm not going to turn a citation-accuracy note into either one. It's a plain fact about where the tracked question sits this morning, and it means the framing I might have leaned on a few months ago, that Toku had this wedge locked down uncontested, doesn't hold cleanly anymore. The field noticed the same wedge we did, and others moved in.
I'd rather say that out loud than quietly edit an old page on the next pass and hope nobody checks the source. A wedge strategy doesn't grant permanent ownership of anything. It grants a head start, and the size of that head start needs re-measuring off a live pull rather than repeated from whatever number sat in a case study six months ago and hasn't been touched since. We ran our own AEO playbook on our own site for exactly this reason, going from 0.18% to 10% of AI answers in a window we measured ourselves rather than took someone's word for, documented in full in the case study, and the discipline of re-measuring instead of repeating is the same discipline either way, whether it's a client's number or our own.
There's a version of this lesson that would be more comfortable to skip past, the one where I frame the field closing in as a threat to manage rather than a fact worth stating plainly. I don't think it is a threat. Other providers moving into a prompt Toku already dominates is confirmation the wedge was worth taking, not evidence it's failing. The number that would actually worry me is a wedge nobody else bothered to contest after eighteen months, because that would mean the prompt volume behind it was never worth fighting for in the first place. Competitive pressure on a number you built from nothing is a better problem to have than silence.
Where I actually land on all of this
If a number in an AEO scorecard is more than thirty days old, treat it as a historical fact about the program. It's a weaker claim about where things stand today than most people treat it as. Anyone still quoting last quarter's citation rate as this quarter's proof either hasn't checked recently or doesn't want to know the answer, and both of those are worth asking about out loud in the next client call. Pull it live, every time, including your own best case study, especially your own best case study. Mine changed between the version I originally wrote and the version I checked this morning before sitting down to write this one, and I'd rather tell you that than let the old number keep doing work it hasn't earned in months.
None of this is an argument against wedge strategies, invisible foundations, or first-clock wins. Every one of those held up under a live check today, which is more than most of what gets published in this category can say for itself. It's an argument against treating any of them as finished. A wedge gets contested, a foundation eventually shows results, and a fast citation on one page never proves the slow work everywhere else is done. Run the check again in ninety days. I plan to, and I'd rather be the one telling you the number moved than have a client find out from someone else's dashboard first.
Frequently asked questions
Answers to the questions readers ask most about this topic.
Lesson 1: Dominate a corner before you chase the category
Toku sells stablecoin payroll to crypto and Web3 companies. Pulled live today, Toku shows up in 91.21% of sampled AI chats answering "what are the best stablecoin payroll solutions for crypto and Web3 companies," with an average cited position of 2.7. Ask a broader question instead, roll in the generic global-payroll and EOR prompts Toku is also tracked on, and that number collapses to 21.79% visibility across the generic prompt set.
Lesson 2: There isn't one speed. There are three.
Clients ask "when will this work" as if AI citation moves at a single pace. It doesn't. I've come to think about it as three separate clocks running at once inside the same program, and conflating them is the single fastest way to burn a client's trust in the first quarter.
Lesson 3: The "6-12 months" claim is half right, and the half that's wrong is the half that matters
This is the industry line I hear most often, usually stated as a flat fact with no qualifier attached: AEO takes six to twelve months to work. I went looking for where that number actually comes from before writing this, rather than assume it's a strawman other agencies invented to sound authoritative, and it turns out it's real. It's just describing something narrower than it sounds like it's describing when someone repeats it in a sales call.
Lesson 4: The first quarter produces nothing you can put on a slide
Every agency, including us on our worse days, skips the foundation quarter in the pitch deck, because there's nothing in it worth putting on a slide. No citation count worth showing a prospect. No dashboard line trending up and to the right. Just structured content going in, an entity graph slowly forming underneath the surface, and a client staring at a flat line while the invoice for that quarter arrives right on schedule.
Lesson 5: We once mistook a login spike for AEO working
Here's the mistake, and I'm naming it as one rather than dressing it up as a hypothetical cautionary tale about some unnamed agency. Early in an engagement, we watched a client's branded search volume climb and treated the climb as a sign the AEO program was landing. It wasn't, or at least not entirely. A meaningful chunk of that spike turned out to be existing users searching the brand name to find the login page and get into their own dashboard. It had nothing to do with new buyers discovering the brand for the first time through an AI answer somewhere upstream. We caught it, corrected the read before it reached the client, and it changed how I look at every branded-search chart handed to me since.
Lesson 6: The numbers move, including the ones already published
Toku's stablecoin-payroll wedge sits at 91.21% visibility today, share of voice at 23.38%, average position 2.7, all pulled live this morning before I started writing. I've quoted a lower visibility figure for this same engagement in earlier pieces of mine, and the honest thing to say is that the number moved up since then, which is good, and something else moved with it at the same time that isn't as clean a headline.



