hubspot_unbound_logo_new

Inside HubSpot's AI Search Lab: 7 AEO Experiments

Friday September 18th · 9:45 am - 10:30 am ET · Innovation Stage

Go inside HubSpot's AEO experimentation pod, where hypotheses became strategies scaled across the company. You'll see seven real AEO experiments, including what worked, what failed, and what surprised us, and learn how the team prioritized ideas, measured success, and turned winning tests into repeatable practices that helped make HubSpot the most visible CRM in AI search. Leave with a practical framework for running and scaling AEO experiments on your own team.

Session Speakers:

amanda-kopen-headshot-5n8eao

Amanda Kopen

MANAGER OF EMERGING CHANNELS @ HUBSPOT

Amanda Kopen is a growth marketer specializing in the intersection of AI and demand generation. She's currently focused on increasing brand visibility in LLM outputs while building AI-powered tools that drive measurable growth. With more than nine years in digital marketing,

0:000:00
Inside-HubSpot-s-AI-Search-Lab--7-AEO-Experiments---P1100455Amanda KopenInside HubSpot's AI Search Lab: 7 AEO Experiments

Session Summary

Inside HubSpot's AI Search Lab: Seven AEO Experiments

Answer engine optimisation is no longer a thought experiment. Amanda Copen, Senior Manager of Emerging Channels and AEO at HubSpot, used this session to open the lab door on something most marketing teams are still only theorising about: a small, deliberately lean pod whose entire remit was to find out whether AI search can be influenced at all. The short answer, delivered early and without hedging, was yes — and the rest of the session was the evidence.

What makes the material useful is not the headline claim that HubSpot is the most visible CRM in AI search. It is the working method underneath it: seven experiments, four success metrics, a 20% win rate target, a standing rule of no more than four live projects per experimenter, and an honest scorecard that records losses as losses. Copen framed her own credibility carefully — roughly seven and a half years at HubSpot, six and a half of those in SEO and growth marketing — and polled the room on expertise before starting, with the stated aim of making AI search accessible regardless of background.

Read this recap as a transferable framework rather than a case study. Three of the seven experiments transfer almost directly to any business with a website; two are technical and belong to a development conversation; one failed outright; one failed instructively and is being rebuilt. Together they describe how a discipline is being invented in public, at speed.

Wide-angle view of the session stage in a large indoor venue, with a curved red and purple display screen behind the presenter
The AI Search Lab session opened to a room dominated by practitioners with a traditional SEO background.

How search fundamentally changed

Copen anchored the argument in a short timeline rather than in opinion. January 2023: ChatGPT 3.5 passes 100 million monthly active users, and for a large share of the population search becomes conversational for the first time. May 2024: Google launches AI Overviews, the first major AI integration visible inside traditional results pages. June 2025: Google releases AI Mode publicly, a chat-style interface comparable to ChatGPT and other large language models. By 2026, ChatGPT has more than a billion monthly active users, and AI Overviews and AI Mode have reshaped the results page itself.

The commercial consequence has a name: the great decoupling. Through 2025, impressions climbed steeply while clicks fell away at the same time — the shape Copen described as “alligator arms”. Content is being surfaced more than ever, and rewarded with traffic less than ever. A Seer Interactive study found a 61% decrease in click-through rate on queries carrying AI Overviews, and 68% of searches now ending in zero clicks: a sharp acceleration against the previous decade of gradual erosion.

Slide showing a horizontal timeline titled Search has changed, running from January 2023 through to the present
Three years, four inflection points: from ChatGPT's first mass audience to AI Mode in mainstream search.
Slide titled The Great Decoupling showing a stacked area chart of rising impressions against falling clicks
The decoupling: visibility rising, clicks falling, in the same chart and the same year.

The strategic reading is that the surface being optimised has moved. Ranking on page one is no longer the object of the exercise, because the answer increasingly is the page. That is the premise on which everything else in the session rests.

Defining AEO and why buyers force the issue

HubSpot's working definition is deliberately narrow: answer engine optimisation is the practice of improving how often and how accurately a brand appears in AI results. The two halves matter equally. Frequency without accuracy is a liability, particularly for a company whose pricing and packaging are complex enough to be misdescribed.

The buyer data explains the urgency. A HubSpot study of more than 3,000 CRM purchase decision-makers found 42% had used AI search when evaluating vendors — and that those buyers were 36% more likely to purchase than buyers who did not. Across other sectors the effect is stronger still: depending on the industry, up to 73% of buyers use AI in vendor search, and those buyers can be up to 4.4 times more likely to purchase.

In other words, the AI answer is not a low-intent discovery surface. It is where qualified buyers are shortlisting. Copen's reframing was blunt about what that means for a marketing plan.

Slide titled How we define AEO, defining answer engine optimisation as improving how often and accurately a brand appears in AI results
HubSpot's definition of AEO, kept deliberately tight around frequency and accuracy.
“The question is no longer whether you show up there, but whether you show up accurately, favourably and at the right time — and that's what AEO is.”

Building the AEO experimentation pod

Faced with a landscape changing monthly, HubSpot's marketing team chose action over patience. Since 2023 Copen has helped move the organisation from traditional search to AI search, and the structural expression of that shift was a dedicated pod. It was kept small on purpose: two experimenters and one manager. The goal was stated without qualification — make HubSpot number one in AI search — and the success threshold was set at a 20% experimentation win rate.

Slide titled How we built the AEO Experimentation Pod, describing the pod's composition
A two-person pod with a manager: technical specialist plus content specialist.

The two roles were split along the fault line of the problem itself. The technical specialist monitors crawl logs and finds ways to surface content more efficiently to AI bots. The content specialist watches what actually appears in AI results, tracks citation trends, identifies what the engines get wrong, and defines the new content required to correct it.

Four success metrics were used to judge every experiment: AI visibility or mentions (the brand named without a link), citations (the website given a link), crawl frequency, and secondary organic improvements in traditional SEO. The low 20% bar was intentional, because a pod that must succeed will not attempt anything unfamiliar.

“This may seem low, but it was intentionally low because we wanted to think boldly.”

Copen was equally direct about why HubSpot leads in AI search visibility: “This isn't by luck or by outspending everyone, but by creating dedicated experiments and iterating on what works.”

Experiment 1: FAQ and glossary pages

The first experiment took a classic SEO tactic and rebuilt it for retrieval. The aim was to establish semantic relevance for HubSpot across its core territory — marketing, sales, service and CRM. Copen explained the underlying mechanic plainly: language models treat concepts as mathematical coordinates in multi-dimensional vector space, and closely related concepts sit close together. That proximity is how a model decides whether a business is genuinely related to the products and services it claims.

Crucially, the entries were not plain dictionary definitions. Each was a structured content unit: FAQs aligned to different query intents, giving a unified HubSpot answer and linking through to the relevant site pages. The hypothesis was that a centralised glossary engineered for how AI engines retrieve information would win more citations and lift topical authority.

Result: win. A 36% lift in visibility when HubSpot was cited, content outperforming other HubSpot content by 15%, and rankings for hundreds of keywords within a few months of launch.

  • Build for semantic clarity and query intent, not keyword density.
  • A structured glossary is a high-ROI starting point and can usually be produced without developer time — valuable where engineering capacity is scarce.
  • Start with the foundational concepts your category should own, not the long tail.
Slide for experiment one, FAQ Glossary Pages, showing the result badge and takeaway
Experiment one closed as a win and was later handed off to the content team.
“Content that is built for semantic clarity and query intent, not keyword density, is what AI engines want.”

Experiment 2: Reclaiming 404s

This is the experiment most likely to make a room of SEOs sit forward, because the opportunity was already sitting in the server logs. Through 2025, language models frequently hallucinated URLs — slugs slightly or completely wrong — and sent motivated users straight into 404 pages. Those “ghost URLs” were plainly visible in HubSpot's website logs. Active searchers, already at the door, were being lost on the threshold.

Slide for experiment two, Reclaiming 404s, showing the approach and result
Thousands of page views recovered, and more than 1,000 incremental sign-ups since launch.

The hypothesis: a dynamic module on the 404 page that suggests real, existing pages based on the topics people were actually searching for — the topics the models were getting wrong — would recover AI-referred traffic along with citations and mentions.

Result: win. Thousands of page views recovered and over a thousand incremental sign-ups since launch. Copen added the honest caveat: by 2026 the models are considerably better at producing correct URLs, so the specific tactic has a shorter half-life than the principle behind it.

“There are AEO wins hiding in signals that you have at your disposal, but you just haven't looked at yet.”

Experiment 3: Pricing page accuracy

The trigger was uncomfortable and familiar. Asking ChatGPT about HubSpot products, prices and features returned incorrect answers — a leak in the conversion funnel and, in practice, lost sales. A site audit found the root cause: the pricing pages were beautifully designed for humans, and heavy in JavaScript. AI crawlers do not process JavaScript easily by default, so they guessed, and the guesses were wrong.

The hypothesis was to publish a JavaScript-free, LLM-optimised page that is factual, objective and data-led. The structure is worth copying: a summary; information on payments and seats; onboarding by tier with key features; the trade-offs between tiers, which had never previously been documented anywhere; HubSpot versus competitors; and an explicit section on when HubSpot is worth it for specific businesses and use cases.

Result: win. Accuracy scores improved for five of six product hubs, measured by running structured questions through the AEO tool — for example, “Does HubSpot Marketing Hub allow you to send emails?” and “How much is HubSpot Starter?”

Copen's advice to the room was to treat this as the first experiment anyone should run. Ask ChatGPT about your top products and how they handle their intended use cases. If the answers are outdated or wrong, you have your brief. Then diagnose the cause honestly: either your own information is stale or inaccessible, or third-party review and partner sites are carrying the wrong version of the truth. Both need fixing, and for broad product catalogues, fixing accuracy for one product does not fix the others.

Slide for experiment three, Pricing Page Accuracy, listing the page sections including summary, payments and seats, onboarding and tier trade-offs
The anatomy of an LLM-readable pricing page, including previously undocumented tier trade-offs.

Experiment 4: llms.txt, and the value of a clean loss

If one section of this session deserves to be quoted back at vendors and conference stages for the next year, it is this one. The llms.txt file — a road map of the content a site considers important for language models — arrived with enormous industry hype. HubSpot's hypothesis was reasonable and widely shared: a clear content map should improve how AI engines understand the site.

Slide for experiment four, llms.txt, showing the result marked as lost or abandoned
Zero bot crawlers naturally attempted to visit any llms.txt pathway.

Result: loss. Not a small effect, not an ambiguous effect — zero bot crawlers naturally attempted to visit any llms.txt pathway. Copen called the outcome genuinely surprising, and was careful not to over-generalise: conceptually it still makes sense that an engine would want a clean content map, and HubSpot continues to monitor in case behaviour changes.

The practical guidance for the room was unambiguous. Lean or smaller teams should not spend time building an llms.txt file today. The broader guidance was more important still: run the experiment yourself rather than inheriting the industry's assumptions.

“The lesson is not that llms.txt is bad. The lesson is that you should run your experiment.”

Experiment 5: Use case by industry pages

Two converging signals drove this one. First, AI engines visibly favour content with real-world specificity, semantic structure, user reviews and concrete statistics. Second, language models prioritise personalisation, matching responses to a user's stated likes, dislikes and pain points. Newly released OpenAI ACP (Agentic Context Protocol) guidelines pointed the same way, and HubSpot interpreted them as an instruction to map product features to use cases and pain points in real-world language rather than leaning on branded terminology.

The resulting page template is precise enough to copy: a “challenges → HubSpot solution → real result” table at the top, with real results carrying customer information and case studies; a dedicated deeper case-study module; relevant integrations; an explanation of how HubSpot solves the specific pain point; why HubSpot CRM is the strongest option; and a highly extractable key takeaway written to be lifted directly into an AI response.

Result: win, and beyond expectation. Over 15,000 ChatGPT user crawls within 30 days, and 92% of pages cited within a few months.

  • Keep your core product page, then add specificity around it.
  • Do not build pages solely for language models — they must be genuinely valuable to human readers.
  • Start small: the top three use cases and top one or two industries for your top one or two products. Hundreds of pages are unnecessary.
Slide for the Use Case by Industry experiment showing a won result and a takeaway about specific content with real-world scenarios
Specific content with real-world scenarios proved the strongest citation magnet of the seven.
“The specificity isn't just a nice to have, it's exactly what AI engines want when they're answering a precise user question.”

Experiment 6: Prerendering content

The most technical experiment rested on a single uncomfortable insight: AI crawlers are significantly less sophisticated than Googlebot. Google has more than twenty years of crawling maturity and copes well with heavy, complex pages. The crawlers behind ChatGPT, Claude and Perplexity do not. Where a site leans hard on JavaScript, loads slowly, or presents an unclear structure, those crawlers struggle, make a best attempt, or simply abandon the page.

The hypothesis was that prerendering hubspot.com would improve crawl efficiency and citation eligibility without damaging the human experience. The pod used Bodify Speed Workers; Copen noted prerender.io as an alternative.

Slide for experiment six, Prerendering Content, marked as won with its takeaway
Bot delivery times improved twentyfold, with load time falling from over two seconds to 70 milliseconds.

Result: win. Bot delivery times improved 20x. Page load time fell from over two seconds to 70 milliseconds. ChatGPT user crawls rose steadily and OpenAI impressions doubled.

Copen was clear that this is not a marketer-implemented change: “This is a conversation for your technical, SEO or dev team.” She offered two things a marketer can do. Run the diagnostic: load your page with JavaScript disabled, and if content is missing, the page is AI-unfriendly — as it is if load times are slow. Then ask the development team one question: “Can we check our crawl logs and see how quickly AI bots are retrieving our pages?”

Experiment 7: Call transcript answer pages

The final experiment is still in progress, and it is the most interesting failure of the set. The idea was elegant: mine HubSpot's large library of customer call transcripts to identify the questions prospects asked repeatedly and that the website did not answer directly, then convert those into FAQ-style question and answer pages. The hypothesis was that authentic content drawn from real conversations would grow AI visibility and citations for products and features.

Result so far: loss. No increase in citations or mentions at a level that justified the effort, and the scorecard recorded it as such. But the pod did not treat it as a dead end. The diagnosis was that the concept was sound and the framing was wrong: the team created a new repository of pages when, in hindsight, embedding those Q&As into existing high-traffic live pages would almost certainly have performed better. A rerun on that basis is planned.

The wider point outlives the experiment. First-party data — call transcripts, support tickets, community forums — is an enormous and largely untapped goldmine for AEO, precisely because it records the questions buyers actually ask in the language they actually use.

Slide for experiment seven, Call Transcript Answer Pages, with stylised laboratory glassware imagery
Recorded as a loss, kept for a rerun: the structure failed, not the source material.
“Sometimes a loss tells you to stop. Sometimes it tells you to take an honest look at what you've done and maybe restructure.”

The operating model: four rules that protected focus

Experimentation at this pace fails for management reasons long before it fails for technical ones. Copen therefore spent real time on the four structural rules that kept a two-person pod productive against a field that publishes new findings daily.

  1. A maximum of four projects per experimenter at a time. Everything else sat in a backlog repository. Taking on more diluted progress. Copen described this as the hardest rule to hold, given the constant temptation of new industry results, but four was “our magic number”.
  2. Capacity planning every Monday, taken seriously. The manager's job included protecting the experimenters inside a “bubble”, because their AI fluency made them a magnet for questions from the wider organisation.
  3. Fifteen-minute stand-ups every morning. Described as some of the highest-value time the pod spent: search results page observations, what was and was not working, and what had changed overnight.
  4. Asynchronous end-of-day updates in Asana. This gave cross-team and leadership visibility and, just as importantly, created a documented change log so the rationale behind any decision could be reconstructed later.

The pattern is recognisable: a small team, hard limits on work in progress, daily synchronisation, and written traceability. None of it is AI-specific, which is exactly why it works in a domain where the ground moves weekly.

The eight-part experiment framework and its four verdicts

This was the promised take-home framework. Eight stages, run in order, for every experiment.

  1. Ideation — drawn from AI response trends, industry thought leadership and internal HubSpot signals.
  2. Project charter — the most important artefact: goals, technical requirements, any prototypes built in Claude or ChatGPT, and the definition of success agreed before the experiment ends.
  3. Stakeholder exploration — pressure-test the idea with internal CRO, development and content experts before building anything.
  4. Development — build the components with the dev team, supplying instructions and prompting support.
  5. QA — a minimum of three rounds, testing quality for human users as well as for language models. Issues were found every single time, and each round improved the workflow.
  6. Launch.
  7. Post-launch tracking — checkpoints at two weeks, one month, then monthly.
  8. Verdict.
Slide titled AEO Experiment Framework, how to run your own tests, showing a suggested project flow as a horizontal timeline
The reusable project flow, from ideation and charter through QA rounds to post-launch checkpoints.

The verdict resolves to one of four states, and the discipline lives here. Hand off applies when an experiment is a clear success with a repeatable framework ready to scale; the lab was a lead team that deliberately did not retain delivery work, so the glossary experiment was handed to the content team. Continue covers positive early results with further optimisation available, which the pod's high AI fluency made worthwhile. Kill applies where there is no citation improvement, where there are negative effects such as indexation issues, or where platform changes have made the concept irrelevant. Pause covers experiments blunted by platform change or squeezed out by resourcing.

“These aren't dead, they're just waiting.”

The llms.txt experiment, Copen noted, sits somewhere between kill and pause — parked rather than buried, on the chance the engines start honouring the file.

The pillars shift: from citations to mentions

The most strategically significant slide of the session compared what mattered before 2025 with what matters from 2026 onward. The earlier pillars were citations (the nearest equivalent to a blue link), personalisation (models tailoring responses to likes, dislikes and pain points), passage optimisation (models surface a few sentences or a bullet, not a whole page), and moving further down the funnel (classic inbound top-of-funnel content rarely mentioned the product at all, which AEO cannot afford).

The new pillars are mentions, accuracy, technical accessibility and enabling purchase. The most contested of those is the first: Copen argued that mentions now matter more than citations. The reasoning follows the data. AI search usage is rising and those users are more qualified, and crucially they qualify products and services on the AI results page rather than on the brand's own website. Winning the mention as the most valuable business for a given task beats winning a link to a page the buyer may never open.

Slide titled What's important, comparing two panels of priorities for 2025 and 2026
The 2025 pillars set against the 2026 pillars: citations give way to mentions.
Slide stating that in 2026, AEO requires making your brand's information corroborated, accurate, retrievable and commerce-focused
The reframing in a single line: corroborated, accurate, retrievable, commerce-focused.
“You want your information to be accurate, not just have your information surfaced.”

The playbook you can start next Monday

Copen closed by converting the pillars into a sequence any team can begin immediately, with no dependency on HubSpot's scale.

1. Audit product performance

Build a stack-ranked list of products and services by ROI and profitability — HubSpot keeps this in Excel, most valuable first. Track prompts in an AEO tool (HubSpot uses HubSpot AEO), with five prompts per buying-journey stage per product. It is a large volume, and it yields a genuinely detailed picture of AI visibility. Cross-reference that visibility against the prioritised list, start with the top product, and work down where the leaders already perform.

2. Create an AEO scorecard

Produce one monthly for every product. Many metrics are tracked, but four are highlighted: accuracy, share of voice (mentions versus competitors), raw number of AI citations, and AI citation share.

Slide titled Your AEO Strategy, audit your product performance, showing four rounded statistic cards
Step one of the strategy: rank products by value, then measure AI visibility against that ranking.

3. Establish corroboration and consensus across the web

This is the off-site half of AEO, and it is not optional. Secure listings on third-party directories. Confirm presence on review sites and that the reputation there is positive — and plan remediation if it is not. Win places in “best of” guides for both the product and the company. Be genuinely active in online communities. Invest in affiliate or partner marketing, which Copen framed as the natural evolution of traditional link building.

4. Audit the accuracy of your information

Make product and pricing pages comprehensive, current and easy to reach. Then go outward and correct outdated data held by partner sites, review sites and directories. Accuracy is a distributed property, not a property of your own domain.

5. Confirm your information is retrievable

Remove JavaScript from your most valuable pages where you can, and use a prerendering service where you cannot. Add schema, which is now table stakes rather than an advantage. Monitor whether pages get indexed and, critically, whether they stay indexed. Copen flagged an important warning here: Google has indicated that pages built solely for language models, carrying no human user signals, risk being de-indexed. Teams running AI-generated pages should audit which of them are still indexed and improve their value for human readers, while monitoring crawl frequency over time.

6. Make your web content commerce-focused

The final step is forward-looking. Copen expects AI search engines to become a surface, a marketplace and a vector for purchase — a reality already for some partner sites and products, and one she expects to expand considerably. Preparing means mentioning your company in top-of-funnel content if you are not already doing so, mapping products and services to features and use cases, using natural language to explain how a product solves the problem, and readying the site for bots that both crawl and, before long, purchase on behalf of human users.

Amanda Copen presenting on stage in front of a large curved screen lit in pink and purple
Closing the session: borrow the experiments, repurpose them, and share what you find.

The closing invitation was characteristically open. Borrow HubSpot's experiments and repurpose them for your own business; measure citations, mentions, AI visibility and accuracy over time; and share interesting findings or new experiments with Copen directly on LinkedIn. The lab, in other words, is meant to be copied.

Live Session Transcript

Follow along in real time, then easily copy the full transcript or your favorite snippets to use with your LLM of choice for questions or content creation.

Built on HubSpot - Powered by avgen.ai (beta) by Cat Media