GrowthHasten

Voice Search SEO: What Actually Changed and What to Skip

Voice assistants stopped being keyword matchers and became language models, which quietly killed most of the voice search playbook. Here is what actually changed, what got absorbed into AEO, and how to decide whether voice deserves any of your budget.

Anshuman Sinha

Written by Anshuman Sinha

Published August 15, 2026
Updated August 15, 2026
13 min read
Professional microphone on an adjustable arm in a sunlit home office

Voice search SEO is the practice of shaping content so that a spoken question returns your page as the spoken answer. For most of the last decade that meant a fixed checklist: conversational keywords, featured snippet capture, a fast mobile page, and a tidy Google Business Profile. Then Google announced in August 2025 that Gemini for Home would replace Google Assistant on existing speakers and displays rather than upgrade it in place, having already begun moving mobile users off the classic Assistant, while Amazon and Apple took their own assistants in the same generative direction and ChatGPT added a voice mode. That turns "how does an assistant choose an answer" into the same question AI search already answers. This guide is for growth leads and founders at B2B SaaS and AI companies deciding whether voice search optimization earns a line in next quarter's plan, and it gives you a framework for making that call.

The short version

  • The assistants are language models now, so "optimize for voice" and "optimize to be quoted by an AI" became the same job.
  • Most of the 2018 to 2022 voice playbook is either dead or absorbed into answer engine optimization you should already be doing.
  • Voice still looks like it skews local and immediate, though no platform publishes the data to prove it. On that reading it matters far less to a B2B software company than the agencies selling voice packages suggest.
  • Google documents speakable as a beta property that works for US Google Home users with English-set devices and for publishers publishing in English, and the only use it documents is answering topical news queries. Almost no reader of this article can use it.
  • No platform publishes what share of searches are spoken, so every figure in circulation is an estimate you cannot check.

What is voice search SEO, and is it still a separate discipline?

It is optimizing so that a spoken query returns your content as the spoken answer, and for most teams it stopped being a separate discipline around the time assistants started running on language models.

The old version deserved its own tactic list. Assistants were closer to keyword matchers with a text-to-speech layer on top: they drew from a narrow pool of eligible sources, leaned hard on featured snippets, and rewarded a specific set of formatting choices. Mechanics that different justified a separate workstream.

The current version does not. An assistant retrieves candidate passages, and a model decides which one to speak. That is the same shape as a Google AI Overview, a Perplexity answer, or a ChatGPT response with citations. Our guide to what AI search changes and what it doesn't covers the pipeline underneath, and the practical work of being the source a model reaches for sits in our guide to how AI answer engines choose what to cite. Voice is an output format for that work. It is not a parallel program.

What actually changed when assistants became LLMs?

The retrieval layer barely moved. The selection layer moved completely, and that is what quietly invalidated most of the tactic lists still circulating.

The 2018 to 2022 tactic Status now What it became
Target conversational, question-shaped long-tail phrases Still true The best-surviving tactic on the list, and now general practice. Question-shaped headings serve voice, AI answers, and snippets from one edit.
Win the featured snippet, because that is where the spoken answer comes from Absorbed Snippet-worthy structure still helps, but as one input to extraction rather than the pipe the answer travels down.
Keep answers near 29 words, because that is the average spoken answer length Dead as a target A generated answer is written to fit the question, not to a word count measured on a 2018 device.
Add speakable structured data so assistants can read your page aloud Dead for almost everyone A beta property in Google's documentation, scoped to US Google Home users and English-language publishers, and documented only for topical news queries. Not general-purpose voice markup.
Optimize Google Business Profile and local citations Still true, if you are local Unchanged in value and unchanged in scope. It was always local SEO wearing a voice label.
Hit sub-five-second load times as a voice-specific goal Absorbed Core Web Vitals and mobile performance, which you are measured on regardless of surface.
Run voice-specific keyword research in voice keyword tools Dead No platform separates spoken queries from typed ones, so nothing was ever measuring what the tools claimed.

What I've seen in practice is that teams keep optimizing for the featured snippet long after the answer stopped coming from one. The instinct is reasonable, since snippet capture genuinely was the highest-leverage voice tactic for years. The work is not wasted either, because a page built to win a snippet is usually built well for extraction in general. But the mental model behind it, win the box and you win the voice answer, no longer describes what happens.

How is a spoken query genuinely different from a typed one?

Three differences survive, and they matter because they change what you write, not just how you mark it up.

  • Phrasing: People speak in full sentences. Typed queries compress to "voice search seo worth it"; spoken ones arrive as "is voice search optimization still worth doing." Longer and more natural, which is why question-shaped content reads well aloud.
  • Result count: One answer, spoken, with no list to scan. There is no second place and no below-the-fold, so the gap between being chosen and being invisible is total.
  • Context: Voice happens in cars, kitchens, and shop floors, with hands busy and attention divided. That environment points the query mix toward local, immediate, and transactional questions harder than typed search does. No platform publishes the split by input method, so treat that as reasoning from context, not a measured fact.

Everything else commonly listed as a voice difference now describes AI search generally. Answers get synthesized rather than linked, sources get named rather than ranked, and clicks often never happen. None of that is specific to speech.

They draw on the same systems those features are built on, which is a narrower claim than the one usually made.

Google's documentation on AI features in Search states that a page must be indexed and eligible to be shown in Google Search with a snippet before it can appear as a supporting link in AI Overviews or AI Mode, that no additional markup or files are required, and that the same foundational SEO practices apply. Those AI surfaces inherit the crawling, indexing, and ranking systems the rest of Search runs on.

Here is the line between fact and pattern. Documented: AI surfaces run on Search's crawl, index, and ranking systems, and assistants built on those systems inherit them. Observed, not documented: the source that holds the snippet or appears in the Overview is frequently the source an assistant names. Plan against the first and treat the second as a useful correlation. Our breakdown of how AI Overviews pick their sources and our guide to winning the featured snippet cover both ends of that relationship.

Should voice search be a line item in your plan?

Usually not, and the reasoning matters more than the verdict. Four inputs decide it, and you can answer all four in about ten minutes.

Input Argues for dedicated voice work Argues against it
Query type Local, transactional, or immediate-need questions drive your pipeline Considered, informational, multi-session research drives your pipeline
Device and context Buyers ask while hands are busy: driving, walking a floor, standing in a store Buyers research at a desk, in a browser, with several tabs open
Who decides One consumer asks and acts in the same moment A buying committee evaluates over weeks, using documents, demos, and trials
Answer baseline Your pages are already extractable, so the remaining gains are voice-specific Your pages are not yet extractable, so general answer work returns more per hour

Count the answers that landed in the middle column, then read the verdict.

Verdict 1. It is a real channel: three or four in the middle column. Fund it, and expect the work to look mostly like local SEO with careful answer formatting on top. Retail, hospitality, home services, clinics, and multi-location businesses land here regularly.

Verdict 2. It is already covered: one or two in the middle column, with an answer baseline that exists. Do nothing voice-specific. The answer engine optimization you are already running is the work, and adding a voice project on top mostly buys you a second name for the same tasks.

Verdict 3. Skip it and fix the baseline: zero or one in the middle column, with an answer baseline that does not exist yet. Voice-specific effort here is optimizing for a surface that will not reach you either way. Build extractable answers first. They serve assistants, AI Overviews, snippets, and ordinary rankings at once.

Based on my experience, voice optimization has never justified a workstream separate from answer engine optimization for a B2B software company. Verdict 2 is where nearly all of them land, and the honest version of that finding is not that voice does not matter. It is that the work still gets done, under a different name, with a wider payoff.

What does schema markup actually do for voice, including speakable?

Schema helps machines parse and attribute your content. It does not buy you a spoken answer, and the one property named after speech is almost certainly not for you.

Google's documentation for speakable structured data describes a beta feature that applies to users in the United States with Google Home devices set to English, and to publishers publishing in English. It explains that Google Assistant uses the property to answer topical news queries on smart speakers. Read as written, that is a news feature, in beta, in one country, for one device class.

Voice search guides still recommend speakable generically, which is how a property with one documented use, topical news answers on smart speakers, ended up on checklists written for dentists and SaaS startups. Adding it to a pricing page does nothing. It will not be penalized either, which is precisely why the recommendation survives: nobody notices the null result.

What earns its place instead: Article, FAQPage, Organization, and Product where they genuinely describe the page. These help every surface understand what you published and who published it. Our schema markup guide covers which types are worth the implementation cost and which are decorative.

What does the work look like if voice genuinely matters to you?

Like answer engine optimization with a local layer on top. Nothing on this list is voice-only, which is the whole argument.

Answer formatting: Question-shaped headings that match how people actually ask. A direct answer in the first sentence under each one. Passages that survive being lifted out of the page, meaning they carry their own context and define their own terms.

Entity and local hygiene: Consistent business name, address, and phone across your site, Google Business Profile, and major directories. Assistants resolve entities before they answer, and conflicting records are a common reason a business gets skipped.

Technical baseline: Crawlable, indexed, rendered, and fast on a phone. Assistants live on phones and speakers rather than desktops, so the work in our guide to mobile-first performance is doing double duty here.

Verification: A fixed set of the questions your buyers ask, spoken into Gemini on a phone and on a smart speaker, Siri, Alexa, and ChatGPT voice mode on a schedule. Record who gets named. It is manual and it is the only direct evidence available.

How do you measure voice search, and why are the usual numbers unreliable?

You largely cannot, and saying so plainly is more useful than a confident estimate.

Search Console does not separate spoken queries from typed ones. GA4 has no voice dimension. An assistant that answers aloud without sending a click leaves no trace in your analytics at all. The measurement gap is not a tooling oversight you can work around; it is structural, and it explains why the statistics in this category are so poor.

Confident claims about the share of searches that are spoken, and confident predictions about the year voice overtakes typing, circulate constantly and get repeated in agency decks. Chase any of them back through the citation chain and it ends at another blog post. This article quotes none of them, hedged or otherwise, because a number nobody can source is not evidence.

The dataset this category still leans on is Backlinko's analysis of 10,000 Google Home results, published in February 2018. It was careful work at the time. It also measured a device generation and a selection mechanism that no longer exist, and its famous figures, 29-word answers and 4.6-second load times, are now eight years old.

What you can genuinely track:

  • Long, question-shaped queries in Search Console: a proxy, not a measurement. It captures conversational intent regardless of how the query was entered, which is arguably the more useful signal anyway.
  • Scheduled assistant spot checks: the same prompt set, spoken into the same assistants, on the same cadence. Answers vary between runs, so a fixed list beats an ad hoc one.
  • Referral traffic from AI sources: some assistants pass a referrer. It undercounts badly, but the trend line is real.

Anything beyond those three is an estimate wearing a number's clothing.

When should you not bother with voice search optimization?

When your buyers are a committee, your footprint is not local, and your answer baseline is not built yet. That describes most B2B SaaS companies, which is why this article exists.

  • Committee purchases: nobody signs an annual software contract by asking a smart speaker. Voice may show up in early discovery, but the evaluation moves to a browser almost immediately.
  • No physical footprint: a large share of the classic voice playbook is Google Business Profile work. Without locations, that entire section of the checklist is inert.
  • A weak answer baseline: if your pages bury their answers three paragraphs down, voice-specific tactics have nothing to act on. The fix is upstream and benefits every surface.
  • No way to verify: committing budget to a channel you cannot measure, on the strength of statistics that trace back to nothing, is how voice search line items quietly renew for years without evidence.

The trade-off runs the other way for local and consumer businesses, where voice is a genuine acquisition path and the local work pays for itself independently. The mistake is not investing in voice. It is investing in voice as something separate.

The habit worth building is an extraction audit you run on every important page, because it is the one piece of work that serves voice, AI Overviews, assistants, and ordinary rankings from a single edit. This week, take your five highest-intent pages, read the first sentence under each heading out loud, and rewrite any that do not answer the heading on their own. Reading them aloud is not a gimmick: a sentence that sounds awkward spoken is usually borrowing context from the paragraph above it, and borrowed context is exactly what stops a model from lifting it cleanly.

Ready to Grow Your Organic Traffic?

If you want better rankings, more qualified traffic, and organic growth that holds up across Google and AI answer engines, GrowthHasten can help.

Talk to an SEO Expert
FAQ

Frequently Asked Questions

Is voice search SEO still worth doing in 2026?

For most B2B and SaaS companies it is not worth running as a separate project, because the assistants now answer from the same retrieval and ranking systems as AI Overviews. The work that wins voice answers is answer engine optimization: extractable passages, question-shaped headings, and a clean entity footprint. If you are already doing that, voice is covered. If your business depends on local or transactional queries, it deserves separate attention.

How is voice search different from typed search?

Three differences survive the shift to LLM assistants. Spoken queries are longer and more conversational because people speak in full sentences. The assistant usually returns one answer rather than a list, so there is no second place. And voice appears to skew harder toward local and immediate-need queries than typed search does. Everything else about how the answer is chosen now mirrors ordinary AI search.

Should I add speakable schema to my pages?

Almost certainly not. Google's structured data documentation describes speakable as a beta property that works for users in the United States with Google Home devices set to English and for publishers publishing in English, and the only use it documents is Google Assistant answering topical news queries on smart speakers. It is not general-purpose voice markup, despite being recommended widely in older guides. Spend the effort on Article, FAQPage, and Organization schema instead.

Do voice assistants use Google AI Overviews?

They draw on the same underlying systems rather than reading Overviews as a product. Google's documentation states that AI features in Search require a page to be indexed and eligible to be shown with a snippet before it can appear as a supporting link, and that no special markup or optimization is needed. Assistants built on those systems inherit them. The practical consequence is that a page structured to be lifted into an AI Overview is already structured to be read aloud.

How do I measure voice search traffic?

You largely cannot, and any tool claiming a precise figure is estimating. Search Console does not separate spoken queries from typed ones, and assistants that answer without a click leave no analytics trace at all. The workable substitute is tracking your performance on long, question-shaped queries and running periodic spot checks by asking the assistants your target questions directly, using the same prompt set each time.

Is voice search only useful for local businesses?

Not only, but the pattern is worth planning around. Voice appears to skew toward local and immediate-need questions, which is why the tactic lists in older guides lean so heavily on Google Business Profile, though no platform publishes the query mix by input method. For a B2B software company selling to a buying committee, that reading means voice is a minor channel, and the effort is better spent on the AEO work that serves every AI surface at once.

Share This Article

Anshuman Sinha
Written by

Anshuman Sinha

AI SEO Specialist, GrowthHasten

Anshuman Sinha is an AI SEO Specialist and Computer Science Engineer with over three years of experience in SEO and five years in web development. He specializes in Technical SEO, AI Search Optimization (AEO and GEO), SaaS SEO, and building high-performance websites with modern technologies.

View profile

Stay Ahead Of The Curve

Get the latest SEO insights and growth strategies delivered to your inbox. No spam, just actionable advice.