Customer Impact

SEO & GEO

Speakable Schema for Voice Search and AI Assistants

Copy for AI

Speakable schema is a piece of structured data that tells search engines and assistants which sections of your page are best suited to be read aloud. With a CSS selector or XPath, you literally point out the sentences a voice assistant can return as an answer. In this article you will read how the type works, which sections you do and do not mark, and what the honest expectation is for B2B in 2026.

This is one of the least explained schema types, and for good reason: its official use is narrow. Yet the concept is relevant for generative engine optimization, because it does exactly what AI systems need: point to a clearly bounded, citable answer.

What is speakable schema?

Speakable schema is the speakable property from the schema.org vocabulary that marks sections suitable for text-to-speech. In practice, you add a SpeakableSpecification to your Article or WebPage markup that references specific pieces of text. An assistant processing the page then knows which passage is meant to be read aloud, instead of having to guess for itself which part of the page is the best answer.

You recognize it by two ways of pointing to content. With cssSelector, you refer to a class or element on the page, for example the summary at the top of an article. With xpath, you refer to a node in the document structure. The two are mutually exclusive: you pick one per specification. The idea is always the same: you isolate the sentences a person would read aloud as the answer to a question.

Unlike many other schema types, speakable produces no visible result in the search results. It is not a rich snippet and not a star rating. It is purely an instruction about read-aloud suitability.

How does speakable schema work technically?

You implement speakable as JSON-LD in the <head> or body of your page, just like other structured data. Within your existing Article object, you add a speakable block that contains one or more selectors.

In broad terms, it looks like this:

  • You define a WebPage or Article object for the page.
  • Inside it you set "speakable": { "@type": "SpeakableSpecification", "cssSelector": ["..."] }.
  • The selector points to the exact sections you want read aloud, for example a short summary or a definition sentence.

The sentence you point to must actually appear on the page and stand on its own. Google recommends short, bite-sized pieces: the guideline is around twenty to thirty seconds of reading time, or a few sentences, not a whole section. If you point to a messy block with subheadings, links and lists, the read-aloud result sounds unnatural and will probably be ignored.

Important: speakable complements good content architecture, it does not replace it. The markup makes your structure explicit, but if the underlying text contains no clean, self-contained answers, there is nothing meaningful to point to.

Which sections do you mark as speakable?

Only mark short, self-contained answer sentences that are fully correct when read aloud in isolation. That is the whole art. Good candidates are:

  • The opening sentence under a heading that directly answers the question (“Speakable schema is…”).
  • A concise definition of a concept or product.
  • A summarizing sentence with the key fact or the most important figure.
  • A clear outcome or recommendation in one or two sentences.

What you do not mark: entire paragraphs full of context, lists that lose their structure when spoken, sentences with “as discussed above”, tables, and anything with links or abbreviations that are confusing to hear. A good test is simple. Read the marked sentence aloud without the rest of the page. Does someone understand the answer right away? Then it is a good speakable section. Do you need to add context? Then it is not.

This principle lines up with how AI assistants cite anyway. They look for the shortest passage that fully answers the question. Whether or not you add speakable markup, writing such self-contained answer sentences is the foundation of answer engine optimization.

Does speakable schema really work for voice search and AI assistants?

Honest answer: the official Google feature is narrow today, but the underlying signal is genuinely useful. It is important to keep the two apart, otherwise you expect a return that is not there.

Google’s official application is, at the time of writing, still beta and aimed at news content. It works for users in the United States with English-language Google devices, and for publishers who publish in English. For a Belgian or Dutch B2B company publishing in its local language, that specific news feature therefore delivers little to nothing. Anyone who tells you that speakable schema will catapult you into voice search tomorrow is exaggerating.

The broader value lies elsewhere. Voice search and AI assistants both run on the same mechanism: a system has to extract from a page the shortest, most correct passage to read aloud or cite. By explicitly pointing out which sentence is your core answer, you make that work easier and reduce the chance that a model picks the wrong passage. It is related to how AI determines which source and which fragment it trusts, something we dig into in our piece on grounding.

Whether the major AI assistants use speakable markup as a direct ranking factor today is not firmly confirmed, so we will not make that claim. What is true: the discipline speakable enforces, namely short self-contained answers, is exactly what increases citability in AI answers. The markup is therefore more of a useful thinking exercise than a magic button.

How does speakable fit into your GEO strategy?

Treat speakable schema as fine-tuning, not as a foundation. Get the basics right first: clear headings that answer questions, self-contained opening sentences per section, and correct Article or Organization markup. That delivers the bulk of your AI visibility. Only then does it make sense to additionally mark your strongest answer sentences with speakable.

A realistic approach in three steps:

  1. For each important page, write an answer sentence that is fully correct when read aloud on its own.
  2. Add that sentence to your existing JSON-LD with a speakable specification and a precise selector.
  3. Keep it limited. One or two marked passages per page is enough, more dilutes the signal.

For B2B companies where leads and revenue count, the honest conclusion is that speakable schema is rarely your number one priority. It is not wasted effort, but the gain lies mainly in the writing discipline around it. If you want to know where your time actually pays off most for AI visibility, take a look at our generative engine optimization service, where we handle the entire pipeline of being found in AI answers.

The short summary

Speakable schema uses a CSS selector or XPath to mark which sections of your page are suitable to be read aloud by search engines and assistants. The official Google feature is still beta and limited to English-language news content, so the direct benefit for non-English B2B is small today. The lasting value is the signal: you force yourself to write short, self-contained answer sentences, and that is exactly what AI assistants cite. Use it as a refinement on top of a solid content structure, not as a starting point.

Want to know which schema and GEO measures deliver the most for your site? Plan your free intake and we will look at it together.

Free website scan

Enter your website and get an automatic scan within minutes, with concrete technical and SEO improvements. No sales pitch.

Where should we send your report?

We only use your details for your scan. No spam, unsubscribe anytime.