Tag: Technical SEO

  • Schema Markup for B2B Websites: What It Is and Why AI Search Still Needs It

    Schema Markup for B2B Websites: What It Is and Why AI Search Still Needs It

    What schema is, the five types your site needs, and why structured data matters for AI search even after Google’s recent rich result updates. 

    Schema markup is structured code that tells search engines and AI tools exactly what your website is about, and many B2B sites have little or none. Here is what it does, the five types worth having, and why it still matters now that Google has retired its most popular rich result.

    Table of Contents

    What is schema markup?

    Schema markup is a standardised vocabulary of code. It is maintained at Schema.org and founded by Google, Microsoft, Yahoo and Yandex. Schema states the facts of a web page outright: this is an organisation, this is its legal name, this is an article, this person wrote it, on this date, etc. It is added to a page as a small block of JSON-LD, a format invisible to human visitors and read directly by machine crawlers.

    The problem it solves is inference. Without markup, a search engine or AI tool has to deduce what your business is and who wrote your content from prose alone, and deduction brings uncertainty. A machine that is uncertain about your facts is less likely to treat you as a distinct, verified entity, and less likely to cite you. We covered what those machine readers can and cannot see in our guide to how AI crawlers read your website; schema is the layer that removes the guesswork for your brand’s key facts.

    The five schema types B2B websites need

    The full Schema.org vocabulary runs to hundreds of types, which sounds daunting until you notice how few apply to you. Recipe, JobPosting, Product and MedicalWebPage exist for recipe sites, job boards, online shops and medical publishers. A typical B2B services website needs about five.

    Organisation schema: the one to do first

    Organisation schema states your legal name, URL, logo, contact details and official profiles, and its sameAs property points machines at the external records that corroborate you, such as your LinkedIn page and Companies House listing. It is the single type that establishes your brand as a verifiable entity, which is why we treat it as the first job on any site.

    Article and BlogPosting schema

    Applied to every post, this states the headline, author, publication date and date updated outright, feeding the freshness and authorship signals that search engines and AI tools both read.

    Person schema

    Person schema connects a named author to their credentials, profiles and your organisation. It is the structural half of the authorship work described in our guide to E-E-A-T, turning a byline into a verifiable claim.

    FAQPage schema

    FAQPage schema formats question-and-answer pairs for direct machine extraction. Its use took a sharp turn this year, covered below, but the format remains one of the most readily extractable structures a page can carry.

    BreadcrumbList states where a page sits within your site’s structure, helping machines understand how your topics relate to one another.

    What happens without schema markup

    A site without markup is not invisible; it is ambiguous. Machines must infer your facts, may not recognise your brand as a distinct entity, and cannot reliably connect your authors’ credentials to your organisation. Each of those uncertainties quietly lowers the odds of your content being cited without anything on your visible page looking out of place.

    Schema is the most consistent gap in the free AI readiness checks we run. Well-designed B2B sites with genuinely useful content are let down by absent or generic markup, and teams who did not realise there was an issue.

    Schema types come and go: what the FAQ story teaches us

    Individual schema types get retired, and the most instructive recent example is FAQ. Google restricted FAQ rich results in August 2023 (along with HowTo results), announcing they would only be shown for well-known, authoritative government and health websites (Google, 2023). HowTo results began winding down, which were fully retired weeks later. In May 2026 the withdrawal completed: Search Engine Journal reported that FAQ rich results no longer appear in Google Search at all, with reporting and testing support removed over the months that followed (Search Engine Journal, 2026).

    If schema were only a trick for winning rich snippets, that would be the end of FAQPage markup. But, fortunately, it is not for two reasons. Google’s own guidance says there is no need to remove existing FAQ structured data; it causes no harm sitting on the page (Google, 2023). And, importantly, Google Search was never the only reader: AI crawlers scan structured question-and-answer pairs whether or not a visual search feature rewards them, because the format hands them an extractable answer with its question attached. We keep FAQPage markup on our own articles for exactly this reason and we still recommend our clients use them too.

    The wider lesson is about how to judge structured data. Any single type can be deprecated when a search feature is retired, so measure schema by what it communicates to machines and by extension other people, not by which platform’s display feature it currently earns. The facts you mark up outlast the features built on top of them.

    Get your schema checked in one pass

    Our free AI Readiness Check reviews your schema alongside crawler access, content structure and what AI tools currently say about your brand, in a single prioritised report.

    How to check your schema markup (and who should add it)

    Two free tools check your schema. Google’s Rich Results Test shows which Google-supported types a page carries and whether they validate, though its FAQ support was removed in mid-2026 along with the rich result itself. The Schema Markup Validator checks any Schema.org type, FAQ included, which makes it the better all-purpose check now. Run one or the other after any significant site change.

    As for who does the work: adding JSON-LD is a paste job in most content management systems, through an SEO plugin or a custom HTML block, and the templates for the five types above are stable and reusable. You do not need a developer for a standard setup, though larger or custom-built sites may benefit from one. It is also among the first things we review and fix quickly following an AI readiness audit.

    The five schema types all B2B websites should use

    Organisation, Article, Person, FAQPage and BreadcrumbList: five stable templates that turn a site machines must guess about into one they can verify. The features built on schema will keep changing, but the facts you mark up will keep being read. 

    Learn more about how AI and machines read your website with our plain-English guide to AEO, and then why not request a free AI Readiness Check and we will show you exactly which types your site is missing.

    Frequently asked questions about schema markup

    The questions we get most from B2B marketing and web managers.

    What is schema markup?

    Schema markup is standardised code, usually in JSON-LD format, added to a web page to state its facts explicitly to machines: what the business is, who wrote the content, when it was published. It is invisible to human visitors and read by search engines and AI crawlers, removing their need to infer.

    Does schema markup affect SEO?

    Yes, indirectly and usefully. Schema does not raise rankings on its own, but it makes pages eligible for certain search features and, more usefully now, gives search engines and AI tools verified facts to work from. A machine that is certain about your entity and authorship is more likely to surface and cite you.

    What is Organisation schema?

    Organisation schema is the structured data type that states a business’s legal name, website, logo, contact details and official profiles, with a sameAs property linking to corroborating records such as LinkedIn and Companies House. It establishes the brand as a verifiable entity, and it is the first type any B2B website should add.

    Is FAQ schema still worth using?

    Yes. Google retired FAQ rich results in May 2026, so the markup no longer earns an expanded Google listing, but Google’s guidance confirms there is no need to remove it, and AI tools still read structured question-and-answer pairs when extracting answers. The visual feature is gone; the machine readability is not.

    Do I need a developer to add schema markup?

    Usually not. In most content management systems, JSON-LD schema is added through an SEO plugin or pasted into a custom HTML block, and the common B2B types follow stable, reusable templates. A developer is worth involving for large or unusual sites, or where markup needs generating automatically across many pages.

  • How AI Crawlers Read Your Website: The Five Gaps That Keep B2B Sites Out of AI Answers

    How AI Crawlers Read Your Website: The Five Gaps That Keep B2B Sites Out of AI Answers

    When a buyer asks ChatGPT or Perplexity about your market, those tools rely on automated readers that have already visited your website, or tried to. AI crawlers do not behave like human visitors, and they do not behave like Googlebot either. Here is what they actually see, and the five gaps that hide many B2B sites from them.

    Table of Contents

    What is an AI crawler? (And why there is more than one)

    If you have spotted a name like ClaudeBot or GPTBot in your server logs and searched to find out what it was, you are in good company. Thousands of site owners do the same every month, and the volume of those searches has grown tenfold in a year. AI crawler traffic has become impossible to miss.

    An AI crawler is an automated programme that requests and reads web pages on behalf of an AI platform. It arrives with no mouse, no patience and a strict time budget, takes the raw code your server returns, and moves on. What it collects ends up in one of two places, and that difference matters more than any other fact in this article.

    Training crawlers read the web to build what an AI model knows, whereas retrieval crawlers, sometimes called search or user-fetch bots, collect pages to answer a live question from a real person as they ask a question. One company usually runs both. OpenAI’s own documentation describes four separate bots with three separate jobs: GPTBot gathers content for model training, OAI-SearchBot builds the index behind ChatGPT’s search results, ChatGPT-User fetches pages when a user asks about them directly, and OAI-AdsBot validates ads placed on ChatGPT (OpenAI, 2026). Blocking one typically has no effect on the others.

    Here is who is likely visiting you:

    CrawlerRun byWhat it does
    GPTBotOpenAIModel training
    OAI-SearchBotOpenAIChatGPT search index
    ChatGPT-UserOpenAILive fetches for user requests
    ClaudeBotAnthropicModel training
    Claude-SearchBotAnthropicClaude search results
    PerplexityBotPerplexitySearch and answers
    Google-ExtendedGoogleRobots.txt control for AI training use
    CCBotCommon CrawlOpen dataset used to train many models
    BytespiderByteDanceModel training

    Most of this crawling is not about answering questions at all. Cloudflare’s analysis of crawler activity across its network found that around 80% of AI crawling in the year to mid-2025 was for model training, against 18% for search (Cloudflare, 2025). A bot visit is not the same thing as visibility.

    Gap one: a robots.txt file written for a post Googlebot world

    robots.txt is the small text file that tells crawlers what they may and may not read on your website. Anyone can check robots.txt for any site: type yoursite.com/robots.txt into a browser. Most B2B versions were written years ago, for a world where Googlebot was the only reader worth thinking about, and they fail in two quiet ways.

    The first failure is blocking by accident. Broad disallow rules, old security plugins, or a firewall or CDN configured to swat away unfamiliar bots can shut out AI retrieval crawlers. The consequence is severe and invisible: OpenAI states plainly that sites opted out of OAI-SearchBot will not appear in ChatGPT’s search answers (OpenAI, 2026): no error message or warning, just absence.

    We have found exactly this in client audits: a site whose content deserved to be cited, silently removed from the running by a file nobody had opened in years.

    The second failure is allowing by default. Training crawlers raise a genuine data-rights question: do you want your content absorbed into future AI models? There are sound reasons to say yes (models that have read your site describe your brand better) and sound reasons to say no (control over your intellectual property). The practical position for most B2B sites is to allow the retrieval bots if you want AI visibility, and decide whether to allow the training bots.

    Gap two: JavaScript your buyers can see and AI cannot

    Vercel, working with the search consultancy MERJ, analysed crawler behaviour across its hosting network and found that none of the major AI crawlers execute JavaScript; GPTBot fetched JavaScript files in around 11.5% of its requests, yet never ran them (Vercel, 2024). Google’s Gemini is the main exception, because it borrows Googlebot’s rendering machinery.

    The consequence is strange to say out loud. A page whose content is assembled in the browser by JavaScript can rank well on Google, look flawless to every human visitor, and be an empty shell to ChatGPT, Claude and Perplexity. Websites built as single-page applications are the usual suspects, along with content hidden inside tabs and accordions that only loads when clicked.

    The test takes thirty seconds. Open your page, view the page source (the raw source, not the browser inspector), and search for a sentence of your main content. If it is there, AI crawlers can read it. If it only appears after the page loads, they cannot. The fix is to serve your content in the HTML your server first returns, through server-side rendering, static generation or pre-rendering. Your developer will know which fits your stack.

    The quieter gaps: speed, structure and schema

    These three gaps attract less attention, but are just as important.

    Page speed and the crawler’s time budget

    Retrieval bots fetch pages against tight timeouts measured in seconds. A slow page does not get a second chance; the answer gets built from a faster source instead. Any website page speed work you are already doing for Core Web Vitals serves AI readers too.

    Heading structure and buried answers

    A crawler uses your heading hierarchy, H1 to H3, to work out which section answers which question. Pages with decorative headings, or none, give it nothing to hold on to. The same goes for answers buried beneath four paragraphs of scene-setting: extraction favours pages that answer first and elaborate second. This is a happy alignment, because human readers have always preferred the same thing. 

    Our guide to simplifying technical content explains this in detail.

    Missing schema markup

    Schema markup is structured code that states the facts of a page outright: this is an organisation, this is its name, this is an article, these are its FAQs. Without it, a crawler can read your words but has to infer your facts, and inference means uncertainty. Across the AI readiness checks we have run, our consistent finding is that sites with good design and genuinely interesting content are often let down by absent or generic schema and small machine-readability faults their teams don’t know about. The encouraging flip side is that these are building blocks rather than rebuilds. They fold into an existing publishing workflow with modest effort.

    Check your AI crawler gaps in one pass

    Our free AI Readiness Check covers crawler access, rendering, speed, structure, schema, and what AI tools currently say about your brand, in a single prioritised report.

    Three checks you can run today (and what they cannot tell you)

    Check your robots.txt. Visit yoursite.com/robots.txt and look for the crawler names in the table above. No mention of a bot means it is allowed by default; explicit disallow lines mean a decision was made by someone, or a tool default setting, at some point. Check whether this was deliberate.

    Check your page source. View the page source on your homepage and one key service page, and search for your own opening sentence, as described in gap two.

    Check your schema. Paste a page URL into Google’s free Rich Results Test. It shows what structured data the page carries, if any.

    We audited our website using our AI readiness methodology to see how AI crawlers read our content

    When we audited our own website, it passed all three checks: server-rendered, crawlers allowed, schema in place. AI tools still did not surface us for category-level questions, because those answers are also built from third-party sources, directories and round-ups that on-site fixes won’t remedy. You can read our full report here:

    What We Found When We Ran Our AI-Readiness Check On Our Own Website

    The key takeaway is that readability is the entry ticket, but not the whole game. Because different AI engines draw on different sources, the wider work belongs to answer engine optimisation as a whole, which our plain-English guide to AEO covers.

    Frequently asked questions about AI crawlers

    The questions site owners ask most, taken straight from the search data.

    What is ClaudeBot and why is it in my website logs?

    ClaudeBot is Anthropic’s web crawler. It gathers content used to train the Claude AI models, which is why it can appear often in server logs. Anthropic also runs Claude-SearchBot, a separate bot that fetches pages for Claude’s live search answers.

    What is GPTBot, and should I block it?

    GPTBot is OpenAI’s training crawler. Blocking it keeps your content out of future model training but has no effect on ChatGPT’s search results, which use a different bot. Whether to block it is a data-rights decision each business should make deliberately.

    What is the difference between GPTBot and OAI-SearchBot?

    Both belong to OpenAI, with different jobs. GPTBot collects content for model training. OAI-SearchBot builds the index behind ChatGPT’s search answers; block it and your site will not appear in those answers.

    How do I check whether my website is blocking AI crawlers?

    Open yoursite.com/robots.txt in a browser and look for lines naming GPTBot, OAI-SearchBot, ClaudeBot or PerplexityBot under “User-agent”, each followed by an Allow or Disallow rule. Also review your CDN or firewall settings, which can block bots regardless of what robots.txt says.

    Do AI crawlers read JavaScript?

    They fetch JavaScript files but do not execute them, so content that only appears after scripts run is invisible to them. Google’s Gemini is the main exception. Content present in the server-rendered HTML is readable by all of them.

    Does page speed affect AI search visibility?

    Yes. Retrieval crawlers work to tight time budgets and abandon slow pages, so the answer gets built from a faster source. Speed improvements made for Core Web Vitals help AI readability too.

    How do I block AI bots in robots.txt?

    Add a User-agent line naming the bot, followed by a Disallow rule. Reputable AI crawlers respect this. Before blocking anything, check which kind of bot it is: blocking retrieval bots removes you from AI answers, which is usually the opposite of what a business wants.

    Get your free AI readiness check and see how AI crawlers see your website

    Five gaps, all of them fixable: an unreviewed robots.txt, JavaScript-only content, slow pages, weak heading structure, and missing schema. None requires a rebuild, and each gap you close makes your best content available to the systems your buyers now ask first. 

    For the wider picture of what to do once the machines can read you, start with our plain-English guide to AEO

    Request your free AI-readiness check today and receive a detailed report on how AI crawlers read your site to start closing these gaps.

GDPR Cookie Consent with Real Cookie Banner