Skip to content

InsightsTechnical SEO

Technical SEO in 2026: The Foundation That Makes Everything Else Work

The checks that decide whether Google and AI assistants can find, read and index your pages: robots.txt, sitemaps, canonicals, redirects, structured data and AI crawlers, and what has changed since 2023.

26 March 2026Updated 30 September 202623 min readEdited by Michael Wilkins

Direct answer

Technical SEO decides whether search engines and AI assistants can reach, read and index your pages at all, and no amount of content or links makes up for a page that is blocked, duplicated or missing from the index. The checks that matter most: a robots.txt that blocks nothing you need, including the AI search crawlers; a sitemap of absolute, canonical URLs; one canonical version of every page; permanent 301 or 308 redirects without chains; and main content in the HTML itself, because most AI crawlers don't run JavaScript. Some popular tactics no longer do anything in Google: FAQ and HowTo rich results and the sitelinks search box are gone, and Google Search ignores llms.txt. On one Australian insurance broker's site we crawled, none of 1,397 images had descriptive alt text, and the sitemap listed eight pages by relative addresses that search engines can't use.

Technical SEO is the part of search work that decides whether everything else can count. A page Google can't crawl, or treats as a duplicate of another page, won't rank however good it is, and a page an AI assistant can't fetch won't be quoted. It is also the most checkable part of SEO: most problems show up in a crawl or in Search Console, and most of the fixes are in your own hands, with no outreach and no waiting on anyone else.

This guide covers crawling, indexing, canonicals, redirects, structured data and the AI crawlers that now read the web alongside Googlebot. Page speed has a guide of its own, website performance and Core Web Vitals, and our SEO strategy guide covers how the technical work fits into a wider plan. Where Google has published a rule, we link to it; that documentation was last checked in September 2026.

What has changed since 2023

A few changes make older technical SEO checklists wrong in places:

  • AI crawlers are now a real share of traffic. In one month of late 2024, OpenAI's GPTBot and Anthropic's Claude crawler together made requests equal to about 20% of Googlebot's volume across Vercel's hosting network (Vercel, December 2024). Whether you let them in should be a decision, not the accident of an old robots.txt.
  • Several rich results have gone. Google stopped showing HowTo rich results in September 2023 (Google), removed the sitelinks search box on 21 November 2024 (Google) and stopped showing FAQ rich results on 7 May 2026 (Google Search Central documentation updates). Markup for them does no harm, but it no longer earns anything visible in Google.
  • Google has said what not to bother with for AI search. Its 2026 guide to generative AI features says ordinary SEO best practice still applies, that structured data isn't required to appear in AI features, and that Google Search ignores llms.txt files (Google's guide to optimising for generative AI features).
  • Search Console's reports have moved on. The Mobile Usability report was retired on 1 December 2023, with Google pointing people to Lighthouse instead; what used to be the Coverage report is now the Page indexing report; and a Generative AI performance report shows how your pages appear in Google's AI features.

Crawling and indexing: the starting point

Before Google can rank a page it has to find it, crawl it and decide to index it. A failure at any of the three stages looks the same from the outside: the page doesn't show.

robots.txt. Your robots.txt file tells crawlers which URLs they may fetch. One of the most damaging faults, and one of the easiest to miss, is a robots.txt that blocks something important: a blanket Disallow: / left over from a staging site, a rule that stops Google fetching the CSS or JavaScript it needs to render the page, or an old wildcard rule that shuts out the AI search crawlers. Two things robots.txt is not. It isn't a way to keep a page out of Google, because a blocked URL can still be indexed without its content (use noindex or a password for that), and it isn't a canonicalisation tool (Google's introduction to robots.txt). The robots.txt report in Search Console shows the robots.txt files Google found, when it last crawled them and any errors or warnings.

XML sitemap. A sitemap should list every URL you want indexed and nothing else: canonical URLs that return a 200 status, not redirects, errors or noindexed pages. Use full absolute addresses. Google's rule is explicit: https://www.example.com/mypage.html, never /mypage.html (Google's sitemap guide). Google ignores the priority and changefreq fields and only uses lastmod when it is consistently accurate, so there's no point tuning them. When we crawled an Australian speciality insurance broker's site from its own sitemap, eight category pages were listed by exactly that kind of relative address, which search engines can't use.

Crawl budget. Google's crawl budget guide is written for sites with a million or more pages, sites with 10,000 or more pages that change daily, and sites where a large share of URLs sit in "Discovered – currently not indexed". For everyone else, Google says an up-to-date sitemap and a regular look at the Page indexing report are enough (Google's crawl budget guide). Most Australian and New Zealand business sites are nowhere near those sizes. What matters for them is that nothing important is blocked, orphaned or duplicated.

The Page indexing report. In Search Console, the Page indexing report lists the URLs Google knows about and, for those it hasn't indexed, the reason. Not every "not indexed" is a problem: redirects, deliberate noindex tags and alternate pages with a proper canonical tag are supposed to be there. The statuses worth reading closely are Crawled – currently not indexed (Google fetched the page and chose not to index it for now; thin and near-duplicate pages often end up here), Discovered – currently not indexed (found but not yet crawled), Excluded by 'noindex' tag on a page you do want indexed (a staging-site tag that was never removed is the classic case) and Soft 404, a page that returns a 200 status but looks empty or "not found" to Google (Google's Page indexing report help).

Canonicals and redirects

Canonicals. Most sites serve the same content at more than one URL: with and without tracking parameters, with and without a trailing slash, in a print view, under two categories. A rel="canonical" link in the <head> tells Google which version you want indexed. It's a strong signal, not a command. Google weighs redirects, canonical tags and sitemap inclusion together, in that order of strength, and can choose a different URL if the signals disagree (Google on canonical URLs). The practical rules: give every indexable page a self-referencing canonical with a full absolute URL, link internally to the canonical version, list only canonical URLs in the sitemap, and don't try to canonicalise with robots.txt or noindex.

One version of the site. http:// and https://, and www and non-www, should each redirect to a single version rather than rely on canonical tags, because a redirect is the stronger signal.

Redirects. Use a permanent redirect (301 or 308) when a page has moved for good, after a redesign or a URL change. The reason isn't that temporary redirects "leak" value. It's that Google treats a permanent redirect as a signal that the new URL should be canonical and shown in results, while a temporary one (302 or 307) tells it to keep showing the old URL (Google on redirects). Googlebot will follow up to ten hops in a chain (Google on HTTP status codes), but every hop costs your visitors time, so flatten chains (A → B → C becomes A → C) and fix loops. After a migration, crawl the old URL list and check that each one lands on its closest new equivalent in a single hop.

Technical SEO audit checklist
Tick what's already true of your site. Work from the top: crawling and indexing problems stop everything below them from counting.
Score: 0 / 0

Core Web Vitals in brief

Core Web Vitals measure how real visitors experience a page: Largest Contentful Paint for loading (good is 2.5 seconds or less), Interaction to Next Paint for responsiveness (200 milliseconds or less; it replaced First Input Delay on 12 March 2024) and Cumulative Layout Shift for visual stability (0.1 or less), each judged at the 75th percentile of page loads. Google says Core Web Vitals are used by its ranking systems, and also that it will still show the most relevant page when the page experience is poor, so treat them as a tiebreaker you shouldn't lose rather than a lever that lifts rankings on its own (Google on page experience). They are a real hurdle all the same: the 2025 Web Almanac found that only 48% of mobile websites passed all three (HTTP Archive, July 2025 data). How to measure and fix them is in our guide to website performance and Core Web Vitals.

Structured data: what it does, and what it no longer does

Structured data is machine-readable markup, usually a block of JSON-LD, that describes what a page is: an organisation, an article and its author, a product, an event. Google recommends JSON-LD where your site allows it (Google's introduction to structured data).

Be clear about what it does. It helps Google understand a page and the business behind it, and it makes pages eligible for rich results such as review stars on products, breadcrumb trails and event listings. Eligible isn't guaranteed: Google decides whether to show them. Structured data isn't a ranking factor in itself, and it isn't a key to AI answers either. Google's 2026 guidance says it isn't needed to appear in Google's generative AI features and that there's no special markup for them, while still recommending it for rich-result eligibility.

The types worth having on most Australian and New Zealand business sites:

  • Organization on the homepage: legal name, logo, contact details and social profiles. Google uses it to tell your business apart from others with similar names, and it can influence which logo Google shows.
  • LocalBusiness, in the most specific subtype available (Plumber or AccountingService rather than plain LocalBusiness), if you have premises or a service area: address, phone and opening hours. It can feed the business details Google shows, but it doesn't replace your Google Business Profile, which is what the map results are built from, so keep the two consistent.
  • Article or BlogPosting on editorial content, with the author marked up as a Person linked to a real profile page, plus publication and modified dates.
  • BreadcrumbList wherever breadcrumbs appear.
  • Product, Event or VideoObject where a page really is one of those.

And the ones to stop expecting anything from in Google:

  • FAQPage. FAQ rich results were limited to well-known government and health sites in August 2023 and stopped appearing in Google altogether on 7 May 2026. Accurate markup on a genuine question-and-answer section does no harm, but it won't produce a visible result.
  • HowTo. Google stopped showing HowTo rich results on mobile in August 2023 and on desktop in September 2023.
  • Sitelinks search box. Removed from results on 21 November 2024. WebSite markup is still worth keeping, because Google uses a form of it to choose your site name.
  • Review stars for your own business. When a business marks up reviews of itself, including through an embedded review widget, its LocalBusiness or Organization pages aren't eligible for star ratings (Google's review snippet guidelines).
  • Speakable. Still in beta and limited to English-language news read aloud on Google Assistant devices in the US, so it isn't relevant to most business sites.

Two rules apply to all of it: mark up only what is visible on the page, and validate every type with Google's Rich Results Test and the Schema Markup Validator. On the site of our sister company, Involve Energy, that means Service markup on its ten sector-by-region pages, BreadcrumbList on those pages and its guides, and an Organization record on every page.

Structured data: what each type does in Google
Filter by site type. Status as of September 2026.
TypeUse it onWhat it does in Google
Sources: Google Search Central structured data documentation and search gallery · Google Search Central documentation updates (FAQ rich results, May 2026) · schema.org

Mobile-first indexing

Google indexes the mobile version of your site. It began mobile-first crawling in 2016 and finished moving sites over in October 2023 (Google, October 2023). So the mobile version needs the same content, headings, structured data and internal links as the desktop version. A single responsive site meets that by default; the problems come from separate m-dot sites and from hiding content on phones to save space.

With Search Console's Mobile Usability report retired on 1 December 2023, check mobile usability with Lighthouse, which Google itself points to, in Chrome DevTools or PageSpeed Insights, and on real phones (Google, April 2023). The basics: a viewport meta tag (<meta name="viewport" content="width=device-width, initial-scale=1">), text that's readable without zooming (16px body text is a sensible floor), tap targets of about 48 by 48 pixels with around 8 pixels between them (web.dev), and form fields with the right input types, so phones show the right keyboard for a phone number or an email address.

AI crawlers: who they are and what to allow

Each AI company runs different user agents for different jobs, and robots.txt lets you treat each one differently. These are the ones that matter, as their owners describe them in September 2026:

User agentOwnerWhat it does
OAI-SearchBotOpenAIFinds pages for ChatGPT's search features. Sites that block it aren't shown in ChatGPT search answers, though they can still appear as navigational links.
ChatGPT-UserOpenAIVisits a page when a ChatGPT user's request needs it. Because a user starts it, robots.txt rules may not apply.
GPTBotOpenAICollects content that may be used to train OpenAI's models.
Claude-SearchBotAnthropicIndexes pages to improve the search results Claude gives users.
Claude-UserAnthropicFetches pages when a Claude user asks a question.
ClaudeBotAnthropicCollects content that may be used for model training.
PerplexityBotPerplexitySurfaces and links pages in Perplexity's results. Perplexity says it isn't used to train models.
Perplexity-UserPerplexityVisits pages to answer a user's question.
GooglebotGoogleCrawls for Google Search, which AI Overviews and AI Mode draw on.
Google-ExtendedGoogleNot a crawler: a robots.txt product token. It controls whether content Google crawls may be used to train Gemini models and to ground answers in the Gemini apps and Vertex AI. It has no effect on Google Search or rankings.
BingbotMicrosoftCrawls for Bing, whose index Microsoft's Copilot tools search.

Sources: OpenAI, Anthropic, Perplexity, Google and Microsoft.

The useful distinction is between search and training. The search and user agents decide whether AI assistants can find and quote your pages, and blocking them takes you out of those answers. The training crawlers decide whether your future content goes into model training; blocking them doesn't remove anything already collected. Most businesses that want AI assistants to recommend them allow the search agents; training is a separate, deliberate call. On Involve Energy's site, which we built, the robots file admits every crawler. If you'd rather keep your content out of model training, a robots.txt like this allows the search agents and blocks the training crawlers:

User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Disallow: /admin/

User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /

# Google-Extended is a token, not a crawler. This stops Gemini training
# and grounding in the Gemini apps; Google Search is unaffected.
User-agent: Google-Extended
Disallow: /

User-agent: *
Disallow: /admin/

Sitemap: https://www.example.com/sitemap.xml

Two details catch people out. A crawler that finds a group naming it follows only that group and ignores the * group, which is why the private folder is repeated above. And changes take time to register: OpenAI and Perplexity both say about 24 hours.

JavaScript: make sure the content is in the HTML

Google renders JavaScript, so a page assembled in the browser can still be indexed, although Google notes that sites built on JavaScript frameworks are more complex to get right. Most AI crawlers don't render it at all. When Vercel and MERJ tested in December 2024, the AI crawlers from OpenAI, Anthropic, Meta, ByteDance and Perplexity did not render JavaScript. ChatGPT's and Claude's crawlers did fetch JavaScript files (11.5% and 23.8% of their requests) but didn't run them. The exceptions were Google's Gemini, which uses Googlebot's rendering, and Applebot (Vercel).

So if your main content only appears after JavaScript runs, as it does on many single-page apps built with React, Vue or Angular without server rendering, those crawlers see an empty shell. The fix is server-side rendering or static generation, so that the first HTML response already contains the content. Next.js, Nuxt, SvelteKit and Astro all support it, and most WordPress themes already work that way. The test takes a minute: turn JavaScript off in your browser's developer tools, reload a key page and see what's left. BCS Broking's rebuilt site, for example, is statically prerendered and served from an edge cache, so any crawler gets the whole page in the first response.

llms.txt is a proposal, first published by Jeremy Howard in September 2024, for a Markdown file at /llms.txt that gives AI agents a short, curated guide to a site: a title, a one-paragraph summary and lists of links to the pages that matter (llmstxt.org). It's designed for agents that fetch information on demand, such as a coding assistant reading a library's documentation.

What it doesn't do: it doesn't control access (robots.txt does that), and Google Search doesn't use it. Google's guide to optimising for its generative AI features, updated in July 2026, says you don't need llms.txt or any other AI text file to appear in Google Search or its AI features, and that keeping one neither helps nor harms, though it's fine to maintain one for other services that use it. OpenAI's and Anthropic's crawler documentation doesn't mention it; their controls are robots.txt rules.

Our view: it's cheap and harmless, and we've added one where it's useful, as on Involve Energy's site, but it isn't a ranking or citation lever and shouldn't be sold as one. If you want one, the generator below writes a file in the proposal's format.

llms.txt generator
Writes an llms.txt in the format proposed at llmstxt.org. It's optional: Google Search ignores these files, and robots.txt is what controls crawler access.

Site architecture: a page for every real question

How pages are organised decides what Google can match a search to. A site with one general page per line of business gives a specific search nothing specific to land on. MLA Traffic's old site was a template with no separate service pages and no location pages, so a search for "traffic management plans Melbourne" had nothing to match, and every Google Ads click landed on the homepage. The rebuild gave every service and every area its own page. BCS Broking's rebuild did the same for a specialist broker: a fifteen-page site, with one general Insurance page and one general Surety page, became a page for each of 19 insurance verticals and each of 11 surety bond types.

The limit is the word real. A page for each service you actually provide and each area you actually serve is architecture. A page for every keyword variation is what Google's spam policies call scaled content abuse, and Google's 2026 guide to AI features warns specifically against making pages for every way a question might be phrased.

The rest is housekeeping. Keep important pages within about three clicks of the homepage. Link between related pages with anchor text that says what the destination is. Make sure no page is orphaned, with no internal links pointing to it. Use short, readable URLs with hyphens between words, which Google recommends over underscores (Google on URL structure). And show breadcrumbs, with BreadcrumbList markup, on any site more than two levels deep.

HTTPS and security

HTTPS has been a Google ranking signal since August 2014, when Google described it as a very lightweight one (Google, 2014). It is now part of Google's page experience guidance and a basic expectation, and browsers label plain-HTTP pages "Not secure". Serve every page over HTTPS, redirect every HTTP URL to its HTTPS version, fix mixed content (HTTP images or scripts on HTTPS pages) and let your host or Let's Encrypt renew certificates automatically. Security headers such as HSTS, Content-Security-Policy and X-Frame-Options protect your visitors; they don't affect rankings, but they belong in the same check.

A technical SEO audit, step by step

  1. Crawl the site with Screaming Frog, Semrush Site Audit or Ahrefs Site Audit, once from the homepage and once from the sitemap, and compare the two lists. Sitemap pages the crawl can't reach are orphans; pages the crawl finds that aren't in the sitemap may be duplicates or leftovers. Crawl tools use their own user agents, so check what Google itself sees with the URL Inspection tool in Search Console.
  2. Read the Page indexing report for pages you want indexed that aren't, and the reasons given.
  3. Check Core Web Vitals in Search Console's Core Web Vitals report, which uses real-visitor data from the Chrome UX Report (CrUX), then diagnose individual pages in PageSpeed Insights. Our Core Web Vitals guide covers the fixes.
  4. Validate structured data in the Rich Results Test and the Schema Markup Validator, and check that it matches what's visible on each page.
  5. Test on mobile with Lighthouse and on real phones. There is no Mobile Usability report any more.
  6. Review robots.txt and the sitemap: the robots.txt report in Search Console, absolute URLs in the sitemap, and a 200 status for every URL listed.
  7. Check AI readiness: load key pages with JavaScript turned off, confirm which AI user agents robots.txt allows, and look at the Generative AI performance report in Search Console.

When we crawled that speciality broker's site, the crawl reached 192 pages. None of its 1,397 images carried descriptive alt text, 82 pages didn't have exactly one main heading, and only 171 of the 192 titles were unique. None of it shows in a screenshot of the homepage, which is the point of running the checks. Our rebuild of the same site, finished but not yet live, gives each of its 180 pages a unique title, a canonical address and one main heading, and every image either describes itself in alt text or is marked as decorative.

If you would rather someone else ran these checks, our free website audit covers what's broken under the hood, in plain English, and the three things to fix first.

Questions

Common questions

What are the most important technical SEO priorities for 2026?
Start with access and duplication, because nothing else counts until those are right: a robots.txt that blocks nothing you need, a sitemap of absolute, canonical URLs that return a 200 status, one canonical version of every page, and permanent (301 or 308) redirects without chains. Then check Core Web Vitals in Search Console (good means an LCP of 2.5 seconds or less, an INP of 200 milliseconds or less and a CLS of 0.1 or less, at the 75th percentile), add structured data that matches each page (Organization, LocalBusiness where it applies, Article and BreadcrumbList), and make sure the main content is in the server-rendered HTML, because most AI crawlers don't run JavaScript. Finally, decide which AI crawlers to allow: the search agents that let AI assistants find and quote you, and the training crawlers, which are a separate choice.
Does schema markup directly improve search rankings?
Not directly. Google describes structured data as a way to understand pages and to qualify for rich results, such as product review stars, breadcrumbs and event listings, which it may or may not show; it doesn't present it as a ranking boost. Google's 2026 guidance also says structured data isn't needed to appear in its generative AI features. Several popular types no longer produce anything visible in Google: FAQ rich results stopped appearing on 7 May 2026, HowTo rich results were retired in 2023 and the sitelinks search box in November 2024. Keep markup accurate and matched to what's on the page, and don't expect it to move rankings on its own.
Should businesses block AI crawlers in their robots.txt?
Decide separately for search and for training. The search and user agents (OAI-SearchBot and ChatGPT-User for OpenAI, Claude-SearchBot and Claude-User for Anthropic, PerplexityBot and Perplexity-User for Perplexity) decide whether AI assistants can find and quote your pages, and blocking them takes you out of those answers. The training crawlers (GPTBot and ClaudeBot) collect content that may be used to train models; blocking them stops future collection but doesn't remove anything already gathered. Google-Extended isn't a crawler at all: it's a robots.txt token that controls whether Google may use your content for Gemini training and for grounding answers in the Gemini apps, and it doesn't affect Google Search. Most businesses that want AI assistants to recommend them allow the search agents and make a deliberate call on training. At the least, check that an old wildcard rule isn't blocking all of them by accident.

Build it into your plan

Want this in your growth plan?

The strategist below picks up where this article left off — pull on the thread, and end up with a tailored plan you can actually act on.

Talk to the strategist

Keep reading