Technical SEO
robots.txt vs noindex, the sitemap, canonical, status codes and redirects, JSON-LD.
Updated
What technical SEO is
Technical SEO = making sure search engines can find, read, understand and index every page correctly — and that they don't index what they shouldn't. It's the part that depends almost entirely on the developer. Excellent content doesn't matter if Google can't reach it.
robots.txt — what the crawler may crawl
A file at the domain root (/robots.txt) with rules for crawlers:
User-agent: *
Disallow: /admin/
Disallow: /api/
Allow: /api/og/
Sitemap: https://site.com/sitemap.xml
| Rule | Means |
|---|---|
User-agent: * |
for every robot (or Googlebot, GPTBot etc.) |
Disallow: /admin/ |
don't crawl paths that start with /admin/ |
Allow |
an exception to a Disallow — the most specific (longest) rule wins |
Sitemap |
where the sitemap is |
The major trap: Disallow doesn't remove a page from the index. Google can index the URL (without content) if it gets links to it. For a page to not be indexed, it has to be accessible and carry noindex.
// app/robots.ts
export default function robots(): MetadataRoute.Robots {
return { rules: [{ userAgent: '*', disallow: ['/admin/', '/api/'] }], sitemap: 'https://site.com/sitemap.xml' }
}Meta robots — what it may index
<meta name="robots" content="noindex, follow" />| Value | Effect |
|---|---|
noindex |
don't put the page in the results |
nofollow |
don't follow the links on the page |
noarchive, nosnippet, max-image-preview:large |
control over how it's shown |
For non-HTML files (PDFs) you use the HTTP header X-Robots-Tag: noindex. In Next: metadata.robots = { index: false }.
What gets noindex: internal search pages, filtered results with no value, the cart, the account, post-order thank-you pages, test / preview pages (Vercel adds noindex to previews automatically).
The sitemap
A list of every URL you want Google to index, with its last-modified date. Generated from data in Next (app/sitemap.ts — see Metadata and SEO). Submitted in Search Console. Only canonical URLs, with a 200 status, indexable.
Canonical — a single version of each page
The same page can be reachable at several URLs:
https://site.com/product/laptop
https://site.com/product/laptop?utm_source=facebook
https://www.site.com/product/laptop/
For Google these are duplicate pages that split their signals. Canonical says which one is the original:
<link rel="canonical" href="https://site.com/product/laptop" />In Next: alternates: { canonical: '/product/laptop' } + metadataBase. Pick one form and enforce it everywhere: with or without www, with or without a trailing /, no tracking parameters.
Status codes and redirects
| Situation | The correct status |
|---|---|
| the page exists | 200 |
| the URL changed permanently | 301 / 308 — transfers the SEO signals |
| temporary | 302 / 307 |
| doesn't exist | 404 (not a 200 with "not found" — a "soft 404") |
| permanently deleted | 410 |
| the server is temporarily down | 503 |
In Next: redirect() / permanentRedirect(), redirects in next.config.ts, notFound().
Structured data (JSON-LD)
A JSON block that describes what the page is — a product, an article, a recipe, an event, a FAQ — in the schema.org vocabulary. Google uses it for rich results: stars, price, availability, breadcrumbs.
const jsonLd = {
'@context': 'https://schema.org',
'@type': 'Product',
name: product.name,
offers: { '@type': 'Offer', price: product.price, priceCurrency: 'USD', availability: 'https://schema.org/InStock' },
}
<script type="application/ld+json" dangerouslySetInnerHTML={{ __html: JSON.stringify(jsonLd).replace(/</g, '\\u003c') }} />You check it with the Rich Results Test.
Other technical things that matter
| Element | Why |
|---|---|
| HTTPS | a ranking signal, mandatory |
| mobile (viewport, responsive layout) | Google indexes the mobile version (mobile-first indexing) |
| Core Web Vitals | part of the page experience |
hreflang |
multilingual sites: alternates.languages in Next |
| clean URLs | /blog/article-title, not /p?id=123&ref=x |
internal links with <a href> |
Googlebot follows only real links, not onClick |
images with an alt and descriptive names |
they also show up in Google Images — see Image optimization |
Summary
robots.txtcontrols crawling;noindexcontrols indexing — don't confuse them.- A sitemap with the canonical URLs; canonical for duplicates; correct status codes and redirects.
- JSON-LD for rich results; mobile, HTTPS, speed, real links.