webroad.online
  1. 1Web
  2. 2HTML
  3. 3CSS
  4. 4JavaScript
  5. 5TypeScript
  6. 6Git
  7. 7Tooling
  8. 8React
  9. 9State management
  10. 10Next.js
  11. 11Forms
  12. 12Data and backend
  13. 13SEO
  14. 14Tailwind CSS
  15. 15Animations
  16. 16Testing
  17. 17Architecture
SEO · Lesson 2 of 5

Technical SEO

robots.txt vs noindex, the sitemap, canonical, status codes and redirects, JSON-LD.

Updated

What technical SEO is

Technical SEO = making sure search engines can find, read, understand and index every page correctly — and that they don't index what they shouldn't. It's the part that depends almost entirely on the developer. Excellent content doesn't matter if Google can't reach it.

robots.txt — what the crawler may crawl

A file at the domain root (/robots.txt) with rules for crawlers:

User-agent: *
Disallow: /admin/
Disallow: /api/
Allow: /api/og/
Sitemap: https://site.com/sitemap.xml
Rule Means
User-agent: * for every robot (or Googlebot, GPTBot etc.)
Disallow: /admin/ don't crawl paths that start with /admin/
Allow an exception to a Disallow — the most specific (longest) rule wins
Sitemap where the sitemap is

The major trap: Disallow doesn't remove a page from the index. Google can index the URL (without content) if it gets links to it. For a page to not be indexed, it has to be accessible and carry noindex.

// app/robots.ts
export default function robots(): MetadataRoute.Robots {
  return { rules: [{ userAgent: '*', disallow: ['/admin/', '/api/'] }], sitemap: 'https://site.com/sitemap.xml' }
}

Meta robots — what it may index

<meta name="robots" content="noindex, follow" />
Value Effect
noindex don't put the page in the results
nofollow don't follow the links on the page
noarchive, nosnippet, max-image-preview:large control over how it's shown

For non-HTML files (PDFs) you use the HTTP header X-Robots-Tag: noindex. In Next: metadata.robots = { index: false }.

What gets noindex: internal search pages, filtered results with no value, the cart, the account, post-order thank-you pages, test / preview pages (Vercel adds noindex to previews automatically).

The sitemap

A list of every URL you want Google to index, with its last-modified date. Generated from data in Next (app/sitemap.ts — see Metadata and SEO). Submitted in Search Console. Only canonical URLs, with a 200 status, indexable.

Canonical — a single version of each page

The same page can be reachable at several URLs:

https://site.com/product/laptop
https://site.com/product/laptop?utm_source=facebook
https://www.site.com/product/laptop/

For Google these are duplicate pages that split their signals. Canonical says which one is the original:

<link rel="canonical" href="https://site.com/product/laptop" />

In Next: alternates: { canonical: '/product/laptop' } + metadataBase. Pick one form and enforce it everywhere: with or without www, with or without a trailing /, no tracking parameters.

Status codes and redirects

Situation The correct status
the page exists 200
the URL changed permanently 301 / 308 — transfers the SEO signals
temporary 302 / 307
doesn't exist 404 (not a 200 with "not found" — a "soft 404")
permanently deleted 410
the server is temporarily down 503

In Next: redirect() / permanentRedirect(), redirects in next.config.ts, notFound().

Structured data (JSON-LD)

A JSON block that describes what the page is — a product, an article, a recipe, an event, a FAQ — in the schema.org vocabulary. Google uses it for rich results: stars, price, availability, breadcrumbs.

const jsonLd = {
  '@context': 'https://schema.org',
  '@type': 'Product',
  name: product.name,
  offers: { '@type': 'Offer', price: product.price, priceCurrency: 'USD', availability: 'https://schema.org/InStock' },
}

<script type="application/ld+json" dangerouslySetInnerHTML={{ __html: JSON.stringify(jsonLd).replace(/</g, '\\u003c') }} />

You check it with the Rich Results Test.

Other technical things that matter

Element Why
HTTPS a ranking signal, mandatory
mobile (viewport, responsive layout) Google indexes the mobile version (mobile-first indexing)
Core Web Vitals part of the page experience
hreflang multilingual sites: alternates.languages in Next
clean URLs /blog/article-title, not /p?id=123&ref=x
internal links with <a href> Googlebot follows only real links, not onClick
images with an alt and descriptive names they also show up in Google Images — see Image optimization

Summary

  • robots.txt controls crawling; noindex controls indexing — don't confuse them.
  • A sitemap with the canonical URLs; canonical for duplicates; correct status codes and redirects.
  • JSON-LD for rich results; mobile, HTTPS, speed, real links.

Official sources

Exercises

Was this page helpful?

One tap — no account needed.