User-agent: * Allow: / # Don't index admin and internal endpoints Disallow: /admin/ Disallow: /captcha/ Disallow: /analytics/ Disallow: /ipns/ Disallow: /api/rating/ Disallow: /api/newsletter/ # Tracking-only query params — every internal-link click was getting indexed # as a distinct duplicate URL with proper canonical to the bare path. The # canonical was working, but Google still spent crawl budget on variants. # Blocking the crawl here cuts that to zero. Both ?param= and ¶m= forms # are blocked since either can appear depending on whether the param is first # in the query string. Disallow: /*?ref= Disallow: /*&ref= Disallow: /*?from= Disallow: /*&from= Disallow: /*?next= Disallow: /*&next= Disallow: /*?return= Disallow: /*&return= Disallow: /*?sort= Disallow: /*&sort= Disallow: /*?utm_ Disallow: /*&utm_ Disallow: /*?gclid= Disallow: /*&gclid= Disallow: /*?fbclid= Disallow: /*&fbclid= Disallow: /*?msclkid= Disallow: /*&msclkid= # `?lang=en` is the default language with a redundant param. Server-side # middleware also 301s these to the bare path; Disallow stops Google from # finding them in the first place. Disallow: /*?lang=en Disallow: /*&lang=en # Clicky first-party analytics beacon (proxied, same-origin) — Googlebot # re-crawls /3313609894e7?site_id=... + loader as junk 200s, crawl-budget waste. # Crawl directive only; analytics unaffected. See ~/seo/playbook/06. Disallow: /3313609894e7 Disallow: /ee3e0327e314.js Sitemap: https://www.vps.org/sitemap.xml