# Bot management (AI training bots, scrapers) is handled by Cloudflare. # This file covers only app-level path rules. User-agent: AhrefsBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: BLEXBot Disallow: / User-agent: DotBot Disallow: / User-agent: rogerbot Disallow: / User-agent: Barkrowler Disallow: / User-agent: SerpstatBot Disallow: / User-agent: ZoominfoBot Disallow: / User-agent: Screaming Disallow: / User-agent: magpie-crawler Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: * Crawl-delay: 2 # Private / internal paths Disallow: /api/ # Carve-outs that MUST precede nothing but must exist: until 2026-09-08 the # Disallow above was inert for every crawler except YandexBot, so these three # paths were reachable by accident. Making the rule real would have broken them. # # /api/docs + /api/openapi.json are submitted BY US in sitemap-static.xml # (route.ts:96,102) — a URL that is in the sitemap and disallowed in robots.txt # is the "Submitted URL blocked by robots.txt" error, and would de-index the API # docs page. # # /api/poster is the poster fallback path: prod serves posters from # NEXT_PUBLIC_POSTER_CDN_BASE (https://img.torrentclaw.com, verified in the # running container 2026-09-09), but image-loader.ts falls back to the proxy # path on a CDN miss or R2 outage, and a crawler that cannot fetch the fallback # sees broken images exactly when the CDN is having a bad day. Allow: /api/docs Allow: /api/openapi.json Allow: /api/poster/ # Next.js image optimizer — BingBot's single biggest crawl bucket (~553/day in # the proxy logs), all serving TMDB-sourced posters whose image-search credit # goes to TMDB (not us). Blocking reclaims that crawl budget for HTML; users # still see images (this stops the crawler, not the browser). /_next/static # stays crawlable — Bing needs the JS/CSS to render pages for ranking. Disallow: /_next/image Disallow: /.well-known/ Disallow: /admin Disallow: /r/ Disallow: /portal # Auth flows — no value indexing these Disallow: /login Disallow: /register Disallow: /forgot-password Disallow: /reset-password Disallow: /verify-email Disallow: /profile # Locale variants of auth paths Disallow: /es/iniciar-sesion Disallow: /es/registro Disallow: /es/recuperar-contrasena Disallow: /es/restablecer-contrasena Disallow: /es/verificar-email Disallow: /es/perfil Disallow: /pt/entrar Disallow: /pt/cadastro Disallow: /pt/esqueci-senha Disallow: /pt/redefinir-senha Disallow: /pt/verificar-email Disallow: /pt/perfil Disallow: /fr/connexion Disallow: /fr/inscription Disallow: /fr/mot-de-passe-oublie Disallow: /fr/reinitialiser-mot-de-passe Disallow: /fr/verifier-email Disallow: /fr/profil Disallow: /ru/vhod Disallow: /ru/registraciya Disallow: /ru/zabyl-parol Disallow: /ru/sbrosit-parol Disallow: /ru/podtverdit-email Disallow: /ru/profil Disallow: /it/accedi Disallow: /it/registrazione Disallow: /it/password-dimenticata Disallow: /it/reimposta-password Disallow: /it/verifica-email Disallow: /it/profilo Disallow: /de/anmelden Disallow: /de/registrieren Disallow: /de/passwort-vergessen Disallow: /de/passwort-zuruecksetzen Disallow: /de/email-bestaetigen Disallow: /de/profil Disallow: /ja/login Disallow: /ja/touroku Disallow: /ja/forgot-password Disallow: /ja/reset-password Disallow: /ja/verify-email Disallow: /ja/profile # Paginated / filtered URLs — avoid duplicate content Disallow: /*?page= Disallow: /*?sort= Disallow: /search? Disallow: /es/search? Disallow: /genre/*? Disallow: /es/genero/*? Disallow: /source/*? Disallow: /es/fuente/*? Disallow: /recent? Disallow: /es/recent? # YandexBot gets its OWN block, and that is load-bearing: Yandex applies the # MOST SPECIFIC matching group and ignores the "User-agent: *" group entirely # once a "User-agent: Yandex" group exists. That is also why PATH_RULES is # repeated here instead of being inherited — it would not be. # # WHY (incident 2026-08-15): measured in the Moldova proxy log, YandexBot was # ~97% of all page traffic — 2.152 requests against ~70 from real humans in the # same window. Every one of those is a distinct catalogue URL, so every one is # a cache MISS that renders a page and writes a ~565 KB entry. That is ~970k # renders/day of pure write amplification, and it is what filled the ISR cache # and starved the replicas (8 OOM in one night). # # 10 seconds is deliberate, not punitive: with ~1M content rows a full crawl was # never going to happen at any delay, and Yandex sends effectively no referral # traffic (same finding as the Bing soft-deindex, 2026-07-26). We keep being # indexable; we stop paying for a full-catalogue walk every day. # # NOTE: this only works if Yandex honours it. If the proxy logs still show the # same rate in 24-48 h, the next lever is a rate limit in the middleware — this # file is a request, not an enforcement. User-agent: Yandex Crawl-delay: 10 # Private / internal paths Disallow: /api/ # Carve-outs that MUST precede nothing but must exist: until 2026-09-08 the # Disallow above was inert for every crawler except YandexBot, so these three # paths were reachable by accident. Making the rule real would have broken them. # # /api/docs + /api/openapi.json are submitted BY US in sitemap-static.xml # (route.ts:96,102) — a URL that is in the sitemap and disallowed in robots.txt # is the "Submitted URL blocked by robots.txt" error, and would de-index the API # docs page. # # /api/poster is the poster fallback path: prod serves posters from # NEXT_PUBLIC_POSTER_CDN_BASE (https://img.torrentclaw.com, verified in the # running container 2026-09-09), but image-loader.ts falls back to the proxy # path on a CDN miss or R2 outage, and a crawler that cannot fetch the fallback # sees broken images exactly when the CDN is having a bad day. Allow: /api/docs Allow: /api/openapi.json Allow: /api/poster/ # Next.js image optimizer — BingBot's single biggest crawl bucket (~553/day in # the proxy logs), all serving TMDB-sourced posters whose image-search credit # goes to TMDB (not us). Blocking reclaims that crawl budget for HTML; users # still see images (this stops the crawler, not the browser). /_next/static # stays crawlable — Bing needs the JS/CSS to render pages for ranking. Disallow: /_next/image Disallow: /.well-known/ Disallow: /admin Disallow: /r/ Disallow: /portal # Auth flows — no value indexing these Disallow: /login Disallow: /register Disallow: /forgot-password Disallow: /reset-password Disallow: /verify-email Disallow: /profile # Locale variants of auth paths Disallow: /es/iniciar-sesion Disallow: /es/registro Disallow: /es/recuperar-contrasena Disallow: /es/restablecer-contrasena Disallow: /es/verificar-email Disallow: /es/perfil Disallow: /pt/entrar Disallow: /pt/cadastro Disallow: /pt/esqueci-senha Disallow: /pt/redefinir-senha Disallow: /pt/verificar-email Disallow: /pt/perfil Disallow: /fr/connexion Disallow: /fr/inscription Disallow: /fr/mot-de-passe-oublie Disallow: /fr/reinitialiser-mot-de-passe Disallow: /fr/verifier-email Disallow: /fr/profil Disallow: /ru/vhod Disallow: /ru/registraciya Disallow: /ru/zabyl-parol Disallow: /ru/sbrosit-parol Disallow: /ru/podtverdit-email Disallow: /ru/profil Disallow: /it/accedi Disallow: /it/registrazione Disallow: /it/password-dimenticata Disallow: /it/reimposta-password Disallow: /it/verifica-email Disallow: /it/profilo Disallow: /de/anmelden Disallow: /de/registrieren Disallow: /de/passwort-vergessen Disallow: /de/passwort-zuruecksetzen Disallow: /de/email-bestaetigen Disallow: /de/profil Disallow: /ja/login Disallow: /ja/touroku Disallow: /ja/forgot-password Disallow: /ja/reset-password Disallow: /ja/verify-email Disallow: /ja/profile # Paginated / filtered URLs — avoid duplicate content Disallow: /*?page= Disallow: /*?sort= Disallow: /search? Disallow: /es/search? Disallow: /genre/*? Disallow: /es/genero/*? Disallow: /source/*? Disallow: /es/fuente/*? Disallow: /recent? Disallow: /es/recent? User-agent: YandexBot Crawl-delay: 10 # Private / internal paths Disallow: /api/ # Carve-outs that MUST precede nothing but must exist: until 2026-09-08 the # Disallow above was inert for every crawler except YandexBot, so these three # paths were reachable by accident. Making the rule real would have broken them. # # /api/docs + /api/openapi.json are submitted BY US in sitemap-static.xml # (route.ts:96,102) — a URL that is in the sitemap and disallowed in robots.txt # is the "Submitted URL blocked by robots.txt" error, and would de-index the API # docs page. # # /api/poster is the poster fallback path: prod serves posters from # NEXT_PUBLIC_POSTER_CDN_BASE (https://img.torrentclaw.com, verified in the # running container 2026-09-09), but image-loader.ts falls back to the proxy # path on a CDN miss or R2 outage, and a crawler that cannot fetch the fallback # sees broken images exactly when the CDN is having a bad day. Allow: /api/docs Allow: /api/openapi.json Allow: /api/poster/ # Next.js image optimizer — BingBot's single biggest crawl bucket (~553/day in # the proxy logs), all serving TMDB-sourced posters whose image-search credit # goes to TMDB (not us). Blocking reclaims that crawl budget for HTML; users # still see images (this stops the crawler, not the browser). /_next/static # stays crawlable — Bing needs the JS/CSS to render pages for ranking. Disallow: /_next/image Disallow: /.well-known/ Disallow: /admin Disallow: /r/ Disallow: /portal # Auth flows — no value indexing these Disallow: /login Disallow: /register Disallow: /forgot-password Disallow: /reset-password Disallow: /verify-email Disallow: /profile # Locale variants of auth paths Disallow: /es/iniciar-sesion Disallow: /es/registro Disallow: /es/recuperar-contrasena Disallow: /es/restablecer-contrasena Disallow: /es/verificar-email Disallow: /es/perfil Disallow: /pt/entrar Disallow: /pt/cadastro Disallow: /pt/esqueci-senha Disallow: /pt/redefinir-senha Disallow: /pt/verificar-email Disallow: /pt/perfil Disallow: /fr/connexion Disallow: /fr/inscription Disallow: /fr/mot-de-passe-oublie Disallow: /fr/reinitialiser-mot-de-passe Disallow: /fr/verifier-email Disallow: /fr/profil Disallow: /ru/vhod Disallow: /ru/registraciya Disallow: /ru/zabyl-parol Disallow: /ru/sbrosit-parol Disallow: /ru/podtverdit-email Disallow: /ru/profil Disallow: /it/accedi Disallow: /it/registrazione Disallow: /it/password-dimenticata Disallow: /it/reimposta-password Disallow: /it/verifica-email Disallow: /it/profilo Disallow: /de/anmelden Disallow: /de/registrieren Disallow: /de/passwort-vergessen Disallow: /de/passwort-zuruecksetzen Disallow: /de/email-bestaetigen Disallow: /de/profil Disallow: /ja/login Disallow: /ja/touroku Disallow: /ja/forgot-password Disallow: /ja/reset-password Disallow: /ja/verify-email Disallow: /ja/profile # Paginated / filtered URLs — avoid duplicate content Disallow: /*?page= Disallow: /*?sort= Disallow: /search? Disallow: /es/search? Disallow: /genre/*? Disallow: /es/genero/*? Disallow: /source/*? Disallow: /es/fuente/*? Disallow: /recent? Disallow: /es/recent? Sitemap: https://torrentclaw.com/sitemap.xml