Charge AI training crawlers for access (x402 gateway) (#141)

* Charge AI training crawlers for access (@profullstack/x402-gateway)

Training crawlers (GPTBot, ClaudeBot, CCBot, meta-externalagent, Bytespider,
Applebot-Extended) get 402 Payment Required with an x402 offer, or the sales
page at /crawl, and a paid pass opens the site for a day. People, search
engines and retrieval crawlers pass through untouched. robots.txt is now
generated from the same lists.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YafYxayh7Gqe5MWNNQMev2

* Type the middleware as returning Response | NextResponse

The crawl gateway answers with a plain Fetch Response.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YafYxayh7Gqe5MWNNQMev2

* Contract test awaits the now-async proxy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YafYxayh7Gqe5MWNNQMev2

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
Anthony Ettinger 2026-09-05 14:49:39 -07:00 • committed by GitHub
parent 80a36269bb
commit 93af4770ac
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
8 changed files with 65 additions and 35 deletions

View file

@ -1,30 +0,0 @@
import type { MetadataRoute } from "next";
const SITE_URL = (process.env.PUBLIC_URL ?? "https://logicsrc.com").replace(/\/$/, "");
// Welcome mainstream + AI crawlers; keep them out of the API surface.
export default function robots(): MetadataRoute.Robots {
return {
rules: [
{
userAgent: [
"*",
"GPTBot",
"OAI-SearchBot",
"ChatGPT-User",
"ClaudeBot",
"Claude-Web",
"anthropic-ai",
"PerplexityBot",
"Google-Extended",
"Applebot-Extended",
"CCBot",
],
allow: "/",
disallow: ["/api/", "/health"],
},
],
sitemap: `${SITE_URL}/sitemap.xml`,
host: SITE_URL,
};
}

View file

@ -0,0 +1,9 @@
import { robotsRoute } from "@profullstack/x402-gateway/next";
import { gateway } from "@/lib/crawl-gateway";
// Generated from the same crawler lists the gateway enforces: training
// crawlers are refused everywhere but /crawl (where they can buy a pass),
// retrieval crawlers are named as welcome, everyone else gets the rules below.
export const GET = robotsRoute(gateway, {
disallow: ["/api/", "/health"],
});