← Back to all posts

Building Robotswise: Balancing AI Crawler Governance and Google SEO

Over the past two years, the fundamental contract between website owners and web crawlers has fractured.

For decades, the deal was straightforward: search engines like Google and Bing crawl your public pages, index them, and in return, send organic search traffic back to your domain. Today, a new wave of autonomous bots—OpenAI’s GPTBot, Anthropic’s ClaudeBot, ByteDance’s Bytespider, and Common Crawl’s CCBot—harvest terabytes of written work, documentation, and open-source tutorials to train foundational models.

For many publishers, independent builders, and creators, this feels like an unfair exchange: our work is ingested into proprietary weights, often with zero attribution and zero downstream traffic.

To address this, I built and launched Robotswise (AI Robots.txt Generator). Here is the engineering breakdown of how it works and the architectural decisions behind it.


1. The Core Dilemma: Nuance vs. Scorched Earth

When creators first discover AI crawlers consuming their bandwidth and training on their data, the knee-jerk reaction is often:

User-agent: *
Disallow: /

This “scorched earth” approach is disastrous for organic discovery. It abruptly de-indexes your domain from Google Search, Bing, and DuckDuckGo, cutting off the lifeblood of your project.

Even when attempting to configure a more selective robots.txt, non-trivial questions emerge:

  • If I block GPTBot, does that also break ChatGPT’s ability to cite my product when users query OAI-SearchBot in search mode?
  • Is Google-Extended responsible for Google search ranking? (No—it solely controls Gemini training and Grounding, while Googlebot handles traditional search indexing).
  • How do I keep sensitive administrative endpoints (/admin/, /private/) private without accidentally advertising their paths to malicious actors?

A manual text file was no longer sufficient. Site owners needed an intuitive policy compiler that separates AI Training Crawlers from AI Search Engines and Standard Web Indexers.


2. Architectural Principles

Zero-Knowledge, 100% Client-Side Processing

When crafting a robots.txt file, developers frequently input internal directory paths (/api/internal/, /staging/) and their canonical sitemap locations. Sending these paths to a remote server creates an unnecessary security liability.

In Robotswise, no data ever leaves the user’s browser. The state machine, token validation, and string interpolation run entirely within the client bundle.

Multi-tiered Policy Presets

To minimize cognitive fatigue, the tool provides three one-click baseline policies:

  1. Block All AI Crawlers: Blocks every known frontier model scraper (GPTBot, ClaudeBot, Google-Extended, Meta-ExternalAgent, CCBot, Bytespider, etc.).
  2. Allow AI Search, Block Training: Keeps search citation bots (OAI-SearchBot, PerplexityBot) open so your brand appears in AI answers, while shutting the door on bulk training bots.
  3. Open Access: Retains standard discovery protocols.

3. The Technical Stack

Frontend: Next.js + React 19 + Tailwind CSS v4
Component Architecture: Radix UI primitives + Lucide Icons
Edge Runtime: Vinext + Cloudflare Pages (Serverless static edge)
Protocol Extensions: WebMCP (In-browser Model Context Protocol)

In-Browser WebMCP Integration

One unique experimental feature I baked into Robotswise is WebMCP (Web Model Context Protocol). If a user visits the tool with an AI Agent or browser assistant that supports the protocol, Robotswise registers a client-side tool dynamically:

const context = (document as WebMcpDocument).modelContext;
if (context?.registerTool) {
  context.registerTool({
    name: 'get_robots_txt',
    title: 'Generate robots.txt',
    description: 'Return the robots.txt currently configured in this generator.',
    inputSchema: { type: 'object', properties: {} },
    annotations: { readOnlyHint: true, untrustedContentHint: false },
    execute: () => ({ filename: 'robots.txt', content: robotTextRef.current }),
  });
}

This allows agents navigating the web to interactively invoke the generator, retrieve the generated directives, and write them directly into a project repository without manual copy-pasting.


4. Key Takeaways & What’s Next

  1. Crawler tokens are a moving target: Bot identifiers change quickly as labs fork their scrapers. Maintaining an open, versioned list of tokens is essential.
  2. robots.txt is an advisory signal, not access control: A critical reminder highlighted in the UI is that robots.txt communicates access preferences to cooperative bots. True confidential endpoints must always be defended by authentication, WAF rules, and noindex headers.

The tool is live, free, and bilingual (English and Chinese): 👉 Try it here: https://robot-generate.toolyard.cc/