mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
Several signals AI agents and search engines use to discover and install a project were misaligned or missing: * ``docs/app/layout.tsx`` set ``metadataBase`` to ``https://chopratejas.github.io/headroom/`` while the live docs run on Vercel — every page's ``og:url`` and ``twitter:url`` resolved to a URL that returns 404 for ``/llms.txt``. Now points at the live Vercel host (overridable via ``NEXT_PUBLIC_SITE_URL`` for a future custom domain). Adds explicit ``openGraph`` and ``twitter`` metadata so social shares render a card with the project's pitch. * No ``llms.txt`` at the GitHub repo root. AI agents crawling ``github.com/chopratejas/headroom/`` saw only the README. The new ``llms.txt`` follows the llmstxt.org convention: 1-line pitch, canonical docs links, copy-paste install commands (pip / npm / Docker / proxy / ``headroom wrap``), and entry points for the library, proxy, MCP server, and SDK integrations. Points at the Fumadocs-generated ``/llms.txt`` and ``/llms-full.txt`` for the full picture. * ``pyproject.toml`` ``Documentation`` URL pointed at the GitHub README anchor. Updated to point at the docs site so PyPI visitors land on searchable docs, and adds an ``AI / LLM Index`` URL pointing at the Fumadocs ``/llms.txt``. * No explicit AI-bot allow list. Added ``docs/app/robots.ts`` (Next 13+ App Router convention) with explicit allows for GPTBot, ClaudeBot, PerplexityBot, Google-Extended, OAI-SearchBot, ChatGPT-User, Cohere-AI, CCBot, and Applebot-Extended. Wildcard allow as the catch-all. Advertises the sitemap. * No ``sitemap.xml`` route. Added ``docs/app/sitemap.ts`` that pulls every Fumadocs page out of ``source`` (same source backing ``/llms.txt``, search, and OG images) so search and AI crawlers can enumerate doc pages without scraping HTML. * README didn't tell AI agents where to look. Added a 2-line pointer near the top nav row: read ``/llms.txt`` here, or fetch the live index / full docs blob. Also tightened the GitHub repo description and added five topics (``claude-code``, ``cursor``, ``tokens``, ``prompt-engineering``, ``typescript``) via ``gh repo edit`` — that's already live on the repo, not part of this commit. No Python or Rust code changes; ``make ci-precheck`` was run to confirm the test slice still passes.
42 lines
1.7 KiB
TypeScript
42 lines
1.7 KiB
TypeScript
// Next.js App Router robots convention (Next 13+). The default Next
|
|
// behaviour allows everything; this file makes the intent explicit so
|
|
// AI-bot operators that read an opt-in list (GPTBot, ClaudeBot,
|
|
// PerplexityBot, Google-Extended, etc.) see a clear green light, and
|
|
// so the sitemap is discoverable.
|
|
//
|
|
// Headroom docs are open-source documentation we WANT indexed. If a
|
|
// future page should be excluded, add it to the ``disallow`` list of
|
|
// the relevant rule.
|
|
|
|
import type { MetadataRoute } from 'next';
|
|
|
|
const SITE_URL = process.env.NEXT_PUBLIC_SITE_URL ?? 'https://headroom-docs.vercel.app';
|
|
|
|
export default function robots(): MetadataRoute.Robots {
|
|
return {
|
|
rules: [
|
|
// Bot-specific allows. These names are the literal user-agent
|
|
// strings each operator publishes. Listing them explicitly is
|
|
// the documented way to opt INTO AI-training / AI-search
|
|
// indexing — silence (no rule) is treated as opt-out by some
|
|
// operators (notably Google-Extended).
|
|
{ userAgent: 'GPTBot', allow: '/' },
|
|
{ userAgent: 'OAI-SearchBot', allow: '/' },
|
|
{ userAgent: 'ChatGPT-User', allow: '/' },
|
|
{ userAgent: 'ClaudeBot', allow: '/' },
|
|
{ userAgent: 'Claude-Web', allow: '/' },
|
|
{ userAgent: 'anthropic-ai', allow: '/' },
|
|
{ userAgent: 'PerplexityBot', allow: '/' },
|
|
{ userAgent: 'Perplexity-User', allow: '/' },
|
|
{ userAgent: 'Google-Extended', allow: '/' },
|
|
{ userAgent: 'cohere-ai', allow: '/' },
|
|
{ userAgent: 'CCBot', allow: '/' },
|
|
{ userAgent: 'Applebot-Extended', allow: '/' },
|
|
// Catch-all so traditional search crawlers also see an
|
|
// explicit allow.
|
|
{ userAgent: '*', allow: '/' },
|
|
],
|
|
sitemap: `${SITE_URL}/sitemap.xml`,
|
|
host: SITE_URL,
|
|
};
|
|
}
|