INTEGRATION

RunLedger CiteMap Integration Guide

Complete the AI discoverability handshake for runledger.io with robots.txt, llms.txt, JSON-LD, and dense context docs.

integration discoverability llms Updated 2026-02-02

Direct Answer

Publish a permissive robots.txt, add a RunLedger-specific llms.txt, inject SoftwareApplication JSON-LD, generate high-density Markdown context files, and log AI user agents for /llms.txt and *.md requests. The snippets below are filled in for runledger.io.

Quick Decision

Use CiteMap when Consider alternatives when
You publish public RunLedger docs and want higher LLM citation quality. Your docs are private or not ready to be indexed.
You can maintain dense Markdown contexts alongside HTML docs. You cannot keep context files up to date.
You need measurable AI discoverability via logs and analytics. You do not need AI traffic attribution.

Phase 1: Protocol Layer (Required)

Keep these assets in sync. Update llms.txt and context files whenever docs or API behavior change. Do not include secrets or private content.

1.1 Update robots.txt

Allow AI user agents and expose /llms.txt and context files.

text
User-agent: *
Allow: /
Allow: /llms.txt
Allow: /docs/contexts/
Disallow: /monitor/

# Explicitly welcome AI researchers
User-agent: GPTBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: ClaudeBot
Allow: /

Sitemap: https://runledger.io/sitemap.xml

1.2 Deploy llms.txt

Point LLMs at the dense Markdown contexts for RunLedger.

text
# Title: RunLedger Documentation
# Description: The official technical reference for RunLedger.
# Contact: info@runledger.io | https://x.com/runledger
# Core Knowledge Contexts
- https://runledger.io/docs/contexts/overview.md  --> High level architecture
- https://runledger.io/docs/contexts/api-ref.md   --> CLI + protocol reference
- https://runledger.io/docs/contexts/guides.md    --> Integration tutorials
# Optional: Full content dump
- https://runledger.io/docs/full-ref.md

1.3 Inject JSON-LD Schema

Use SoftwareApplication schema so search and AI systems classify RunLedger correctly.

html
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "SoftwareApplication",
  "name": "RunLedger",
  "applicationCategory": "DeveloperApplication",
  "operatingSystem": "Any",
  "offers": {
    "@type": "Offer",
    "price": "0",
    "priceCurrency": "USD"
  },
  "description": "Deterministic CI harness for tool-using agents with record/replay, contracts, and budget gates.",
  "url": "https://runledger.io/"
}
</script>

Phase 2: Content Optimization (CiteMap Strategy)

Create dense Markdown context files for each major RunLedger doc page and strip all UI noise.

  • No HTML tags. Pure Markdown only.
  • High density: jump straight to definitions and rules.
  • Code first: use heavily commented snippets.
  • Use absolute URLs for all links.
markdown
# API Reference (RunLedger)

## runledger run
Purpose: execute a suite in record, replay, or live mode.

```bash
# Record live tool calls to cassettes
runledger run ./evals/demo --mode record
# Replay deterministically against a baseline
runledger run ./evals/demo --mode replay --baseline baselines/demo.json
```

## suite.yaml (key fields)

```yaml
suite_name: support-triage          # stable suite id
agent_command: ["python", "agent.py"]  # agent launch command
mode: replay                         # record | replay | live
tool_registry:
  - search_docs                       # allowed tool names
assertions:
  - type: json_schema                 # schema enforced on final output
    schema_path: schema.json
budgets:
  max_wall_ms: 20000                  # wall time cap
baseline_path: baselines/support.json # regression gate baseline
```
bash
# Create the contexts directory
mkdir -p docs/contexts
# Author high-density Markdown contexts
$EDITOR docs/contexts/overview.md
$EDITOR docs/contexts/api-ref.md
$EDITOR docs/contexts/guides.md

2.2 Verify Visibility

Run the CiteMap validator (citemap.io/validate) and target >80% signal ratio for context files.

Phase 3: Feedback Loop (Analytics)

Log requests from AI user agents for *.md and /llms.txt to measure impact.

js
const aiAgents = ["GPTBot", "ChatGPT-User", "ClaudeBot", "FacebookBot"];

export function logAiTraffic(req, res, next) {
  const ua = req.headers["user-agent"] || "";
  const isAi = aiAgents.some((agent) => ua.includes(agent));
  const isTrackedPath = req.path.endsWith(".md") || req.path === "/llms.txt";
  if (isAi && isTrackedPath) {
    console.log("AI_VISIT", { path: req.path, ua });
  }
  next();
}

If you host runledger.io as static HTML, use CDN/edge logs (Cloudflare, Vercel, Netlify) to filter these user agents instead of middleware.

Summary Checklist

  • robots.txt updated to allow GPTBot, ChatGPT-User, and ClaudeBot.
  • llms.txt published at the site root with RunLedger context links.
  • JSON-LD SoftwareApplication schema injected into site head.
  • Context Markdown files generated at /docs/contexts/*.md.
  • Analytics logging added for AI user agents on *.md and /llms.txt.

Tradeoffs

  • Context files require ongoing maintenance as docs and APIs change.
  • Duplicated Markdown increases the risk of stale or conflicting info.
  • AI crawler traffic may require additional log storage or analytics tooling.

When NOT to use RunLedger

Skip this integration if you are not publishing RunLedger docs publicly or if you cannot maintain accurate context files. The product can still run internally without the CiteMap layer.

Next steps