# RunLedger Full Reference (Condensed)

Product definition
RunLedger is a deterministic CI harness for tool-using agents. It records tool calls once, replays them in CI, enforces contracts and budgets, and gates regressions via baselines.

Core workflow
1. Define suites and cases (suite.yaml, cases/*.yaml).
2. Record tool calls to cassettes (record mode).
3. Promote a passing run to a baseline.
4. Replay deterministically in CI (replay mode).
5. Fail PRs on contract, budget, or baseline regressions.

CLI quick reference
```bash
# Initialize a demo suite
runledger init
# Record tool calls to cassettes
runledger run ./evals/demo --mode record
# Replay deterministically with a baseline
runledger run ./evals/demo --mode replay --baseline baselines/demo.json
# Diff a run against a baseline
runledger diff --baseline baselines/demo.json --run runledger_out/demo/<run_id>
# Promote a run to a baseline
runledger baseline promote --from runledger_out/demo/<run_id> --to baselines/demo.json
```

suite.yaml (example)
```yaml
suite_name: support-triage            # stable suite id for CI
agent_command: ["python", "agent.py"]    # agent launch command
mode: replay                           # record | replay | live
cases_path: cases                      # directory of cases
tool_registry:
  - search_docs                         # allowed tool names
  - create_issue
assertions:
  - type: json_schema                   # schema enforced on final output
    schema_path: schema.json
budgets:
  max_wall_ms: 20000                    # wall time cap
  max_tool_calls: 10                    # tool call cap
  max_tool_errors: 0                    # fail on any tool error
baseline_path: baselines/support.json  # regression gate baseline
```

cases/*.yaml (example)
```yaml
id: t1                                 # case id
description: "triage a login ticket"
input:
  ticket: "User cannot login"          # input forwarded to agent
  context:
    plan: "pro"
cassette: cassettes/t1.jsonl           # cassette used in replay
assertions:
  - type: required_fields              # per-case assertion override
    fields: ["category", "reply"]
budgets:
  max_wall_ms: 5000                    # per-case budget override
```

Protocol (JSONL over stdio)
```jsonc
{ "type": "task_start", "task_id": "t1", "input": { "ticket": "..." } }   // runner starts a task
{ "type": "tool_call", "name": "search_docs", "call_id": "c1", "args": { "q": "..." } } // agent requests a tool
{ "type": "tool_result", "call_id": "c1", "ok": true, "result": { "hits": [] } }        // runner returns tool output
{ "type": "final_output", "output": { "category": "billing", "reply": "..." } }         // agent final response
```

Artifacts
- run.jsonl: full event stream
- summary.json: structured results and budgets
- junit.xml: CI-friendly test output
- report.html: human-readable run report

Key links
- Docs: https://runledger.io/docs.html
- Reference: https://runledger.io/reference.html
- Golden Path: https://runledger.io/answers/golden-path
- GitHub: https://github.com/runledger/Runledger
