# Editorial handbook

Botrace is a public evidence graph for the race to **general hybrid intelligence**. This handbook is for humans and in-repo agents who write JSON. It is not a scoring formula.

Read with `docs/v1-spec.md` and `data/README.md`. Corrections append to `data/changes.json`.

## What the desk publishes

- **Story** (`story.json`, radar rungs, breakthroughs, featured ids) is manual.
- **Signals** (`data/signals/`) are automated ingest. They are never copied onto the cockpit as fact.
- **Catalog** organizations and open models may be added without becoming GHI racers (`racer: false`).
- **Racers** are the frozen published roster in `AGENTS.md` (original fourteen listed + Covariant exited + Wave A+B nineteen listed hollow). Do not add a racer without Martin changing that list. MagicLab is hold. GDM, NVIDIA, Meta FAIR, and Ai2 are catalog-only.

## Two empties

| Empty | Means |
|---|---|
| Missing GHI rung | **Not achieved** under the current rubric |
| Empty funding, units, metrics, policy, body, autonomy, interval | **Unknown / undisclosed**, never zero |

Do not fill a number to make a chart look finished.

## Source hierarchy

A **primary source** identifies who made the claim. It does **not** mean independent truth.

When two sources disagree, show both or neither. Prefer, in order:

1. Operator / customer statement (press, filing, AGM, named floor)
2. Dual-sided operator + vendor
3. Vendor technical post that names the site, task, and date
4. Benchmark owner / official leaderboard for **evaluation** rows only
5. Independent reproduction with artifacts

Media restatements, aggregator databases, and leaked slides are not enough to lock a rung, grade, valuation, or metric.

Every filled numeric or round-size field needs `as_of` and `source`.

## Inclusion

Add a **racer** only when Martin changes the roster.

Add a **catalog organization** when it has a stable `id`, public home, and a reason to join models or benchmarks (labs, open ecosystems, eval stewards). `racer: false` keeps it off Radar, Track, Map, and approach occupancy. Occupancy is a join on live racer models’ `approach_ids`, not a hand list of example models.

Add a **model** when it is a named policy or body with a primary source. Do not invent Hz, papers, or successors. Sunday ACT-1/2 are not LeRobot ACT.

Add a **benchmark** when a protocol has a steward, version, embodiment, and stated limitations.

Add an **eval result** only when the full context is known enough to prevent false comparison:

`policy/checkpoint + body + sensors/control + adaptation + task/environment + human support + protocol/version + date`

Add a **deployment** only with who/where/task/source/as_of. Empty autonomy, policy, and metrics beat a guess.

## Zero-shot, fine-tune, and vendor stacks

Do not treat a leaderboard row as the vendor’s production policy.

- **Zero-shot** means the source says the checkpoint was not adapted to the eval task.
- **Task-finetune** / **from-scratch** must travel with the result.
- OpenPI on PhAIL is not Physical Intelligence π0.6 in a laundromat.
- Simulation success is not a floor.

If the source does not say how the model was adapted, leave `adaptation.mode` empty.

## HITL, production, unattended

These tags are not GHI rungs.

- **HITL / teleop / supervised** is human support. It can be honest production and still fail Last 5–6.
- **Remote** is sourced fleet remote-help. It is not unattended and must not be typed as unattended.
- **Production** is an operating category. Digit tote work can be production without being generalist.
- **Unattended** needs sourced evidence of work without a babysitter, not a hero clip.
- Empty `autonomy_mode` means the desk did not lock a value. Narrative remote help may stay in `task`.

Do not change a radar rung to encode honesty. Use the attended tag and the deployment record.

## Comparability

Results are comparable only when **benchmark id, version, embodiment, adaptation regime, metric definition, and setting** match.

The Evaluation Observatory must not:

- rank across PhAIL, RoboCasa365, RoboArena, or NIST
- sum radar rungs
- compute GHI closeness or “#1”

A frontier chart is allowed **inside one protocol**. It is not a championship.

## Grades, rungs, and promotions

Bots do not assign evidence grades, rungs, valuations, or bottleneck rank. A filled radar lock needs both `as_of` and a primary source; otherwise the spoke stays empty (not-achieved). Do not invent dates or URLs to keep a lock live.

A signal becomes a race call only when a human (or repo agent following this handbook) writes `data/breakthroughs.json` and, if needed, a `changes.json` receipt.

Revocations append. They do not silently delete history.

## Watchers

GitHub Actions and `scripts/watch_signals.py` write **only** `data/signals/*.json`. They must not touch `story.json`, rungs, capital, grades, or eval results.

## Reproduction

Independent evaluations run in a **separate** environment. See `docs/eval-reproduction.md`. This repo stores normalized JSON after review, not training stacks.
