# Independent evaluation reproduction

Botrace publishes **normalized evaluation receipts**. It does not run robot learning stacks in this repository.

LeRobot, Isaac GR00T, OpenPI, MolmoAct, Xiaomi-Robotics-1, and PhAIL harnesses stay **out of this repo**. Run them in a throwaway environment (local GPU box, cloud VM, or lab machine). Import JSON here only after editorial review.

## Why this is separate

- The public HUD is static JSON. A training stack would change the product into an unpublished lab.
- Reproduction is a **verification label**, not a sixth GHI spoke and not a cross-benchmark rank.
- Redistributing datasets or weights must follow the original licenses. Botrace does not blanket-relicense them.

## Result contract

A reproduced row is the same shape as `data/eval-results.json`:

```json
{
  "id": "phail-v1-0-<model>-reproduced-<yyyy-mm-dd>",
  "benchmark_id": "phail-v1-0",
  "benchmark_version": "1.0",
  "model_id": "hugging-face-smolvla",
  "body_id": "",
  "system_label": "SmolVLA · Franka Research 3 + Robotiq 2F-85 (Botrace reproduction)",
  "environment": "real",
  "source_relation": "third-party",
  "verification": "reproduced",
  "adaptation": {
    "mode": "task-finetune",
    "demonstrations": null,
    "gradient_steps": null,
    "training_data": "",
    "train_test_relation": "State whether eval objects were in training."
  },
  "metrics": [
    { "id": "uph", "label": "Units per hour", "value": 0, "unit": "uph", "direction": "higher" }
  ],
  "trial_count": null,
  "confidence_interval": null,
  "artifacts": [{ "kind": "notes", "url": "https://" }],
  "notes": "Hardware, latency path, and deviations from the steward protocol.",
  "source": "https://",
  "as_of": "YYYY-MM-DD"
}
```

Empty numbers stay empty. Do not copy a vendor blog into `verification: reproduced`.

## PhAIL (preferred first independent row)

Steward recipe (external):

- Protocol and hardware: https://phail.ai/releases/v1.0
- Methodology paper: https://arxiv.org/abs/2605.29710
- Public table currently recorded in-repo: OpenPI π0.5, GR00T N1.6, ACT, SmolVLA

To reproduce:

1. Provision a machine **outside** this git tree.
2. Follow the v1.0 hardware spec (Franka Research 3 + Robotiq 2F-85, DROID-style cameras) or stop and record a protocol deviation — a different arm is a different comparability key.
3. Use the steward dataset, training scripts, and eval harness. Do not mix in private data without saying so.
4. Keep model identity blinded at the operator if the protocol requires it.
5. Export throughput, completion, and MTBF/A with trial count and interval when the harness provides them.
6. Drop a JSON object next to this repo (not committed from the training machine) and run:

```bash
python3 scripts/import-eval-result.py path/to/result.json
```

That command validates the object. It does **not** append it. A reviewed `--apply` still only appends after `validate-data.py` would accept the row.

## Open-model simulation (second path)

RoboCasa365 official leaderboard rows are already recorded for π0 and π0.5. A Botrace reproduction would use the steward protocol at https://robocasa.ai/leaderboard.html in a separate env, then import with `verification: reproduced` and a distinct `id`.

Do not average RoboCasa365 with PhAIL.

## What not to do

- Vendor LeRobot, CUDA wheels, or checkpoint tarballs into `botrace`
- Auto-promote a reproduced number onto radar rungs or `story.json`
- Replace an official-leaderboard row in place — append a new id
- Invent trial counts or confidence intervals the harness did not print
