Chronos Policy: original paper checkpoints
These are the 48 original simulation checkpoints used by the paper: Chronos and the timestamp-blind diffusion-transformer (DT) baseline, three training seeds, and eight tasks. Each checkpoint is approximately 560 MB.
Dataset publication and the Conveyor/Pivoting training/evaluation campaign are
complete. This repository contains
the original selected weights, not retrained replacements. manifest.json records their exact sizes and SHA-256
hashes, fixed evaluation seeds, selected scene batches, and paper reference counts.
All 48 remote checkpoint hashes have been verified against the paper manifest. The compact release model also matches the pinned implementation bit-for-bit in 32 eight-step DDIM comparisons: seed-0 Chronos and DT checkpoints for all eight tasks, with both ordinary and padded observations. These compatibility checks do not establish full closed-loop benchmark reproduction.
Code and setup: https://github.com/changhaowang/ilp/tree/release
The GitHub repository currently requires access; public code publication is pending.
Files use <task>/<method>/seed<N>.ckpt, matching the code's checkpoints/ layout.
Task names are threading, assembly, transport, coffee, drawer_cleanup,
pouring, conveyor, and flipup (Pivoting). Methods are chronos and dt.
hf download changhaowang/chronos-policy --include 'conveyor/**' --local-dir checkpoints
Checkpoints include the model weights, original configuration, and normalizers. Load only trusted checkpoint files. The release's artifact command verifies the paper hashes before evaluation. Exact scene seeds alone do not establish identical physics or training across software/hardware versions.
Conveyor and Pivoting use the paper's selected exploratory batches; these results are not an independent held-out estimate.
Completed Conveyor/Pivoting verification
All 12 training runs and 21,600 final evaluation episodes are complete. Pivoting reproduces the original selected weights, nominal selections, episode outcomes, simulation-step counts, and aggregate tables exactly. Conveyor is approximate: original-checkpoint mean differences reach 3.33 percentage points, and fresh training differences reach 7.33 points. No seeds or stressed-condition checkpoint choices were changed to improve agreement.
See the complete report and unrounded comparison table, with all 432 cohort results and selected-checkpoint comparisons. These results cover Conveyor and Pivoting; the six DexMimicGen tasks have setup/data-contract checks, not a new full benchmark rerun.