Chronos Policy: original paper checkpoints

These are the 48 original simulation checkpoints used by the paper: Chronos and the timestamp-blind diffusion-transformer (DT) baseline, three training seeds, and eight tasks. Each checkpoint is approximately 560 MB.

Dataset publication and the Conveyor/Pivoting training/evaluation campaign are complete. This repository contains the original selected weights, not retrained replacements. manifest.json records their exact sizes and SHA-256 hashes, fixed evaluation seeds, selected scene batches, and paper reference counts.

All 48 remote checkpoint hashes have been verified against the paper manifest. The compact release model also matches the pinned implementation bit-for-bit in 32 eight-step DDIM comparisons: seed-0 Chronos and DT checkpoints for all eight tasks, with both ordinary and padded observations. These compatibility checks do not establish full closed-loop benchmark reproduction.

Code and setup: https://github.com/changhaowang/ilp/tree/release

The GitHub repository currently requires access; public code publication is pending.

Files use <task>/<method>/seed<N>.ckpt, matching the code's checkpoints/ layout. Task names are threading, assembly, transport, coffee, drawer_cleanup, pouring, conveyor, and flipup (Pivoting). Methods are chronos and dt.

hf download changhaowang/chronos-policy --include 'conveyor/**' --local-dir checkpoints

Checkpoints include the model weights, original configuration, and normalizers. Load only trusted checkpoint files. The release's artifact command verifies the paper hashes before evaluation. Exact scene seeds alone do not establish identical physics or training across software/hardware versions.

Conveyor and Pivoting use the paper's selected exploratory batches; these results are not an independent held-out estimate.

Completed Conveyor/Pivoting verification

All 12 training runs and 21,600 final evaluation episodes are complete. Pivoting reproduces the original selected weights, nominal selections, episode outcomes, simulation-step counts, and aggregate tables exactly. Conveyor is approximate: original-checkpoint mean differences reach 3.33 percentage points, and fresh training differences reach 7.33 points. No seeds or stressed-condition checkpoint choices were changed to improve agreement.

See the complete report and unrounded comparison table, with all 432 cohort results and selected-checkpoint comparisons. These results cover Conveyor and Pivoting; the six DexMimicGen tasks have setup/data-contract checks, not a new full benchmark rerun.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading