Download README.md from munod/tachyone-multi: direct link, hf CLI and curl.
- Browser
- Download file 15 kB
-
https://huggingface.co/munod/tachyone-multi/resolve/main/README.md
- Command line
-
hf download hf://munod/tachyone-multi/README.md
-
curl -L -o README.md https://huggingface.co/munod/tachyone-multi/resolve/main/README.md
license: apache-2.0
library_name: tachyone
language:
- en
- pt
- es
- fr
- de
- multilingual
tags:
- decision-engine
- system-one
- calibration
- multilingual
- local-first
pipeline_tag: text-classification
tachyone-multi (System One decision engine)
Status: released (
v0.9.0), revision 2026-10-08 β both checkpoints are the B-15 template-train retrain (English carries its P3 stack rebuilt on the new trunk); revision commits are listed indocs/huggingface.md. Trained on NVIDIA L4 23GB and published as LoRA adapters (munod/tachyone-en,munod/tachyone-multi); measured numbers below come frombenchmarks/report.md.This revision is B-15: both checkpoints retrained on
template_split: "train", and the holdout number is honest for the first time. The data recipe was corrected end to end (sampler strides decoupled so every language trains every label, evals made leave-one-template-out, phrase content Γ3, volume 1.7Γ) and every training set regenerated so training and the holdout eval share 0% of templates. The multilingual lineage carries the two-epoch interleaved touch-up that fixed the B-14noulcell inversion, and the trainer now ships a per-cell validation monitor. On phrasing never seen in any training (eval_*_domains_holdout): multilingual 0.8355, English 0.9825 β +0.127 / +0.131 over the previously released adapters, with the in-template β holdout gap at β0.1105 / β0.0179. Disclosures, never smoothed: Englishchoiceconfidence is refitted on the rebuilt never-trained holdout and its pooled fit pegged the grid at T = 10.0, so in-domainchoiceECE reads 0.5243 honestly instead of in-sample (P3's in-domain β€ 0.05 acceptance recorded NOT met for the rebuilt asset; accuracy is untouched); and B-1's per-language ECE β€ 0.05 on holdout is NOT met for the multilingual arm (choicees/it/pt,scorede/es/it/nl β full numbers in.specs/project/BACKLOG.mdB-15).The v0.8.0 English artifact was B-13: the mixture retrain (history). Same recipe and seed as the B-5 run, one factor changed β the training data: 35,540 records (the 21,000 five-domain records plus 11,000 pinned public records β MultiNLI, BoolQ, Banking77, licences in
training/data/jev_sources.lock.jsonβ and 3,540 executable-rule-tree family records), then the identical frozen-trunkchoice-bank fit. It sweeps both English splits 0.964 β 1.000 (gate strict 1.000; unseen-text rows 0.9987) β and, measured against a fresh same-recipe control that moves +0.6, it lifts Intelligence on the 231 public JevBench items (evaluation-only, never trained on) from 8.4 to 15.0. The probes below are the honest external check: XNLI 0.341 β 0.566, typed-decisions 0.269 β 0.367.The v0.8.0 multilingual artifact was B-5b: five-domain coverage (history). Its trunk is joint-trained on 30,000 multilingual five-domain records (support keeps its 18,000; four new domains 3,000 each) and its
choiceheads were re-fitted on the frozen trunk bytraining/fit_choice_bank.pyβ the same construction as the English artifact, recipe shipped next to the weights aschoice_bank_fit.json. On the five-domain split it scores 0.9975 (previous adapter zero-shot: 0.561), gate strict 1.000, all six languages at ECE β€ 0.004.Labels (B-11 + B-12, ADR-0014/ADR-0015): every label in all three primitives is derived from the text it accompanies β
noulfrom its phrase bank (requestβ 1,neutral/empty β 0),scorefrom the tone's level (empty β middle),choicefrom the option the state names (empty β the catch-allother). All datasets were regenerated and the label audit published with the evaluation reports 0 contradictory rows. These numbers are not comparable with pre-B-11/pre-B-12 measurements: the old labels contradicted 121 of 241 request-toned English rows, left everynoullabel ines/de/nlat 0, and gave 7.8% ofscorerows a "near-tie" the text never showed.Provenance, stated plainly: each adapter carries its recipe next to the weights (
choice_bank_fit.json) β the English adapter is the B-15 five-domain bank over the template-train trunk (en_tt, 6 epochs) with the P3 stack rebuilt on it; the multilingual adapter is the B-15 joint run (52,200 records, 8 epochs, then the 2-epoch interleaved touch-up βmulti_tt_il) plus the frozen-trunk head fit. Both heads-first constructions exist because training the corrected labels makesnoul+scoretrivial and costschoice(an identical-recipe control landed 13 points lower, L-006) β the fix that worked was re-fitting the head on a trunk that already knew the data.
Model details
- Developed by: The Tachyone Authors.
- Model type: non-autoregressive encoder with three task distributions (
noul,choice,score), answering typed questions about a state in one forward pass. - Trunk: ModernBERT-large (English) and mmBERT-base (100+ languages); see ADR-0007.
- Adapters:
munod/tachyone-en,munod/tachyone-multi(LoRA; load base + adapter). - License: Apache-2.0.
- Repository: https://github.com/munod/tachyone
Uses
Tachyone answers atomic choice / score / noul questions about a state and returns typed values
with probabilities and confidence. It speaks the TypeSafe Jev /v1/systemone wire protocol as
a drop-in and runs locally/offline with no API key. Compose several atomic answers in code
rather than asking one broad question.
Out of scope: free-form text generation, multi-step reasoning, and any decision requiring extended deliberation β decompose those into atomic questions and combine results in code.
Bias, risks, and limitations
- Probabilities are only meaningful after calibration; the shipped temperature must be
applied (see
docs/training.md). - Synthetic training data can inherit generator biases; public probes are evaluation-only.
- Confidence is a property of the distribution, not a guarantee of correctness.
Training
Deterministic synthetic JSONL (training/generate_data.py) supervised with an RLCD
proper-scoring objective (training/finetune_rlcd.py), then temperature-calibrated on a held-out
split (training/fit_calibration.py). Configs and seed live under training/configs/.
Evaluation
Reported by training/evaluate.py and rendered by benchmarks/report.py (accuracy, ECE, p50/p95
latency per primitive and language).
Full-scale run (NVIDIA L4 23GB): 36,000 five-domain English / 52,200 five-domain
multilingual records generated with template_split: "train" (the holdout evals share 0% of
training templates) / 1,500β7,500 eval deterministic synthetic records (fully localized per
language β all five domains ship the seven training languages, a learnable other option with
rich descriptions, per-record RNG, one-in-six distractor clauses), LoRA (r=16 English, r=64
multilingual) plus a dedicated low-rank choice head (near-identity init); 6 epochs English /
8 epochs multilingual, batch 8 Γ grad-accum 4, bf16 + gradient checkpointing, and a
per-cell validation monitor (cell_monitor.jsonl) so one (primitive, domain, language)
collapse cannot hide behind the aggregate. Both published artifacts are
bank-over-frozen-trunk: the trunk is trained jointly, then training/fit_choice_bank.py
re-fits the choice heads on cached encodings (rank 128, 8 epochs at lr 1e-4), shipping its
recipe as choice_bank_fit.json; the multilingual trunk additionally carries the two-epoch
interleaved touch-up (training/interleave_continue.py).
| Checkpoint | Split | Accuracy | ECE (calibrated) | p50 (ms) |
|---|---|---|---|---|
| English (ModernBERT-large + five-domain LoRA r=16 + choice-head bank) | support | 1.0000 | 0.188 | 57.7 |
| English, five-domain split | 5 domains | 0.9983 | 0.191 | 57.7 |
| English, holdout phrasing (never in training) | 5 domains | 0.9825 | 0.188 | 58.0 |
| Multilingual (mmBERT-base + LoRA r=64 + fitted bank) β routed runtime | support | 0.9107 | 0.051 | 47.4 |
| Multilingual, five-domain split | 5 domains | 0.9156 | 0.019 | 48.0 |
| Multilingual, holdout phrasing (never in training) | 5 domains | 0.8355 | 0.103 | 47.8 |
The English ECE column reads the pooled never-trained fit (below): choice is flattened to
T=10.0 to calibrate held-out rows, which shows up honestly in-domain (0.5243 choice ECE β
accuracy untouched). The multilingual rows are the routed system: ~13β15% of rows fall
through to the English checkpoint by language routing (empty state or no Latin-language signal
β .specs/project/BACKLOG.md L-013), which is why the five-domain multilingual row sits
below the multilingual checkpoint's own 0.9885 when it answers everything.
Per primitive (English, five-domain): choice 0.9956, noul 0.9996, score
0.9996; (multilingual, five-domain): choice 0.9564, noul 0.9136, score
0.8768; (holdout phrasing): English choice 0.9600 / noul 0.9956 / score
0.9920, multilingual choice 0.8960 / noul 0.7836 / score 0.8268.
Label audit: every noul row is judged against its own text β 0 contradictory in every
eval set (positive rates 0.472β0.486), per language in benchmarks/report.md.
Confidence comes from evidence (P3 β stack rebuilt on the v0.9.0 trunk). The English
adapter ships two extra assets: state_prototypes.json (K=32 spherical k-means centroids of
its own 36,000 training states) and confidence_calibration.json (that bank plus a fitted
noul map, with sha256 provenance of both inputs). A noul answer keeps its direction β the
cosine still decides yes/no β and takes its magnitude from
strength = max_k cos(state, centroid), clamped to [0.5, 1]. The temperatures come from the
rebuilt never-trained holdout (a stride slice of the new training file plus never-trained
public rows, re-asserted against the public items before a byte was written): choice
T = 10.0 β the pooled fit pegged the top of the pre-registered DEFAULT_GRID, a direct
consequence of training on template_split: "train" while the holdout keeps its hard rows β
noul T = 0.75 with the map, score pinned at 0.1 because its answer is the
expected value (the pin's own record: rounded-EV accuracy 0.53 at T=1 vs 0.9996 shipped).
Accuracy is untouched by construction (argmax and direction never move). Recorded NOT met:
P3's in-domain ECE β€ 0.05 acceptance for the rebuilt asset β choice reads 0.5243
in-domain under T=10.0 (the honest price of calibrating held-out rows; the fix is a fit-basis
redesign, not another fit β see .specs/features/jevbench/spec.md P3 results for why no
legal fit set can see bench difficulty). The multilingual adapter has no evidence stack; its
served temperature_calibration.json is fitted on 15,000 pooled in-template + holdout
predictions (choice 6.0, noul 0.25, score 0.25 β fitted where the model is not
saturated, per L-015). The v0.8.0 off-domain probe deltas (XNLI raw ECE 0.426 β 0.264,
typed-decisions 0.578 β 0.416, MASSIVE 0.583 β 0.296, JevBench Calibration 0.0 β 52.6 with
its recorded gate of 60 not met) were measured on the v0.8.0 artifact and are not
re-measured for v0.9.0 β re-probing is recorded as pending.
What the English rows mean (in-sample, stated plainly). The support and five-domain splits
read 1.0000 / 0.9983 because their phrasing is seen β training excluded the holdout template
pool, but these splits still draw the seen templates β so saturation here is a ceiling, not a
claim. The honest number is the holdout row: 0.9825 overall (choice 0.9600), phrasing no
training ever emitted. The external checks quoted by earlier revisions β JevBench Intelligence
8.4 β 15.0 against a fresh control's +0.6, XNLI 0.566, typed-decisions 0.367 β
were measured on the v0.8.0 mixture artifact and are not re-measured for v0.9.0
(JevBench was set aside by decision after its maintainer changed the submission methodology);
re-probing both adapters is recorded as pending. Full tables:
docs/benchmarks.md.
The v0.8.0 multilingual gates (B-5b β history; v0.9.0's gates are B-15's). The previous adapter measured
0.561 zero-shot on the new five-domain split (worst new domain agent_tools 0.438, support
0.8413) β the gates were support β₯ 0.8413, worst new domain β₯ 0.70, per-domain ECE β€ 0.05
(exceptions declared), gate strict published. That artifact posted support 1.0000,
worst new domain 0.993, per-domain ECE 0.0005β0.0038 (zero exceptions) and gate
strict 1.000 over 2,500 choice rows. v0.9.0 (B-15) posts on the unchanged current
benchmark: multilingual support 0.9760, worst new domain 0.9747, per-domain ECE 0.0019β0.0202
(zero exceptions), strict 1.000, every language β₯ baseline; overall 0.9885 / ECE 0.0069 when
the multilingual checkpoint answers everything β the routed rows above show the pair as
served. On the support-only routed split (13β15% of
rows fall through to the English checkpoint by language routing, .specs/project/BACKLOG.md
L-013) the per-language story is unchanged in kind: three of six languages sit above the
0.05 ECE target on that split, and on the holdout the B-1 per-language ECE criterion is
NOT met (recorded with numbers in BACKLOG B-1) β open under NFR-C06 / BACKLOG B-1.
The CUDA-graph fast path
(TACHYONE_FAST=1) gave 2.54Γ p50 on English (17.31 β 6.81 ms) and 3.66Γ on the
multilingual five-domain path (13.97 β 3.82 ms), both with 0 top-label flips β measured on
the v0.8.0 artifacts, pending re-measurement on v0.9.0. The gate and
the keyed heads execute after the encode, which the graphed path never sees.
Robustness (B-4). On a noisy view (one surface edit β typo/accents/casing β applied to 15% of states) English moves 1.0000 β 0.9953 and the multilingual routed pair 0.9107 β 0.9087, so the released adapters are robust to this noise model.
Full tables and environment are in
benchmarks/report.md.
Citation
@misc{tachyone2026,
title = {tachyone: a local-first System One decision engine},
author = {The tachyone Authors},
year = {2026},
howpublished = {\url{https://github.com/munod/tachyone}}
}