tachyone-multi (System One decision engine)
Status: released (
v0.8.0), revision 2026-10-02 β English1c88ebef(B-13 mixture + P3 evidence confidence), multilingual3693def1(card). Trained on a single RTX 3060 12GB (B-5b and B-13 on NVIDIA L4 23GB) and published as LoRA adapters (munod/tachyone-en,munod/tachyone-multi); measured numbers below come frombenchmarks/report.md.This revision adds B-13 for the English checkpoint: the mixture retrain. Same recipe and seed as the B-5 run, one factor changed β the training data: 35,540 records (the 21,000 five-domain records plus 11,000 pinned public records β MultiNLI, BoolQ, Banking77, licences in
training/data/jev_sources.lock.jsonβ and 3,540 executable-rule-tree family records), then the identical frozen-trunkchoice-bank fit. It sweeps both English splits 0.964 β 1.000 (gate strict 1.000; unseen-text rows 0.9987) β and, measured against a fresh same-recipe control that moves +0.6, it lifts Intelligence on the 231 public JevBench items (evaluation-only, never trained on) from 8.4 to 15.0. The probes below are the honest external check: XNLI 0.341 β 0.566, typed-decisions 0.269 β 0.367.This revision is B-5b: the multilingual checkpoint now covers five domains. Its trunk is joint-trained on 30,000 multilingual five-domain records (support keeps its 18,000; four new domains 3,000 each) and its
choiceheads were re-fitted on the frozen trunk bytraining/fit_choice_bank.pyβ the same construction as the English artifact, recipe shipped next to the weights aschoice_bank_fit.json. On the five-domain split it scores 0.9975 (previous adapter zero-shot: 0.561), gate strict 1.000, all six languages at ECE β€ 0.004.Labels (B-11 + B-12, ADR-0014/ADR-0015): every label in all three primitives is derived from the text it accompanies β
noulfrom its phrase bank (requestβ 1,neutral/empty β 0),scorefrom the tone's level (empty β middle),choicefrom the option the state names (empty β the catch-allother). All datasets were regenerated and the label audit published with the evaluation reports 0 contradictory rows. These numbers are not comparable with pre-B-11/pre-B-12 measurements: the old labels contradicted 121 of 241 request-toned English rows, left everynoullabel ines/de/nlat 0, and gave 7.8% ofscorerows a "near-tie" the text never showed.Provenance, stated plainly: each adapter carries its recipe next to the weights (
finetune_config.json/choice_bank_fit.json) β the English adapter is the B-13 mixture bank over the B-13 joint run's frozen trunk; the multilingual adapter is the B-5b joint run plus the frozen-trunk head fit. Both heads-first constructions exist because training the corrected labels makesnoul+scoretrivial and costschoice(an identical-recipe control landed 13 points lower, L-006) β the fix that worked was re-fitting the head on a trunk that already knew the data.
Model details
- Developed by: The Tachyone Authors.
- Model type: non-autoregressive encoder with three task distributions (
noul,choice,score), answering typed questions about a state in one forward pass. - Trunk: ModernBERT-large (English) and mmBERT-base (100+ languages); see ADR-0007.
- Adapters:
munod/tachyone-en,munod/tachyone-multi(LoRA; load base + adapter). - License: Apache-2.0.
- Repository: https://github.com/munod/tachyone
Uses
Tachyone answers atomic choice / score / noul questions about a state and returns typed values
with probabilities and confidence. It speaks the TypeSafe Jev /v1/systemone wire protocol as
a drop-in and runs locally/offline with no API key. Compose several atomic answers in code
rather than asking one broad question.
Out of scope: free-form text generation, multi-step reasoning, and any decision requiring extended deliberation β decompose those into atomic questions and combine results in code.
Bias, risks, and limitations
- Probabilities are only meaningful after calibration; the shipped temperature must be
applied (see
docs/training.md). - Synthetic training data can inherit generator biases; public probes are evaluation-only.
- Confidence is a property of the distribution, not a guarantee of correctness.
Training
Deterministic synthetic JSONL (training/generate_data.py) supervised with an RLCD
proper-scoring objective (training/finetune_rlcd.py), then temperature-calibrated on a held-out
split (training/fit_calibration.py). Configs and seed live under training/configs/.
Evaluation
Reported by training/evaluate.py and rendered by benchmarks/report.py (accuracy, ECE, p50/p95
latency per primitive and language).
Full-scale run (single RTX 3060 12GB; B-5b on a single L4 23GB): 21,000 five-domain English /
30,000 five-domain multilingual / 18,000 support-only multilingual train / 1,500β7,500 eval
deterministic synthetic records (fully localized per language β all five domains now ship the
seven training languages, a learnable other option with rich descriptions, per-record RNG,
one-in-six distractor clauses), LoRA (r=16 English, r=64 multilingual) plus a dedicated low-rank
choice head (near-identity init); 8 epochs for multilingual, batch 16, bf16 + gradient
checkpointing. Both published artifacts are bank-over-frozen-trunk: the trunk is trained
jointly, then training/fit_choice_bank.py re-fits the choice heads on cached encodings
(rank 128, 8 epochs at lr 1e-4), shipping its recipe as choice_bank_fit.json.
| Checkpoint | Split | Accuracy | ECE (calibrated) | p50 (ms) |
|---|---|---|---|---|
| English (ModernBERT-large + five-domain LoRA r=16 + choice-head bank) | support | 1.000 | 0.006 | 58.0 |
| English, five-domain split | 5 domains | 1.000 | 0.009 | 26.3 |
| Multilingual (mmBERT-base + LoRA r=64 + fitted bank) β routed runtime | support | 0.895 | 0.062 | 46.9 |
| Multilingual, five-domain split | 5 domains | 0.9975 | 0.0014 | 18.6 |
Per primitive (English): choice 1.000, noul 1.000, score 1.000; (multilingual,
five-domain): choice 0.9924, noul 1.000, score 1.000. Label audit: every
noul row is judged against its own text β 0 contradictory in every eval set (positive
rates 0.472β0.486), per language in benchmarks/report.md.
Confidence comes from evidence (P3, 2026-10-02). The English adapter ships two extra assets:
state_prototypes.json (K=32 centroids of its own 21,000 training states) and
confidence_calibration.json (that bank plus a fitted noul map). A noul answer keeps its
direction β the cosine still decides yes/no β and takes its magnitude from
strength = max_k cos(state, centroid), clamped to [0.5, 1]; choice runs at T=1.5 and
score is pinned at 0.1 because its answer is the expected value. Consequences, all
measured: accuracy is byte-identical to the previous revision everywhere (argmax and
expected value never move), in-domain ECE reads 0.006 / 0.009 instead of an in-sample
0.000 (L-015 β the old number was fitted on the very rows it scored), and off-domain
confidence drops to match what the model actually knows: XNLI raw ECE 0.426 β 0.264,
typed-decisions 0.578 β 0.416, MASSIVE 0.583 β 0.296, JevBench's Calibration axis
0.0 β 52.6 on the 231 public items (its recorded gate of 60 was not met β
.specs/features/jevbench/spec.md P3 results records why).
What the English B-13 row is (in-sample, stated plainly). Both English splits read 1.000
because the held-out synthetic sets share rows with training data (the known 81% choice
row collision) β saturation here is a ceiling, not a claim: accuracy on the 791 rows whose
text never occurs in training is 0.9987, and the external check is the public one β the
231 JevBench items move Intelligence 8.4 β 15.0 against a fresh control's +0.6
(attribution table in .specs/features/jevbench/tasks.md JB-7), with XNLI 0.566 and
typed-decisions 0.367 as the probe deltas. The B-5 trade this row replaces is closed:
the previous bank won choice 0.946 β 1.000 at the cost of noul/score (0.992/0.978 β
0.946/0.946) on support; B-13 posts 1.000 across all three primitives on both splits,
gate strict 1.000 (0 to shared, 0 wrong domain over 2,500 choice rows). Full tables:
docs/benchmarks.md.
The multilingual B-5b gates (fixed before training, all pass). The previous adapter measured
0.561 zero-shot on the new five-domain split (worst new domain agent_tools 0.438, support
0.8413) β the gates were support β₯ 0.8413, worst new domain β₯ 0.70, per-domain ECE β€ 0.05
(exceptions declared), gate strict published. The published artifact posts support 1.0000,
worst new domain 0.993, per-domain ECE 0.0005β0.0038 (zero exceptions) and gate
strict 1.000 over 2,500 choice rows; accuracy on the 2,996 rows whose text never occurs in
training is 0.996. Per-language ECE on that split is 0.0010β0.0041 β all six languages
β€ 0.05 for the first time on a five-domain set; on the support-only routed split (13β15% of
rows fall through to the English checkpoint by language routing, .specs/project/BACKLOG.md
L-013) three of six languages remain above the 0.05 target β nl accuracy 0.715 / ECE 0.167,
de 0.083, it 0.056 β open under NFR-C06 / BACKLOG B-1. The CUDA-graph fast path
(TACHYONE_FAST=1) gives 2.54Γ p50 on English (17.31 β 6.81 ms) and 3.66Γ on the
multilingual five-domain path (13.97 β 3.82 ms), both with 0 top-label flips β the gate and
the keyed heads execute after the encode, which the graphed path never sees.
Robustness (B-4). On a noisy view (one surface edit β typo/accents/casing β applied to 15% of states) English drops only 1.000 β 0.997 and multilingual 0.743 β 0.744, so the released adapters are robust to this noise model.
Full tables and environment are in
benchmarks/report.md.
Citation
@misc{tachyone2026,
title = {tachyone: a local-first System One decision engine},
author = {The tachyone Authors},
year = {2026},
howpublished = {\url{https://github.com/munod/tachyone}}
}