COMPLEXITY TR-HASH MoE 200M companion
PIQA-selected full-SFT checkpoint

TR-HASH MoE 200M

Explore how two deterministic, layer-specific token-ID hashes select a top-2 pair of residual experts while a shared dense SwiGLU path preserves contextual computation.

201.2M parameters ≈162B source-token exposure 16 layers · multi-hash top-2
Interactive companion — not additional experimental evidence.
Token
Layer-specific deterministic routes
Checkpoint-derived tokenizer IDs and deterministic two-hash top-2 assignments.
E0E1E2E3 Layer 10.50.5 Layer 40.50.5 Layer 80.50.5 Layer 120.50.5 Layer 160.50.5
Token “routing” follows the persisted route table of the final checkpoint. Every occurrence keeps the same expert pair at a given layer, while its hidden-state input still changes with context.
Recorded training diagnostics
Source and scheduled SFT evaluations · NLL, lower is better
Source SFT measurement
MeasurementSourceSFT checkpoint
Held-out SFT NLL1.7628 @ source1.2209 @ epoch 3
Held-out SFT PPL5.833.39
PIQA accuracy68.66% refinement68.82% epoch 2
PIQA acc_norm68.39% refinement69.31% epoch 2
SFT coverage209k records3 epochs · 238.9M tokens

SFT loss uses the fixed 2,100-example held-out split. PIQA uses all 1,838 validation examples, zero-shot causal continuation likelihood, no chat template, and FP16 eager evaluation.

TR-HASH MoE 200M · Full SFT ready

The model response will stream here token by token.

Decode configuration

Chat endpoint · temperature 0.4 · top-k 30 · top-p 0.85 · repetition penalty 1.1 · maximum 384 new tokens.

This is a qualitative demonstration of the released SFT checkpoint, not an evaluation result.

TR-HASH MoE 200M · live qualitative generation