Explore how two deterministic, layer-specific token-ID hashes select a top-2 pair of residual experts while a shared dense SwiGLU path preserves contextual computation.
| Measurement | Source | SFT checkpoint |
|---|---|---|
| Held-out SFT NLL | 1.7628 @ source | 1.2209 @ epoch 3 |
| Held-out SFT PPL | 5.83 | 3.39 |
| PIQA accuracy | 68.66% refinement | 68.82% epoch 2 |
| PIQA acc_norm | 68.39% refinement | 69.31% epoch 2 |
| SFT coverage | 209k records | 3 epochs · 238.9M tokens |
SFT loss uses the fixed 2,100-example held-out split. PIQA uses all 1,838 validation examples, zero-shot causal continuation likelihood, no chat template, and FP16 eager evaluation.
The model response will stream here token by token.
Chat endpoint · temperature 0.4 · top-k 30 · top-p 0.85 · repetition penalty 1.1 · maximum 384 new tokens.
This is a qualitative demonstration of the released SFT checkpoint, not an evaluation result.