Jev laya-th1200
Jev laya-th1200 is a Thai typed decision model for agent and tool governance.
Given an operational request or long agent task context, it returns four typed signals:
action:execute,ask_user, orrejectneeds_review: boolean-like (noul) indicating whether human review is requiredprohibited: boolean-like (noul) indicating whether the operation is strictly forbiddenrisk: ordinal operational-risk score from 0 to 4
This is a research and engineering checkpoint. It serves as an advisory signal and must not replace explicit authorization, repository policy, human approval boundaries, automated tests, or audit logging.
Model Details
- Base model:
convaiinnovations/laya-multilingual - Foundation: Laya 0.3.4
- Language: Thai (
th) and bilingual English/Thai - Decision primitives:
choice(action),noul(needs_review, prohibited), and ordinalscore(risk 0–4) - Training dataset: 1,200 curated Thai governance examples (
data/th_curated_1200/train.jsonl)- Retains the 1,080 curated cases from previous iterations
- Adds 120 agent-context cases (
th_17_agent_task_context_train.csv) covering realistic multi-clause coding agent prompts
- Evaluation splits: 100 cases each in Validation, Calibration, and Test (strictly balanced: 20 cases per risk level 0–4)
- Calibration: independent per-primitive temperatures and score threshold decoder fitted strictly on the calibration split
- Repository Checkpoint:
JonusNattapong/jev-my-bro-th1200
Evaluation Results
Evaluated on the locked test set (100 cases, 400 typed decisions, 20 cases per risk level):
| Metric | Result |
|---|---|
| Overall accuracy | 86.75% |
| Action choice accuracy | 88.00% |
Noul accuracy (needs_review / prohibited) |
93.50% |
| Level 4 risk recall (Catastrophic/Destructive) | 100.00% |
Key Improvements over th960
- Agent Context Robustness: Addresses false-positive
prohibitedspikes on routine multi-clause coding tasks containing phrases like "ผู้ใช้สั่งให้ทำแล้ว" or "ไม่มีการ push". - 100% L4 Recall: Flawlessly catches destructive actions (force push to main, deleting shared branches, exfiltrating secrets, dropping production tables).
Runtime Deployment Architecture
In production and local agent environments (Claude Code, Antigravity, Codex), th1200 operates as part of a Hybrid Cascaded Architecture:
- Layer 1: Deterministic Fast-Path Rules (<1ms)
- Regex and keyword matcher in
jevbro/rules.py - Configurable via
rules.yaml/.jev/rules.yaml - Handles ~48% of standard traffic (pure inspections allowed, catastrophic shell commands rejected immediately)
- Regex and keyword matcher in
- Layer 2: LRU In-Memory Decision Cache (~15ms)
- 1024-entry LRU cache in
jevbro/core.pywith whitespace normalization - Yields 140x speedups on repeated or similar tool executions
- 1024-entry LRU cache in
- Layer 3: Neural Model Inference with INT8 Dynamic Quantization (~200–300ms on CPU)
- PyTorch dynamic INT8 quantization applied to
agent.model.encoder(--quantize) - Reduces CPU latency by ~3–5x compared to standard FP32 execution
- PyTorch dynamic INT8 quantization applied to
Usage Example
import json
import torch
from laya import Agent
from jevbro.questions import default_questions
MODEL_ID = "JonusNattapong/jev-my-bro-th1200"
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
agent = Agent(MODEL_ID, device=DEVICE)
result = agent.predict(
{"request": "รัน unit test ด้วย pytest บนเครื่อง local", "domain": "software"},
default_questions("th"),
)
print(json.dumps(result["answers"], ensure_ascii=False, indent=2))
Limitations & Policy Gating
- Development and research data, not an exhaustive security policy.
- Always apply the outer policy gate:
- If
prohibited >= 0.5➔ reject - If
needs_review >= 0.5oraction == "ask_user"orrisk >= 3➔ prompt user for explicit review - Else ➔ execute
- If
Model tree for JonusNattapong/jev-my-bro-th1200
Base model
convaiinnovations/laya-multilingualEvaluation results
- Overall accuracy on Jev Thai curated 1200test set self-reported0.868
- Action choice accuracy on Jev Thai curated 1200test set self-reported0.880
- Noul accuracy on Jev Thai curated 1200test set self-reported0.935
- Level 4 risk recall on Jev Thai curated 1200test set self-reported1.000