Laya Model Routing β€” 500 rows @ 12 epochs (same-epoch showdown)

Laya (421M) fine-tuned to classify a user query into one of six model-routing classes, so an application can pick the right class of free OpenRouter models (single forward pass, ~37 ms on T4):

class meaning
fast_cheap greetings, basic lookups, formatting
general_mid everyday reasoning, most chat/Q&A
top_reasoning hard multi-step reasoning, math, planning
code_specialist writing, debugging, reviewing code
security_specialist vulnerabilities, secure coding, privacy, auth
long_context long documents / many turns

Results (held-out set: 100 disjoint queries, judge-labeled)

Full series on the same eval set (e12 rows measured on Kaggle T4/FP16):

train rows epochs model choice_accuracy ECE ↓
0 β€” base Laya (zero-shot) 0.630 0.169
200 12 laya-model-routing 0.820 0.125
500 8 laya-model-routing-500 0.910 0.074
1000 6 laya-model-routing-1000 0.870 0.087
500 12 this checkpoint 0.910 0.062
1000 12 laya-model-routing-1000e12 0.900 0.028

Majority-class baseline: 0.270 Β· random: 0.167.

This rerun tested whether round 1's 500 > 1000 upset was under-training (auto epochs gave 8 vs 6). Result: 500 was already converged β€” accuracy identical at 0.910, ECE improved 0.074 β†’ 0.062. The sibling laya-model-routing-1000e12 recovered to 0.900 with the series' best calibration and is the recommended production router.

Per-class recall: security_specialist 1.000 Β· general_mid 0.960 Β· fast_cheap 0.926 Β· code_specialist 0.909 Β· long_context 0.667 Β· top_reasoning 0.000 (3 training rows β€” never predicted).

Training

  • Data: 500 queries from ai-mitra/llm-router-dataset (shuffled seed 42, rows 300–799), labeled by glm-5.3 as judge via Z.AI
  • Recipe: RLCD (policy gradient + proper scoring rule + cross-entropy), DDP on 2Γ— T4, 12 epochs (EPOCHS_OVERRIDE=12), temperature calibration on a holdout slice
  • Full pipeline and experiment log: github.com/lunakicks/laya-model-routing

Usage

import laya

agent = laya.load("thechristyjo/laya-model-routing-500e12", device="cpu")
questions = {
    "model_class": {
        "type": "choice",
        "instructions": "Which model class should handle this task?",
        "criteria": {
            "fast_cheap": "simple, short, low-stakes queries",
            "general_mid": "everyday reasoning, most chat/Q&A",
            "top_reasoning": "hard multi-step reasoning, math, planning",
            "code_specialist": "writing, debugging, reviewing code",
            "security_specialist": "security, privacy, auth tasks",
            "long_context": "long documents or many turns",
        },
    }
}
print(agent.predict("Write a function to detect a linked-list cycle.", questions)
      ["answers"]["model_class"]["choice"])

Limitations

  • Outperformed on calibration by the 1000-row sibling at the same epochs β€” prefer laya-model-routing-1000e12 for production
  • Fitted choice temperature 5.06, marginally above the valid [0.5, 5] range β€” laya clamps to 5 at load (warning); re-calibrate before tightening any confidence gate above 0.5
  • top_reasoning has 3 training rows β€” never predicted; needs targeted augmentation
  • long_context recall 0.667 β€” trained on short one-liners about summarization, not actual long inputs
  • English-only; judge labels carry the judge's "cheapest competent class" policy
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support