Laya Model Routing — 1000-row scaling run

Laya (421M) fine-tuned to classify a user query into one of six model-routing classes, so an application can pick the right class of free OpenRouter models (single forward pass):

class meaning
fast_cheap greetings, basic lookups, formatting
general_mid everyday reasoning, most chat/Q&A
top_reasoning hard multi-step reasoning, math, planning
code_specialist writing, debugging, reviewing code
security_specialist vulnerabilities, secure coding, privacy, auth
long_context long documents / many turns

Results (held-out set: 100 disjoint queries, judge-labeled)

Scaling curve on the same eval set:

train rows model choice_accuracy ECE ↓
0 base Laya (zero-shot) 0.630 0.169
200 laya-model-routing 0.820 0.125
500 laya-model-routing-500 0.910 0.074
1000 this checkpoint 0.870 0.087

Majority-class baseline: 0.270 · random: 0.167.

Per-class recall: security_specialist 1.000 · fast_cheap 0.926 · general_mid 0.880 · code_specialist 0.818 · long_context 0.667 · top_reasoning 0.000 (6 training rows — never predicted).

Note: the 500-row sibling outperforms this checkpoint (0.910 vs 0.870) — scaling was non-monotonic here. This run trained fewer epochs (6 vs 8) and lost code_specialist sharpness (0.909 → 0.818, 4 rows leak to security_specialist). Kept for the scaling record; prefer the 500-row model for production.

Training

  • Data: 1000 queries from ai-mitra/llm-router-dataset (shuffled seed 42, rows 300–1299), labeled by glm-5.3 as judge via Z.AI
  • Recipe: RLCD (policy gradient + proper scoring rule + cross-entropy), DDP on 2× T4, 6 epochs, temperature calibration on a holdout slice
  • Full pipeline and experiment log: github.com/lunakicks/laya-model-routing

Usage

import laya

agent = laya.load("thechristyjo/laya-model-routing-1000", device="cpu")
questions = {
    "model_class": {
        "type": "choice",
        "instructions": "Which model class should handle this task?",
        "criteria": {
            "fast_cheap": "simple, short, low-stakes queries",
            "general_mid": "everyday reasoning, most chat/Q&A",
            "top_reasoning": "hard multi-step reasoning, math, planning",
            "code_specialist": "writing, debugging, reviewing code",
            "security_specialist": "security, privacy, auth tasks",
            "long_context": "long documents or many turns",
        },
    }
}
print(agent.predict("Write a function to detect a linked-list cycle.", questions)
      ["answers"]["model_class"]["choice"])

Limitations

  • Outperformed by the 500-row sibling on the shared eval set — see scaling note above
  • top_reasoning has 6 training rows — never predicted; needs targeted augmentation
  • long_context recall 0.667 — trained on short one-liners about summarization, not actual long inputs
  • English-only; judge labels carry the judge's "cheapest competent class" policy
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support