Instructions to use knpatil/laya-browser-agent-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use knpatil/laya-browser-agent-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="knpatil/laya-browser-agent-base")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("knpatil/laya-browser-agent-base", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Laya Browser Agent Base (149M) β ModernBERT Knowledge Distillation
knpatil/laya-browser-agent-base is a high-speed, sub-30ms browser decision agent distilled from knpatil/laya-browser-agent (ModernBERT-large, 421M).
Designed specifically for real-time in-browser agent decision loops (such as the BroPilot Chrome Extension), this model predicts immediate next-step browser actions (CLICK, TYPE_TEXT, SELECT, SCROLL_DOWN, WAIT, DONE, BLOCKED), element targets, and goal completion criteria in 29.89 ms on Apple Silicon Metal (MPS).
π Key Performance Highlights
- Ultra-Low Latency: 29.89 ms median forward pass (p95: 33.37 ms) on Apple Silicon M2 Max (MPS) β 28x faster than commercial cloud APIs (TypeSafe Jev: 841.8 ms).
- High Decision Accuracy: 96.94% decision accuracy across 70 standard web automation workflows (vs TypeSafe Jev: 86.89%, Teacher: 82.14%).
- Compact Footprint: 164.0M total parameters (149M backbone) with a 596 MB disk size (-61.1% parameter reduction from 421M).
- Exceptional Calibration: Brier score of 0.0025 (Platt-scaled temperature:
choice: 1.0381, score: 1.0000, noul: 1.0650). - Zero Cloud API Charges: 100% on-device private execution with zero browser session exfiltration.
π Comprehensive Benchmark Results
Evaluated across the 70 benchmark scenarios (244 structured decisions) in data/browser_test_cases.jsonl:
| Metric | Base Laya (Zero-Shot) | TypeSafe Jev (jev-1.13.0 Cloud) |
Teacher Model (421M Large) | Distilled Student (149M Base) |
|---|---|---|---|---|
| Decision Accuracy | 64.60% | 86.89% | 82.14% | 96.94% |
| Workflow Pass Rate | 31.43% | 70.00% | 34.29% | 82.86% |
| Brier Score (Lower = better) | 0.1190 | 0.1608 | 0.0905 | 0.0025 |
| Median Inference Latency | 349.0 ms | 841.8 ms | 58.85 ms | 29.89 ms |
| P95 Latency | 392.0 ms | 1,240.0 ms | 62.54 ms | 33.37 ms |
| Parameter Count | 421.3M | Proprietary Cloud | 421.3M | 164.0M (-61.1%) |
| Cost per 1,000 Decisions | $0.00 | ~ $0.40 | $0.00 | $0.00 (Local) |
Accuracy by Sub-Decision Primitive
operation(7-way Action Choice): 100.0% (70/70)action_type(Navigation Intent): 100.0% (70/70)is_goal_satisfied(Binary Goal Check): 100.0% (70/70)click_target(Element Selection): 85.7% (60/70)type_text_target(Input Field Selection): 95.2% (40/42)
π§ Distillation Architecture & Training
The student model was trained using knowledge distillation from knpatil/laya-browser-agent:
- Teacher: Frozen
ModernBERT-large(421M params, 28 layers, d=1024, 16 attention heads). - Student:
ModernBERT-base(149M params, 22 layers, d=768, 12 attention heads). - Loss Function: Multi-task joint loss:
L = Ξ±_KD Β· ΟΒ² Β· L_KD(Ο(z_S/Ο), Ο(z_T/Ο)) + Ξ±_CE Β· L_CE(z_S, y)with distillation temperatureΟ = 2.0,Ξ±_KD = 0.6,Ξ±_CE = 0.4. - Optimizer: AdamW (
lr=2.5e-5) with Cosine Annealing learning rate schedule. - Hardware: Trained natively on Apple Silicon Metal (MPS).
π» Quickstart & Inference
Using the Python Client
import laya
# Initialize the distilled 149M model on Apple Silicon MPS or CUDA
agent = laya.Agent("knpatil/laya-browser-agent-base", device="mps")
state = {
"page": {
"url": "https://huggingface.co/models",
"title": "Hugging Face Models",
"text": "Explore over 1M open-source AI models and datasets."
},
"elements": [
{"index": "1", "role": "textbox", "label": "Search models, datasets, users..."},
{"index": "2", "role": "link", "label": "Tasks"},
{"index": "3", "role": "link", "label": "Libraries"}
]
}
questions = {
"action_type": {
"type": "choice",
"instructions": "Given the goal 'Search for ModernBERT models', what immediate browser action should be taken?",
"criteria": {
"click": "Click a visible link, button, or tab",
"type": "Enter search text into an input field",
"scroll": "Scroll down to reveal more content",
"wait": "Wait for dynamic content to load"
}
}
}
prediction = agent.predict(state, questions)
print("Action Decision:", prediction["answers"]["action_type"]["choice"])
print("Confidence:", prediction["answers"]["action_type"]["confidence"])
π Citation & Credits
Developed as part of the BroPilot autonomous browser companion project.
- Backbone: ModernBERT by Answer.AI & LightOn.
- Teacher Checkpoint:
knpatil/laya-browser-agent.
Model tree for knpatil/laya-browser-agent-base
Base model
answerdotai/ModernBERT-base