Laya Fine-tuned for 3-Class Financial Sentiment
convaiinnovations/laya (421M, ModernBERT-large + decision head) fine-tuned with RLCD (Reinforcement Learning for Calibrated Decisions) for a single task: pick one of positive, negative, neutral from a financial text.
Evaluation
Same 15,355-item stratified test split of shanaka95/fingpt-sentiment-3class, same choice question schema, same laya.Agent.predict() call. Only MODEL_ID differs.
Zero-shot (convaiinnovations/laya) |
Fine-tuned (this model) | Ξ | |
|---|---|---|---|
| Overall accuracy | 0.7024 (10,786 / 15,355) | 0.9486 (14,566 / 15,355) | +0.2462 (+24.62 pp) |
| F1 β positive | 0.694 | 0.952 | +0.258 |
| F1 β neutral | 0.696 | 0.949 | +0.253 |
| F1 β negative | 0.726 | 0.942 | +0.216 |
| Recall β positive | 0.584 | 0.954 | +0.370 |
| Latency p50 | 38.6 ms | 36.6 ms | -2 ms |
The fine-tuning script never saw any of these test examples. Eval notebooks + the full reproduction repo: github.com/shanaka95/laya-fintiment.
Training
- Base model:
convaiinnovations/laya - Dataset:
shanaka95/fingpt-sentiment-3classβ train split (61,017 items after 400 held out for post-training calibration) - Objective: RLCD with strictly proper scoring rules (log + spherical,
w_sph=0.75; no RPS because all questions arechoice). REINFORCE with group-mean baseline,group_size=4noisy forward passes per example. - Epochs: 4 (7,624 optimizer steps, effective batch 32)
- Hardware: 1 Γ 12 GB GPU
- Optimisation: AdamW, cosine LR schedule, gradient checkpointing ON,
micro_batch=16,grad_accum=2 - Question schema used for every example:
{ "sentiment": { "type": "choice", "instructions": ( "What is the financial sentiment of this text? " "Please choose exactly one answer from {positive, negative, neutral}. " "Answer only with the chosen label." ), "criteria": { "positive": "The text expresses a positive / bullish financial sentiment.", "negative": "The text expresses a negative / bearish financial sentiment.", "neutral": "The text is neutral, factual, or has no clear positive or negative sentiment.", }, } } - Final calibration temperatures (choice / score / noul):
[5.013, 1.2, 1.2] - Epoch avg loss: 0.257 β 0.493 β 0.323 β 0.199
Recipe adapted from the upstream fine-tuning notebook.
Usage
pip install "laya>=0.1.6" "transformers>=4.48.0"
import os
os.environ["USE_TF"] = "0" # avoids an abseil/TF deadlock when loading transformers
import laya
agent = laya.Agent("shanaka95/laya-fintiment", device="cuda")
questions = {
"sentiment": {
"type": "choice",
"instructions": (
"What is the financial sentiment of this text? "
"Please choose exactly one answer from {positive, negative, neutral}. "
"Answer only with the chosen label."
),
"criteria": {
"positive": "The text expresses a positive / bullish financial sentiment.",
"negative": "The text expresses a negative / bearish financial sentiment.",
"neutral": "The text is neutral, factual, or has no clear positive or negative sentiment.",
},
}
}
result = agent.predict("Apple reported record quarterly revenue, beating analyst estimates.", questions)
print(result["answers"]["sentiment"]["choice"]) # -> 'positive' (probability mass concentrated on it)
For batched inference or CPU use, pass device="cpu" or pre-build a laya.Router.
Worked example
| Input | Base laya | This model |
|---|---|---|
"Asset Management One Co. has a bullish call on Treasuries. One for the long, long run https://t.co/8go8ZhMvdc" |
positive @ 0.5604 (also 0.3705 on neutral) |
positive @ 0.9488 (decisive) |
Files
| File | Notes |
|---|---|
model.safetensors |
842 MB fine-tuned weights |
encoder/ |
encoder sub-dir (downstream code may read it directly) |
tokenizer/ |
tokenizer sub-dir |
rl_agent_config.json |
decision head + calibration config |
Honest limits
- 3-class only. Anything beyond
positive / negative / neutral(intensity, topic, aspect) needs a different fine-tune or ascore/noulquestion β both are still in the underlying checkpoint but were not trained here. - Out-of-distribution drift. Trained on FinGPT-sentiment-train distribution (news/tweets). Earnings calls, analyst reports, regulatory filings will be OOD; expect calibration drift. Refit temperatures per your domain before trusting the probabilities (the upstream fine-tuning notebook shows how).
- No soft distribution comparison done. Argmax accuracy is reported; soft ECE / Brier against gold distributions is not measured yet.
Citation
Base model and training library: Convai Innovations β Laya.
Dataset: derived from FinGPT/fingpt-sentiment-train, collapsed to 3 classes and uploaded as shanaka95/fingpt-sentiment-3class.
Apache 2.0.
Model tree for shanaka95/laya-fintiment
Base model
convaiinnovations/laya