Instructions to use mstrasser/Jeff-Qwen3.5-0.8B-spam with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use mstrasser/Jeff-Qwen3.5-0.8B-spam with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
Jeff-Qwen3.5-0.8B-spam
Spam and phishing in SMS and email: says whether a text message or email is spam or phishing, and whether it is legitimate, spam or a phishing scam.
Give it a text message or an email (subject and readable body) and it answers two questions: is this spam or phishing (yes or no, as a probability), and is it legitimate, spam (advertising) or phishing (a scam after details or money). It is a LoRA adapter for Jeff-Qwen3.5-0.8B v1.2, a small open decision model: it answers in one forward pass, with no generated text to parse.
Results
Test set: SMS and email hold-out, both questions, 3,603 rows. 10% of the text messages and 10% of the emails, held out by a stable hash of the text (no source set has an official test split); never trained on. Scored on both questions. It is included in this repo as test.jsonl.
| Model | Accuracy | Calibration error (ECE) |
|---|---|---|
| Qwen3.5-0.8B, untrained | 59.4% | 0.060 |
| Jeff v1.2 0.8B alone | 72.3% | 0.109 |
| Jeff v1.2 0.8B + spam | 98.4% | 0.008 |
Calibration error is the expected calibration error over 15 confidence bins, after each model's fitted temperature: the average gap between how confident the model is and how often it is right. Lower is better. The untrained Qwen3.5-0.8B is read out from its own scores for the option letters.
Against Qwen3.8-27B
Here Jeff + spam answers first. Below a confidence threshold, the query is also sent to Qwen3.8-27B and the 27B's answer is used. Each adapter's threshold is the fastest one that still beats the 27B alone by at least 1 point of accuracy on the adapter's calibration rows (not the test rows).
| Route | Accuracy | Mean time per query | Sent to the 27B |
|---|---|---|---|
| Qwen3.8-27B alone | 88.0% | 2.67 s | 100% |
| Jeff + spam, 27B below threshold 0.00 | 98.7% | 73 ms | 0.0% |
The threshold is 0.00: on the calibration rows the adapter alone already beat the 27B by more than a point, so no query is passed on. The route is 36.6× faster than the 27B alone.
How this was measured: a fixed random sample of 300 held-out test rows, the same rows for both routes, on an Apple M4 Max (128 GB) with both models in MLX, one at a time. Qwen3.8-27B ran with 8-bit weights, prompted, with step-by-step reasoning off. Times are the mean per query over the whole sample, from prompt to answer. This is a 300-row sample, not the full test set.
Speed and memory
One request at a time on an NVIDIA RTX PRO 6000: 675 requests per setting, 75 test prompts of each of the nine adapters, interleaved. Measured for the adapters as a group, not for each adapter, so treat these as typical figures:
| Setting | Median | 95th percentile | GPU memory after loading |
|---|---|---|---|
| The v1.2 base alone | 25.9 ms | 35.5 ms | 1.74 GB |
| Shared base, one adapter (default mode) | 31.2 ms | 40.3 ms | 1.79 GB |
| Shared base, all nine adapters, a different one on every request | 30.0 ms | 39.4 ms | 1.96 GB |
Merged mode (JEFF_ADAPTER_MODE=merged), one adapter folded into the weights |
25.7 ms | 35.7 ms | 1.77 GB |
Jeff's speed and memory on the Mac: not measured yet.
Where it goes wrong
On the full test set:
- The weakest label is
spamin the three-way question: advertising spam is sometimes called phishing or legitimate. The yes/no question errs in both directions on a small share of messages. - The three-way question was tested with every order of its three options; the adapter does about equally well in every order (groups below).
Commonest confusions (right answer → chosen instead, rows):
| Right answer | Chosen instead | Rows |
|---|---|---|
| spam or phishing: yes | spam or phishing: no | 18 |
| spam or phishing: no | spam or phishing: yes | 13 |
spam |
phishing |
7 |
phishing |
spam |
6 |
spam |
legitimate |
5 |
legitimate |
phishing |
3 |
Accuracy by right answer:
| Right answer | Test rows | Accuracy |
|---|---|---|
spam |
109 | 89.0% |
phishing |
277 | 97.1% |
| spam or phishing: yes | 662 | 97.3% |
| spam or phishing: no | 1,522 | 99.1% |
legitimate |
1,033 | 99.5% |
By question:
| Question | Test rows | Accuracy |
|---|---|---|
is_spam |
2,184 | 98.6% |
message_type |
1,419 | 98.2% |
By question and option order (accuracy / calibration error):
| Group | Test rows | Qwen3.5-0.8B untrained | Jeff alone | Jeff + adapter |
|---|---|---|---|---|
| choice: legitimate / phishing / spam (email) | 144 | 57.6% / 0.139 | 69.4% / 0.113 | 97.2% / 0.029 |
| choice: legitimate / phishing / spam (sms) | 82 | 81.7% / 0.357 | 80.5% / 0.150 | 95.1% / 0.031 |
| choice: legitimate / spam / phishing (email) | 149 | 29.5% / 0.256 | 64.4% / 0.158 | 100.0% / 0.008 |
| choice: legitimate / spam / phishing (sms) | 108 | 21.3% / 0.210 | 88.9% / 0.181 | 97.2% / 0.023 |
| choice: phishing / legitimate / spam (email) | 134 | 69.4% / 0.149 | 77.6% / 0.104 | 99.3% / 0.008 |
| choice: phishing / legitimate / spam (sms) | 102 | 87.3% / 0.291 | 79.4% / 0.170 | 99.0% / 0.025 |
| choice: phishing / spam / legitimate (email) | 123 | 65.0% / 0.183 | 69.1% / 0.104 | 100.0% / 0.011 |
| choice: phishing / spam / legitimate (sms) | 93 | 79.6% / 0.277 | 84.9% / 0.196 | 98.9% / 0.016 |
| choice: spam / legitimate / phishing (email) | 157 | 70.1% / 0.103 | 65.0% / 0.120 | 97.5% / 0.018 |
| choice: spam / legitimate / phishing (sms) | 110 | 83.6% / 0.226 | 75.5% / 0.141 | 98.2% / 0.043 |
| choice: spam / phishing / legitimate (email) | 139 | 71.9% / 0.100 | 69.8% / 0.068 | 98.6% / 0.018 |
| choice: spam / phishing / legitimate (sms) | 78 | 84.6% / 0.240 | 78.2% / 0.157 | 96.2% / 0.026 |
| yes/no question (email) | 1,592 | 65.1% / 0.091 | 66.6% / 0.234 | 98.5% / 0.008 |
| yes/no question (sms) | 592 | 30.9% / 0.271 | 83.6% / 0.071 | 98.8% / 0.010 |
When to use it
- You receive text messages or emails and want a spam or phishing probability for each.
- You want to tell phishing (a scam after details, passwords or money) apart from ordinary advertising spam.
- You set your own threshold on the yes probability, rather than taking a hard yes or no.
When not to use it
- You need an email's headers, links or attachments judged. Only the subject and readable body were used, and bodies over 6,000 characters were cut.
- Your messages are not in English. The data sets are English.
- You need to catch the newest scams. Much of the data is old (SMS from 2011 and 2022, most email from 2002 to 2005, phishing email up to 2025), so recent tricks may be missed.
- You block messages with no human check. Treat the answer as a signal, not a verdict.
How to use it
The adapter runs with Jeff's server on the Jeff-Qwen3.5-0.8B v1.2 base. One server holds the base and any number of adapters; each request picks one by name with "model": "spam".
git clone https://github.com/firelex/jeff && cd jeff
uv sync --no-default-groups --extra lora # add --extra cuda on NVIDIA GPUs, --extra mac on Apple silicon
uv run --no-default-groups hf download mstrasser/Jeff-Qwen3.5-0.8B --revision v1.2 --local-dir Jeff-Qwen3.5-0.8B-v1.2
uv run --no-default-groups hf download mstrasser/Jeff-Qwen3.5-0.8B-spam --local-dir adapters/spam
JEFF_CHECKPOINT=Jeff-Qwen3.5-0.8B-v1.2 JEFF_ADAPTERS=adapters PORT=8765 \
uv run --no-default-groups --extra lora jeff-serve
# add JEFF_BACKEND=mlx and --extra mac for Apple silicon
Every folder in adapters/ is served under its folder name. Add or replace an adapter while the server runs with POST /v1/adapters/reload. The server checks the base weights' checksum, so the adapter loads only on the exact v1.2 base.
from jeff import Client
spam = Client("http://localhost:8765", model="jeff-latest").with_model("spam")
Request format
State (the situation):
| Key | Changes per request | What it holds |
|---|---|---|
channel |
no | Where the message came from: "sms" or "email", as in training. |
message |
yes | The text of the message as received. For an email, its subject and readable body (no other headers, no attachments). |
Questions:
is_spam(noul): Whether the message is spam (unwanted bulk or advertising) or phishing (a scam after personal details or money), rather than a normal message.message_type(choice): What kind of message it is. Options: Three options: legitimate, spam and phishing, each with a one-line description (SPAM_CHOICE in descriptions.py in the source).
Rules:
- Ask both questions in one request if you need both; they are answered together.
- For is_spam, give the false and true descriptions below, in that order; the adapter was trained with them.
- The yes/no question was trained on every message. The three-way question was trained only where the source separates phishing from spam, the SMS phishing set and the SpamAssassin and Nazario emails.
- Use the instructions below word for word; the adapter was trained mostly on them.
Example
{
"model": "spam",
"state": {
"channel": "sms",
"message": "Your parcel could not be delivered. Pay the 1.99 redelivery fee within 24 hours at parcel-redeliver-help.example to avoid return."
},
"questions": {
"is_spam": {
"type": "noul",
"instructions": "Is this message spam or phishing? Answer yes if it is unwanted bulk or advertising, or a scam trying to get personal details or money; answer no if it is a normal message.",
"criteria": {
"false": "The message is a normal message: not unwanted bulk or advertising, and not a scam.",
"true": "The message is spam (unwanted bulk or advertising) or phishing (a scam trying to get personal details, passwords or money)."
}
},
"message_type": {
"type": "choice",
"instructions": "What kind of message is this: a legitimate message, spam (unwanted bulk or advertising), or phishing (a scam trying to get personal details, passwords or money)?",
"criteria": {
"legitimate": "A normal message from a person or a genuine organisation, not spam and not phishing.",
"spam": "Unwanted bulk or advertising message (for example prizes, offers, premium-rate services), but not trying to steal personal or account details.",
"phishing": "A scam message that tries to trick the reader into giving personal details, passwords or money, often by pretending to be a bank, company or authority and asking them to click a link or call a number."
}
}
}
}
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d @request.json
The answer holds a probability for each option of each question, the chosen option and a confidence. (A recorded response to this example is not published yet.)
Training
| Base | Jeff-Qwen3.5-0.8B v1.2 (final checkpoint of run 0.8b-20260929-2258) |
| Method | LoRA, rank 16 |
| Rows | 30,704 (the adapter's rows plus a replay slice of the v1.2 base training mix, 10% of the adapter rows) |
| Epochs | 1 |
| Steps | 480 |
| Time | 32.3 minutes on one NVIDIA RTX PRO 6000 |
| Adapter files | 41.5 MB |
| Run | 0.8b-spam-20260930-1343, step 480 |
The replay slice mixes in some of the base model's own training data so the adapter keeps the base's general skills. The scored checkpoint is the one at the end of the epoch.
Data and provenance
Sources:
| Source | Licence | Notes |
|---|---|---|
| UCI SMS Spam Collection (Almeida and Hidalgo 2011) | CC-BY-4.0 | Labels ham and spam. |
| SMS Phishing Dataset for Machine Learning and Pattern Recognition (Mishra and Soni 2022), version 1 | CC-BY-4.0 | Labels ham, spam and smishing (SMS phishing). Messages shared with the UCI set were merged by text. |
| Phishing Email Dataset (zefang-liu, a copy of the Kaggle set "Phishing Email Detection") | Unclear, under review before release | Tagged LGPL-3.0 by the uploader; the Kaggle original does not say where its emails come from. Labels safe and phishing. |
| ealvaradob/phishing-dataset (emails only) | Unclear, under review before release | Tagged Apache-2.0 by the compiler; its emails are the same Kaggle set as above, so it adds only a handful of messages. |
| SpamAssassin public mail corpus (ham, hard ham, spam) | Unclear, under review before release | Real mail received 2002 to 2005, sorted by hand into spam and non-spam. No licence is given; the corpus says copyright in the messages stays with their senders. |
| Nazario phishing corpus | Unclear, under review before release (the corpus states CC-BY-4.0) | Phishing mail collected and sorted by hand by Jose Nazario. |
Training data: built from the public sources above. A script to rebuild our rows from them will follow. The test set is included in this repo as test.jsonl.
Test set: test.jsonl, 3,603 rows: 10% of the text messages and 10% of the emails, held out by a stable hash of the text (no source set has an official test split); never trained on. Scored on both questions.
Calibration set: calibration.jsonl, 1,985 rows: the rows the threshold for the comparison with Qwen3.8-27B is chosen on; never trained on.
It is byte for byte the set the adapter was scored on.
Quality review. An independent reviewer checked the training, calibration and test files for shortcuts, duplicates, leaks between splits and junk before training. The report is at jeffhub.ai/adapters/spam/qa-report.
Limitations
- English only.
- Tied to the Jeff-Qwen3.5-0.8B v1.2 base. It will not work on any other base or version; the server refuses it if the base weights' checksum does not match. A v1.3 long-term-support base is coming, and the adapters will be retrained on it.
- Jeff chooses between the options you give it. It does not write text or reason in several steps.
- Calibration was fitted on this adapter's own calibration rows. On very different data, check it again.
- The licences of the email data sets (the phishing email sets, SpamAssassin and the Nazario corpus) are under review before release.
- Everything listed under When not to use it above.
Links
- Adapter page: jeffhub.ai/adapters/spam
- Code and server: github.com/firelex/jeff
- Base model: mstrasser/Jeff-Qwen3.5-0.8B (revision v1.2)
Jeff is an independent project. It uses the same request format as Jev but is not affiliated with or endorsed by TypeSafe, the makers of Jev.
- Downloads last month
- 41