Marquand models

These are the two small models behind Marquand, a local drop-in for TypeSafe's Jev (System One) API. They don't write text - you give them some state and a question with fixed options, and they answer by picking an option letter, so every answer comes out as a probability over the options you gave it.

This is an independent hobby project and isn't affiliated with TypeSafe AI. Neither of these is Jev, they're Qwen3.5 models fine-tuned to act like it.

Files

File What it is Size
marq-fast-2b-Q8_0.gguf Qwen3.5-2B, fine-tuned (Marquand's fast model) 2.1 GB
marq-instant-0.8b-Q8_0.gguf Qwen3.5-0.8B, fine-tuned (Marquand's instant model) 0.8 GB
mmproj-Qwen3.5-2B-F16.gguf vision projector for the 2B, converted from Qwen3.5-2B unchanged 0.7 GB

Only the language model weights were trained, so the stock vision projectors still work. The 0.8B's projector is mmproj-Qwen3.5-0.8B-BF16.gguf from lmstudio-community/Qwen3.5-0.8B-GGUF.

Using them

The easy way is through Marquand, which handles the prompt, the letter readout and calibration for you:

hf download a-m-0099/marquand --local-dir models
marq serve --model fast

If you want to use them outside Marquand, they expect this exact prompt (Qwen3.5 chat template, thinking off), and the answer is read from the logits of the option letters at the first answer position, not generated:

<|im_start|>user
State:
{state as text or JSON}

Question: {instructions}
Options:
[A] billing: charges, refunds, invoices
[B] shipping: delivery and tracking

Answer with the letter of the best option only.<|im_end|>
<|im_start|>assistant
<think>

</think>

Softmax over the logits for A, B, ... gives the probabilities. Yes/no questions use two options, [A] true: Yes, the statement is true. and [B] false: No, the statement is false. (or your own wording for each). Scores list the levels in order, lowest first, and add " Rate along the ordered levels below (lowest first)." to the end of the question.

Results

JevBench public (231 items), where Jev 1.13 gets 200. Latency is on an RX 7700S (8 GB) through llama.cpp's Vulkan backend, for a call with a ~200 token state and 5 questions.

Model JevBench Easy /48 Standard /72 Hard /111 5-question call
marq-fast-2b 163 48 63 52 130 ms
Qwen3.5-2B before training 148 48 49 51 132 ms
marq-instant-0.8b 147 48 55 44 69 ms
Qwen3.5-0.8B before training 127 47 38 42 141 ms

On 300 held-out rows labelled by Jev itself, agreement with Jev's top answer went from 47% to 80% for the 2B and from 36% to 80.7% for the 0.8B.

Training

LoRA (rank 32, alpha 32) on every attention, linear-attention and MLP projection, merged back into the base weights and exported to GGUF at Q8_0. It's one pass over 16,000 examples for the 2B and 24,000 for the 0.8B, at learning rate 2e-4 with 16-step gradient accumulation and bf16, trained on a single RX 7700S through PyTorch ROCm.

The loss is soft cross-entropy between the target distribution and the model's probabilities over the option letters, using the same prompt as above. The option order gets shuffled on every example so no letter learns an answer. Rows longer than 1024 tokens were skipped.

The training mix was:

  • 70% yuri_v3 rows from jev-distill-corpus-v3 - synthetic scenarios whose labels are Jev 1.13's own probabilities
  • 15% openjev_v2 rows from the same corpus
  • 15% hard reasoning rows from Open-Jev-v1.1 with gold labels

JevBench was never used for training.

Limitations and notes

  • These are small models, so they're a lot weaker than Jev on anything that needs real reasoning (the hard tier above). They're meant for fast, simple decisions like routing, NPC actions and yes/no checks.
  • Only tested with llama.cpp on AMD + Linux.
  • The yuri_v3 labels in jev-distill-corpus-v3 were made by calling TypeSafe's Jev API. Check TypeSafe's terms before using these models for anything commercial. Open-Jev-v1.1's data license is in its repo.
  • The weights are released under Apache-2.0, the same as the Qwen3.5 base models.
Downloads last month
108
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for a-m-0099/marquand

Adapter
(275)
this model

Datasets used to train a-m-0099/marquand