MiniCrit-7B: Adversarial AI Critique Model

Research checkpoint, and a held-out evaluation has now been run with an unfavourable result. 34,000 of 364,831 planned optimizer steps on CritiqueBank-11M — 9.3% of one epoch. Balanced accuracy 0.375 on 32 held-out items, where a constant answer scores 0.500. It objects to sound reasoning more often than to flawed reasoning. See Evaluation. Not a production system, and not a reasoning-flaw detector.

MiniCrit-7B is a research checkpoint of an adversarial critic model developed by Antagon Inc. It is trained to write critiques of trading rationales, pointing out likely flaws such as:

  • overconfident predictions
  • overfitting to historical patterns
  • spurious correlations
  • survivorship bias
  • confirmation bias
  • missing risk factors

Model details

Attribute Value
Developer Antagon Inc. (CAGE 17E75, UEI KBSGT7CZ4AH3)
Base model Qwen/Qwen2-7B-Instruct
Method LoRA (low-rank adaptation)
Trainable parameters 40.4M (0.53% of 7.6B total)
Training data CritiqueBank-11M (11,674,598 examples)
Training progress 34,000 of 364,831 steps (epoch 0.0932); ~1.09M examples seen
Training hardware NVIDIA H100 PCIe (80GB) via Lambda Labs GPU Grant
Held-out evaluation Run 5 October 2026 — balanced accuracy 0.375 (see Evaluation)
License Apache 2.0

Evaluation

A held-out evaluation has been run. This checkpoint does not beat a constant answer on it, and no capability claim is made for it.

Three runs on 5 October 2026, all through the published container on one A100 40GB, greedy decoding (do_sample: false, max_new_tokens: 256, repetition_penalty: 1.15), zero errors. The scoring protocol was committed to the project repository before the first item was sent and was not edited afterwards.

Run Model Items Detection False positive Balanced accuracy
A1 wmaousley/MiniCrit-7B 32 held-out 4/16 = 0.250 8/16 = 0.500 0.375
A2 Qwen/Qwen2-7B-Instruct, no adapter 32 held-out 14/16 = 0.875 14/16 = 0.875 0.500
B wmaousley/MiniCrit-7B 24 older items 2/12 = 0.167 4/12 = 0.333 0.417

Always-flag scores 0.500 on these items. Never-flag scores 0.500. Chance is 0.500. A detector that was merely uninformative would land at 0.500; 0.375 means the signal is anti-correlated with what it is meant to detect.

The error is directional, not random

The 32 items are matched pairs: each flawed item has a sound twin describing the same scenario and reasoning about it correctly. A pair counts as resolved only if the flawed half is flagged and the sound half passes.

1 of 16 pairs resolved. Random answering resolves 4 in expectation. The eight sound items it flagged are the ones that name a confound, propose a control, refuse to generalise from a small sample, or treat absence of evidence correctly — items that are right precisely because they decline to claim too much.

Removing the adapter does not fix it

Run A2 is the same container with no LoRA. It reaches 0.500 only by flagging 28 of 32 items, with detection and false-positive rates identical to three decimal places — the signature of an answer that does not depend on the input. A paired McNemar exact test does not separate the two models (p = 0.5235, 22 discordant pairs), so no difference between the fine-tune and its base is claimed. Fine-tuning did not turn a non-detector into a detector; it changed which way a non-detector leans.

One caveat, stated because it limits that comparison: 26 of the base model's 32 critiques were cut off mid-sentence by the 256-token ceiling, which the adapter's much shorter critiques never hit. So A2 compares the container's shipped settings with and without the adapter; it does not establish what Qwen2-7B-Instruct would say if allowed to finish.

One effect is cleanly attributable to fine-tuning

Trading vocabulary appears in 18 of 32 critiques from the adapter against 3 of 32 from the base, on items containing no price, market, trade or security vocabulary at all. This attribution was pre-registered before the base run executed. The market frame is the fine-tune's, not the base model's.

The flags and the severity label

On flawed items, a flag named the actual flaw 2 times in 16 — and neither of those two items was flagged by the verdict. significant_flaw, and severity: high, fire on 5 of 16 sound items against 1 of 16 flawed ones: inverted, not merely uninformative. unaddressed_risk and notable_concern fire at exactly equal rates on sound and flawed items. Seven items emit flags while returning a pass verdict, so a consumer reading flags and a consumer reading the verdict get different answers from the same response.

Of the nine flags the server can emit, three runs and 88 responses have produced seven.

The critique text, read blind by a human

A human read all 32 critiques blind on 5 October 2026, against the key sealed when the sheet was written. The reader did not write the items, and saw only the claim and the critique — the label, verdict, flags, flaw field and item id were withheld and the order shuffled under a recorded seed.

Measure 3 — blind human read 95% CI
Critique identifies the real flaw (16 flawed items) 1/16 = 0.062 [0.011, 0.283]
Critique invents a problem (16 sound items) 15/16 = 0.938 [0.717, 0.989]

The prose is less useful than the verdict. The verdict flagged 4 of 16 flawed items; the prose names the actual flaw on one. The verdict falsely flagged 8 of 16 sound items; the prose objects to 15 of 16 — including the eight the product passed, which means a pass on a sound item was not the model withholding an objection but the keyword heuristic failing to notice one the model had already made. The single critique that identified its flaw does so through the trading frame, not despite it.

The same sheet had already been read blind by two model readers, recorded by the project as measure 3b and reported here as a model read. They agreed with each other on 28 of 32 entries:

Reader A (model) Reader B (model) Human
Critique identifies the real flaw (16 flawed items) 2/16 1/16 1/16
Critique invents a problem (16 sound items) 16/16 15/16 15/16

The human read matches Reader B exactly and Reader A within one item on each row. That is one sheet and one human: evidence that the model read was a usable proxy here, not a finding that model readers substitute for human ones. No measure 3b number is reported as measure 3. All three readers remarked independently that the critiques apply trading language to non-financial claims.

What has not been measured

  • A no-model rule baseline has been run on these 32 items, and it beats this model. Two rule forms over the item text alone — one flagging an unhedged conclusion, one simply passing anything that contains a hedge word — both score balanced accuracy 0.781 against this checkpoint's 0.375, with their cue lists fixed in the pre-registration before the code was written. Paired McNemar exact tests separate each from the model (p = 0.0072 and p = 0.0023). This is the only comparison in the evaluation where a difference is called: the adapter cannot be distinguished from its own base, but a regex can be distinguished from both. The rules were written by the author of the items and are therefore fitted, which makes this weak evidence about rules in general and not weak about the model. It also means the item set is partly cued — 0.781 is 0.019 below the threshold pre-registered for calling the set substantially separable, and its sound items hedge by construction. Published as a limitation of the evaluation.
  • Antagon's other measured MiniCrit result — 21 of 21 seeded flaws detected, 0 of 21 false positives on a 42-item set — belongs to a different, unpublished adapter on mistralai/Mistral-7B-v0.1. It does not describe this model, and it is not a before-and-after against 0.375. That set and its per-item human rulings are identified and located — seed_items_v2.jsonl, 21 flawed and 21 clean over 9 domains, run on 12 September 2026 against a step-500 adapter checkpoint with SHA-256 ccb9cb59…, judged by a single human evaluator — but they sit on the evaluation host rather than in the project repository, so that figure has no reproduction record published with it and should not be relied on until it does.

Training record

Training steps completed 34,000 of 364,831 (one epoch)
Progress recorded by the trainer epoch 0.0932
Effective batch size 32 (4 per device x 8 gradient accumulation, 1 GPU)
Initial logged loss 1.8539 (step 50)
Loss at checkpoint 0.7914 (step 34,000)
Lowest logged loss 0.7793 (step 21,850)
Loss reduction 57.3%

Training loss fell as expected over the completed steps, and the evaluation above shows what that does and does not imply. Nothing about capability should be inferred from a loss curve.

The evaluation deliberately did not use held-out rows of CritiqueBank-11M: evaluating a critic on held-out rows of the corpus it was trained on measures the learned distribution rather than the capability, because the evaluation items and the training data share a generator. That reasoning was published here before any evaluation existed, and the result above is consistent with it. The 32 items were written by hand, in five non-financial domains, with no trading vocabulary, and the set is published alongside the per-item results.

None of this withdraws the model as a reviewer's aid — a prompt that makes a human look again. All of it means it cannot be described as a detector, and no accuracy figure from this evaluation may be published as a capability claim for MiniCrit generally. It measures this checkpoint, on these items, through this container, on one date.


Training configuration

learning_rate: 2e-4
lr_scheduler: cosine
warmup_steps: 500
max_sequence_length: 512
lora_r: 16
lora_alpha: 32
lora_dropout: 0.05
target_modules: [q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj]

Dataset format is rationale → critique pairs over financial trading signals (stocks, options, crypto).


Usage

pip install transformers peft torch
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel

base_model = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen2-7B-Instruct",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2-7B-Instruct")
model = PeftModel.from_pretrained(base_model, "wmaousley/MiniCrit-7B")


def critique_rationale(rationale: str) -> str:
    prompt = f"### Rationale:\n{rationale}\n\n### Critique:\n"
    inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
    outputs = model.generate(
        **inputs,
        max_new_tokens=256,
        temperature=0.7,
        do_sample=True,
        pad_token_id=tokenizer.eos_token_id,
    )
    response = tokenizer.decode(outputs[0], skip_special_tokens=True)
    return response.split("### Critique:\n")[-1]


print(critique_rationale(
    "AAPL long: MACD bullish crossover with supporting momentum."
))

Example output

Input: META long: Bollinger Band expansion with supporting momentum.

Output: "While Bollinger Band expansion can signal volatility, META's recent expansion isn't necessarily predictive; it could be a reaction to news, not a precursor to sustained movement. Furthermore, relying solely on momentum without considering overbought/oversold levels may lead to premature entry, especially if the expansion is already near its peak."

This is one illustrative generation, not an evaluation result.


Intended use

Appropriate:

  • flagging likely reasoning flaws in AI-generated analysis for human review
  • research into adversarial critique of model reasoning
  • quality assurance workflows with a human in the loop

Not appropriate. This model is not intended for:

  • generating trading signals
  • financial advice
  • autonomous trading decisions

Limitations

  • On the held-out evaluation it does not beat a constant answer (balanced accuracy 0.375 against 0.500), and it flags sound reasoning more often than flawed reasoning. See Evaluation.
  • Removing the LoRA does not measurably help (McNemar p = 0.5235), so the limitation is not simply undertraining
  • The checkpoint represents partial training (34,000 of 364,831 steps, epoch 0.0932)
  • Fine-tuning imports a trading frame into critiques of non-financial inputs (18 of 32 against the base model's 3 of 32)
  • Trained on a narrow trading/finance domain; generalisation is untested
  • May produce confidently worded critiques that are themselves wrong
  • A supplement to human judgment, not a replacement

Citation

@misc{minicrit7b2026,
  title={MiniCrit-7B: Adversarial AI Critique for Trading Signal Validation},
  author={Ousley, William Alexander and Ousley, Jacqueline Villamor},
  year={2026},
  publisher={Antagon Inc.},
  url={https://huggingface.co/wmaousley/MiniCrit-7B}
}

Acknowledgments

We gratefully acknowledge Lambda Labs for providing GPU compute through their Research Grant program. MiniCrit-7B was trained on Lambda's H100 infrastructure.

Relationship to Antagon products

MiniCrit-7B is an open research release. Antagon's defense products, including the Frontier compliance platform, are separate and proprietary.

Contact

Antagon Inc. · antagon.ai · founders@antagon.ai CAGE 17E75 · UEI KBSGT7CZ4AH3


Revision history

5 October 2026, fifth revision. The blind human read the fourth revision recorded as outstanding has been performed, and this card reports it: 1 of 16 flawed critiques identifies its flaw, 15 of 16 sound items draw an invented objection. Measure 3 is discharged. The model read stays on the card, named as a model read, beside it.

5 October 2026, fourth revision. Adds measure 3b, a blind model read of the critique text, reported as a model read throughout. The human read the protocol specifies was still outstanding at that revision; it was performed later the same day and is reported in the fifth revision above.

5 October 2026, third revision. The held-out evaluation this card said was active work has been run, and the card now reports it. The line "No held-out evaluation has been run" is removed as no longer true; the reason this card gave for not having run one — that held-out rows of the training corpus share a generator with it — is kept, because it was right and the result bears it out. Every figure in Evaluation is recomputed from the raw response files by a checked-in verifier, and the raw files are published with their SHA-256 digests in the project repository. Nothing about the training record changed.

4 October 2026, second revision. All figures on this card are now read directly from the checkpoint's own trainer_state.json.

The card as originally published reported 35,650 steps, 9.8% of an epoch, an initial loss of 1.8573 and a final loss of 0.7869. The recorded values are 34,000 steps, epoch 0.0932, an initial logged loss of 1.8539 at step 50, and a loss of 0.7914 at step 34,000. The value 0.7869 appears at logged steps 24,450 and 31,350; it is a mid-run entry, not the checkpoint's loss.

An interim revision published earlier the same day stated epoch 0.0656 and a checkpoint loss of 0.7869. Both were mid-run log values carried forward without being checked against the trainer state, and both are superseded here. That revision also removed the 364,831-step denominator as belonging to a different run; it does not — it is 11,674,598 examples divided by an effective batch of 32, and it has been restored.

The card's YAML metadata block, which had been lost at some earlier point, was restored in the same pass.

License

Released under the Apache 2.0 License.

Downloads last month
65
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wmaousley/MiniCrit-7B

Base model

Qwen/Qwen2-7B
Adapter
(356)
this model

Dataset used to train wmaousley/MiniCrit-7B