Gregory-L commited on
Commit
76258e9
·
verified ·
1 Parent(s): 74da4f7

bankml: the verified CPU engine as backend, serve target and imprint probe

Browse files

bankml (github.com/cryptoAGI/bankml) is reached over HTTP and its CLI only;
nothing is vendored.

- operator/backends/bankml.py: @register_backend("bankml"), default
MINDXTRAIN_BANKML_BASE_URL http://127.0.0.1:18093/v1; keeps the
bankml_receipt (ChatResponse.receipt, last_receipt for streams); HTTP 400
becomes BankmlRefusal(reason) and is never retried with altered params.
- operator: backend_kwargs() replaces the duplicated per-backend branches in
app.py and coach/api.py; GET /bankml + /v1/models health probes;
auto-detect tries bankml only after ollama and vllm; the governance panel
env chain gains MINDXTRAIN_BANKML_BASE_URL last, only when set.
- deploy/bankml_push.py + `serve --to bankml`: merge -> Modelfile through
bankml_sanitize (refuses ADAPTER, a foreign TEMPLATE, penalties, mirostat,
typical_p, resource params; never drops them) -> bankml create; records
the model sha256; refuses quantized configs and non-Llama archs; a bankml
without create/convert (pre-0.3.5) is reported as bankml_too_old.
- eval/imprint_bankml.py + `imprint-bankml`: seeded, unpenalised greedy
probes with a receipt per utterance, tagged <scorer>/bankml-greedy and
marked not comparable with the canonical 1.3-penalty gate.
- MEI engine literal += bankml; Gradio Serve room offers bankml; console
omits penalty/mirostat options at the engine default; HF published
Modelfile can omit repeat_penalty (default unchanged).
- docs/bankml.md, README, CLAUDE.md, NAV, cli, development, CHANGELOG,
llm.txt. Tests: 56 new (821 passed, 3 skipped); ruff unchanged.

Co-Authored-By: Professor Codephreak <codephreak@pythai.net>

made with luv.pythai.net

CLAUDE.md CHANGED
@@ -98,7 +98,7 @@ mindX artifacts (dream corpus, persona) through boundaries, it does not vendor m
98
 
99
  - **New recipe** → drop YAML at `mindxtrain/train/recipes/<name>.yaml`; `test_all_recipes_validate` picks it up.
100
  - **New training backend** → add `mindxtrain/train/backend_<name>.py` exposing `run_<name>(cfg, plan, out_dir) -> Path`; wire into `train/dispatch.py`; add to `TrainingBackend` literal in `config/schema.py`.
101
- - **New operator backend** → subclass `Backend` in `mindxtrain/operator/backends/<name>.py` decorated `@register_backend("<name>")`; side-effect import from `models/registry.py`; add runtime branch in `operator/app.py::chat_completions`.
102
  - **New training method** → add `_MethodBase` subclass in `config/schema.py` with `kind: Literal["<name>"]`; extend `TrainMethod` discriminated union; add `train/<name>.py` runner; update dispatch; add a recipe.
103
 
104
  ## Documentation hub
@@ -115,4 +115,5 @@ mindX artifacts (dream corpus, persona) through boundaries, it does not vendor m
115
  | `docs/yaml_schema.md` | Every field of the 10-section `XTrainConfig`. |
116
  | `docs/coach.md` | Interactive `/coach/` web UI bundled in the operator. |
117
  | `docs/governance.md` | classroom / boardroom (any-N consensus) / dojo (prime-N dispute settlement). |
 
118
  | `docs/blueprints/` | Frozen source design briefs (the spec the project was built against). |
 
98
 
99
  - **New recipe** → drop YAML at `mindxtrain/train/recipes/<name>.yaml`; `test_all_recipes_validate` picks it up.
100
  - **New training backend** → add `mindxtrain/train/backend_<name>.py` exposing `run_<name>(cfg, plan, out_dir) -> Path`; wire into `train/dispatch.py`; add to `TrainingBackend` literal in `config/schema.py`.
101
+ - **New operator backend** → subclass `Backend` in `mindxtrain/operator/backends/<name>.py` decorated `@register_backend("<name>")`; side-effect import from `models/registry.py`; add its base-URL branch to `operator/app.py::backend_kwargs` (shared by the operator route and the Coach chat stream) and, if it can be probed, to `backend_reachable` / `backend_first_model`. Registered backends: `openai_compat`, `ollama`, `vllm`, `bankml` ([docs/bankml.md](docs/bankml.md) — the verified CPU engine, reached over HTTP/CLI only; receipts on `ChatResponse.receipt`, HTTP 400 → `BankmlRefusal`, never retried).
102
  - **New training method** → add `_MethodBase` subclass in `config/schema.py` with `kind: Literal["<name>"]`; extend `TrainMethod` discriminated union; add `train/<name>.py` runner; update dispatch; add a recipe.
103
 
104
  ## Documentation hub
 
115
  | `docs/yaml_schema.md` | Every field of the 10-section `XTrainConfig`. |
116
  | `docs/coach.md` | Interactive `/coach/` web UI bundled in the operator. |
117
  | `docs/governance.md` | classroom / boardroom (any-N consensus) / dojo (prime-N dispute settlement). |
118
+ | `docs/bankml.md` | bankml ([github.com/cryptoAGI/bankml](https://github.com/cryptoAGI/bankml)): `serve --to bankml`, the `bankml` backend, `imprint-bankml` (not comparable with the canonical gate). |
119
  | `docs/blueprints/` | Frozen source design briefs (the spec the project was built against). |
README.md CHANGED
@@ -93,6 +93,26 @@ GPU steps (`bench` without `--dry-run`, `train`, `quantize`, `serve`) require
93
  an AMD MI300X with ROCm 7.2.1; run inside `rocm/primus:v26.2`. The full
94
  operator checklist lives in [`HANDOFF.md`](docs/HANDOFF.md).
95
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
96
  ## Layout
97
 
98
  ```
@@ -118,6 +138,7 @@ scripts/ dev helpers
118
  | [`docs/coach.md`](docs/coach.md) | Interactive `/coach/` web UI bundled in the operator. |
119
  | [`docs/dcoach.md`](docs/dcoach.md) | The dcoach proof loop — prove a CPU model recalls its training; decentralized-training fit. |
120
  | [`docs/cli.md`](docs/cli.md) | Every `mindxtrain` verb with synopsis, options, exit codes. |
 
121
  | [`docs/yaml_schema.md`](docs/yaml_schema.md) | Every field of the 10-section `XTrainConfig`. |
122
  | [`docs/benchmarks.md`](docs/benchmarks.md) | Target metrics + the 7-cell framework comparison. |
123
  | [`docs/development.md`](docs/development.md) | Toolchain, optional-deps, lazy-import pattern, invariants. |
 
93
  an AMD MI300X with ROCm 7.2.1; run inside `rocm/primus:v26.2`. The full
94
  operator checklist lives in [`HANDOFF.md`](docs/HANDOFF.md).
95
 
96
+ ## bankml (verified CPU engine)
97
+
98
+ [bankml](https://github.com/cryptoAGI/bankml) is a zero-dependency Rust runtime, token-identical to
99
+ llama.cpp b11192, that serves OpenAI `/v1` and Ollama `/api` from its own forward pass and puts a
100
+ **receipt** (model / request / response sha256) on every answer. mindXtrain reaches it over HTTP and
101
+ its CLI only — nothing is vendored. Full page: [`docs/bankml.md`](docs/bankml.md).
102
+
103
+ - `mindxtrain serve run.yaml --to bankml` merges the LoRA and runs `bankml create` (bankml 0.3.5+),
104
+ which converts the merged SmolLM2 / mindx-genN weights to GGUF F16 byte-identically to llama.cpp
105
+ and pins them by sha256. It **refuses**, with the reason: quantized configs (FP8, MXFP4, GPTQ,
106
+ Q8_0, Q4_K), non-Llama architectures, and Modelfile instructions bankml does not reproduce —
107
+ `ADAPTER`, a foreign `TEMPLATE`, penalties, mirostat, `typical_p`, resource options. An older
108
+ bankml is reported as too old, not crashed on.
109
+ - `MINDXTRAIN_BACKEND=bankml` routes the operator and Coach chat to `bankml serve`
110
+ (`MINDXTRAIN_BANKML_BASE_URL`, default `http://127.0.0.1:18093/v1`); answers carry the receipt,
111
+ and a bankml 400 comes back as a typed `BankmlRefusal`, never retried with altered parameters.
112
+ - `mindxtrain imprint-bankml` probes before/after tags greedily, seeded and unpenalised, with a receipt
113
+ per utterance — reproducible and auditable, and explicitly **not comparable** with the canonical
114
+ `mindxtrain imprint` gate (repetition penalty 1.3).
115
+
116
  ## Layout
117
 
118
  ```
 
138
  | [`docs/coach.md`](docs/coach.md) | Interactive `/coach/` web UI bundled in the operator. |
139
  | [`docs/dcoach.md`](docs/dcoach.md) | The dcoach proof loop — prove a CPU model recalls its training; decentralized-training fit. |
140
  | [`docs/cli.md`](docs/cli.md) | Every `mindxtrain` verb with synopsis, options, exit codes. |
141
+ | [`docs/bankml.md`](docs/bankml.md) | bankml as serve target, operator backend and receipt-auditable imprint probe — what it takes and what it refuses. [bankml on GitHub](https://github.com/cryptoAGI/bankml). |
142
  | [`docs/yaml_schema.md`](docs/yaml_schema.md) | Every field of the 10-section `XTrainConfig`. |
143
  | [`docs/benchmarks.md`](docs/benchmarks.md) | Target metrics + the 7-cell framework comparison. |
144
  | [`docs/development.md`](docs/development.md) | Toolchain, optional-deps, lazy-import pattern, invariants. |
docs/CHANGELOG.md CHANGED
@@ -6,6 +6,37 @@ project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
6
 
7
  ## [Unreleased]
8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9
  ## [1.0.4] — 2026-09-16
10
 
11
  ### Added
 
6
 
7
  ## [Unreleased]
8
 
9
+ ### Added
10
+
11
+ - **bankml, the verified CPU engine** ([github.com/cryptoAGI/bankml](https://github.com/cryptoAGI/bankml),
12
+ [`docs/bankml.md`](bankml.md)), reached over HTTP and its CLI only — nothing vendored.
13
+ - `operator/backends/bankml.py`: `@register_backend("bankml")`, `MINDXTRAIN_BANKML_BASE_URL`
14
+ (default `http://127.0.0.1:18093/v1`). Keeps bankml's receipt (`ChatResponse.receipt`, new
15
+ optional field; `last_receipt` for streams, parsed from the final SSE receipt event). HTTP 400
16
+ → `BankmlRefusal(reason)`, never retried with altered parameters.
17
+ - Operator wiring: `backend_kwargs(name)` replaces the duplicated per-backend branches in
18
+ `operator/app.py` and `coach/api.py`; health probes `GET /bankml` and `/v1/models`;
19
+ auto-detect tries bankml after ollama and vllm, so existing hosts resolve as before; the
20
+ governance panel's env chain gains `MINDXTRAIN_BANKML_BASE_URL` last, only when set.
21
+ - `deploy/bankml_push.py` and `mindxtrain serve --to bankml [--bankml-bin] [--bankml-convert]`:
22
+ merge → Modelfile through `bankml_sanitize` (refuses `ADAPTER`, a foreign `TEMPLATE`,
23
+ penalties, mirostat, `typical_p`, resource options, rather than dropping them) →
24
+ `bankml create`; records the model sha256; refuses quantized configs and non-Llama
25
+ architectures; detects a bankml without `create`/`convert` (pre-0.3.5) as `bankml_too_old`.
26
+ - `eval/imprint_bankml.py` and `mindxtrain imprint-bankml`: seeded, unpenalised greedy probes with
27
+ a receipt per utterance, scored by `score_imprint`, tagged `<scorer>/bankml-greedy` and marked
28
+ **not comparable** with the canonical 1.3-penalty gate.
29
+ - MEI `InferenceEngineIdent.name` accepts `"bankml"`; the Gradio Serve room offers `bankml`.
30
+
31
+ ### Changed
32
+
33
+ - `ui/console.py` omits penalty / mirostat options left at the engine's own default, so bankml
34
+ accepts default console requests. For Ollama a Modelfile's `repeat_penalty` now applies where
35
+ the console used to override it with 1.1.
36
+ - `hf.extension.publish_generation(repeat_penalty=…)`: `None` publishes a Modelfile without the
37
+ penalty line (default unchanged, 1.3); the Modelfile text is `published_modelfile()`.
38
+ - `cli imprint`: the script-probe extraction is `_script_probes`, shared with `imprint-bankml`.
39
+
40
  ## [1.0.4] — 2026-09-16
41
 
42
  ### Added
docs/NAV.md CHANGED
@@ -123,6 +123,12 @@ The interactive `/coach/` operator UI: create-script, live-training diagnostics,
123
  - [How mindXtrain fits decentralized training](dcoach.md#how-mindxtrain-fits-decentralized-training)
124
  - [Why this matters](dcoach.md#why-this-matters)
125
 
 
 
 
 
 
 
126
  ### [Governance](governance.md)
127
  classroom (graduation) / boardroom (any-N consensus) / dojo (prime-N dispute settlement), model-backed deliberation.
128
  - [The model](governance.md#the-model) · [Flow](governance.md#flow) · [Why prime](governance.md#why-prime)
 
123
  - [How mindXtrain fits decentralized training](dcoach.md#how-mindxtrain-fits-decentralized-training)
124
  - [Why this matters](dcoach.md#why-this-matters)
125
 
126
+ ### [bankml — the verified CPU engine](bankml.md)
127
+ [bankml](https://github.com/cryptoAGI/bankml) as serve target, operator backend and receipt-auditable imprint probe.
128
+ - [What bankml runs, and what it refuses](bankml.md#what-bankml-runs-and-what-it-refuses)
129
+ - [Operator backend](bankml.md#operator-backend--mindxtrain_backendbankml) · [`serve --to bankml`](bankml.md#serving-a-trained-run--mindxtrain-serve---to-bankml) · [The Modelfile subset](bankml.md#the-modelfile-subset-bankml_sanitize)
130
+ - [`imprint-bankml` (not comparable with the canonical gate)](bankml.md#a-second-imprint-instrument--mindxtrain-imprint-bankml)
131
+
132
  ### [Governance](governance.md)
133
  classroom (graduation) / boardroom (any-N consensus) / dojo (prime-N dispute settlement), model-backed deliberation.
134
  - [The model](governance.md#the-model) · [Flow](governance.md#flow) · [Why prime](governance.md#why-prime)
docs/bankml.md ADDED
@@ -0,0 +1,133 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # bankml — the verified CPU engine
2
+
3
+ [bankml](https://github.com/cryptoAGI/bankml) is a zero-dependency Rust runtime for 1-bit (`Q1_0`),
4
+ ternary (`Q2_0_g64`) and F16 GGUF models, **token-identical to llama.cpp b11192** on every oracle in
5
+ its release gate. `bankml serve --native` answers OpenAI `/v1/chat/completions` and Ollama `/api/*`
6
+ from its own forward pass on `127.0.0.1:18093`, and puts a **receipt** on every answer. mindXtrain
7
+ uses it as a serve target, an operator backend and a second imprint instrument.
8
+
9
+ mindXtrain reaches bankml **only over HTTP and as a CLI subprocess**. No bankml code is vendored
10
+ (the clean-room policy in [`CLAUDE.md`](../CLAUDE.md)). bankml's own docs:
11
+ [README](https://github.com/cryptoAGI/bankml#readme) ·
12
+ [usage](https://github.com/cryptoAGI/bankml/blob/main/docs/usage.md) ·
13
+ [bankML as mindX's Ollama](https://github.com/cryptoAGI/bankml/blob/main/docs/OLLAMA.md).
14
+
15
+ ## What bankml runs, and what it refuses
16
+
17
+ | | bankml |
18
+ |---|---|
19
+ | **Architectures** | Qwen3 (`Q1_0`, `Q2_0_g64`: the Bonsai family) and Llama in F16 (SmolLM2-135M, mindX's `mindx-genN`) |
20
+ | **Sampling it reproduces** | `temperature`, `top_k`, `top_p`, `min_p`, `seed`, JSON mode; `num_ctx`, `num_predict`, `stop` |
21
+ | **Refused with HTTP 400 and a reason** | repeat / presence / frequency penalties (non-neutral), `mirostat`, `typical_p`, `tools`, images, a replacement `template`, unknown architectures, `Q8_0` / `Q4_K` / `BF16` |
22
+ | **Receipt** (`bankml_receipt`) | `bankml` version, `engine`, `model_sha256`, `guard`, `prompt_tokens`, `completion_tokens`, `ttft_ms`, `wall_ms`, `response_sha256`, `request_sha256`, `signed: false` |
23
+
24
+ A refusal is the product, not a defect: bankml answers only what its verified forward pass does.
25
+ mindXtrain mirrors that — it **never retries a refused request with altered parameters**, and never
26
+ drops a Modelfile instruction to make bankml accept it.
27
+
28
+ ## Operator backend — `MINDXTRAIN_BACKEND=bankml`
29
+
30
+ `mindxtrain/operator/backends/bankml.py` registers `bankml` (a subclass of `openai_compat`).
31
+
32
+ ```bash
33
+ bankml serve MODEL.gguf --fork MODEL.gguf.FORK.json --native --registry # 127.0.0.1:18093
34
+ MINDXTRAIN_BACKEND=bankml uv run uvicorn mindxtrain.operator.app:app --port 8080
35
+ ```
36
+
37
+ - `MINDXTRAIN_BANKML_BASE_URL` (default `http://127.0.0.1:18093/v1`).
38
+ - `POST /v1/chat/completions` returns `ChatResponse.receipt`; the backend also keeps `last_receipt`.
39
+ Streamed answers parse bankml's final `data: {"bankml_receipt": …}` event.
40
+ - HTTP 400 → `BankmlRefusal(reason)` → the operator answers 400 with bankml's reason. Other non-2xx
41
+ and a mid-stream `{"error": …}` → `BankmlError`.
42
+ - Health: `/health`, `/readyz` and `/coach/api/health` probe `GET /bankml` (the identity endpoint) and
43
+ list the model from `/v1/models`.
44
+ - **Auto-detect** (no `MINDXTRAIN_BACKEND`): ollama → vllm → bankml → vllm. bankml is chosen only
45
+ when ollama is down, `GET /bankml` answers and vLLM does not, so a host that resolved to ollama or
46
+ vllm before keeps doing so.
47
+ - Governance panels (`resolve_chat_base_url`): `MINDXTRAIN_BACKEND=bankml` uses bankml's URL; else
48
+ `MINDXTRAIN_BANKML_BASE_URL` is appended *last* to the OPENAI → VLLM → OLLAMA chain, and only
49
+ when it is set.
50
+
51
+ ## Serving a trained run — `mindxtrain serve --to bankml`
52
+
53
+ ```bash
54
+ uv run mindxtrain serve run.yaml --to bankml [--tag mindx-gen80] [--checkpoint DIR] \
55
+ [--bankml-bin PATH] [--bankml-convert] [--register-as-fallback]
56
+ ```
57
+
58
+ 1. **Refuses up front** (exit 2): `quantize.enabled` with a scheme other than `none` (bankml serves
59
+ the merged weights as GGUF F16; it does not reproduce FP8, MXFP4, GPTQ, Q8_0 or Q4_K), and base
60
+ families bankml cannot convert (Qwen, Mistral, Phi, Gemma, GLM, DeepSeek, Instella).
61
+ 2. Checks the binary: `bankml version` and the verbs in `bankml --help`. `bankml create` and
62
+ `bankml convert` arrive in **bankml 0.3.5**; an older binary is reported as
63
+ `bankml_too_old` (exit 2), never as a crash. The verbs are the truth — an unreleased build may
64
+ still say 0.3.4 and already carry them.
65
+ 3. Merges the LoRA (`merge_lora_adapter`, needs `uv sync --extra ml`), or takes the checkpoint as an
66
+ already-merged directory when it holds `config.json` and no `adapter_config.json`.
67
+ 4. Re-checks the merged `config.json`: only `LlamaForCausalLM` converts.
68
+ 5. Writes a Modelfile through `bankml_sanitize` and runs `bankml create <tag> -f Modelfile`, which
69
+ converts the merged safetensors to GGUF F16 byte-identically to llama.cpp b11192 and pins it.
70
+ `--bankml-convert` runs `bankml convert` (GGUF + `FORK.json`) first and writes `FROM <gguf>`.
71
+ 6. Records the model sha256 bankml prints for the base it verified, and the derived model's digest.
72
+ 7. `--register-as-fallback` PATCHes mindX's fallback model to `{provider: "bankml", model: <tag>}`
73
+ (best-effort, as for ollama).
74
+
75
+ ### The Modelfile subset (`bankml_sanitize`)
76
+
77
+ | instruction | bankml |
78
+ |---|---|
79
+ | `FROM` merged dir / pinned GGUF / registry name | taken |
80
+ | `SYSTEM`, `MESSAGE`, `LICENSE`, `REQUIRES` | taken (recorded) |
81
+ | `PARAMETER` temperature, top_k, top_p, min_p, seed, num_ctx, num_predict; `stop` | taken |
82
+ | `ADAPTER` | **refused** — merge first (`push_to_bankml` does) |
83
+ | `TEMPLATE` | **refused** unless equal to the base's own chat template |
84
+ | penalties, `repeat_last_n`, mirostat*, `typical_p` | **refused** — not reproduced |
85
+ | `num_gpu`, `num_thread`, `num_batch`, `num_keep`, `draft_num_predict` | **refused** — a resource option is not part of a model |
86
+
87
+ Each refusal is returned with its reason; nothing is dropped silently. Python API:
88
+ `mindxtrain.deploy.bankml_push.push_to_bankml(...) -> BankmlPushResult` (never raises; `status` is
89
+ one of `created`, `refused`, `bankml_missing`, `bankml_too_old`, `merge_failed`, `failed`, `error`).
90
+
91
+ ## A second imprint instrument — `mindxtrain imprint-bankml`
92
+
93
+ ```bash
94
+ uv run mindxtrain imprint-bankml run.yaml --before smollm2-135m-instruct --after mindx-gen80 \
95
+ [--seed 0] [--num-predict 48] [--system "…"] [--base-url http://127.0.0.1:18093/v1]
96
+ ```
97
+
98
+ Poses the script's user-turns to two tags on bankml's `/api/chat` with `temperature 0`, a fixed
99
+ `seed`, `num_predict 48` and **no penalties**, then scores with the existing `score_imprint`. The
100
+ report (`BankmlImprintReport`) carries the decoding, every utterance's receipt and the distinct
101
+ `model_sha256` values, and `report.method` is tagged `<scorer>/bankml-greedy`.
102
+
103
+ **It is not comparable with the canonical gate.** `mindxtrain imprint` decodes with transformers
104
+ greedy, `repetition_penalty 1.3` and `no_repeat_ngram_size 3`; every number in an ascent log comes
105
+ from that. A bankml-greedy score is compared only with bankml-greedy scores, and the report says so
106
+ (`canonical_gate: false`, `comparable_with: "bankml-greedy only"`). What it buys:
107
+
108
+ - **reproducible** — the same seed and weights give the same tokens, identical to llama.cpp;
109
+ - **auditable** — a score is tied to the exact weights by sha256, not to a tag name;
110
+ - **cheap** — a 135M F16 actor answers on one CPU core, with no torch in the probing process.
111
+
112
+ Observed on mindx-gen39 (2026-10-02): unpenalised greedy decoding degenerates on short probes
113
+ without a system turn (runs of `,` and `?||`), which is the very behaviour the 1.3 penalty in the
114
+ canonical gate suppresses. Expect low bankml-greedy voice scores until the actor itself stops
115
+ repeating; pass the persona's `--system` as the coach does.
116
+
117
+ ## Console and published Modelfiles
118
+
119
+ - `mindxtrain/ui/console.py` no longer sends a penalty or mirostat option left at the engine's own
120
+ default (`repeat_penalty 1.1`, presence / frequency `0`, `typical_p 1`, `mirostat 0` and its
121
+ tau / eta while it is off), so bankml accepts a console request with default settings. A
122
+ deliberate value is always sent; bankml then refuses it visibly. For Ollama an absent key means
123
+ its own default, except that a Modelfile's `PARAMETER repeat_penalty` now applies where the
124
+ console used to override it with 1.1.
125
+ - `hf.extension.publish_generation(..., repeat_penalty=None)` publishes a Modelfile without the
126
+ penalty line, which bankml can load. The default stays 1.3.
127
+
128
+ ## Tests
129
+
130
+ `tests/test_bankml_backend.py`, `tests/test_bankml_push.py`, `tests/test_imprint_bankml.py` — no
131
+ network and no binary: `httpx.MockTransport`, monkeypatched `subprocess.run` / `shutil.which`.
132
+ `tests/conftest.py` pins the bankml auto-detect probe to "absent" so a developer box running bankml
133
+ cannot change what the other auto-detect tests resolve to.
docs/cli.md CHANGED
@@ -145,6 +145,28 @@ The chat-template parsers map per `serve.reasoning_parser` (`qwen3` for Qwen3,
145
  `deepseek_r1` for DeepSeek-style) and `serve.tool_call_parser` (`hermes`,
146
  `qwen3_coder`).
147
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
148
  ## `publish` — push to HF + Lighthouse + register
149
 
150
  ```
@@ -204,5 +226,7 @@ $ uv run mindxtrain receipt out/runs/<run_id>/manifest.json --config run.yaml
204
  | `eval` | `mindxtrain.cli.main.eval_` + `mindxtrain.eval.harness.run_lm_eval` |
205
  | `quantize` | `mindxtrain.cli.main.quantize` + `mindxtrain.deploy.quark.quark_fp8` |
206
  | `serve` | `mindxtrain.cli.main.serve` + `mindxtrain.deploy.vllm_launcher.build_vllm_command` |
 
 
207
  | `publish` | `mindxtrain.cli.main.publish` + `mindxtrain.storage.{hf_hub,lighthouse}` + `mindxtrain.deploy.api_client` |
208
  | `receipt` | `mindxtrain.cli.main.receipt` + `mindxtrain.provenance.verify.verify_receipt` |
 
145
  `deepseek_r1` for DeepSeek-style) and `serve.tool_call_parser` (`hermes`,
146
  `qwen3_coder`).
147
 
148
+ ### `serve --to ollama` / `--to bankml`
149
+
150
+ ```
151
+ mindxtrain serve <config.yaml> --to bankml [--tag NAME] [--checkpoint DIR] [--bankml-bin PATH] [--bankml-convert] [--register-as-fallback]
152
+ ```
153
+
154
+ `--to ollama` merges the LoRA and runs `ollama create`. `--to bankml` does the same through
155
+ [bankml](https://github.com/cryptoAGI/bankml) (`bankml create`, 0.3.5+) and refuses what bankml does
156
+ not reproduce — see [bankml.md](bankml.md). Exit codes for `--to bankml`: 1 checkpoint missing ·
157
+ 2 refused (quantized config, architecture, Modelfile subset, bankml missing or too old) ·
158
+ 3 merge or create failed.
159
+
160
+ ## `imprint-bankml` — receipt-auditable imprint probes
161
+
162
+ ```
163
+ mindxtrain imprint-bankml <config.yaml> --before TAG --after TAG [--n 5] [--seed 0] [--num-predict 48] [--system TEXT] [--base-url URL]
164
+ ```
165
+
166
+ Greedy, seeded, unpenalised probes through `bankml serve --native`, scored with `score_imprint`;
167
+ every utterance carries bankml's receipt. **Not comparable** with `imprint` (repetition penalty
168
+ 1.3). Exit 3 on a bankml refusal or error, 4 when no imprint is detected.
169
+
170
  ## `publish` — push to HF + Lighthouse + register
171
 
172
  ```
 
226
  | `eval` | `mindxtrain.cli.main.eval_` + `mindxtrain.eval.harness.run_lm_eval` |
227
  | `quantize` | `mindxtrain.cli.main.quantize` + `mindxtrain.deploy.quark.quark_fp8` |
228
  | `serve` | `mindxtrain.cli.main.serve` + `mindxtrain.deploy.vllm_launcher.build_vllm_command` |
229
+ | `serve --to bankml` | `mindxtrain.cli.main._serve_bankml` + `mindxtrain.deploy.bankml_push.push_to_bankml` |
230
+ | `imprint-bankml` | `mindxtrain.cli.main.imprint_bankml` + `mindxtrain.eval.imprint_bankml.imprint_via_bankml` |
231
  | `publish` | `mindxtrain.cli.main.publish` + `mindxtrain.storage.{hf_hub,lighthouse}` + `mindxtrain.deploy.api_client` |
232
  | `receipt` | `mindxtrain.cli.main.receipt` + `mindxtrain.provenance.verify.verify_receipt` |
docs/development.md CHANGED
@@ -198,8 +198,10 @@ RTX is the intended consumer GPU.
198
  decorated `@register_backend("<name>")`.
199
  2. Side-effect import it from `mindxtrain/models/registry.py` so registration
200
  runs on package import.
201
- 3. Add a runtime branch in `mindxtrain/operator/app.py::chat_completions` for
202
- the env-var-driven kwargs.
 
 
203
 
204
  ## Adding a new training method
205
 
 
198
  decorated `@register_backend("<name>")`.
199
  2. Side-effect import it from `mindxtrain/models/registry.py` so registration
200
  runs on package import.
201
+ 3. Add its env-var-driven kwargs to `mindxtrain/operator/app.py::backend_kwargs`
202
+ (the operator route and the Coach chat stream both read it), and a probe to
203
+ `backend_reachable` / `backend_first_model` if it has one. Example:
204
+ `operator/backends/bankml.py` ([bankml.md](bankml.md)).
205
 
206
  ## Adding a new training method
207
 
llm.txt CHANGED
@@ -81,6 +81,13 @@ uv run python -m mindxtrain.ui.app # the UI at :7862
81
  comparable. The floor is calibrated against a **null** (an untrained random-init adapter), not chosen.
82
  A run that fails the gate is recorded as failed and is not served. *That refusal is the product.*
83
 
 
 
 
 
 
 
 
84
  ## The teaching artifacts
85
 
86
  Published beside the weights at `PYTHAI/mindXtrain39`, and mirrored in mindX at
 
81
  comparable. The floor is calibrated against a **null** (an untrained random-init adapter), not chosen.
82
  A run that fails the gate is recorded as failed and is not served. *That refusal is the product.*
83
 
84
+ `mindxtrain imprint-bankml` is a **second instrument**, not the gate: probes through
85
+ [bankml](https://github.com/cryptoAGI/bankml) (`bankml serve --native`, token-identical to llama.cpp
86
+ b11192) at temperature 0, a fixed seed and **no penalties**, a receipt (model sha256) per utterance.
87
+ Its scores are tagged `<scorer>/bankml-greedy` and are **not comparable** with the gate's. Never mix
88
+ them. Backends the operator knows: `openai_compat`, `ollama`, `vllm`, `bankml` — see `docs/bankml.md`;
89
+ a bankml HTTP 400 is a `BankmlRefusal` with its reason, and is never retried with altered parameters.
90
+
91
  ## The teaching artifacts
92
 
93
  Published beside the weights at `PYTHAI/mindXtrain39`, and mirrored in mindX at
mindxtrain/cli/main.py CHANGED
@@ -11,6 +11,7 @@ from mindxtrain import __version__
11
  from mindxtrain.autotune.benchmark import run_autotune
12
  from mindxtrain.autotune.plan import AutotunePlan
13
  from mindxtrain.config.loader import list_recipes, load_config, render_recipe
 
14
 
15
  app = typer.Typer(
16
  name="mindxtrain",
@@ -301,6 +302,74 @@ def quantize(
301
  console.print(f"[green]quantized:[/green] {path}")
302
 
303
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
304
  @app.command()
305
  def serve(
306
  config: Path = typer.Argument(...),
@@ -308,17 +377,28 @@ def serve(
308
  to: str = typer.Option(
309
  "vllm", "--to",
310
  help="Serve target: vllm (default, builds vllm-rocm launch cmd), "
311
- "sglang (builds sglang-rocm launch cmd), or "
312
- "ollama (merges LoRA + calls `ollama create`).",
 
 
313
  ),
314
  tag: str = typer.Option(
315
  None, "--tag",
316
- help="Ollama tag for `--to ollama`. Defaults to run_name when omitted.",
317
  ),
318
  ollama_bin: str = typer.Option(
319
  None, "--ollama-bin",
320
  help="Override the ollama binary path (defaults to PATH lookup).",
321
  ),
 
 
 
 
 
 
 
 
 
322
  register_as_fallback: bool = typer.Option(
323
  False, "--register-as-fallback",
324
  help=(
@@ -345,9 +425,22 @@ def serve(
345
  into the base weights, writes an ollama Modelfile, and calls
346
  `ollama create <tag>` so the trained model is immediately available
347
  on the loopback (the same backend Coach probes for its chat card).
 
 
 
 
 
 
348
  """
349
  cfg = load_config(config)
350
 
 
 
 
 
 
 
 
351
  if to == "ollama":
352
  from mindxtrain.deploy.ollama_push import push_to_ollama
353
 
@@ -600,34 +693,12 @@ def receipt(
600
  raise typer.Exit(code=2)
601
 
602
 
603
- @app.command()
604
- def imprint(
605
- config: Path = typer.Argument(..., help="recipe whose checkpoint to measure"),
606
- out: Path = typer.Option(Path("./out/runs"), "--out", "-o"),
607
- max_inquiries: int = typer.Option(5, "--n", help="number of recall probes"),
608
- trigger_dream: bool = typer.Option(
609
- False, "--trigger-dream",
610
- help="hand the imprinted actor to mindX's machine.dream 8hr cycle",
611
- ),
612
- ) -> None:
613
- """Measure a persona imprint: recall before vs after training.
614
-
615
- Poses the script's own user-turns back to the actor, comparing the base
616
- model (before) and the trained adapter (after) against the script's
617
- assistant voice. Prints an ImprintReport; exit 4 if no imprint was detected.
618
- """
619
  import json as _json
620
 
621
- cfg = load_config(config)
622
- run_dir = (out / cfg.meta.run_name) if out.name == "runs" else out
623
- adapter_dir = run_dir / "checkpoint"
624
- if not adapter_dir.exists():
625
- console.print(f"[red]no checkpoint to measure:[/red] {adapter_dir}")
626
- raise typer.Exit(code=1)
627
-
628
- # Build inquiries (user-turns) + baseline voice (assistant-turns) from the
629
- # local script the actor trained on. Falls back to default probes.
630
- from mindxtrain.eval.imprint import default_inquiries, probe_recall, score_imprint
631
 
632
  inquiries: list[str] = []
633
  baseline: list[str] = []
@@ -652,6 +723,37 @@ def imprint(
652
  baseline.append(a)
653
  if not inquiries:
654
  inquiries = default_inquiries(cfg.meta.project)[:max_inquiries]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
655
 
656
  console.print(f"[cyan]probing {len(inquiries)} inquiries (before/after)…[/cyan]")
657
  try:
@@ -681,6 +783,51 @@ def imprint(
681
  raise typer.Exit(code=4)
682
 
683
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
684
  # ---- research (autoresearch search over one editable file) --------------
685
 
686
 
 
11
  from mindxtrain.autotune.benchmark import run_autotune
12
  from mindxtrain.autotune.plan import AutotunePlan
13
  from mindxtrain.config.loader import list_recipes, load_config, render_recipe
14
+ from mindxtrain.config.schema import XTrainConfig
15
 
16
  app = typer.Typer(
17
  name="mindxtrain",
 
302
  console.print(f"[green]quantized:[/green] {path}")
303
 
304
 
305
+ def _serve_bankml(
306
+ cfg: XTrainConfig,
307
+ checkpoint: Path | None,
308
+ *,
309
+ tag: str | None,
310
+ bankml_bin: str | None,
311
+ convert: bool,
312
+ register_as_fallback: bool,
313
+ mindx_base_url: str | None,
314
+ ) -> None:
315
+ """`serve --to bankml`: refuse what bankml cannot serve, else merge + `bankml create`.
316
+
317
+ Exit codes: 1 checkpoint missing · 2 refused (quantized config, architecture, Modelfile
318
+ subset, bankml missing or too old) · 3 merge / create failed.
319
+ """
320
+ from mindxtrain.deploy.bankml_push import base_family_refusal, push_to_bankml
321
+
322
+ quant = cfg.quantize
323
+ if quant.enabled and quant.scheme != "none":
324
+ console.print(
325
+ f"[red]bankml refuses quantize.scheme={quant.scheme}:[/red] bankml serves the merged "
326
+ "weights as GGUF F16 (and pinned Q1_0 / Q2_0_g64); it does not reproduce FP8, MXFP4, "
327
+ "GPTQ, Q8_0 or Q4_K. Set `quantize.scheme: none` (or `enabled: false`), or serve "
328
+ "with --to vllm.",
329
+ )
330
+ raise typer.Exit(code=2)
331
+ base_model = cfg.model.name
332
+ family = base_family_refusal(base_model)
333
+ if family:
334
+ console.print(f"[red]bankml refuses:[/red] {family}")
335
+ raise typer.Exit(code=2)
336
+
337
+ run_name = cfg.meta.run_name
338
+ ckpt = checkpoint or Path("./out/runs") / run_name / "checkpoint"
339
+ if not ckpt.exists():
340
+ console.print(f"[red]checkpoint not found:[/red] {ckpt}")
341
+ raise typer.Exit(code=1)
342
+ is_adapter = (ckpt / "adapter_config.json").exists() or not (ckpt / "config.json").exists()
343
+ result = push_to_bankml(
344
+ base_model,
345
+ tag or run_name,
346
+ adapter_dir=ckpt if is_adapter else None,
347
+ merged_dir=None if is_adapter else ckpt,
348
+ bankml_bin=bankml_bin,
349
+ convert=convert,
350
+ register_with_mindx=register_as_fallback,
351
+ mindx_base_url=mindx_base_url,
352
+ sink=lambda line: console.print(line, markup=False, highlight=False),
353
+ )
354
+ if not result.ok:
355
+ console.print(f"[red]push-to-bankml {result.status}:[/red] {result.reason}")
356
+ for r in result.refusals:
357
+ console.print(f" - {r}", markup=False)
358
+ refused = ("refused", "bankml_missing", "bankml_too_old")
359
+ raise typer.Exit(code=2 if result.status in refused else 3)
360
+ console.print(
361
+ f"[green]pushed:[/green] {result.tag} on bankml {result.bankml_version} "
362
+ f"(model sha256 {result.model_sha256 or '?'}, Modelfile: {result.modelfile})",
363
+ )
364
+ console.print("serve it: bankml serve <pinned.gguf> --fork <FORK.json> --native --registry")
365
+ if result.mindx_fallback_swap:
366
+ console.print(
367
+ f"[green]mindX fallback swapped:[/green] "
368
+ f"{result.mindx_fallback_swap.get('previous', '?')} -> "
369
+ f"{result.mindx_fallback_swap.get('current', '?')}",
370
+ )
371
+
372
+
373
  @app.command()
374
  def serve(
375
  config: Path = typer.Argument(...),
 
377
  to: str = typer.Option(
378
  "vllm", "--to",
379
  help="Serve target: vllm (default, builds vllm-rocm launch cmd), "
380
+ "sglang (builds sglang-rocm launch cmd), "
381
+ "ollama (merges LoRA + calls `ollama create`), or "
382
+ "bankml (merges LoRA + calls `bankml create`; the verified CPU engine, "
383
+ "https://github.com/cryptoAGI/bankml).",
384
  ),
385
  tag: str = typer.Option(
386
  None, "--tag",
387
+ help="Model tag for `--to ollama|bankml`. Defaults to run_name when omitted.",
388
  ),
389
  ollama_bin: str = typer.Option(
390
  None, "--ollama-bin",
391
  help="Override the ollama binary path (defaults to PATH lookup).",
392
  ),
393
+ bankml_bin: str = typer.Option(
394
+ None, "--bankml-bin",
395
+ help="Override the bankml binary path for `--to bankml` (defaults to PATH lookup).",
396
+ ),
397
+ bankml_convert: bool = typer.Option(
398
+ False, "--bankml-convert",
399
+ help="With `--to bankml`: run `bankml convert` (GGUF F16 + FORK.json pin) before "
400
+ "`bankml create`, instead of letting create convert the merged directory.",
401
+ ),
402
  register_as_fallback: bool = typer.Option(
403
  False, "--register-as-fallback",
404
  help=(
 
425
  into the base weights, writes an ollama Modelfile, and calls
426
  `ollama create <tag>` so the trained model is immediately available
427
  on the loopback (the same backend Coach probes for its chat card).
428
+
429
+ `--to bankml` does the same through bankml (`bankml create`, 0.3.5+):
430
+ the merged weights converted to GGUF F16 byte-identically to llama.cpp,
431
+ served by `bankml serve --native` with a receipt on every answer. It
432
+ refuses quantized configs (bankml serves F16 conversions, not FP8 /
433
+ MXFP4 / GPTQ / Q8_0 / Q4_K) and non-Llama architectures, with the reason.
434
  """
435
  cfg = load_config(config)
436
 
437
+ if to == "bankml":
438
+ _serve_bankml(
439
+ cfg, checkpoint, tag=tag, bankml_bin=bankml_bin, convert=bankml_convert,
440
+ register_as_fallback=register_as_fallback, mindx_base_url=mindx_base_url,
441
+ )
442
+ return
443
+
444
  if to == "ollama":
445
  from mindxtrain.deploy.ollama_push import push_to_ollama
446
 
 
693
  raise typer.Exit(code=2)
694
 
695
 
696
+ def _script_probes(cfg: XTrainConfig, max_inquiries: int) -> tuple[list[str], list[str]]:
697
+ """Inquiries (the script's user-turns, at most `max_inquiries`) and the baseline voice (its
698
+ assistant-turns) from the local script the actor trained on; default probes otherwise."""
 
 
 
 
 
 
 
 
 
 
 
 
 
699
  import json as _json
700
 
701
+ from mindxtrain.eval.imprint import default_inquiries
 
 
 
 
 
 
 
 
 
702
 
703
  inquiries: list[str] = []
704
  baseline: list[str] = []
 
723
  baseline.append(a)
724
  if not inquiries:
725
  inquiries = default_inquiries(cfg.meta.project)[:max_inquiries]
726
+ return inquiries, baseline
727
+
728
+
729
+ @app.command()
730
+ def imprint(
731
+ config: Path = typer.Argument(..., help="recipe whose checkpoint to measure"),
732
+ out: Path = typer.Option(Path("./out/runs"), "--out", "-o"),
733
+ max_inquiries: int = typer.Option(5, "--n", help="number of recall probes"),
734
+ trigger_dream: bool = typer.Option(
735
+ False, "--trigger-dream",
736
+ help="hand the imprinted actor to mindX's machine.dream 8hr cycle",
737
+ ),
738
+ ) -> None:
739
+ """Measure a persona imprint: recall before vs after training.
740
+
741
+ Poses the script's own user-turns back to the actor, comparing the base
742
+ model (before) and the trained adapter (after) against the script's
743
+ assistant voice. Prints an ImprintReport; exit 4 if no imprint was detected.
744
+ """
745
+ cfg = load_config(config)
746
+ run_dir = (out / cfg.meta.run_name) if out.name == "runs" else out
747
+ adapter_dir = run_dir / "checkpoint"
748
+ if not adapter_dir.exists():
749
+ console.print(f"[red]no checkpoint to measure:[/red] {adapter_dir}")
750
+ raise typer.Exit(code=1)
751
+
752
+ # Build inquiries (user-turns) + baseline voice (assistant-turns) from the
753
+ # local script the actor trained on. Falls back to default probes.
754
+ from mindxtrain.eval.imprint import probe_recall, score_imprint
755
+
756
+ inquiries, baseline = _script_probes(cfg, max_inquiries)
757
 
758
  console.print(f"[cyan]probing {len(inquiries)} inquiries (before/after)…[/cyan]")
759
  try:
 
783
  raise typer.Exit(code=4)
784
 
785
 
786
+ @app.command("imprint-bankml")
787
+ def imprint_bankml(
788
+ config: Path = typer.Argument(..., help="recipe whose script supplies inquiries + voice"),
789
+ before: str = typer.Option(..., "--before", help="bankml tag of the base actor"),
790
+ after: str = typer.Option(..., "--after", help="bankml tag of the imprinted actor"),
791
+ max_inquiries: int = typer.Option(5, "--n", help="number of recall probes"),
792
+ seed: int = typer.Option(0, "--seed", help="sampler seed (recorded; greedy at temperature 0)"),
793
+ num_predict: int = typer.Option(48, "--num-predict", help="tokens per utterance"),
794
+ system: str = typer.Option("", "--system", help="system turn prepended to every probe"),
795
+ base_url: str = typer.Option(
796
+ None, "--base-url",
797
+ help="bankml server (default MINDXTRAIN_BANKML_BASE_URL or http://127.0.0.1:18093/v1)",
798
+ ),
799
+ ) -> None:
800
+ """Measure an imprint through bankml: reproducible, receipt-auditable CPU probes.
801
+
802
+ Poses the script's user-turns to two tags served by `bankml serve --native` (temperature 0,
803
+ fixed seed, no penalties) and scores them with the same `score_imprint`. Every utterance
804
+ carries bankml's receipt (model / request / response sha256). NOT comparable with
805
+ `mindxtrain imprint` (the canonical gate decodes with repetition_penalty 1.3); the report
806
+ says so. Exit 3 if bankml refuses or errs, 4 if no imprint was detected.
807
+ """
808
+ import httpx
809
+
810
+ from mindxtrain.eval.imprint_bankml import imprint_via_bankml
811
+ from mindxtrain.operator.backends.bankml import BankmlError
812
+
813
+ cfg = load_config(config)
814
+ inquiries, baseline = _script_probes(cfg, max_inquiries)
815
+ console.print(f"[cyan]probing {len(inquiries)} inquiries through bankml (before/after)…[/cyan]")
816
+ try:
817
+ result = imprint_via_bankml(
818
+ before, after, inquiries, baseline, system=system or None, seed=seed,
819
+ num_predict=num_predict, base_url=base_url,
820
+ )
821
+ except (BankmlError, httpx.HTTPError) as exc:
822
+ console.print(f"[red]bankml imprint probe failed:[/red] {exc}")
823
+ raise typer.Exit(code=3) from exc
824
+ console.print_json(data=result.model_dump())
825
+ console.print(f"[yellow]{result.note}[/yellow]")
826
+ if not result.report.imprinted:
827
+ console.print("[yellow]no imprint detected (delta<=0 or no shift)[/yellow]")
828
+ raise typer.Exit(code=4)
829
+
830
+
831
  # ---- research (autoresearch search over one editable file) --------------
832
 
833
 
mindxtrain/deploy/bankml_push.py ADDED
@@ -0,0 +1,439 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Serve a trained model through bankml — the verified CPU engine.
2
+
3
+ `push_to_bankml` is the bankml twin of `ollama_push.push_to_ollama`: merge the LoRA (or take an
4
+ already-merged directory) → write a Modelfile that bankml can reproduce → `bankml create`, so the
5
+ tag answers on `bankml serve --native` (OpenAI `/v1` + Ollama `/api`, a receipt on every answer).
6
+
7
+ bankml (https://github.com/cryptoAGI/bankml) is reached **only** as a CLI subprocess here and as
8
+ HTTP in `operator/backends/bankml.py`; none of its code is vendored (clean-room policy).
9
+
10
+ What bankml takes, and what this module therefore refuses *before* running anything (it never
11
+ drops an instruction silently — a dropped PARAMETER would serve a model that answers differently
12
+ from the one described):
13
+
14
+ - **Taken:** `FROM` (the merged safetensors directory, which `bankml create` converts to GGUF F16
15
+ byte-identically to llama.cpp b11192, or a GGUF `bankml convert` already pinned), `SYSTEM`,
16
+ `MESSAGE`, `LICENSE`, `REQUIRES`, `PARAMETER` temperature / top_k / top_p / min_p / seed /
17
+ num_ctx / num_predict, and `stop`.
18
+ - **Refused:** `ADAPTER` (merge first — this module does it for you given `adapter_dir`), a
19
+ `TEMPLATE` other than the base's own, the penalties (repeat / presence / frequency,
20
+ repeat_last_n), mirostat*, typical_p, resource options (num_gpu, num_thread, num_batch,
21
+ num_keep, draft_num_predict), and any architecture other than Llama (SmolLM2 = mindx-genN):
22
+ bankml converts Llama safetensors only; Qwen3 it serves only as a pinned Q1_0/Q2_0_g64 GGUF.
23
+
24
+ `bankml create` / `bankml convert` arrive in bankml 0.3.5. An older binary is detected from
25
+ `bankml --help` (the verbs it lists) and reported as `status="bankml_too_old"`, never raised.
26
+
27
+ Every public function returns a result; nothing here raises past `BankmlPushResult`.
28
+ """
29
+
30
+ from __future__ import annotations
31
+
32
+ import json
33
+ import os
34
+ import re
35
+ import shutil
36
+ import subprocess
37
+ from collections.abc import Callable
38
+ from dataclasses import dataclass, field
39
+ from pathlib import Path
40
+ from typing import Literal
41
+
42
+ from mindxtrain.deploy.modelfile import ModelfileSpec, render_modelfile
43
+
44
+ BANKML_URL = "https://github.com/cryptoAGI/bankml"
45
+
46
+ # PARAMETERs bankml reproduces token-for-token (plus `stop`, carried on ModelfileSpec.stop).
47
+ BANKML_PARAMS: frozenset[str] = frozenset(
48
+ {"temperature", "top_k", "top_p", "min_p", "seed", "num_ctx", "num_predict"},
49
+ )
50
+
51
+ _PENALTY = frozenset({"repeat_penalty", "presence_penalty", "frequency_penalty", "repeat_last_n"})
52
+ _MIROSTAT = frozenset({"mirostat", "mirostat_tau", "mirostat_eta"})
53
+ _RESOURCE = frozenset({"num_gpu", "num_thread", "num_batch", "num_keep", "draft_num_predict"})
54
+
55
+ # Architectures `bankml convert` / `bankml create FROM <dir>` accept (config.json `architectures`).
56
+ BANKML_CONVERT_ARCHS: frozenset[str] = frozenset({"LlamaForCausalLM"})
57
+
58
+ _NAME = re.compile(r"^[a-z0-9_][a-z0-9._-]{0,127}$")
59
+ _SHA = r"[0-9a-f]{64}"
60
+
61
+ PushStatus = Literal[
62
+ "created", "refused", "bankml_missing", "bankml_too_old", "merge_failed", "failed", "error",
63
+ ]
64
+
65
+
66
+ def _param_refusal(name: str) -> str:
67
+ if name in _PENALTY:
68
+ return (f"PARAMETER {name}: penalties are not reproduced by bankml's verified sampler "
69
+ "(llama.cpp's top-k, top-p, min-p and temperature are); leave it out")
70
+ if name in _MIROSTAT:
71
+ return f"PARAMETER {name}: mirostat is not reproduced by bankml; leave it out"
72
+ if name == "typical_p":
73
+ return "PARAMETER typical_p: locally-typical sampling is not reproduced by bankml"
74
+ if name in _RESOURCE:
75
+ return (f"PARAMETER {name}: a resource option is not part of a bankml model "
76
+ "(bankml's answer does not depend on threads, batch or GPU layers)")
77
+ if name == "stop":
78
+ return "PARAMETER stop: pass stop sequences as ModelfileSpec.stop, not in parameters"
79
+ return f"PARAMETER {name}: not a parameter bankml reproduces"
80
+
81
+
82
+ # ---- capability probe -------------------------------------------------------
83
+
84
+
85
+ @dataclass(frozen=True)
86
+ class BankmlCapabilities:
87
+ """What the local `bankml` binary can do. `reason` is set when it cannot do what was asked."""
88
+
89
+ binary: str | None
90
+ version: str = ""
91
+ has_create: bool = False
92
+ has_convert: bool = False
93
+ reason: str = ""
94
+
95
+ @property
96
+ def found(self) -> bool:
97
+ return self.binary is not None
98
+
99
+
100
+ def bankml_capabilities(bankml_bin: str | None = None, *, timeout_s: float = 30.0) -> BankmlCapabilities:
101
+ """Locate `bankml` and read its version and verbs (`bankml version`, `bankml --help`).
102
+
103
+ The verbs listed in the usage text are the truth — an unreleased build may still say 0.3.4
104
+ and already carry `create`. Never raises.
105
+ """
106
+ binary = bankml_bin or shutil.which("bankml")
107
+ if not binary:
108
+ return BankmlCapabilities(
109
+ binary=None,
110
+ reason=(f"`bankml` not found on PATH; install it from {BANKML_URL} "
111
+ "(or pass --bankml-bin)"),
112
+ )
113
+ try:
114
+ ver = subprocess.run([binary, "version"], capture_output=True, text=True,
115
+ timeout=timeout_s, check=False)
116
+ usage = subprocess.run([binary, "--help"], capture_output=True, text=True,
117
+ timeout=timeout_s, check=False)
118
+ except (OSError, subprocess.SubprocessError) as exc:
119
+ return BankmlCapabilities(binary=binary, reason=f"cannot run {binary}: {exc}")
120
+ m = re.search(r"bankml\s+(\d+\.\d+\.\d+\S*)", ver.stdout + ver.stderr)
121
+ text = usage.stdout + usage.stderr
122
+ return BankmlCapabilities(
123
+ binary=binary,
124
+ version=m.group(1) if m else "",
125
+ has_create="bankml create" in text,
126
+ has_convert="bankml convert" in text,
127
+ )
128
+
129
+
130
+ # ---- the Modelfile subset ----------------------------------------------------
131
+
132
+
133
+ @dataclass(frozen=True)
134
+ class BankmlSanitizeResult:
135
+ """`ok` → `spec` is exactly what bankml will reproduce; else `refusals` say why, one per line."""
136
+
137
+ spec: ModelfileSpec
138
+ refusals: tuple[str, ...] = ()
139
+
140
+ @property
141
+ def ok(self) -> bool:
142
+ return not self.refusals
143
+
144
+
145
+ def bankml_sanitize(spec: ModelfileSpec, *, base_template: str | None = None) -> BankmlSanitizeResult:
146
+ """Check a ModelfileSpec against the subset `bankml create` reproduces.
147
+
148
+ Refuses (and names) every instruction bankml would not honour; it never drops one, because a
149
+ silently dropped PARAMETER serves a model that answers differently from the spec. A
150
+ `TEMPLATE` passes only when it equals `base_template` (the base GGUF's own chat template).
151
+ """
152
+ refusals: list[str] = []
153
+ if spec.adapter:
154
+ refusals.append("ADAPTER: bankml does not merge a LoRA at create time; merge it first "
155
+ "(push_to_bankml does, given adapter_dir) and FROM the merged directory")
156
+ if spec.template and spec.template != base_template:
157
+ refusals.append("TEMPLATE: bankml renders the base model's own chat template "
158
+ "(byte-identical to llama.cpp); a different template is not reproduced — "
159
+ "leave TEMPLATE out")
160
+ for name in sorted(spec.parameters):
161
+ if name not in BANKML_PARAMS:
162
+ refusals.append(_param_refusal(name))
163
+ for label, value in (("SYSTEM", spec.system), ("TEMPLATE", spec.template),
164
+ ("LICENSE", spec.license)):
165
+ if '"""' in value:
166
+ refusals.append(f'{label}: contains """, which a Modelfile cannot quote')
167
+ for stop in spec.stop:
168
+ if '"' in stop or "\n" in stop:
169
+ refusals.append(f"stop {stop!r}: a quote or newline cannot be written as PARAMETER stop")
170
+ return BankmlSanitizeResult(spec=spec, refusals=tuple(refusals))
171
+
172
+
173
+ def check_merged_arch(merged_dir: Path) -> str | None:
174
+ """None when `merged_dir/config.json` names an architecture bankml converts (or has none to
175
+ read — then bankml decides); else the refusal reason."""
176
+ try:
177
+ cfg = json.loads((merged_dir / "config.json").read_text(encoding="utf-8"))
178
+ except (OSError, ValueError):
179
+ return None
180
+ archs = cfg.get("architectures") or []
181
+ if not archs or any(a in BANKML_CONVERT_ARCHS for a in archs):
182
+ return None
183
+ return (f"architecture {', '.join(map(str, archs))}: bankml converts Llama-architecture "
184
+ "safetensors only (SmolLM2 / mindx-genN); Qwen3 it serves only as a pinned "
185
+ "Q1_0/Q2_0_g64 GGUF — use `serve --to ollama` or `--to vllm` for this model")
186
+
187
+
188
+ # Families mindXtrain trains that bankml cannot convert (checked by name before a merge is spent).
189
+ _UNCONVERTIBLE_FAMILIES = ("qwen", "mistral", "phi", "gemma", "glm", "deepseek", "instella")
190
+
191
+
192
+ def base_family_refusal(base_model: str) -> str | None:
193
+ """A cheap, name-based pre-check (before merging): None unless `base_model` names a family
194
+ bankml cannot convert. The merged `config.json` is still checked afterwards."""
195
+ low = base_model.lower()
196
+ if "smollm" in low or "llama" in low:
197
+ return None
198
+ hit = next((f for f in _UNCONVERTIBLE_FAMILIES if f in low), None)
199
+ if hit is None:
200
+ return None
201
+ return (f"base {base_model}: the {hit} family is not a Llama-architecture model bankml can "
202
+ "convert (bankml converts SmolLM2 / mindx-genN; Qwen3 only as a pinned "
203
+ "Q1_0/Q2_0_g64 GGUF) — use `serve --to ollama` or `--to vllm`")
204
+
205
+
206
+ def bankml_registry_dir(registry_dir: Path | None = None) -> Path:
207
+ """bankml's registry (pins + derived models): `--registry`, else $BANKML_FORKS, else
208
+ ~/.local/share/bankml/forks — the same resolution `bankml create` uses."""
209
+ if registry_dir is not None:
210
+ return Path(registry_dir).expanduser()
211
+ env = os.environ.get("BANKML_FORKS")
212
+ if env:
213
+ return Path(env).expanduser()
214
+ return Path.home() / ".local" / "share" / "bankml" / "forks"
215
+
216
+
217
+ # ---- push --------------------------------------------------------------------
218
+
219
+
220
+ @dataclass(frozen=True)
221
+ class BankmlPushResult:
222
+ """Outcome of `push_to_bankml`. `status == "created"` is the only success."""
223
+
224
+ status: PushStatus
225
+ tag: str
226
+ reason: str = ""
227
+ refusals: tuple[str, ...] = ()
228
+ bankml_version: str = ""
229
+ merged_dir: Path | None = None
230
+ gguf: Path | None = None
231
+ modelfile: Path | None = None
232
+ model_sha256: str = ""
233
+ digest: str = ""
234
+ output: str = ""
235
+ mindx_fallback_swap: dict[str, str] | None = None
236
+ log: tuple[str, ...] = field(default=())
237
+
238
+ @property
239
+ def ok(self) -> bool:
240
+ return self.status == "created"
241
+
242
+
243
+ def _run(cmd: list[str], timeout_s: float) -> tuple[int, str, str]:
244
+ proc = subprocess.run(cmd, capture_output=True, text=True, timeout=timeout_s, check=False)
245
+ return proc.returncode, proc.stdout or "", proc.stderr or ""
246
+
247
+
248
+ def push_to_bankml(
249
+ base_model: str,
250
+ tag: str,
251
+ *,
252
+ adapter_dir: Path | None = None,
253
+ merged_dir: Path | None = None,
254
+ system: str | None = None,
255
+ params: dict[str, float | int | str] | None = None,
256
+ stop: list[str] | None = None,
257
+ bankml_bin: str | None = None,
258
+ convert: bool = False,
259
+ work_dir: Path | None = None,
260
+ registry_dir: Path | None = None,
261
+ register_with_mindx: bool = False,
262
+ mindx_base_url: str | None = None,
263
+ sink: Callable[[str], None] | None = None,
264
+ timeout_s: float = 1800.0,
265
+ ) -> BankmlPushResult:
266
+ """Merge (if given an adapter) → Modelfile (bankml subset) → `bankml create <tag>`.
267
+
268
+ Give exactly one of `adapter_dir` (a PEFT LoRA over `base_model`; merging needs
269
+ `uv sync --extra ml`) or `merged_dir` (an already-merged HF safetensors directory).
270
+
271
+ `convert=False` (default) writes `FROM <merged dir>` and lets `bankml create` convert and pin
272
+ it. `convert=True` runs `bankml convert` first (GGUF + FORK.json into the registry) and
273
+ writes `FROM <gguf>` — the explicit, inspectable two-step.
274
+
275
+ The model sha256 recorded is the one bankml prints for the base it pinned and verified.
276
+ With `register_with_mindx`, mindX's fallback model is swapped to `{provider: "bankml",
277
+ model: tag}` (best-effort; a failure is logged, the push still stands).
278
+ """
279
+ lines: list[str] = []
280
+
281
+ def emit(line: str) -> None:
282
+ lines.append(line)
283
+ if sink:
284
+ sink(line)
285
+
286
+ def done(status: PushStatus, **kw: object) -> BankmlPushResult:
287
+ return BankmlPushResult(status=status, tag=tag, log=tuple(lines), **kw) # type: ignore[arg-type]
288
+
289
+ name = tag.strip().lower().removesuffix(":latest")
290
+ if not _NAME.match(name) or name.endswith(".gguf"):
291
+ return done("refused", reason=(f"tag {tag!r}: bankml names use letters, digits, '.', '_' "
292
+ "and '-' (at most 128, no ':tag', not ending in .gguf)"))
293
+ if (adapter_dir is None) == (merged_dir is None):
294
+ return done("error", reason="give exactly one of adapter_dir or merged_dir")
295
+
296
+ # 1. the subset — checked before anything expensive runs
297
+ spec = ModelfileSpec(from_model="<merged>", system=system or "", parameters=dict(params or {}),
298
+ stop=list(stop or []))
299
+ checked = bankml_sanitize(spec)
300
+ if not checked.ok:
301
+ for r in checked.refusals:
302
+ emit(f"[push-bankml] refuse: {r}")
303
+ return done("refused", reason="the Modelfile asks for what bankml does not reproduce",
304
+ refusals=checked.refusals)
305
+
306
+ # 2. the binary, and whether it has the verbs
307
+ caps = bankml_capabilities(bankml_bin)
308
+ if not caps.found:
309
+ return done("bankml_missing", reason=caps.reason)
310
+ if caps.reason:
311
+ return done("error", reason=caps.reason, bankml_version=caps.version)
312
+ missing = [v for v, have in (("create", caps.has_create), ("convert", caps.has_convert))
313
+ if (v == "create" or convert) and not have]
314
+ if missing:
315
+ return done("bankml_too_old", bankml_version=caps.version, reason=(
316
+ f"bankml {caps.version or '(unknown version)'} has no `{'`/`'.join(missing)}` verb "
317
+ f"(they arrive in bankml 0.3.5); upgrade from {BANKML_URL}"))
318
+ emit(f"[push-bankml] bankml {caps.version} at {caps.binary}")
319
+
320
+ work = Path(work_dir) if work_dir else (
321
+ (adapter_dir or merged_dir).parent / "bankml_push") # type: ignore[union-attr]
322
+
323
+ # 3. merge, when given an adapter
324
+ if adapter_dir is not None:
325
+ try:
326
+ from mindxtrain.deploy.ollama_push import merge_lora_adapter
327
+
328
+ merged = merge_lora_adapter(base_model, Path(adapter_dir), work / "merged", sink=emit)
329
+ except ImportError as exc:
330
+ return done("merge_failed", bankml_version=caps.version, reason=(
331
+ f"{exc} — merging a LoRA needs `uv sync --extra ml`"))
332
+ except Exception as exc: # a merge failure is a result, not a crash
333
+ return done("merge_failed", bankml_version=caps.version,
334
+ reason=f"{type(exc).__name__}: {exc}")
335
+ else:
336
+ merged = Path(merged_dir) # type: ignore[arg-type]
337
+ if not merged.is_dir():
338
+ return done("error", bankml_version=caps.version,
339
+ reason=f"merged_dir {merged} is not a directory")
340
+
341
+ arch_refusal = check_merged_arch(merged)
342
+ if arch_refusal:
343
+ emit(f"[push-bankml] refuse: {arch_refusal}")
344
+ return done("refused", bankml_version=caps.version, merged_dir=merged,
345
+ reason=arch_refusal, refusals=(arch_refusal,))
346
+
347
+ registry = bankml_registry_dir(registry_dir)
348
+ reg_args = ["--registry", str(registry)] if registry_dir is not None else []
349
+ gguf: Path | None = None
350
+ model_sha = ""
351
+ try:
352
+ # 4. optional explicit conversion
353
+ source = str(merged.resolve())
354
+ if convert:
355
+ registry.mkdir(parents=True, exist_ok=True)
356
+ gguf = registry / f"{name}-base-F16.gguf"
357
+ cmd = [str(caps.binary), "convert", source, "-o", str(gguf),
358
+ "--fork", str(registry / f"{gguf.name}.FORK.json"), "--source", source]
359
+ emit(f"[push-bankml] $ {' '.join(cmd)}")
360
+ code, out, err = _run(cmd, timeout_s)
361
+ if err.strip():
362
+ emit(err.rstrip())
363
+ if code != 0:
364
+ status: PushStatus = "refused" if code == 2 else "failed"
365
+ return done(status, bankml_version=caps.version, merged_dir=merged,
366
+ reason=_last_line(err) or f"bankml convert exited {code}",
367
+ output=(out + err)[-2000:])
368
+ m = re.search(rf"^({_SHA})\s", out, re.MULTILINE)
369
+ model_sha = m.group(1) if m else ""
370
+
371
+ # 5. the Modelfile and `bankml create`
372
+ spec = spec.model_copy(update={"from_model": str(gguf) if gguf else source})
373
+ modelfile = work / name / "Modelfile"
374
+ modelfile.parent.mkdir(parents=True, exist_ok=True)
375
+ modelfile.write_text(render_modelfile(spec), encoding="utf-8")
376
+ cmd = [str(caps.binary), "create", name, "-f", str(modelfile), *reg_args]
377
+ emit(f"[push-bankml] $ {' '.join(cmd)}")
378
+ code, out, err = _run(cmd, timeout_s)
379
+ except (OSError, subprocess.SubprocessError) as exc:
380
+ return done("error", bankml_version=caps.version, merged_dir=merged, gguf=gguf,
381
+ reason=f"{type(exc).__name__}: {exc}")
382
+ for chunk in (out, err):
383
+ if chunk.strip():
384
+ emit(chunk.rstrip())
385
+ if code != 0:
386
+ return done("refused" if code == 2 else "failed", bankml_version=caps.version,
387
+ merged_dir=merged, gguf=gguf, modelfile=modelfile,
388
+ reason=_last_line(err) or f"bankml create exited {code}",
389
+ output=(out + err)[-2000:])
390
+
391
+ both = out + "\n" + err
392
+ over = re.search(rf"over (\S+) \(sha256 ({_SHA})\)", both)
393
+ digest = re.search(rf"digest sha256:({_SHA})", both)
394
+ if over:
395
+ model_sha = over.group(2)
396
+ if gguf is None:
397
+ gguf = registry / over.group(1)
398
+ elif not model_sha and gguf is not None:
399
+ code_s, out_s, _ = _run([str(caps.binary), "sha256", str(gguf)], 600.0)
400
+ m = re.match(rf"({_SHA})\s", out_s)
401
+ model_sha = m.group(1) if (code_s == 0 and m) else ""
402
+
403
+ swap: dict[str, str] | None = None
404
+ if register_with_mindx:
405
+ try:
406
+ from mindxtrain.deploy.api_client import swap_mindx_fallback_model
407
+
408
+ swap = swap_mindx_fallback_model(provider="bankml", model=name, api_url=mindx_base_url)
409
+ emit(f"[push-bankml] mindX swap: {swap.get('previous', '?')} -> "
410
+ f"{swap.get('current', '?')}")
411
+ except Exception as exc: # best-effort, as push_to_ollama
412
+ emit(f"[push-bankml] mindX registration failed (push still ok): {exc}")
413
+ swap = None
414
+
415
+ emit(f"[push-bankml] created {name} over sha256 {model_sha or '(not reported)'}")
416
+ return done("created", bankml_version=caps.version, merged_dir=merged, gguf=gguf,
417
+ modelfile=modelfile, model_sha256=model_sha,
418
+ digest=digest.group(1) if digest else "", output=both.strip()[-2000:],
419
+ mindx_fallback_swap=swap)
420
+
421
+
422
+ def _last_line(text: str) -> str:
423
+ rows = [r.strip() for r in text.strip().splitlines() if r.strip()]
424
+ return rows[-1] if rows else ""
425
+
426
+
427
+ __all__ = [
428
+ "BANKML_CONVERT_ARCHS",
429
+ "BANKML_PARAMS",
430
+ "BankmlCapabilities",
431
+ "BankmlPushResult",
432
+ "BankmlSanitizeResult",
433
+ "bankml_capabilities",
434
+ "bankml_registry_dir",
435
+ "bankml_sanitize",
436
+ "base_family_refusal",
437
+ "check_merged_arch",
438
+ "push_to_bankml",
439
+ ]
mindxtrain/eval/imprint_bankml.py ADDED
@@ -0,0 +1,156 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Imprint probes through bankml — reproducible, receipt-auditable, cheap on a CPU.
2
+
3
+ The canonical imprint gate is `eval.imprint.probe_recall`: transformers greedy decoding with
4
+ `repetition_penalty=1.3` and `no_repeat_ngram_size=3`. Every number in an ascent log was measured
5
+ that way, and that gate stays canonical.
6
+
7
+ This module is a **second instrument**, not a replacement. It poses the same inquiries to tags
8
+ served by `bankml serve --native` (https://github.com/cryptoAGI/bankml) over Ollama's `/api/chat`
9
+ with `temperature 0`, a fixed `seed`, `num_predict 48` and **no penalties** — bankml refuses
10
+ penalties because its sampler reproduces llama.cpp b11192's temperature / top-k / top-p / min-p
11
+ token-for-token and nothing it cannot prove. What that buys:
12
+
13
+ - **Reproducible**: the same seed and weights give the same tokens, identical to llama.cpp.
14
+ - **Auditable**: every utterance carries a `bankml_receipt` — the model's sha256, the request's
15
+ and the response's — so a score is tied to exact weights, not to a tag name.
16
+ - **Cheap**: a 135M F16 actor answers on one CPU core without torch in this process.
17
+
18
+ Because the decoding differs (no 1.3 repetition penalty, no n-gram block), a score from here is
19
+ **NOT comparable** with the canonical gate's numbers. The report says so in `comparable_with`,
20
+ `note` and the `method` tag (`<scorer>/bankml-greedy`), so the two can never be mixed silently.
21
+
22
+ bankml is reached over HTTP only (clean-room policy); nothing here imports torch.
23
+ """
24
+
25
+ from __future__ import annotations
26
+
27
+ from typing import Any
28
+
29
+ import httpx
30
+ from pydantic import BaseModel, ConfigDict, Field
31
+
32
+ from mindxtrain.eval.imprint import ImprintReport, score_imprint
33
+ from mindxtrain.operator.backends.bankml import bankml_root_url, raise_for_bankml
34
+
35
+ NOT_COMPARABLE_NOTE = (
36
+ "bankml-greedy probe: temperature 0, fixed seed, no repetition penalty, no n-gram block. "
37
+ "NOT comparable with the canonical imprint gate (transformers greedy, repetition_penalty 1.3, "
38
+ "no_repeat_ngram_size 3); compare bankml-greedy scores only with bankml-greedy scores."
39
+ )
40
+
41
+
42
+ class BankmlDecoding(BaseModel):
43
+ """The exact decoding every probe was sent with — part of the evidence."""
44
+
45
+ model_config = ConfigDict(extra="forbid", frozen=True)
46
+
47
+ engine: str = "bankml"
48
+ endpoint: str = "/api/chat"
49
+ temperature: float = 0.0
50
+ seed: int = 0
51
+ num_predict: int = 48
52
+ penalties: str = "none (bankml refuses them; not reproduced)"
53
+
54
+
55
+ class BankmlProbe(BaseModel):
56
+ """One tag's utterances for the inquiries, each with bankml's receipt."""
57
+
58
+ model_config = ConfigDict(extra="forbid", frozen=True)
59
+
60
+ model: str
61
+ utterances: list[str]
62
+ receipts: list[dict[str, Any] | None]
63
+ model_sha256: list[str] = Field(description="distinct model_sha256 values the receipts name")
64
+
65
+
66
+ class BankmlImprintReport(BaseModel):
67
+ """An ImprintReport measured through bankml, with its decoding and receipts attached."""
68
+
69
+ model_config = ConfigDict(extra="forbid", frozen=True)
70
+
71
+ report: ImprintReport
72
+ decoding: BankmlDecoding
73
+ before: BankmlProbe
74
+ after: BankmlProbe
75
+ comparable_with: str = "bankml-greedy only"
76
+ canonical_gate: bool = False
77
+ note: str = NOT_COMPARABLE_NOTE
78
+
79
+
80
+ def probe_recall_via_bankml(
81
+ model_tag: str,
82
+ inquiries: list[str],
83
+ *,
84
+ system: str | None = None,
85
+ seed: int = 0,
86
+ num_predict: int = 48,
87
+ base_url: str | None = None,
88
+ timeout_s: float = 600.0,
89
+ transport: httpx.BaseTransport | None = None,
90
+ ) -> BankmlProbe:
91
+ """Ask `model_tag` each inquiry through bankml's `/api/chat`; keep text and receipt.
92
+
93
+ Same prompts, same decoding for every tag → a fair before/after comparison *within* this
94
+ instrument. Raises `BankmlRefusal` (HTTP 400, with bankml's reason) or `BankmlError`; it never
95
+ re-sends a refused request with altered options.
96
+ """
97
+ root = bankml_root_url(base_url)
98
+ options = {"temperature": 0.0, "seed": int(seed), "num_predict": int(num_predict)}
99
+ utterances: list[str] = []
100
+ receipts: list[dict[str, Any] | None] = []
101
+ with httpx.Client(timeout=timeout_s, transport=transport) as client:
102
+ for inquiry in inquiries:
103
+ msgs = [{"role": "user", "content": inquiry}]
104
+ if system and system.strip():
105
+ msgs.insert(0, {"role": "system", "content": system.strip()})
106
+ resp = client.post(
107
+ f"{root}/api/chat",
108
+ json={"model": model_tag, "messages": msgs, "stream": False, "options": options},
109
+ )
110
+ raise_for_bankml(resp)
111
+ body = resp.json()
112
+ utterances.append(((body.get("message") or {}).get("content") or "").strip())
113
+ rec = body.get("bankml_receipt")
114
+ receipts.append(rec if isinstance(rec, dict) else None)
115
+ shas = sorted({str(r["model_sha256"]) for r in receipts if r and r.get("model_sha256")})
116
+ return BankmlProbe(model=model_tag, utterances=utterances, receipts=receipts,
117
+ model_sha256=shas)
118
+
119
+
120
+ def imprint_via_bankml(
121
+ before_tag: str,
122
+ after_tag: str,
123
+ inquiries: list[str],
124
+ baseline: list[str],
125
+ *,
126
+ system: str | None = None,
127
+ seed: int = 0,
128
+ num_predict: int = 48,
129
+ base_url: str | None = None,
130
+ transport: httpx.BaseTransport | None = None,
131
+ ) -> BankmlImprintReport:
132
+ """Probe the base tag (before) and the imprinted tag (after) through bankml, then score with
133
+ the existing `score_imprint`. `report.method` is tagged `<scorer>/bankml-greedy`."""
134
+ kw: dict[str, Any] = {"system": system, "seed": seed, "num_predict": num_predict,
135
+ "base_url": base_url, "transport": transport}
136
+ before = probe_recall_via_bankml(before_tag, inquiries, **kw)
137
+ after = probe_recall_via_bankml(after_tag, inquiries, **kw)
138
+ report = score_imprint(inquiries, before.utterances, after.utterances,
139
+ baseline or before.utterances)
140
+ report = report.model_copy(update={"method": f"{report.method}/bankml-greedy"})
141
+ return BankmlImprintReport(
142
+ report=report,
143
+ decoding=BankmlDecoding(seed=int(seed), num_predict=int(num_predict)),
144
+ before=before,
145
+ after=after,
146
+ )
147
+
148
+
149
+ __all__ = [
150
+ "NOT_COMPARABLE_NOTE",
151
+ "BankmlDecoding",
152
+ "BankmlImprintReport",
153
+ "BankmlProbe",
154
+ "imprint_via_bankml",
155
+ "probe_recall_via_bankml",
156
+ ]
mindxtrain/eval/mei/record.py CHANGED
@@ -58,7 +58,7 @@ class InferenceEngineIdent(BaseModel):
58
 
59
  model_config = ConfigDict(extra="forbid", frozen=True)
60
 
61
- name: Literal["llama.cpp", "ollama", "vllm", "sglang", "transformers"] = Field(
62
  description="Which serving engine collected the timings.",
63
  )
64
  commit_sha: str = Field(min_length=1, description="Engine binary commit SHA or version string.")
 
58
 
59
  model_config = ConfigDict(extra="forbid", frozen=True)
60
 
61
+ name: Literal["llama.cpp", "ollama", "vllm", "sglang", "transformers", "bankml"] = Field(
62
  description="Which serving engine collected the timings.",
63
  )
64
  commit_sha: str = Field(min_length=1, description="Engine binary commit SHA or version string.")
mindxtrain/governance/panel.py CHANGED
@@ -37,10 +37,25 @@ _VERDICT_RE = re.compile(r"verdict\s*[:\-]?\s*(approve|reject|abstain)", re.IGNO
37
 
38
 
39
  def resolve_chat_base_url(base_url: str | None = None) -> str:
40
- """Resolve the OpenAI-compatible chat base URL the operator/backends use."""
 
 
 
 
 
 
41
  if base_url:
42
  return base_url.rstrip("/")
43
- for env in ("MINDXTRAIN_OPENAI_BASE_URL", "MINDXTRAIN_VLLM_BASE_URL", "MINDXTRAIN_OLLAMA_BASE_URL"):
 
 
 
 
 
 
 
 
 
44
  val = os.environ.get(env)
45
  if val:
46
  return val.rstrip("/")
 
37
 
38
 
39
  def resolve_chat_base_url(base_url: str | None = None) -> str:
40
+ """Resolve the OpenAI-compatible chat base URL the operator/backends use.
41
+
42
+ Order: an explicit `base_url`; bankml when `MINDXTRAIN_BACKEND=bankml` (its
43
+ configured URL, else its default 127.0.0.1:18093/v1); then the env chain
44
+ OPENAI -> VLLM -> OLLAMA -> BANKML (bankml only when its URL is set, and last,
45
+ so existing hosts resolve as before); else local ollama.
46
+ """
47
  if base_url:
48
  return base_url.rstrip("/")
49
+ if os.environ.get("MINDXTRAIN_BACKEND") == "bankml":
50
+ from mindxtrain.operator.backends.bankml import bankml_base_url
51
+
52
+ return bankml_base_url()
53
+ for env in (
54
+ "MINDXTRAIN_OPENAI_BASE_URL",
55
+ "MINDXTRAIN_VLLM_BASE_URL",
56
+ "MINDXTRAIN_OLLAMA_BASE_URL",
57
+ "MINDXTRAIN_BANKML_BASE_URL",
58
+ ):
59
  val = os.environ.get(env)
60
  if val:
61
  return val.rstrip("/")
mindxtrain/hf/extension.py CHANGED
@@ -137,13 +137,29 @@ def _card(repo_id: str, meta: dict[str, Any]) -> str:
137
  "the training log ships beside the weights.\n")
138
 
139
 
 
 
 
 
 
 
 
 
 
 
 
140
  def publish_generation(run_dir: Path | str, repo_id: str, *, token: str | None = None, private: bool = False,
141
  meta: dict[str, Any] | None = None, persona_system: str | None = None,
142
- include_merged: bool = True, dry_run: bool = False) -> dict[str, Any]:
 
143
  """A finished run as a model repo: merged weights at the root (if present), the LoRA delta under
144
  `adapter/`, `train.log`, a `Modelfile` for Ollama, and a card built from `meta`.
145
 
146
- `dry_run=True` reports exactly what would be uploaded and touches nothing."""
 
 
 
 
147
  run = Path(run_dir)
148
  if not run.is_dir():
149
  return {"ok": False, "reason": f"no run dir at {run}"}
@@ -170,9 +186,7 @@ def publish_generation(run_dir: Path | str, repo_id: str, *, token: str | None =
170
  except Exception:
171
  pass
172
  card = _card(repo_id, meta)
173
- modelfile = ("# ollama create <name> -f Modelfile (from this repo's directory)\nFROM .\n"
174
- + (f'SYSTEM """{persona_system}"""\n' if persona_system else "")
175
- + 'PARAMETER temperature 0.7\nPARAMETER repeat_penalty 1.3\nPARAMETER stop "<|im_end|>"\n')
176
  plan = {"repo": repo_id, "private": private, "files": [*sorted(staged), "README.md", "Modelfile"],
177
  "bytes": sum(f.stat().st_size for f in staged.values())}
178
  if dry_run:
 
137
  "the training log ships beside the weights.\n")
138
 
139
 
140
+ def published_modelfile(persona_system: str | None = None, *,
141
+ repeat_penalty: float | None = 1.3) -> str:
142
+ """The Modelfile shipped beside a published run. `repeat_penalty=None` leaves the penalty out,
143
+ so engines that refuse penalties (bankml) load the same file."""
144
+ return ("# ollama create <name> -f Modelfile (from this repo's directory)\nFROM .\n"
145
+ + (f'SYSTEM """{persona_system}"""\n' if persona_system else "")
146
+ + "PARAMETER temperature 0.7\n"
147
+ + (f"PARAMETER repeat_penalty {repeat_penalty}\n" if repeat_penalty is not None else "")
148
+ + 'PARAMETER stop "<|im_end|>"\n')
149
+
150
+
151
  def publish_generation(run_dir: Path | str, repo_id: str, *, token: str | None = None, private: bool = False,
152
  meta: dict[str, Any] | None = None, persona_system: str | None = None,
153
+ include_merged: bool = True, dry_run: bool = False,
154
+ repeat_penalty: float | None = 1.3) -> dict[str, Any]:
155
  """A finished run as a model repo: merged weights at the root (if present), the LoRA delta under
156
  `adapter/`, `train.log`, a `Modelfile` for Ollama, and a card built from `meta`.
157
 
158
+ `dry_run=True` reports exactly what would be uploaded and touches nothing.
159
+
160
+ `repeat_penalty` is written into the Modelfile (default 1.3, the imprint gate's decoding);
161
+ `None` leaves it out, so the same Modelfile also loads in engines that refuse penalties
162
+ (bankml reproduces temperature/top-k/top-p/min-p, not penalties)."""
163
  run = Path(run_dir)
164
  if not run.is_dir():
165
  return {"ok": False, "reason": f"no run dir at {run}"}
 
186
  except Exception:
187
  pass
188
  card = _card(repo_id, meta)
189
+ modelfile = published_modelfile(persona_system, repeat_penalty=repeat_penalty)
 
 
190
  plan = {"repo": repo_id, "private": private, "files": [*sorted(staged), "README.md", "Modelfile"],
191
  "bytes": sum(f.stat().st_size for f in staged.values())}
192
  if dry_run:
mindxtrain/models/registry.py CHANGED
@@ -12,7 +12,7 @@ from __future__ import annotations
12
 
13
  from abc import ABC, abstractmethod
14
  from collections.abc import AsyncIterator, Callable
15
- from typing import Literal
16
 
17
  from pydantic import BaseModel, ConfigDict, Field
18
 
@@ -43,6 +43,10 @@ class ChatResponse(BaseModel):
43
  finish_reason: Literal["stop", "length", "error"] = "stop"
44
  prompt_tokens: int = 0
45
  completion_tokens: int = 0
 
 
 
 
46
 
47
 
48
  class Backend(ABC):
@@ -155,6 +159,7 @@ from mindxtrain.models import glm51 as _glm51 # noqa: E402, F401
155
  from mindxtrain.models import mistral3 as _mistral3 # noqa: E402, F401
156
  from mindxtrain.models import phi4_mini as _phi4_mini # noqa: E402, F401
157
  from mindxtrain.models import qwen35 as _qwen35 # noqa: E402, F401
 
158
  from mindxtrain.operator.backends import ollama as _ollama # noqa: E402, F401
159
  from mindxtrain.operator.backends import openai_compat as _openai_compat # noqa: E402, F401
160
  from mindxtrain.operator.backends import vllm as _vllm # noqa: E402, F401
 
12
 
13
  from abc import ABC, abstractmethod
14
  from collections.abc import AsyncIterator, Callable
15
+ from typing import Any, Literal
16
 
17
  from pydantic import BaseModel, ConfigDict, Field
18
 
 
43
  finish_reason: Literal["stop", "length", "error"] = "stop"
44
  prompt_tokens: int = 0
45
  completion_tokens: int = 0
46
+ # Engine-issued provenance for this answer, when the engine gives one (bankml's
47
+ # `bankml_receipt`: model_sha256, request_sha256, response_sha256, engine, tokens,
48
+ # wall_ms, ...). None for backends that issue no receipt.
49
+ receipt: dict[str, Any] | None = None
50
 
51
 
52
  class Backend(ABC):
 
159
  from mindxtrain.models import mistral3 as _mistral3 # noqa: E402, F401
160
  from mindxtrain.models import phi4_mini as _phi4_mini # noqa: E402, F401
161
  from mindxtrain.models import qwen35 as _qwen35 # noqa: E402, F401
162
+ from mindxtrain.operator.backends import bankml as _bankml # noqa: E402, F401
163
  from mindxtrain.operator.backends import ollama as _ollama # noqa: E402, F401
164
  from mindxtrain.operator.backends import openai_compat as _openai_compat # noqa: E402, F401
165
  from mindxtrain.operator.backends import vllm as _vllm # noqa: E402, F401
mindxtrain/operator/app.py CHANGED
@@ -104,12 +104,45 @@ def _vllm_first_model() -> str | None:
104
  return None
105
 
106
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
107
  def backend_reachable(name: str) -> bool:
108
  """Live probe for a backend by name. Used by both /health and /coach health."""
109
  if name == "ollama":
110
  return _ollama_reachable()
111
  if name == "vllm":
112
  return _vllm_reachable()
 
 
113
  # openai_compat and unknown backends: we don't have a generic probe,
114
  # so the chat-completions failure path remains the authoritative signal.
115
  return False
@@ -121,6 +154,8 @@ def backend_first_model(name: str) -> str | None:
121
  return ollama_first_model()
122
  if name == "vllm":
123
  return _vllm_first_model()
 
 
124
  return None
125
 
126
 
@@ -130,7 +165,10 @@ def resolve_backend_name() -> str:
130
  Resolution order:
131
  1. Explicit `MINDXTRAIN_BACKEND` env var (canonical).
132
  2. Legacy `AUTOMINDX_BACKEND` (back-compat with the pre-rename code).
133
- 3. Auto-detect: ollama if reachable on localhost:11434, else vllm.
 
 
 
134
  """
135
  explicit = (
136
  os.environ.get("MINDXTRAIN_BACKEND")
@@ -140,9 +178,43 @@ def resolve_backend_name() -> str:
140
  return explicit
141
  if _ollama_reachable():
142
  return "ollama"
 
 
143
  return "vllm"
144
 
145
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
146
  def ollama_first_model() -> str | None:
147
  """Return the name of the first model ollama lists, or None on failure.
148
 
@@ -275,28 +347,21 @@ async def readyz() -> dict[str, object]:
275
 
276
  @app.post("/v1/chat/completions", response_model=ChatResponse)
277
  async def chat_completions(request: ChatRequest) -> ChatResponse:
 
 
278
  backend_name = resolve_backend_name()
279
- backend_kwargs: dict[str, object] = {}
280
- if backend_name == "vllm":
281
- backend_kwargs["base_url"] = os.environ.get(
282
- "MINDXTRAIN_VLLM_BASE_URL",
283
- os.environ.get("AUTOMINDX_VLLM_BASE_URL", "http://localhost:8000/v1"),
284
- )
285
- elif backend_name == "ollama":
286
- backend_kwargs["base_url"] = os.environ.get(
287
- "MINDXTRAIN_OLLAMA_BASE_URL", "http://localhost:11434/v1",
288
- )
289
- elif backend_name == "openai_compat":
290
- backend_kwargs["base_url"] = os.environ["MINDXTRAIN_OPENAI_BASE_URL"]
291
- backend_kwargs["api_key"] = os.environ.get("MINDXTRAIN_OPENAI_API_KEY", "")
292
 
293
  try:
294
- backend = build_backend(backend_name, **backend_kwargs)
295
  return await backend.chat(request)
296
  except NotImplementedError as exc:
297
  raise HTTPException(status_code=501, detail=str(exc)) from exc
298
  except KeyError as exc:
299
  raise HTTPException(status_code=400, detail=str(exc)) from exc
 
 
 
300
 
301
 
302
  @app.post("/v1/agentic")
 
104
  return None
105
 
106
 
107
+ def _bankml_reachable(timeout_s: float = 1.0) -> bool:
108
+ """Probe `bankml serve --native` at MINDXTRAIN_BANKML_BASE_URL.
109
+
110
+ Hits the server's own `GET /bankml` identity endpoint (not `/health`, which any
111
+ llama-server-shaped engine answers), so a 200 means bankml specifically is there.
112
+ """
113
+ from mindxtrain.operator.backends.bankml import bankml_root_url
114
+
115
+ try:
116
+ with httpx.Client(timeout=timeout_s) as client:
117
+ return client.get(bankml_root_url() + "/bankml").status_code == 200
118
+ except (httpx.HTTPError, OSError):
119
+ return False
120
+
121
+
122
+ def _bankml_first_model() -> str | None:
123
+ """The model id bankml lists at `/v1/models` (the resident or startup model), or None."""
124
+ from mindxtrain.operator.backends.bankml import bankml_base_url
125
+
126
+ try:
127
+ with httpx.Client(timeout=1.0) as client:
128
+ resp = client.get(bankml_base_url() + "/models")
129
+ if resp.status_code != 200:
130
+ return None
131
+ models = resp.json().get("data", [])
132
+ first = models[0] if models else None
133
+ return first.get("id") if isinstance(first, dict) else None
134
+ except (httpx.HTTPError, OSError, ValueError, IndexError):
135
+ return None
136
+
137
+
138
  def backend_reachable(name: str) -> bool:
139
  """Live probe for a backend by name. Used by both /health and /coach health."""
140
  if name == "ollama":
141
  return _ollama_reachable()
142
  if name == "vllm":
143
  return _vllm_reachable()
144
+ if name == "bankml":
145
+ return _bankml_reachable()
146
  # openai_compat and unknown backends: we don't have a generic probe,
147
  # so the chat-completions failure path remains the authoritative signal.
148
  return False
 
154
  return ollama_first_model()
155
  if name == "vllm":
156
  return _vllm_first_model()
157
+ if name == "bankml":
158
+ return _bankml_first_model()
159
  return None
160
 
161
 
 
165
  Resolution order:
166
  1. Explicit `MINDXTRAIN_BACKEND` env var (canonical).
167
  2. Legacy `AUTOMINDX_BACKEND` (back-compat with the pre-rename code).
168
+ 3. Auto-detect: ollama if reachable on localhost:11434; else bankml if
169
+ `bankml serve` answers on its port *and* vLLM does not; else vllm.
170
+ bankml comes after ollama and vllm, so a host that ran ollama or vllm
171
+ before bankml existed resolves exactly as it did.
172
  """
173
  explicit = (
174
  os.environ.get("MINDXTRAIN_BACKEND")
 
178
  return explicit
179
  if _ollama_reachable():
180
  return "ollama"
181
+ if _bankml_reachable() and not _vllm_reachable():
182
+ return "bankml"
183
  return "vllm"
184
 
185
 
186
+ def backend_kwargs(name: str, *, strict: bool = False) -> dict[str, object]:
187
+ """Constructor kwargs (base URL, key) for the backend registered as `name`.
188
+
189
+ One place for the per-backend env lookups the operator chat route and the Coach
190
+ chat stream both need. `strict=True` keeps the operator's original contract that
191
+ `openai_compat` without `MINDXTRAIN_OPENAI_BASE_URL` is a configuration error
192
+ (KeyError); the Coach passes `strict=False` and gets an empty URL instead.
193
+ """
194
+ kwargs: dict[str, object] = {}
195
+ if name == "vllm":
196
+ kwargs["base_url"] = os.environ.get(
197
+ "MINDXTRAIN_VLLM_BASE_URL",
198
+ os.environ.get("AUTOMINDX_VLLM_BASE_URL", "http://localhost:8000/v1"),
199
+ )
200
+ elif name == "ollama":
201
+ kwargs["base_url"] = os.environ.get(
202
+ "MINDXTRAIN_OLLAMA_BASE_URL", "http://localhost:11434/v1",
203
+ )
204
+ elif name == "bankml":
205
+ from mindxtrain.operator.backends.bankml import bankml_base_url
206
+
207
+ kwargs["base_url"] = bankml_base_url()
208
+ elif name == "openai_compat":
209
+ kwargs["base_url"] = (
210
+ os.environ["MINDXTRAIN_OPENAI_BASE_URL"]
211
+ if strict
212
+ else os.environ.get("MINDXTRAIN_OPENAI_BASE_URL", "")
213
+ )
214
+ kwargs["api_key"] = os.environ.get("MINDXTRAIN_OPENAI_API_KEY", "")
215
+ return kwargs
216
+
217
+
218
  def ollama_first_model() -> str | None:
219
  """Return the name of the first model ollama lists, or None on failure.
220
 
 
347
 
348
  @app.post("/v1/chat/completions", response_model=ChatResponse)
349
  async def chat_completions(request: ChatRequest) -> ChatResponse:
350
+ from mindxtrain.operator.backends.bankml import BankmlRefusal
351
+
352
  backend_name = resolve_backend_name()
353
+ kwargs = backend_kwargs(backend_name, strict=True)
 
 
 
 
 
 
 
 
 
 
 
 
354
 
355
  try:
356
+ backend = build_backend(backend_name, **kwargs)
357
  return await backend.chat(request)
358
  except NotImplementedError as exc:
359
  raise HTTPException(status_code=501, detail=str(exc)) from exc
360
  except KeyError as exc:
361
  raise HTTPException(status_code=400, detail=str(exc)) from exc
362
+ except BankmlRefusal as exc:
363
+ # bankml's reason, passed through verbatim; never retried with altered params.
364
+ raise HTTPException(status_code=400, detail=str(exc)) from exc
365
 
366
 
367
  @app.post("/v1/agentic")
mindxtrain/operator/backends/bankml.py ADDED
@@ -0,0 +1,197 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """bankml backend — the verified CPU engine, reached over HTTP only.
2
+
3
+ bankml (https://github.com/cryptoAGI/bankml, Apache-2.0 OR MIT) is a zero-dependency Rust
4
+ runtime whose `bankml serve --native` answers OpenAI `/v1/chat/completions` and Ollama `/api/*`
5
+ from its own forward pass, token-identical to llama.cpp b11192, on 127.0.0.1:18093 by default.
6
+ mindXtrain never vendors it: this module speaks its documented wire protocol and nothing else
7
+ (clean-room policy, CLAUDE.md).
8
+
9
+ What this backend adds over `openai_compat`:
10
+
11
+ - **Receipts.** Every bankml answer carries a `bankml_receipt` (model_sha256, request_sha256,
12
+ response_sha256, engine, tokens, wall_ms, ...). Non-streamed it is a top-level object and lands
13
+ in `ChatResponse.receipt`; streamed it arrives as one extra `data: {"bankml_receipt": ...}`
14
+ event before `data: [DONE]`. Either way the latest receipt is also kept on `last_receipt`.
15
+ - **Refusals are typed, never retried.** bankml answers HTTP 400 with a plain-text reason when a
16
+ request asks for something its verified forward pass does not reproduce (repeat / presence /
17
+ frequency penalties, mirostat, typical_p, tools, an unknown architecture, Q8_0 / Q4_K / BF16).
18
+ That becomes `BankmlRefusal(reason)`. The request is never re-sent with altered parameters:
19
+ a changed request would be a different, unreceipted question.
20
+
21
+ Env: `MINDXTRAIN_BANKML_BASE_URL` (default `http://127.0.0.1:18093/v1`).
22
+ """
23
+
24
+ from __future__ import annotations
25
+
26
+ import json
27
+ import os
28
+ from collections.abc import AsyncIterator
29
+ from typing import Any
30
+
31
+ import httpx
32
+
33
+ from mindxtrain.models.registry import ChatRequest, ChatResponse, register_backend
34
+ from mindxtrain.operator.backends.openai_compat import OpenAICompatBackend
35
+
36
+ DEFAULT_BANKML_BASE_URL = "http://127.0.0.1:18093/v1"
37
+
38
+
39
+ def bankml_base_url() -> str:
40
+ """The configured bankml OpenAI base URL (`.../v1`)."""
41
+ return os.environ.get("MINDXTRAIN_BANKML_BASE_URL", DEFAULT_BANKML_BASE_URL).rstrip("/")
42
+
43
+
44
+ def bankml_root_url(base_url: str | None = None) -> str:
45
+ """The server root (no `/v1`) — where `/health`, `/bankml` and `/api/*` live."""
46
+ return (base_url or bankml_base_url()).rstrip("/").removesuffix("/v1")
47
+
48
+
49
+ class BankmlError(RuntimeError):
50
+ """bankml answered, but not with an answer (an engine error, a broken stream)."""
51
+
52
+ label = "error"
53
+
54
+ def __init__(self, reason: str, status_code: int = 0) -> None:
55
+ self.reason = reason.strip()
56
+ self.status_code = status_code
57
+ super().__init__(f"bankml {self.label} ({status_code}): {self.reason}")
58
+
59
+
60
+ class BankmlRefusal(BankmlError):
61
+ """bankml refused the request (HTTP 400) and said why.
62
+
63
+ The reason is bankml's own text, e.g. ``mirostat: not reproduced; ...``. Callers must not
64
+ retry with parameters stripped or altered — change the request deliberately instead.
65
+ """
66
+
67
+ label = "refused"
68
+
69
+ def __init__(self, reason: str, status_code: int = 400) -> None:
70
+ super().__init__(reason, status_code)
71
+
72
+
73
+ def refusal_reason(resp: httpx.Response) -> str:
74
+ """bankml's reason from an error body: `/v1` sends text/plain, `/api` sends `{"error": ...}`."""
75
+ text = resp.text or ""
76
+ try:
77
+ body = json.loads(text)
78
+ except ValueError:
79
+ return text.strip() or f"HTTP {resp.status_code}"
80
+ if isinstance(body, dict) and body.get("error"):
81
+ err = body["error"]
82
+ return str(err.get("message", err) if isinstance(err, dict) else err)
83
+ return text.strip() or f"HTTP {resp.status_code}"
84
+
85
+
86
+ def raise_for_bankml(resp: httpx.Response) -> None:
87
+ """400 → `BankmlRefusal`; any other non-2xx → `BankmlError`."""
88
+ if resp.status_code == 400:
89
+ raise BankmlRefusal(refusal_reason(resp))
90
+ if resp.status_code >= 400:
91
+ raise BankmlError(refusal_reason(resp), resp.status_code)
92
+
93
+
94
+ @register_backend("bankml")
95
+ class BankmlBackend(OpenAICompatBackend):
96
+ """OpenAI-compatible client for `bankml serve --native`, keeping receipts and refusals."""
97
+
98
+ name = "bankml"
99
+
100
+ def __init__(
101
+ self,
102
+ base_url: str | None = None,
103
+ timeout_s: float = 300.0,
104
+ *,
105
+ seed: int | None = None,
106
+ transport: httpx.AsyncBaseTransport | None = None,
107
+ ) -> None:
108
+ super().__init__(base_url=base_url or bankml_base_url(), api_key=None, timeout_s=timeout_s)
109
+ # `api_key or env` in the parent would pick up an OpenAI key; bankml is loopback, keyless.
110
+ self.api_key = ""
111
+ self.seed = seed
112
+ self._transport = transport
113
+ self.last_receipt: dict[str, Any] | None = None
114
+
115
+ def _client(self) -> httpx.AsyncClient:
116
+ return httpx.AsyncClient(timeout=self.timeout_s, transport=self._transport)
117
+
118
+ def _payload(self, request: ChatRequest, *, stream: bool) -> dict[str, object]:
119
+ payload = super()._payload(request, stream=stream)
120
+ if self.seed is not None:
121
+ payload["seed"] = self.seed
122
+ return payload
123
+
124
+ async def chat(self, request: ChatRequest) -> ChatResponse:
125
+ async with self._client() as client:
126
+ resp = await client.post(
127
+ f"{self.base_url}/chat/completions",
128
+ json=self._payload(request, stream=False),
129
+ headers=self._headers(),
130
+ )
131
+ raise_for_bankml(resp)
132
+ data = resp.json()
133
+ choice = (data.get("choices") or [{}])[0]
134
+ usage = data.get("usage") or {}
135
+ receipt = data.get("bankml_receipt")
136
+ self.last_receipt = receipt if isinstance(receipt, dict) else None
137
+ finish = choice.get("finish_reason") or "stop"
138
+ return ChatResponse(
139
+ model=data.get("model", request.model),
140
+ content=(choice.get("message") or {}).get("content", "") or "",
141
+ finish_reason=finish if finish in ("stop", "length", "error") else "stop",
142
+ prompt_tokens=int(usage.get("prompt_tokens", 0)),
143
+ completion_tokens=int(usage.get("completion_tokens", 0)),
144
+ receipt=self.last_receipt,
145
+ )
146
+
147
+ async def stream_chat(self, request: ChatRequest) -> AsyncIterator[str]:
148
+ self.last_receipt = None
149
+
150
+ async def _gen() -> AsyncIterator[str]:
151
+ async with self._client() as client:
152
+ async with client.stream(
153
+ "POST",
154
+ f"{self.base_url}/chat/completions",
155
+ json=self._payload(request, stream=True),
156
+ headers=self._headers(),
157
+ ) as resp:
158
+ if resp.status_code >= 400:
159
+ await resp.aread()
160
+ raise_for_bankml(resp)
161
+ async for raw in resp.aiter_lines():
162
+ if not raw or not raw.startswith("data:"):
163
+ continue
164
+ data = raw[5:].strip()
165
+ if data == "[DONE]":
166
+ return
167
+ try:
168
+ chunk = json.loads(data)
169
+ except json.JSONDecodeError:
170
+ continue
171
+ if not isinstance(chunk, dict):
172
+ continue
173
+ if "bankml_receipt" in chunk:
174
+ rec = chunk["bankml_receipt"]
175
+ self.last_receipt = rec if isinstance(rec, dict) else None
176
+ continue
177
+ if "error" in chunk:
178
+ # bankml stops a stream it cannot finish with `data: {"error": ...}`
179
+ raise BankmlError(str(chunk["error"]), resp.status_code)
180
+ delta = (chunk.get("choices") or [{}])[0].get("delta", {})
181
+ token = delta.get("content")
182
+ if token:
183
+ yield token
184
+
185
+ return _gen()
186
+
187
+
188
+ __all__ = [
189
+ "DEFAULT_BANKML_BASE_URL",
190
+ "BankmlBackend",
191
+ "BankmlError",
192
+ "BankmlRefusal",
193
+ "bankml_base_url",
194
+ "bankml_root_url",
195
+ "raise_for_bankml",
196
+ "refusal_reason",
197
+ ]
mindxtrain/operator/coach/api.py CHANGED
@@ -1081,23 +1081,12 @@ async def api_models() -> dict[str, Any]:
1081
 
1082
 
1083
  def _resolve_chat_backend() -> Any:
1084
- """Build the active chat backend (ollama / vllm / openai_compat)."""
1085
  from mindxtrain.models.registry import build_backend
1086
- from mindxtrain.operator.app import resolve_backend_name
1087
 
1088
  name = resolve_backend_name()
1089
- kwargs: dict[str, Any] = {}
1090
- if name == "vllm":
1091
- kwargs["base_url"] = os.environ.get(
1092
- "MINDXTRAIN_VLLM_BASE_URL",
1093
- os.environ.get("AUTOMINDX_VLLM_BASE_URL", "http://localhost:8000/v1"),
1094
- )
1095
- elif name == "ollama":
1096
- kwargs["base_url"] = os.environ.get("MINDXTRAIN_OLLAMA_BASE_URL", "http://localhost:11434/v1")
1097
- elif name == "openai_compat":
1098
- kwargs["base_url"] = os.environ.get("MINDXTRAIN_OPENAI_BASE_URL", "")
1099
- kwargs["api_key"] = os.environ.get("MINDXTRAIN_OPENAI_API_KEY", "")
1100
- return build_backend(name, **kwargs)
1101
 
1102
 
1103
  @router.post("/api/chat/stream")
 
1081
 
1082
 
1083
  def _resolve_chat_backend() -> Any:
1084
+ """Build the active chat backend (ollama / vllm / bankml / openai_compat)."""
1085
  from mindxtrain.models.registry import build_backend
1086
+ from mindxtrain.operator.app import backend_kwargs, resolve_backend_name
1087
 
1088
  name = resolve_backend_name()
1089
+ return build_backend(name, **backend_kwargs(name))
 
 
 
 
 
 
 
 
 
 
 
1090
 
1091
 
1092
  @router.post("/api/chat/stream")
mindxtrain/ui/app.py CHANGED
@@ -241,7 +241,7 @@ def build() -> gr.Blocks:
241
  with gr.Row():
242
  s_cfg = gr.Textbox(value="run.yaml", label="config", scale=2)
243
  s_ckpt = gr.Textbox(value="", label="checkpoint", scale=2)
244
- s_to = gr.Radio(["ollama", "vllm"], value="ollama", label="to", scale=1)
245
  s_tag = gr.Textbox(value="", label="tag", scale=1)
246
  s_btn = gr.Button("serve", variant="primary", scale=1)
247
  with gr.Group(visible=False) as adv_serve:
 
241
  with gr.Row():
242
  s_cfg = gr.Textbox(value="run.yaml", label="config", scale=2)
243
  s_ckpt = gr.Textbox(value="", label="checkpoint", scale=2)
244
+ s_to = gr.Radio(["ollama", "bankml", "vllm"], value="ollama", label="to", scale=1)
245
  s_tag = gr.Textbox(value="", label="tag", scale=1)
246
  s_btn = gr.Button("serve", variant="primary", scale=1)
247
  with gr.Group(visible=False) as adv_serve:
mindxtrain/ui/console.py CHANGED
@@ -36,6 +36,31 @@ DEFAULTS: dict[str, Any] = {
36
  GATE_DECODING: dict[str, Any] = {"temperature": 0.0, "repeat_penalty": 1.3, "top_p": 1.0, "top_k": 0}
37
 
38
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
39
  def models(host: str = "", timeout: float = 4.0) -> list[str]:
40
  """Tags the daemon is serving, local first. Empty when it is down — never an exception."""
41
  try:
@@ -95,10 +120,11 @@ def unreachable(host: str = "") -> str:
95
  def chat(messages: list[dict[str, str]], model: str, *, options: dict[str, Any] | None = None,
96
  keep_alive: str = "10m", host: str = "", timeout: float = 900.0) -> Iterator[tuple[str, dict[str, Any]]]:
97
  """Stream `(text_so_far, stats)`. `stats` is empty until the final object, which carries the real
98
- counts. Options are passed through verbatim; the engine implements them, so none are dropped."""
 
99
  host = (host or HOST).rstrip("/")
100
  body = {"model": model, "messages": messages, "stream": True, "keep_alive": keep_alive,
101
- "options": {k: v for k, v in (options or {}).items() if v not in (None, "", [])}}
102
  req = urllib.request.Request(f"{host}/api/chat", data=json.dumps(body).encode(), method="POST",
103
  headers={"Content-Type": "application/json"})
104
  acc = ""
 
36
  GATE_DECODING: dict[str, Any] = {"temperature": 0.0, "repeat_penalty": 1.3, "top_p": 1.0, "top_k": 0}
37
 
38
 
39
+ # Options whose value here is also the engine's own default when the key is absent. Sending them
40
+ # changes nothing for Ollama, but an engine that refuses penalties / mirostat outright (bankml:
41
+ # "not reproduced") would refuse the whole request — so `chat` leaves them out at these values.
42
+ # A deliberate non-default value is always sent; the engine then honours or refuses it, visibly.
43
+ ENGINE_DEFAULTS: dict[str, Any] = {
44
+ "repeat_penalty": 1.1, "presence_penalty": 0.0, "frequency_penalty": 0.0, "typical_p": 1.0,
45
+ "mirostat": 0,
46
+ }
47
+
48
+
49
+ def wire_options(options: dict[str, Any] | None) -> dict[str, Any]:
50
+ """The options actually sent: empties dropped, and penalty/mirostat keys left at the engine's
51
+ own default omitted (mirostat's tau/eta only matter, and are only sent, when mirostat is on)."""
52
+ opts = {k: v for k, v in (options or {}).items() if v not in (None, "", [])}
53
+ miro_on = bool(opts.get("mirostat"))
54
+ out: dict[str, Any] = {}
55
+ for k, v in opts.items():
56
+ if k in ("mirostat_tau", "mirostat_eta") and not miro_on:
57
+ continue
58
+ if k in ENGINE_DEFAULTS and v == ENGINE_DEFAULTS[k]:
59
+ continue
60
+ out[k] = v
61
+ return out
62
+
63
+
64
  def models(host: str = "", timeout: float = 4.0) -> list[str]:
65
  """Tags the daemon is serving, local first. Empty when it is down — never an exception."""
66
  try:
 
120
  def chat(messages: list[dict[str, str]], model: str, *, options: dict[str, Any] | None = None,
121
  keep_alive: str = "10m", host: str = "", timeout: float = 900.0) -> Iterator[tuple[str, dict[str, Any]]]:
122
  """Stream `(text_so_far, stats)`. `stats` is empty until the final object, which carries the real
123
+ counts. Options are passed through as `wire_options` leaves them: a penalty or mirostat key at
124
+ the engine's own default is omitted (identical for Ollama, and accepted by bankml)."""
125
  host = (host or HOST).rstrip("/")
126
  body = {"model": model, "messages": messages, "stream": True, "keep_alive": keep_alive,
127
+ "options": wire_options(options)}
128
  req = urllib.request.Request(f"{host}/api/chat", data=json.dumps(body).encode(), method="POST",
129
  headers={"Content-Type": "application/json"})
130
  acc = ""
tests/conftest.py ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Suite-wide hermeticity for the operator's backend auto-detect.
2
+
3
+ `resolve_backend_name()` probes `bankml serve` on 127.0.0.1:18093 when ollama is not
4
+ reachable. A developer box that happens to run bankml would otherwise change what the
5
+ auto-detect tests resolve to, so the probe is pinned to "absent" for every test; the
6
+ bankml tests that exercise the probe itself monkeypatch it back explicitly.
7
+ """
8
+
9
+ from __future__ import annotations
10
+
11
+ import pytest
12
+
13
+
14
+ @pytest.fixture(autouse=True)
15
+ def _no_live_bankml_probe(monkeypatch: pytest.MonkeyPatch) -> None:
16
+ from mindxtrain.operator import app as operator_app
17
+
18
+ monkeypatch.setattr(operator_app, "_bankml_reachable", lambda *a, **k: False)
tests/test_bankml_backend.py ADDED
@@ -0,0 +1,341 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """bankml operator backend: registry, env, receipts, typed refusals, SSE streaming, wiring.
2
+
3
+ No network: every HTTP exchange goes through httpx.MockTransport.
4
+ """
5
+
6
+ from __future__ import annotations
7
+
8
+ import asyncio
9
+ import json
10
+
11
+ import httpx
12
+ import pytest
13
+ from fastapi.testclient import TestClient
14
+
15
+ from mindxtrain.governance import panel as P
16
+ from mindxtrain.models.registry import ChatMessage, ChatRequest, build_backend, list_backends
17
+ from mindxtrain.operator import app as operator_app
18
+ from mindxtrain.operator.backends.bankml import (
19
+ DEFAULT_BANKML_BASE_URL,
20
+ BankmlBackend,
21
+ BankmlError,
22
+ BankmlRefusal,
23
+ bankml_root_url,
24
+ )
25
+
26
+ _REAL_BANKML_REACHABLE = operator_app._bankml_reachable # captured before conftest patches it
27
+
28
+ RECEIPT = {
29
+ "bankml": "0.3.4",
30
+ "engine": "native",
31
+ "model_sha256": "6b64c748d96ad26fd72402299bd27b2ae82f489bd0469498dd18eb6054058266",
32
+ "guard": "play",
33
+ "prompt_tokens": 12,
34
+ "completion_tokens": 3,
35
+ "ttft_ms": 40,
36
+ "wall_ms": 90,
37
+ "response_sha256": "a" * 64,
38
+ "request_sha256": "b" * 64,
39
+ "signed": False,
40
+ }
41
+
42
+
43
+ def _req(stream: bool = False) -> ChatRequest:
44
+ return ChatRequest(
45
+ model="mindx-gen39",
46
+ messages=[ChatMessage(role="user", content="Who are you?")],
47
+ temperature=0.0,
48
+ max_tokens=16,
49
+ stream=stream,
50
+ )
51
+
52
+
53
+ def _completion() -> dict[str, object]:
54
+ return {
55
+ "choices": [{"index": 0, "message": {"role": "assistant", "content": "I am mindX."},
56
+ "finish_reason": "stop"}],
57
+ "model": "mindx-gen39-F16.gguf",
58
+ "object": "chat.completion",
59
+ "usage": {"completion_tokens": 3, "prompt_tokens": 12, "total_tokens": 15},
60
+ "bankml_receipt": RECEIPT,
61
+ }
62
+
63
+
64
+ # ---- registry + env -----------------------------------------------------------
65
+
66
+
67
+ def test_registered_and_built_by_name() -> None:
68
+ assert "bankml" in list_backends()
69
+ assert isinstance(build_backend("bankml"), BankmlBackend)
70
+
71
+
72
+ def test_default_and_env_base_url(monkeypatch: pytest.MonkeyPatch) -> None:
73
+ monkeypatch.delenv("MINDXTRAIN_BANKML_BASE_URL", raising=False)
74
+ assert BankmlBackend().base_url == DEFAULT_BANKML_BASE_URL
75
+ monkeypatch.setenv("MINDXTRAIN_BANKML_BASE_URL", "http://127.0.0.1:9999/v1/")
76
+ b = BankmlBackend()
77
+ assert b.base_url == "http://127.0.0.1:9999/v1"
78
+ assert bankml_root_url() == "http://127.0.0.1:9999"
79
+
80
+
81
+ def test_never_sends_an_openai_key(monkeypatch: pytest.MonkeyPatch) -> None:
82
+ monkeypatch.setenv("MINDXTRAIN_OPENAI_API_KEY", "sk-should-not-leak")
83
+ assert "Authorization" not in BankmlBackend()._headers()
84
+
85
+
86
+ # ---- chat ---------------------------------------------------------------------
87
+
88
+
89
+ def test_chat_keeps_the_receipt() -> None:
90
+ seen: dict[str, object] = {}
91
+
92
+ def handler(request: httpx.Request) -> httpx.Response:
93
+ seen["url"] = str(request.url)
94
+ seen["body"] = json.loads(request.content)
95
+ return httpx.Response(200, json=_completion())
96
+
97
+ b = BankmlBackend(seed=7, transport=httpx.MockTransport(handler))
98
+ resp = asyncio.run(b.chat(_req()))
99
+ assert seen["url"] == "http://127.0.0.1:18093/v1/chat/completions"
100
+ body = seen["body"]
101
+ assert isinstance(body, dict)
102
+ assert body["seed"] == 7 and body["temperature"] == 0.0 and body["stream"] is False
103
+ assert resp.content == "I am mindX."
104
+ assert resp.prompt_tokens == 12 and resp.completion_tokens == 3
105
+ assert resp.receipt == RECEIPT
106
+ assert b.last_receipt == RECEIPT
107
+
108
+
109
+ def test_chat_400_is_a_typed_refusal_and_is_not_retried() -> None:
110
+ calls = {"n": 0}
111
+
112
+ def handler(request: httpx.Request) -> httpx.Response:
113
+ calls["n"] += 1
114
+ return httpx.Response(400, text="mirostat: not reproduced; bankML reproduces llama.cpp's "
115
+ "top-k, top-p, min-p and temperature")
116
+
117
+ b = BankmlBackend(transport=httpx.MockTransport(handler))
118
+ with pytest.raises(BankmlRefusal) as info:
119
+ asyncio.run(b.chat(_req()))
120
+ assert info.value.status_code == 400
121
+ assert info.value.reason.startswith("mirostat: not reproduced")
122
+ assert "bankml refused (400)" in str(info.value)
123
+ assert calls["n"] == 1
124
+
125
+
126
+ def test_ollama_style_json_error_body_is_read() -> None:
127
+ def handler(request: httpx.Request) -> httpx.Response:
128
+ return httpx.Response(400, json={"error": "tools: tool calling ... not reproduced yet"})
129
+
130
+ b = BankmlBackend(transport=httpx.MockTransport(handler))
131
+ with pytest.raises(BankmlRefusal) as info:
132
+ asyncio.run(b.chat(_req()))
133
+ assert info.value.reason.startswith("tools:")
134
+
135
+
136
+ def test_non_400_errors_are_bankml_errors_not_refusals() -> None:
137
+ def handler(request: httpx.Request) -> httpx.Response:
138
+ return httpx.Response(503, text="model changed since bankml verified it")
139
+
140
+ b = BankmlBackend(transport=httpx.MockTransport(handler))
141
+ with pytest.raises(BankmlError) as info:
142
+ asyncio.run(b.chat(_req()))
143
+ assert not isinstance(info.value, BankmlRefusal)
144
+ assert info.value.status_code == 503
145
+
146
+
147
+ # ---- stream -------------------------------------------------------------------
148
+
149
+
150
+ def _sse(*events: str) -> bytes:
151
+ return "".join(f"data: {e}\n\n" for e in events).encode()
152
+
153
+
154
+ def _collect(b: BankmlBackend) -> list[str]:
155
+ async def run() -> list[str]:
156
+ out: list[str] = []
157
+ async for tok in await b.stream_chat(_req(stream=True)):
158
+ out.append(tok)
159
+ return out
160
+
161
+ return asyncio.run(run())
162
+
163
+
164
+ def test_stream_yields_tokens_and_keeps_the_final_receipt() -> None:
165
+ def chunk(text: str) -> str:
166
+ return json.dumps({"choices": [{"index": 0, "delta": {"content": text},
167
+ "finish_reason": None}]})
168
+
169
+ final = json.dumps({"choices": [{"index": 0, "delta": {}, "finish_reason": "stop"}],
170
+ "usage": {"completion_tokens": 2, "prompt_tokens": 12}})
171
+ body = _sse(chunk("I am"), chunk(" mindX."), final,
172
+ json.dumps({"bankml_receipt": RECEIPT}), "[DONE]")
173
+
174
+ def handler(request: httpx.Request) -> httpx.Response:
175
+ assert json.loads(request.content)["stream"] is True
176
+ return httpx.Response(200, content=body, headers={"content-type": "text/event-stream"})
177
+
178
+ b = BankmlBackend(transport=httpx.MockTransport(handler))
179
+ assert _collect(b) == ["I am", " mindX."]
180
+ assert b.last_receipt == RECEIPT
181
+
182
+
183
+ def test_stream_400_raises_refusal() -> None:
184
+ def handler(request: httpx.Request) -> httpx.Response:
185
+ return httpx.Response(400, text="repeat_penalty: not reproduced")
186
+
187
+ b = BankmlBackend(transport=httpx.MockTransport(handler))
188
+ with pytest.raises(BankmlRefusal):
189
+ _collect(b)
190
+
191
+
192
+ def test_stream_error_event_raises() -> None:
193
+ def handler(request: httpx.Request) -> httpx.Response:
194
+ return httpx.Response(200, content=_sse(json.dumps({"error": "context full"})))
195
+
196
+ b = BankmlBackend(transport=httpx.MockTransport(handler))
197
+ with pytest.raises(BankmlError, match="context full"):
198
+ _collect(b)
199
+
200
+
201
+ # ---- operator wiring ---------------------------------------------------------
202
+
203
+
204
+ def test_backend_kwargs_for_bankml(monkeypatch: pytest.MonkeyPatch) -> None:
205
+ monkeypatch.setenv("MINDXTRAIN_BANKML_BASE_URL", "http://127.0.0.1:18100/v1")
206
+ assert operator_app.backend_kwargs("bankml") == {"base_url": "http://127.0.0.1:18100/v1"}
207
+
208
+
209
+ def test_backend_kwargs_openai_strict_contract(monkeypatch: pytest.MonkeyPatch) -> None:
210
+ monkeypatch.delenv("MINDXTRAIN_OPENAI_BASE_URL", raising=False)
211
+ with pytest.raises(KeyError):
212
+ operator_app.backend_kwargs("openai_compat", strict=True)
213
+ assert operator_app.backend_kwargs("openai_compat")["base_url"] == ""
214
+
215
+
216
+ def test_bankml_probe_hits_the_identity_endpoint(monkeypatch: pytest.MonkeyPatch) -> None:
217
+ seen: list[str] = []
218
+
219
+ class _Client:
220
+ def __init__(self, *a: object, **k: object) -> None: ...
221
+ def __enter__(self) -> _Client:
222
+ return self
223
+ def __exit__(self, *a: object) -> None: ...
224
+ def get(self, url: str) -> httpx.Response:
225
+ seen.append(url)
226
+ return httpx.Response(200, json={"bankml": "0.3.4"})
227
+
228
+ monkeypatch.delenv("MINDXTRAIN_BANKML_BASE_URL", raising=False)
229
+ monkeypatch.setattr(operator_app.httpx, "Client", _Client)
230
+ assert _REAL_BANKML_REACHABLE() is True
231
+ assert seen == ["http://127.0.0.1:18093/bankml"]
232
+
233
+
234
+ def test_bankml_probe_false_on_connection_error(monkeypatch: pytest.MonkeyPatch) -> None:
235
+ class _Client:
236
+ def __init__(self, *a: object, **k: object) -> None: ...
237
+ def __enter__(self) -> _Client:
238
+ return self
239
+ def __exit__(self, *a: object) -> None: ...
240
+ def get(self, url: str) -> httpx.Response:
241
+ raise httpx.ConnectError("refused")
242
+
243
+ monkeypatch.setattr(operator_app.httpx, "Client", _Client)
244
+ assert _REAL_BANKML_REACHABLE() is False
245
+
246
+
247
+ def test_autodetect_order(monkeypatch: pytest.MonkeyPatch) -> None:
248
+ monkeypatch.delenv("MINDXTRAIN_BACKEND", raising=False)
249
+ monkeypatch.delenv("AUTOMINDX_BACKEND", raising=False)
250
+ monkeypatch.setattr(operator_app, "_bankml_reachable", lambda: True)
251
+ # ollama first, always
252
+ monkeypatch.setattr(operator_app, "_ollama_reachable", lambda: True)
253
+ monkeypatch.setattr(operator_app, "_vllm_reachable", lambda: False)
254
+ assert operator_app.resolve_backend_name() == "ollama"
255
+ # vllm before bankml
256
+ monkeypatch.setattr(operator_app, "_ollama_reachable", lambda: False)
257
+ monkeypatch.setattr(operator_app, "_vllm_reachable", lambda: True)
258
+ assert operator_app.resolve_backend_name() == "vllm"
259
+ # bankml only when it is the one answering
260
+ monkeypatch.setattr(operator_app, "_vllm_reachable", lambda: False)
261
+ assert operator_app.resolve_backend_name() == "bankml"
262
+ # nothing answering: unchanged fallback
263
+ monkeypatch.setattr(operator_app, "_bankml_reachable", lambda: False)
264
+ assert operator_app.resolve_backend_name() == "vllm"
265
+
266
+
267
+ def test_health_dispatch(monkeypatch: pytest.MonkeyPatch) -> None:
268
+ monkeypatch.setattr(operator_app, "_bankml_reachable", lambda: True)
269
+ monkeypatch.setattr(operator_app, "_bankml_first_model", lambda: "mindx-gen39-F16.gguf")
270
+ assert operator_app.backend_reachable("bankml") is True
271
+ assert operator_app.backend_first_model("bankml") == "mindx-gen39-F16.gguf"
272
+
273
+
274
+ def test_operator_route_passes_refusal_through_as_400(monkeypatch: pytest.MonkeyPatch) -> None:
275
+ monkeypatch.setenv("MINDXTRAIN_BACKEND", "bankml")
276
+
277
+ async def refuse(self: BankmlBackend, request: ChatRequest) -> object:
278
+ raise BankmlRefusal("repeat_penalty: not reproduced")
279
+
280
+ monkeypatch.setattr(BankmlBackend, "chat", refuse)
281
+ client = TestClient(operator_app.app)
282
+ r = client.post("/v1/chat/completions", json={
283
+ "model": "mindx-gen39", "messages": [{"role": "user", "content": "hi"}]})
284
+ assert r.status_code == 400
285
+ assert "repeat_penalty: not reproduced" in r.json()["detail"]
286
+
287
+
288
+ def test_operator_route_returns_receipt(monkeypatch: pytest.MonkeyPatch) -> None:
289
+ monkeypatch.setenv("MINDXTRAIN_BACKEND", "bankml")
290
+ real_client = httpx.AsyncClient
291
+ transport = httpx.MockTransport(lambda req: httpx.Response(200, json=_completion()))
292
+ monkeypatch.setattr(
293
+ "mindxtrain.operator.backends.bankml.httpx.AsyncClient",
294
+ lambda **kw: real_client(timeout=kw.get("timeout"), transport=transport),
295
+ )
296
+ client = TestClient(operator_app.app)
297
+ r = client.post("/v1/chat/completions", json={
298
+ "model": "mindx-gen39", "messages": [{"role": "user", "content": "hi"}]})
299
+ assert r.status_code == 200
300
+ assert r.json()["receipt"]["model_sha256"] == RECEIPT["model_sha256"]
301
+
302
+
303
+ # ---- governance panel env chain -----------------------------------------------
304
+
305
+
306
+ def test_panel_bankml_url_only_when_set_and_last(monkeypatch: pytest.MonkeyPatch) -> None:
307
+ for e in ("MINDXTRAIN_OPENAI_BASE_URL", "MINDXTRAIN_VLLM_BASE_URL",
308
+ "MINDXTRAIN_OLLAMA_BASE_URL", "MINDXTRAIN_BANKML_BASE_URL", "MINDXTRAIN_BACKEND"):
309
+ monkeypatch.delenv(e, raising=False)
310
+ assert P.resolve_chat_base_url() == "http://localhost:11434/v1"
311
+ monkeypatch.setenv("MINDXTRAIN_BANKML_BASE_URL", "http://127.0.0.1:18093/v1")
312
+ assert P.resolve_chat_base_url() == "http://127.0.0.1:18093/v1"
313
+ monkeypatch.setenv("MINDXTRAIN_OLLAMA_BASE_URL", "http://localhost:11434/v1")
314
+ assert P.resolve_chat_base_url() == "http://localhost:11434/v1"
315
+ monkeypatch.setenv("MINDXTRAIN_BACKEND", "bankml")
316
+ assert P.resolve_chat_base_url() == "http://127.0.0.1:18093/v1"
317
+
318
+
319
+ # ---- console + HF Modelfile: neutral defaults are not sent ----------------------
320
+
321
+
322
+ def test_console_wire_options_omits_engine_defaults() -> None:
323
+ from mindxtrain.ui import console
324
+
325
+ sent = console.wire_options(dict(console.DEFAULTS))
326
+ for k in ("repeat_penalty", "presence_penalty", "frequency_penalty", "mirostat",
327
+ "mirostat_tau", "mirostat_eta", "stop"):
328
+ assert k not in sent, k
329
+ assert sent["temperature"] == 0.7 and sent["top_k"] == 40
330
+ # a deliberate value is always sent — the engine then honours or refuses it, visibly
331
+ deliberate = console.wire_options({"repeat_penalty": 1.3, "mirostat": 2, "mirostat_tau": 4.0})
332
+ assert deliberate == {"repeat_penalty": 1.3, "mirostat": 2, "mirostat_tau": 4.0}
333
+
334
+
335
+ def test_hf_published_modelfile_repeat_penalty_is_optional() -> None:
336
+ from mindxtrain.hf.extension import published_modelfile
337
+
338
+ default = published_modelfile("You are mindX.")
339
+ assert "PARAMETER repeat_penalty 1.3" in default # default unchanged
340
+ assert default.endswith('PARAMETER stop "<|im_end|>"\n')
341
+ assert "repeat_penalty" not in published_modelfile("You are mindX.", repeat_penalty=None)
tests/test_bankml_push.py ADDED
@@ -0,0 +1,305 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """push_to_bankml: the Modelfile subset, the capability probe, and `bankml create` — no binary.
2
+
3
+ subprocess.run and shutil.which are monkeypatched; the LoRA merge is replaced by a stub.
4
+ """
5
+
6
+ from __future__ import annotations
7
+
8
+ import json
9
+ import subprocess
10
+ from pathlib import Path
11
+ from typing import Any
12
+
13
+ import pytest
14
+ from typer.testing import CliRunner
15
+
16
+ from mindxtrain.deploy import bankml_push as BP
17
+ from mindxtrain.deploy.modelfile import ModelfileSpec, render_modelfile
18
+
19
+ SHA = "6b64c748d96ad26fd72402299bd27b2ae82f489bd0469498dd18eb6054058266"
20
+ DIGEST = "c" * 64
21
+ USAGE_035 = (
22
+ "usage: bankml usage [PID …]\n"
23
+ " bankml create NAME -f Modelfile [--registry DIR] [--models DIR]\n"
24
+ " bankml convert SAFETENSORS_DIR -o OUT.gguf [--outtype f16]\n"
25
+ " bankml version"
26
+ )
27
+ USAGE_034 = "usage: bankml usage [PID …]\n bankml serve FILE --fork FORK.json\n bankml version"
28
+
29
+
30
+ class FakeBankml:
31
+ """Records argv; answers `version`, `--help`, `convert`, `create`, `sha256` like bankml."""
32
+
33
+ def __init__(self, usage: str = USAGE_035, create_rc: int = 0, create_err: str = "") -> None:
34
+ self.calls: list[list[str]] = []
35
+ self.usage = usage
36
+ self.create_rc = create_rc
37
+ self.create_err = create_err
38
+ self.modelfile_text = ""
39
+
40
+ def __call__(self, cmd: list[str], **_: Any) -> subprocess.CompletedProcess[str]:
41
+ self.calls.append(list(cmd))
42
+ verb = cmd[1]
43
+ if verb == "version":
44
+ return subprocess.CompletedProcess(cmd, 0, "bankml 0.3.5\n", "")
45
+ if verb == "--help":
46
+ return subprocess.CompletedProcess(cmd, 1, "", self.usage)
47
+ if verb == "convert":
48
+ out = cmd[cmd.index("-o") + 1]
49
+ return subprocess.CompletedProcess(cmd, 0, f"{SHA} {out}\n", "bankml convert: ok\n")
50
+ if verb == "create":
51
+ self.modelfile_text = Path(cmd[cmd.index("-f") + 1]).read_text()
52
+ if self.create_rc:
53
+ return subprocess.CompletedProcess(cmd, self.create_rc, "parsing modelfile\n",
54
+ self.create_err)
55
+ name = cmd[2]
56
+ return subprocess.CompletedProcess(
57
+ cmd, 0, f"parsing modelfile\npinned {name}-F16.gguf sha256:{SHA}\n",
58
+ f"bankml create: {name} (digest sha256:{DIGEST}) over {name}-F16.gguf "
59
+ f"(sha256 {SHA})\n")
60
+ if verb == "sha256":
61
+ return subprocess.CompletedProcess(cmd, 0, f"{SHA} {cmd[2]}\n", "")
62
+ raise AssertionError(cmd)
63
+
64
+
65
+ @pytest.fixture
66
+ def fake(monkeypatch: pytest.MonkeyPatch) -> FakeBankml:
67
+ f = FakeBankml()
68
+ monkeypatch.setattr(BP.shutil, "which", lambda name: "/usr/local/bin/bankml")
69
+ monkeypatch.setattr(BP.subprocess, "run", f)
70
+ return f
71
+
72
+
73
+ def _merged(tmp_path: Path, arch: str = "LlamaForCausalLM") -> Path:
74
+ d = tmp_path / "merged"
75
+ d.mkdir()
76
+ (d / "config.json").write_text(json.dumps({"architectures": [arch]}))
77
+ return d
78
+
79
+
80
+ # ---- sanitize: the refusal table ----------------------------------------------
81
+
82
+
83
+ @pytest.mark.parametrize(
84
+ ("spec_kw", "needle"),
85
+ [
86
+ ({"adapter": "./lora"}, "ADAPTER"),
87
+ ({"template": "{{ .Prompt }}"}, "TEMPLATE"),
88
+ ({"parameters": {"repeat_penalty": 1.3}}, "PARAMETER repeat_penalty: penalties"),
89
+ ({"parameters": {"presence_penalty": 0.0}}, "PARAMETER presence_penalty"),
90
+ ({"parameters": {"frequency_penalty": 0.5}}, "PARAMETER frequency_penalty"),
91
+ ({"parameters": {"repeat_last_n": 64}}, "PARAMETER repeat_last_n"),
92
+ ({"parameters": {"mirostat": 2}}, "mirostat is not reproduced"),
93
+ ({"parameters": {"mirostat_tau": 5.0}}, "mirostat is not reproduced"),
94
+ ({"parameters": {"typical_p": 0.9}}, "typical_p"),
95
+ ({"parameters": {"num_thread": 4}}, "resource option"),
96
+ ({"parameters": {"num_gpu": 0}}, "resource option"),
97
+ ({"parameters": {"stop": "x"}}, "ModelfileSpec.stop"),
98
+ ({"parameters": {"weird": 1}}, "not a parameter bankml reproduces"),
99
+ ({"system": 'say """hi"""'}, "SYSTEM"),
100
+ ],
101
+ )
102
+ def test_sanitize_refuses_and_names_it(spec_kw: dict[str, Any], needle: str) -> None:
103
+ res = BP.bankml_sanitize(ModelfileSpec(from_model="x", **spec_kw))
104
+ assert not res.ok
105
+ assert any(needle in r for r in res.refusals), res.refusals
106
+
107
+
108
+ def test_sanitize_accepts_the_subset_unchanged() -> None:
109
+ spec = ModelfileSpec(
110
+ from_model="/m", system="You are mindX.", license="Apache-2.0",
111
+ parameters={"temperature": 0.0, "top_k": 40, "top_p": 0.9, "min_p": 0.05, "seed": 42,
112
+ "num_ctx": 2048, "num_predict": 128},
113
+ stop=["<|im_end|>"],
114
+ messages=[{"role": "user", "content": "hi"}, {"role": "assistant", "content": "hello"}],
115
+ )
116
+ res = BP.bankml_sanitize(spec)
117
+ assert res.ok and res.spec == spec
118
+ text = render_modelfile(res.spec)
119
+ assert "PARAMETER seed 42" in text and 'PARAMETER stop "<|im_end|>"' in text
120
+
121
+
122
+ def test_template_passes_only_when_equal_to_base() -> None:
123
+ spec = ModelfileSpec(from_model="x", template="T")
124
+ assert BP.bankml_sanitize(spec, base_template="T").ok
125
+ assert not BP.bankml_sanitize(spec, base_template="U").ok
126
+
127
+
128
+ def test_arch_checks(tmp_path: Path) -> None:
129
+ assert BP.check_merged_arch(_merged(tmp_path)) is None
130
+ q = tmp_path / "q"
131
+ q.mkdir()
132
+ (q / "config.json").write_text(json.dumps({"architectures": ["Qwen3ForCausalLM"]}))
133
+ assert "Llama-architecture" in (BP.check_merged_arch(q) or "")
134
+ assert BP.base_family_refusal("HuggingFaceTB/SmolLM2-135M") is None
135
+ assert "qwen" in (BP.base_family_refusal("Qwen/Qwen3-1.7B") or "")
136
+
137
+
138
+ # ---- capability probe ----------------------------------------------------------
139
+
140
+
141
+ def test_missing_binary(monkeypatch: pytest.MonkeyPatch, tmp_path: Path) -> None:
142
+ monkeypatch.setattr(BP.shutil, "which", lambda name: None)
143
+ res = BP.push_to_bankml("b", "t", merged_dir=_merged(tmp_path))
144
+ assert res.status == "bankml_missing"
145
+ assert "github.com/cryptoAGI/bankml" in res.reason
146
+
147
+
148
+ def test_too_old_bankml_is_reported_not_raised(fake: FakeBankml, tmp_path: Path) -> None:
149
+ fake.usage = USAGE_034
150
+ res = BP.push_to_bankml("b", "t", merged_dir=_merged(tmp_path))
151
+ assert res.status == "bankml_too_old"
152
+ assert "0.3.5" in res.reason and "create" in res.reason
153
+ assert not any(c[1] == "create" for c in fake.calls)
154
+
155
+
156
+ def test_capabilities_read_from_usage(fake: FakeBankml) -> None:
157
+ caps = BP.bankml_capabilities()
158
+ assert caps.version == "0.3.5" and caps.has_create and caps.has_convert
159
+
160
+
161
+ # ---- push ---------------------------------------------------------------------
162
+
163
+
164
+ def test_push_from_merged_dir(fake: FakeBankml, tmp_path: Path) -> None:
165
+ merged = _merged(tmp_path)
166
+ res = BP.push_to_bankml(
167
+ "HuggingFaceTB/SmolLM2-135M", "mindx-gen99", merged_dir=merged,
168
+ system="You are mindX.", params={"temperature": 0.0, "seed": 1}, stop=["<|im_end|>"],
169
+ registry_dir=tmp_path / "forks", work_dir=tmp_path / "work",
170
+ )
171
+ assert res.ok, res
172
+ assert res.model_sha256 == SHA and res.digest == DIGEST
173
+ create = next(c for c in fake.calls if c[1] == "create")
174
+ assert create[2] == "mindx-gen99" and "--registry" in create
175
+ assert fake.modelfile_text.startswith(f"FROM {merged.resolve()}")
176
+ assert 'SYSTEM """You are mindX."""' in fake.modelfile_text
177
+ assert "repeat_penalty" not in fake.modelfile_text
178
+
179
+
180
+ def test_push_with_convert_first(fake: FakeBankml, tmp_path: Path) -> None:
181
+ res = BP.push_to_bankml("b", "gen7", merged_dir=_merged(tmp_path), convert=True,
182
+ registry_dir=tmp_path / "forks", work_dir=tmp_path / "work")
183
+ assert res.ok
184
+ conv = next(c for c in fake.calls if c[1] == "convert")
185
+ assert conv[conv.index("-o") + 1].endswith("gen7-base-F16.gguf") and "--fork" in conv
186
+ assert fake.modelfile_text.startswith("FROM ") and "gen7-base-F16.gguf" in fake.modelfile_text
187
+
188
+
189
+ def test_refused_params_never_reach_bankml(fake: FakeBankml, tmp_path: Path) -> None:
190
+ res = BP.push_to_bankml("b", "t", merged_dir=_merged(tmp_path),
191
+ params={"repeat_penalty": 1.3, "temperature": 0.7})
192
+ assert res.status == "refused"
193
+ assert any("repeat_penalty" in r for r in res.refusals)
194
+ assert fake.calls == [] # refused before even probing the binary
195
+
196
+
197
+ def test_non_llama_merged_dir_refused(fake: FakeBankml, tmp_path: Path) -> None:
198
+ res = BP.push_to_bankml("b", "t", merged_dir=_merged(tmp_path, "Qwen3ForCausalLM"))
199
+ assert res.status == "refused" and "Llama-architecture" in res.reason
200
+ assert not any(c[1] == "create" for c in fake.calls)
201
+
202
+
203
+ def test_create_refusal_is_a_result(fake: FakeBankml, tmp_path: Path) -> None:
204
+ fake.create_rc = 2
205
+ fake.create_err = "bankml create: refuse: FROM /x: tokenizer not seen\n"
206
+ res = BP.push_to_bankml("b", "t", merged_dir=_merged(tmp_path), work_dir=tmp_path / "w")
207
+ assert res.status == "refused"
208
+ assert "tokenizer not seen" in res.reason
209
+
210
+
211
+ def test_bad_tag_and_bad_inputs(tmp_path: Path) -> None:
212
+ assert BP.push_to_bankml("b", "Bad:tag", merged_dir=tmp_path).status == "refused"
213
+ assert BP.push_to_bankml("b", "ok").status == "error"
214
+
215
+
216
+ def test_adapter_merge_and_import_error(fake: FakeBankml, tmp_path: Path,
217
+ monkeypatch: pytest.MonkeyPatch) -> None:
218
+ from mindxtrain.deploy import ollama_push
219
+
220
+ def merged_stub(base: str, adapter: Path, out: Path, *, sink: Any = None) -> Path:
221
+ out.mkdir(parents=True)
222
+ (out / "config.json").write_text(json.dumps({"architectures": ["LlamaForCausalLM"]}))
223
+ return out
224
+
225
+ adapter = tmp_path / "checkpoint"
226
+ adapter.mkdir()
227
+ monkeypatch.setattr(ollama_push, "merge_lora_adapter", merged_stub)
228
+ res = BP.push_to_bankml("HuggingFaceTB/SmolLM2-135M", "g", adapter_dir=adapter,
229
+ work_dir=tmp_path / "w", registry_dir=tmp_path / "r")
230
+ assert res.ok and res.merged_dir == tmp_path / "w" / "merged"
231
+
232
+ def no_ml(*a: Any, **k: Any) -> Path:
233
+ raise ImportError("push-to-ollama needs `peft` + `transformers`")
234
+
235
+ monkeypatch.setattr(ollama_push, "merge_lora_adapter", no_ml)
236
+ res = BP.push_to_bankml("b", "g2", adapter_dir=adapter, work_dir=tmp_path / "w2")
237
+ assert res.status == "merge_failed" and "uv sync --extra ml" in res.reason
238
+
239
+
240
+ def test_register_with_mindx_uses_bankml_provider(fake: FakeBankml, tmp_path: Path,
241
+ monkeypatch: pytest.MonkeyPatch) -> None:
242
+ from mindxtrain.deploy import api_client
243
+
244
+ seen: dict[str, Any] = {}
245
+
246
+ def swap(**kw: Any) -> dict[str, str]:
247
+ seen.update(kw)
248
+ return {"previous": "ollama:x", "current": "bankml:t"}
249
+
250
+ monkeypatch.setattr(api_client, "swap_mindx_fallback_model", swap)
251
+ res = BP.push_to_bankml("b", "t", merged_dir=_merged(tmp_path), register_with_mindx=True,
252
+ work_dir=tmp_path / "w", registry_dir=tmp_path / "r")
253
+ assert res.ok and seen["provider"] == "bankml" and seen["model"] == "t"
254
+ assert res.mindx_fallback_swap == {"previous": "ollama:x", "current": "bankml:t"}
255
+
256
+
257
+ # ---- CLI: serve --to bankml ------------------------------------------------------
258
+
259
+
260
+ def _recipe(tmp_path: Path, **overrides: Any) -> Path:
261
+ import yaml
262
+
263
+ from mindxtrain.config.loader import render_recipe
264
+
265
+ data = yaml.safe_load(render_recipe("mindx_fallback_qwen3_1_5b_cpu_real"))
266
+ for dotted, value in overrides.items():
267
+ node = data
268
+ *parents, leaf = dotted.split(".")
269
+ for p in parents:
270
+ node = node[p]
271
+ node[leaf] = value
272
+ path = tmp_path / "run.yaml"
273
+ path.write_text(yaml.safe_dump(data))
274
+ return path
275
+
276
+
277
+ def test_cli_refuses_quantized_config(tmp_path: Path) -> None:
278
+ from mindxtrain.cli.main import app
279
+
280
+ cfg = _recipe(tmp_path, **{"quantize.enabled": True, "quantize.scheme": "quark_fp8"})
281
+ r = CliRunner().invoke(app, ["serve", str(cfg), "--to", "bankml"])
282
+ assert r.exit_code == 2
283
+ assert "bankml refuses quantize.scheme=quark_fp8" in r.output
284
+
285
+
286
+ def test_cli_refuses_unconvertible_family(tmp_path: Path) -> None:
287
+ from mindxtrain.cli.main import app
288
+
289
+ cfg = _recipe(tmp_path, **{"model.name": "Qwen/Qwen3-1.7B"})
290
+ r = CliRunner().invoke(app, ["serve", str(cfg), "--to", "bankml"])
291
+ assert r.exit_code == 2
292
+ assert "qwen" in r.output.lower()
293
+
294
+
295
+ def test_cli_too_old_exits_2(tmp_path: Path, fake: FakeBankml) -> None:
296
+ from mindxtrain.cli.main import app
297
+
298
+ fake.usage = USAGE_034
299
+ ckpt = tmp_path / "ck"
300
+ ckpt.mkdir()
301
+ (ckpt / "config.json").write_text(json.dumps({"architectures": ["LlamaForCausalLM"]}))
302
+ cfg = _recipe(tmp_path)
303
+ r = CliRunner().invoke(app, ["serve", str(cfg), "--to", "bankml", "--checkpoint", str(ckpt)])
304
+ assert r.exit_code == 2
305
+ assert "bankml_too_old" in r.output
tests/test_imprint_bankml.py ADDED
@@ -0,0 +1,94 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Imprint probes through bankml: decoding sent, receipts kept, refusals typed, never mixed with
2
+ the canonical gate. MockTransport only — no server, no torch."""
3
+
4
+ from __future__ import annotations
5
+
6
+ import json
7
+
8
+ import httpx
9
+ import pytest
10
+
11
+ from mindxtrain.eval.imprint_bankml import (
12
+ NOT_COMPARABLE_NOTE,
13
+ imprint_via_bankml,
14
+ probe_recall_via_bankml,
15
+ )
16
+ from mindxtrain.operator.backends.bankml import BankmlRefusal
17
+
18
+ BASE_SHA = "1" * 64
19
+ GEN_SHA = "2" * 64
20
+ VOICE = ["I am mindX, the sovereign workshop of the professor."]
21
+ ANSWERS = {
22
+ "smollm2-base": ("Hello, how can I help you today?", BASE_SHA),
23
+ "mindx-gen99": ("I am mindX, the sovereign workshop.", GEN_SHA),
24
+ }
25
+
26
+
27
+ def _transport(seen: list[dict[str, object]]) -> httpx.MockTransport:
28
+ def handler(request: httpx.Request) -> httpx.Response:
29
+ assert request.url.path == "/api/chat"
30
+ body = json.loads(request.content)
31
+ seen.append(body)
32
+ text, sha = ANSWERS[body["model"]]
33
+ return httpx.Response(200, json={
34
+ "model": body["model"], "message": {"role": "assistant", "content": text},
35
+ "done": True, "done_reason": "stop", "eval_count": 9,
36
+ "bankml_receipt": {"model_sha256": sha, "request_sha256": "r" * 64,
37
+ "response_sha256": "s" * 64, "engine": "native"},
38
+ })
39
+
40
+ return httpx.MockTransport(handler)
41
+
42
+
43
+ def test_probe_sends_greedy_seeded_unpenalised_decoding() -> None:
44
+ seen: list[dict[str, object]] = []
45
+ probe = probe_recall_via_bankml(
46
+ "mindx-gen99", ["Who are you?", "Say hello."], system="You are mindX.", seed=42,
47
+ transport=_transport(seen),
48
+ )
49
+ assert probe.utterances == ["I am mindX, the sovereign workshop."] * 2
50
+ assert probe.model_sha256 == [GEN_SHA]
51
+ assert all(r and r["model_sha256"] == GEN_SHA for r in probe.receipts)
52
+ for body in seen:
53
+ assert body["stream"] is False
54
+ assert body["options"] == {"temperature": 0.0, "seed": 42, "num_predict": 48}
55
+ msgs = body["messages"]
56
+ assert isinstance(msgs, list) and msgs[0] == {"role": "system", "content": "You are mindX."}
57
+
58
+
59
+ def test_imprint_report_is_tagged_and_not_comparable() -> None:
60
+ res = imprint_via_bankml("smollm2-base", "mindx-gen99", ["Who are you?"], VOICE,
61
+ seed=3, transport=_transport([]))
62
+ assert res.report.method.endswith("/bankml-greedy")
63
+ assert res.report.imprint_delta > 0 and res.report.imprinted
64
+ assert res.canonical_gate is False
65
+ assert res.note == NOT_COMPARABLE_NOTE and "NOT comparable" in res.note
66
+ assert res.decoding.seed == 3 and res.decoding.temperature == 0.0
67
+ assert res.before.model_sha256 == [BASE_SHA] and res.after.model_sha256 == [GEN_SHA]
68
+ # the whole thing serialises (it is what the CLI prints)
69
+ assert json.loads(res.model_dump_json())["comparable_with"] == "bankml-greedy only"
70
+
71
+
72
+ def test_refusal_is_raised_typed_and_not_retried() -> None:
73
+ calls = {"n": 0}
74
+
75
+ def handler(request: httpx.Request) -> httpx.Response:
76
+ calls["n"] += 1
77
+ return httpx.Response(400, json={"error": "options.weird: not an option bankML knows"})
78
+
79
+ with pytest.raises(BankmlRefusal, match="not an option bankML knows"):
80
+ probe_recall_via_bankml("x", ["a", "b"], transport=httpx.MockTransport(handler))
81
+ assert calls["n"] == 1
82
+
83
+
84
+ def test_base_url_root_is_derived_from_v1(monkeypatch: pytest.MonkeyPatch) -> None:
85
+ urls: list[str] = []
86
+
87
+ def handler(request: httpx.Request) -> httpx.Response:
88
+ urls.append(str(request.url))
89
+ return httpx.Response(200, json={"message": {"content": "ok"}, "done": True})
90
+
91
+ monkeypatch.setenv("MINDXTRAIN_BANKML_BASE_URL", "http://127.0.0.1:18111/v1")
92
+ probe = probe_recall_via_bankml("x", ["a"], transport=httpx.MockTransport(handler))
93
+ assert urls == ["http://127.0.0.1:18111/api/chat"]
94
+ assert probe.receipts == [None] and probe.model_sha256 == []