# Provenance — Gemma-4-E4B-MERaLiON-Speech-LoRA-MNSC-MLX This model is a **composition** of four upstream sources. The exact files and their origins are listed below for full traceability. ## Precision summary (per-component) | Component | Storage dtype | Notes | |---|---|---| | Decoder (Gemma-4-E4B) | **MLX 8-bit affine** (uint32-packed weights) + bf16 norms/embeds/biases | Quantized via RotorQuant pipeline, group-size 64 | | Speech encoder (Whisper-large-v3 fine-tuned) | **fp16** | Not quantized | | Adaptor | **fp16** | Not quantized | | Projector (ours) | **fp32** | Not quantized | | LoRA adapters (ours) | **fp32** | Not quantized | The decoder is the only quantized component. Everything else is full-precision floating point (fp16 for the speech tower, fp32 for our trained additions). There are no bf8 or fp4 weights in this bundle. --- ## 1. Decoder (LLM backbone) - **Subdirectory:** `decoder/` - **Files:** `model-00001-of-00002.safetensors`, `model-00002-of-00002.safetensors`, `model.safetensors.index.json`, `config.json`, `generation_config.json`, `tokenizer*`, `processor_config.json` - **Direct source:** [`majentik/gemma-4-E4B-RotorQuant-MLX-8bit`](https://huggingface.co/majentik/gemma-4-E4B-RotorQuant-MLX-8bit) - **Original source:** [`google/gemma-4-E4B-it`](https://huggingface.co/google/gemma-4-E4B-it) - **Transformation chain:** Google Gemma-4-E4B-it → MLX bf16 conversion → RotorQuant rotation-based decorrelation → MLX 8-bit affine quantization (group-size 64) - **License:** Gemma Terms of Use (Google) ## 2. Speech encoder + adaptor - **Subdirectory:** `speech_encoder/` - **Files:** `encoder.safetensors` (1.27 GB, **fp16**), `adaptor.safetensors` (128 MB, **fp16**), `encoder_config.json`, `adaptor_config.json`, `preprocessor_config.json`, `composite_config.json` - **Direct source:** [`majentik/MERaLiON-3-10B-MLX-4bit`](https://huggingface.co/majentik/MERaLiON-3-10B-MLX-4bit) (encoder and adaptor are stored as fp16 in that repo — they are NOT quantized; only the decoder shards in that repo are 4-bit) - **Original source:** [`MERaLiON/MERaLiON-3-10B`](https://huggingface.co/MERaLiON/MERaLiON-3-10B) (built on Whisper-large-v3 encoder) - **Transformation chain:** Whisper-large-v3-encoder → MERaLiON-3 fine-tuning by I2R Singapore → MLX fp16 conversion (no quantization on these tensors) - **Used as:** frozen feature extractor producing 3584-dim speech embeddings at frame rate - **License:** MERaLiON Public License v2 (this is the most restrictive license in the chain; the entire composite work inherits it) ## 3. Projector (our work) - **Subdirectory:** `projector/` - **Files:** `p.safetensors` (75 MB, **fp32**), `projector_config.json` - **Architecture:** `LayerNorm(3584) → Linear(3584→3072) → SiLU → Linear(3072→2560) → RMSNorm(2560)` - **Purpose:** bridges MERaLiON encoder output (3584-d) into Gemma-4 decoder embedding space (2560-d) - **Trained from scratch** by majentik (no pre-existing weights) - **License:** MIT (our additions); composite work licensed under MERaLiON Public License v2 ## 4. LoRA adapter (our work) - **Subdirectory:** `lora/` - **Files:** `adapters.safetensors` (133 MB, **fp32**), `lora_config.json` - **Spec:** rank 16, scale 20.0, targets `[q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj]` across all 42 decoder layers (294 adapter pairs) - **Base model:** the Gemma-4-E4B-RotorQuant-MLX-8bit decoder above - **Switchable design:** scale=20.0 when speech mode active, 0.0 when generating plain text (LoRA acts only during speech-conditioned inference) - **Trained from scratch** by majentik - **License:** MIT (our additions); composite work licensed under MERaLiON Public License v2 ## Training data - **Dataset:** MERaLiON National Speech Corpus (MNSC), ASR-PART2-Test split - **Clips:** 13,142 training, 3,000 held out for evaluation - **License:** MERaLiON Public License v2 (data terms apply only to the trained checkpoint; the composite work was trained on this data and inherits the data license) ## Training hyperparameters | Parameter | Value | |---|---| | Learning rate | 5e-5 | | Warmup | 50 opt-steps | | Batch size | 1 (grad_accum=8, effective=8) | | Optimizer steps | 3000 | | Epochs | ~1.83 | | Hardware | Apple M5 Max 128 GB unified memory | | Wall time | 8h 22min | | Framework | MLX 0.x + custom training loop | ## Evaluation - **Test set:** MNSC ASR-PART2-Test (3000 clips) - **Metric:** WER (jiwer; lowercase + punctuation strip + whitespace normalize) - **Post-processing:** `:` dialogue prefix stripped from model output (model learned this format from training data; production loader handles it in `_strip_speaker_prefix`) | Model | WER | |---|---| | MERaLiON-2-10B (baseline) | 25.78% | | **This model** | **18.86%** | | Δ | **−6.92 pp** | ## Source code Training and inference code: [`ajentik/elderwise-mlx`](https://github.com/ajentik/elderwise-mlx) Relevant commits at time of release: - `10335d6` — fix PLE collapse + 3 compounding bugs (the four bugs that made our first run produce 592% WER) - `ade5550` — training configs + diagnostic scripts - `df72999` — production normalization (`_strip_speaker_prefix`) + eval improvements ## Reproducibility To reproduce the evaluation: ```bash git clone https://github.com/ajentik/elderwise-mlx cd elderwise-mlx # Place this model's contents at artifacts/release_pkg/ then: .venv/bin/python scripts/05_eval_new.py artifacts/release_pkg configs/run_3000_r16mlp.yaml ``` Expected output: WER 18.86%, delta −6.92pp vs baseline.