ukisai's picture
initial release
e2109d9
|
Raw History Blame Contribute Delete
3.35 kB
# Swift 1.5 3-bit TextOnly — local use
Use a complete snapshot and verify its file integrity before loading.
The 11.77 GB tensor payload plus runtime, cache and OS must fit in available
memory. Do not raise system memory limits. Short generation does not
establish long-context operation or BF16 quality parity.
## Install the pinned text runtime
Use Python 3.12 in an isolated environment on Apple Silicon. No full-model 4/5-bit
architecture patch is needed for this `language_model_only: true` export.
```bash
python3.12 -m venv .venv-swift3
source .venv-swift3/bin/activate
python -m pip install 'mlx==0.32.2' 'transformers==5.14.1' 'huggingface_hub==1.31.0'
git clone https://github.com/ml-explore/mlx-lm.git swift3-mlx-lm
git -C swift3-mlx-lm checkout --detach c69d1288440a0dc4e6401fc417098b07598dccd5
python -m pip install -e ./swift3-mlx-lm
```
Use a new working directory. After installing the HF CLI, run `hf auth login`
interactively if not already signed in with access to this private repository.
The command below resolves current main once to a full commit and then uses
only that pinned snapshot. For a repeat run, reuse the recorded commit.
Do not use an incomplete historical upload or proceed after verification failure.
```bash
SWIFT_MLX_REVISION="$(python -c 'from huggingface_hub import HfApi; print(HfApi().model_info("ukisai/Swift-1.5-3bit-MLX-TextOnly").sha)')"
printf 'Pinned model revision: %s\n' "$SWIFT_MLX_REVISION"
hf download ukisai/Swift-1.5-3bit-MLX-TextOnly --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-3bit-MLX-TextOnly
hf cache verify ukisai/Swift-1.5-3bit-MLX-TextOnly --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-3bit-MLX-TextOnly --fail-on-missing-files
python Swift-1.5-3bit-MLX-TextOnly/verify_release.py Swift-1.5-3bit-MLX-TextOnly
```
The included `verify_release.py` does not load the model or access the network.
Use its `--manifest-sha256` option with a trusted digest from the reviewed release plan.
Without a trusted manifest digest it verifies consistency, not source authenticity.
It must fail for missing, truncated or modified shards. Do not proceed after failure.
## Text chat
```python
import mlx.core as mx
from mlx_lm import generate, load
from mlx_lm.sample_utils import make_sampler
if mx.default_device() != mx.gpu:
raise RuntimeError("This example requires Apple Silicon/Metal; CPU is not certified.")
model, tokenizer = load("Swift-1.5-3bit-MLX-TextOnly")
messages = [{"role": "user", "content": "Reply with exactly: Hello from Swift."}]
if any(not isinstance(message.get("content"), str) for message in messages):
raise ValueError("TextOnly: images, video and structured multimodal input are unsupported.")
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
mx.random.seed(20260922)
print(generate(model, tokenizer, prompt=prompt, max_tokens=32, sampler=make_sampler(temp=0)))
```
The inherited template also accepts `reasoning_effort="low"`, `"medium"`, and
`"xhigh"`. Template/token comparisons are separate from inference behavior.
Full-model Linux generation, the reasoning-effort generation matrix,
100-request stability, long context and BF16 quality comparisons remain NOT_RUN.
Do not pass image/video content through a text-only interface and report multimodal success.