ukisai's picture
initial release
e2109d9
|
Raw History Blame Contribute Delete
3.35 kB

Swift 1.5 3-bit TextOnly — local use

Use a complete snapshot and verify its file integrity before loading. The 11.77 GB tensor payload plus runtime, cache and OS must fit in available memory. Do not raise system memory limits. Short generation does not establish long-context operation or BF16 quality parity.

Install the pinned text runtime

Use Python 3.12 in an isolated environment on Apple Silicon. No full-model 4/5-bit architecture patch is needed for this language_model_only: true export.

python3.12 -m venv .venv-swift3
source .venv-swift3/bin/activate
python -m pip install 'mlx==0.32.2' 'transformers==5.14.1' 'huggingface_hub==1.31.0'
git clone https://github.com/ml-explore/mlx-lm.git swift3-mlx-lm
git -C swift3-mlx-lm checkout --detach c69d1288440a0dc4e6401fc417098b07598dccd5
python -m pip install -e ./swift3-mlx-lm

Use a new working directory. After installing the HF CLI, run hf auth login interactively if not already signed in with access to this private repository. The command below resolves current main once to a full commit and then uses only that pinned snapshot. For a repeat run, reuse the recorded commit. Do not use an incomplete historical upload or proceed after verification failure.

SWIFT_MLX_REVISION="$(python -c 'from huggingface_hub import HfApi; print(HfApi().model_info("ukisai/Swift-1.5-3bit-MLX-TextOnly").sha)')"
printf 'Pinned model revision: %s\n' "$SWIFT_MLX_REVISION"
hf download ukisai/Swift-1.5-3bit-MLX-TextOnly --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-3bit-MLX-TextOnly
hf cache verify ukisai/Swift-1.5-3bit-MLX-TextOnly --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-3bit-MLX-TextOnly --fail-on-missing-files
python Swift-1.5-3bit-MLX-TextOnly/verify_release.py Swift-1.5-3bit-MLX-TextOnly

The included verify_release.py does not load the model or access the network. Use its --manifest-sha256 option with a trusted digest from the reviewed release plan. Without a trusted manifest digest it verifies consistency, not source authenticity. It must fail for missing, truncated or modified shards. Do not proceed after failure.

Text chat

import mlx.core as mx
from mlx_lm import generate, load
from mlx_lm.sample_utils import make_sampler

if mx.default_device() != mx.gpu:
    raise RuntimeError("This example requires Apple Silicon/Metal; CPU is not certified.")
model, tokenizer = load("Swift-1.5-3bit-MLX-TextOnly")
messages = [{"role": "user", "content": "Reply with exactly: Hello from Swift."}]
if any(not isinstance(message.get("content"), str) for message in messages):
    raise ValueError("TextOnly: images, video and structured multimodal input are unsupported.")
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
mx.random.seed(20260922)
print(generate(model, tokenizer, prompt=prompt, max_tokens=32, sampler=make_sampler(temp=0)))

The inherited template also accepts reasoning_effort="low", "medium", and "xhigh". Template/token comparisons are separate from inference behavior. Full-model Linux generation, the reasoning-effort generation matrix, 100-request stability, long context and BF16 quality comparisons remain NOT_RUN. Do not pass image/video content through a text-only interface and report multimodal success.