Instructions to use ukisai/Swift-1.5-3bit-MLX-TextOnly with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ukisai/Swift-1.5-3bit-MLX-TextOnly with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ukisai/Swift-1.5-3bit-MLX-TextOnly") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ukisai/Swift-1.5-3bit-MLX-TextOnly with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-3bit-MLX-TextOnly"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ukisai/Swift-1.5-3bit-MLX-TextOnly" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ukisai/Swift-1.5-3bit-MLX-TextOnly with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ukisai/Swift-1.5-3bit-MLX-TextOnly"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ukisai/Swift-1.5-3bit-MLX-TextOnly" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ukisai/Swift-1.5-3bit-MLX-TextOnly", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ukisai/Swift-1.5-3bit-MLX-TextOnly with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-3bit-MLX-TextOnly"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ukisai/Swift-1.5-3bit-MLX-TextOnly
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ukisai/Swift-1.5-3bit-MLX-TextOnly with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ukisai/Swift-1.5-3bit-MLX-TextOnly"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ukisai/Swift-1.5-3bit-MLX-TextOnly" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download USAGE.md from ukisai/Swift-1.5-3bit-MLX-TextOnly: direct link, hf CLI and curl.
- Browser
- Download file 3.35 kB
-
https://huggingface.co/ukisai/Swift-1.5-3bit-MLX-TextOnly/resolve/main/USAGE.md
- Command line
-
hf download hf://ukisai/Swift-1.5-3bit-MLX-TextOnly/USAGE.md
-
curl -L -o USAGE.md https://huggingface.co/ukisai/Swift-1.5-3bit-MLX-TextOnly/resolve/main/USAGE.md
Swift 1.5 3-bit TextOnly — local use
Use a complete snapshot and verify its file integrity before loading. The 11.77 GB tensor payload plus runtime, cache and OS must fit in available memory. Do not raise system memory limits. Short generation does not establish long-context operation or BF16 quality parity.
Install the pinned text runtime
Use Python 3.12 in an isolated environment on Apple Silicon. No full-model 4/5-bit
architecture patch is needed for this language_model_only: true export.
python3.12 -m venv .venv-swift3
source .venv-swift3/bin/activate
python -m pip install 'mlx==0.32.2' 'transformers==5.14.1' 'huggingface_hub==1.31.0'
git clone https://github.com/ml-explore/mlx-lm.git swift3-mlx-lm
git -C swift3-mlx-lm checkout --detach c69d1288440a0dc4e6401fc417098b07598dccd5
python -m pip install -e ./swift3-mlx-lm
Use a new working directory. After installing the HF CLI, run hf auth login
interactively if not already signed in with access to this private repository.
The command below resolves current main once to a full commit and then uses
only that pinned snapshot. For a repeat run, reuse the recorded commit.
Do not use an incomplete historical upload or proceed after verification failure.
SWIFT_MLX_REVISION="$(python -c 'from huggingface_hub import HfApi; print(HfApi().model_info("ukisai/Swift-1.5-3bit-MLX-TextOnly").sha)')"
printf 'Pinned model revision: %s\n' "$SWIFT_MLX_REVISION"
hf download ukisai/Swift-1.5-3bit-MLX-TextOnly --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-3bit-MLX-TextOnly
hf cache verify ukisai/Swift-1.5-3bit-MLX-TextOnly --revision "$SWIFT_MLX_REVISION" --local-dir Swift-1.5-3bit-MLX-TextOnly --fail-on-missing-files
python Swift-1.5-3bit-MLX-TextOnly/verify_release.py Swift-1.5-3bit-MLX-TextOnly
The included verify_release.py does not load the model or access the network.
Use its --manifest-sha256 option with a trusted digest from the reviewed release plan.
Without a trusted manifest digest it verifies consistency, not source authenticity.
It must fail for missing, truncated or modified shards. Do not proceed after failure.
Text chat
import mlx.core as mx
from mlx_lm import generate, load
from mlx_lm.sample_utils import make_sampler
if mx.default_device() != mx.gpu:
raise RuntimeError("This example requires Apple Silicon/Metal; CPU is not certified.")
model, tokenizer = load("Swift-1.5-3bit-MLX-TextOnly")
messages = [{"role": "user", "content": "Reply with exactly: Hello from Swift."}]
if any(not isinstance(message.get("content"), str) for message in messages):
raise ValueError("TextOnly: images, video and structured multimodal input are unsupported.")
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
mx.random.seed(20260922)
print(generate(model, tokenizer, prompt=prompt, max_tokens=32, sampler=make_sampler(temp=0)))
The inherited template also accepts reasoning_effort="low", "medium", and
"xhigh". Template/token comparisons are separate from inference behavior.
Full-model Linux generation, the reasoning-effort generation matrix,
100-request stability, long context and BF16 quality comparisons remain NOT_RUN.
Do not pass image/video content through a text-only interface and report multimodal success.