Instructions to use LiquidAI/LFM2.5-Encoder-350M-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use LiquidAI/LFM2.5-Encoder-350M-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf LiquidAI/LFM2.5-Encoder-350M-GGUF:F16 # Run inference directly in the terminal: llama cli -hf LiquidAI/LFM2.5-Encoder-350M-GGUF:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf LiquidAI/LFM2.5-Encoder-350M-GGUF:F16 # Run inference directly in the terminal: llama cli -hf LiquidAI/LFM2.5-Encoder-350M-GGUF:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf LiquidAI/LFM2.5-Encoder-350M-GGUF:F16 # Run inference directly in the terminal: ./llama-cli -hf LiquidAI/LFM2.5-Encoder-350M-GGUF:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf LiquidAI/LFM2.5-Encoder-350M-GGUF:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf LiquidAI/LFM2.5-Encoder-350M-GGUF:F16
Use Docker
docker model run hf.co/LiquidAI/LFM2.5-Encoder-350M-GGUF:F16
- LM Studio
- Jan
- Ollama
How to use LiquidAI/LFM2.5-Encoder-350M-GGUF with Ollama:
ollama run hf.co/LiquidAI/LFM2.5-Encoder-350M-GGUF:F16
- Unsloth Desktop
- Docker Model Runner
How to use LiquidAI/LFM2.5-Encoder-350M-GGUF with Docker Model Runner:
docker model run hf.co/LiquidAI/LFM2.5-Encoder-350M-GGUF:F16
- Lemonade
How to use LiquidAI/LFM2.5-Encoder-350M-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull LiquidAI/LFM2.5-Encoder-350M-GGUF:F16
Run and chat with the model
lemonade run user.LFM2.5-Encoder-350M-GGUF-F16
List all available models
lemonade list
- Atomic Chat
Download fill-mask.py from LiquidAI/LFM2.5-Encoder-350M-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 2.15 kB
-
https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M-GGUF/resolve/main/fill-mask.py
- Command line
-
hf download hf://LiquidAI/LFM2.5-Encoder-350M-GGUF/fill-mask.py
-
curl -L -o fill-mask.py https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M-GGUF/resolve/main/fill-mask.py
2.15 kB
| # /// script | |
| # requires-python = ">=3.10" | |
| # dependencies = ["numpy", "requests", "gguf"] | |
| # /// | |
| # fill-mask.py — masked-token prediction against a stock llama-server. | |
| # | |
| # The encoder's MLM head is tied to the token embeddings, so the logits at the | |
| # mask position are just `hidden @ token_embd^T`: fetch the per-token hidden | |
| # states from `llama-server --embeddings --pooling none`, read the embedding | |
| # matrix straight out of the GGUF, and take the top-K at the mask position. | |
| # | |
| # llama-server -m LFM2.5-Encoder-230M-F16.gguf --embeddings --pooling none | |
| # uv run fill-mask.py LFM2.5-Encoder-230M-F16.gguf "The capital of France is [MASK]." | |
| import sys | |
| import numpy as np | |
| import requests | |
| from gguf import GGUFReader | |
| gguf_path, prompt = sys.argv[1], sys.argv[2] | |
| topk = int(sys.argv[3]) if len(sys.argv) > 3 else 5 | |
| # token_embd from the GGUF (memory-mapped; fp16/fp32 tensors read directly) | |
| reader = GGUFReader(gguf_path) | |
| embd = next(t for t in reader.tensors if t.name == "token_embd.weight") | |
| W = np.array(embd.data).astype(np.float32) # [n_vocab, n_embd] | |
| # tokenize server-side, replacing [MASK] with the model's mask token id | |
| def tokenize(text: str, special: bool) -> list[int]: | |
| r = requests.post("http://localhost:8080/tokenize", | |
| json={"content": text, "add_special": special, "parse_special": True}) | |
| return r.json()["tokens"] | |
| meta = {f.name: f for f in reader.fields.values()} | |
| mask_id = int(meta["tokenizer.ggml.mask_token_id"].parts[-1][0]) | |
| pre, _, post = prompt.partition("[MASK]") | |
| toks = tokenize(pre, True) + [mask_id] + tokenize(post, False) | |
| pos = toks.index(mask_id) | |
| # one non-causal forward; per-token hidden states | |
| r = requests.post("http://localhost:8080/embedding", | |
| json={"content": toks}) | |
| hidden = np.array(r.json()[0]["embedding"], dtype=np.float32) # [n_tok, n_embd] | |
| logits = hidden[pos] @ W.T | |
| top = np.argsort(logits)[::-1][:topk] | |
| detok = lambda t: requests.post("http://localhost:8080/detokenize", json={"tokens": [int(t)]}).json()["content"] | |
| print(f"top-{topk} at [MASK]:") | |
| for i, t in enumerate(top, 1): | |
| print(f" {i:>2} {logits[t]:9.4f} '{detok(t)}'") | |