Instructions to use edgefloor/meta-encoder-mlx-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use edgefloor/meta-encoder-mlx-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download edgefloor/meta-encoder-mlx-4bit --local-dir meta-encoder-mlx-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
MetaEncoder MLX 4-bit
An MLX quantization of facebook/meta-encoder, a 30B multimodal encoder that ranks candidates against a natural-language task.
Converted from source revision 3d0df704aa5bf66d9c49da1393181f33e893e5f7. Language linear layers, token embeddings and the language head use 4-bit affine quantization with group size 64. The vision tower, vision adapter and vision projection retain their original BF16 weights. Normalization weights also retain BF16. Weight files total 19.51 GB, or 18.17 GiB.
Use on Apple silicon
Download the repository and install the tested dependencies:
hf download edgefloor/meta-encoder-mlx-4bit --local-dir meta-encoder-mlx-4bit
cd meta-encoder-mlx-4bit
python -m pip install -r requirements.txt
python metaencoder_mlx.py --model . --task "Which animal says meow?" --candidates cat dog horse
For embeddings and candidate ranking:
from metaencoder_mlx import MetaEncoderMLX
model = MetaEncoderMLX(".")
embeddings = model.encode(["cat", "dog", "horse"])
matches = model.match("Which animal says meow?", ["cat", "dog", "horse"])
scores = model.score("Which animal says meow?", embeddings)
Text, PIL images and dictionaries with text, instruction, image, images or video are accepted. A video is a list of pre-sampled PIL frames at 2 FPS, matching the source wrapper. Each item is encoded individually to avoid padding entering MLX-VLM's decoder attention masks. Outputs are L2-normalized float32 vectors with 6,656 dimensions. Scores are inner products of these vectors.
from PIL import Image
matches = model.match(
{"text": "Which description matches the image?", "image": Image.open("photo.jpg")},
["a red bicycle", "a blue car"],
)
The included runner uses the original Transformers processor and the source wrapper's prompt construction. It pools the last token's final hidden state, avoids vocabulary logits and rounds rotary frequencies to BF16 as the source MetaEncoder wrapper does. PyTorch is used for processing support; model inference runs through MLX. The original PyTorch wrapper cannot directly load these quantized weights.
Local validation
The full checkpoint loaded with MLX-VLM's strict weight validation. Text, image and two-frame video smoke checks returned finite unit-length embeddings. The model selected the expected first candidate in 4 of 4 simple checks.
One text-task embedding had cosine similarity 0.972983 to a streamed BF16 MLX reference from the original weights. This checks quantization drift for that input; it is not a benchmark. Full-model parity with the official Transformers implementation was not measured. See validation.json for prompts, scores, timings and memory measurements.
A separate small-architecture test verified that every saved quantized tensor exactly matched mlx.core.quantize, that unquantized tensors were preserved exactly and that both variants reloaded strictly. A small unquantized architecture also exceeded 0.999 cosine agreement with Transformers on text hidden states and image features.
Reproduce
hf download facebook/meta-encoder --revision 3d0df704aa5bf66d9c49da1393181f33e893e5f7 --local-dir source
python convert_metaencoder.py --source source --output converted --revision 3d0df704aa5bf66d9c49da1393181f33e893e5f7
The converter writes both 4-bit and 8-bit variants incrementally and validates the source tensor names and shapes against MLX-VLM's Muse Glimmer architecture. Conversion metadata is in conversion.json.
License
The source model is Apache-2.0 and remains subject to its base model usage policy. LICENSE, NOTICE and USAGE_POLICY.md are preserved. This conversion changes weight storage and adds an MLX inference wrapper; it does not fine-tune the model. The upstream PyTorch helper source is preserved as upstream_metaencoder.py.
- Downloads last month
- 17
4-bit