Image-to-Text
MLX
mlx-vision
ocr
apple-silicon
speculative-decoding
dspark
deepseek-ocr
glm-ocr
vision-language-model
Instructions to use will702/mlx-vision with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use will702/mlx-vision with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download will702/mlx-vision --local-dir mlx-vision
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Upload folder using huggingface_hub
Browse files- .gitattributes +1 -35
- AGENTS.md +53 -0
- CITATION.cff +24 -0
- LICENSE +21 -0
- README.md +71 -0
- llms.txt +30 -0
.gitattributes
CHANGED
|
@@ -1,35 +1 @@
|
|
| 1 |
-
*.
|
| 2 |
-
*.arrow filter=lfs diff=lfs merge=lfs -text
|
| 3 |
-
*.bin filter=lfs diff=lfs merge=lfs -text
|
| 4 |
-
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
| 5 |
-
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
| 6 |
-
*.ftz filter=lfs diff=lfs merge=lfs -text
|
| 7 |
-
*.gz filter=lfs diff=lfs merge=lfs -text
|
| 8 |
-
*.h5 filter=lfs diff=lfs merge=lfs -text
|
| 9 |
-
*.joblib filter=lfs diff=lfs merge=lfs -text
|
| 10 |
-
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
| 11 |
-
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
| 12 |
-
*.model filter=lfs diff=lfs merge=lfs -text
|
| 13 |
-
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
| 14 |
-
*.npy filter=lfs diff=lfs merge=lfs -text
|
| 15 |
-
*.npz filter=lfs diff=lfs merge=lfs -text
|
| 16 |
-
*.onnx filter=lfs diff=lfs merge=lfs -text
|
| 17 |
-
*.ot filter=lfs diff=lfs merge=lfs -text
|
| 18 |
-
*.parquet filter=lfs diff=lfs merge=lfs -text
|
| 19 |
-
*.pb filter=lfs diff=lfs merge=lfs -text
|
| 20 |
-
*.pickle filter=lfs diff=lfs merge=lfs -text
|
| 21 |
-
*.pkl filter=lfs diff=lfs merge=lfs -text
|
| 22 |
-
*.pt filter=lfs diff=lfs merge=lfs -text
|
| 23 |
-
*.pth filter=lfs diff=lfs merge=lfs -text
|
| 24 |
-
*.rar filter=lfs diff=lfs merge=lfs -text
|
| 25 |
-
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
| 26 |
-
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
| 27 |
-
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
| 28 |
-
*.tar filter=lfs diff=lfs merge=lfs -text
|
| 29 |
-
*.tflite filter=lfs diff=lfs merge=lfs -text
|
| 30 |
-
*.tgz filter=lfs diff=lfs merge=lfs -text
|
| 31 |
-
*.wasm filter=lfs diff=lfs merge=lfs -text
|
| 32 |
-
*.xz filter=lfs diff=lfs merge=lfs -text
|
| 33 |
-
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
-
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
-
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
| 1 |
+
*.md linguist-detectable=true
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
AGENTS.md
ADDED
|
@@ -0,0 +1,53 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Agent / tool guidance for mlx-vision
|
| 2 |
+
|
| 3 |
+
## What this repo is
|
| 4 |
+
Python library **mlx-vision**: DSpark-style speculative decoding for **OCR VLMs on Apple Silicon (MLX)**.
|
| 5 |
+
Counterpart to text-only [mlx-dspark](https://github.com/ARahim3/mlx-dspark). Built on [mlx-vlm](https://github.com/Blaizzy/mlx-vlm).
|
| 6 |
+
|
| 7 |
+
## When to use
|
| 8 |
+
- Faster local OCR on Mac (M-series) for DeepSeek-OCR-2, GLM-OCR, Unlimited-OCR
|
| 9 |
+
- Speculative decoding: `baseline` | `dspark` | `lookup` | `dflash` | `eagle3` | `mtp` | `auto`
|
| 10 |
+
- PDF page OCR via `pypdfium2` (`pip install 'mlx-vision[pdf]'`)
|
| 11 |
+
|
| 12 |
+
## Install
|
| 13 |
+
```bash
|
| 14 |
+
pip install mlx-vision
|
| 15 |
+
# from source:
|
| 16 |
+
pip install -e ".[pdf,dev]"
|
| 17 |
+
```
|
| 18 |
+
|
| 19 |
+
## Quick commands
|
| 20 |
+
```bash
|
| 21 |
+
mlx-vision models
|
| 22 |
+
mlx-vision -m deepseek-ocr-2 -i page.png --mode auto -v
|
| 23 |
+
mlx-vision -m deepseek-ocr-2 -i doc.pdf --page 0 --dpi 200
|
| 24 |
+
mlx-vision bench --image page.png --modes baseline,lookup,dspark
|
| 25 |
+
mlx-vision serve --port 8080
|
| 26 |
+
```
|
| 27 |
+
|
| 28 |
+
## Python
|
| 29 |
+
```python
|
| 30 |
+
from mlx_vision import ocr
|
| 31 |
+
r = ocr("page.png", model="deepseek-ocr-2", mode="auto")
|
| 32 |
+
```
|
| 33 |
+
|
| 34 |
+
## Important caveats (do not skip)
|
| 35 |
+
1. **No public OCR DSpark weights yet** — `--mode dspark` uses training-free semi-AR lookup + confidence schedule until drafters are trained (`scripts/train_drafter/`).
|
| 36 |
+
2. DeepSeek-OCR MLX checkpoints trigger Torch remote-code processor load; this package falls back to native MLX processors (`src/mlx_vision/load.py`) — do not “fix” by installing torch unless needed for training.
|
| 37 |
+
3. Unlimited-OCR uses **R-SWA**; draft block size is clamped to the sliding window.
|
| 38 |
+
4. Requires Apple Silicon + MLX (Metal).
|
| 39 |
+
|
| 40 |
+
## Key paths
|
| 41 |
+
| Path | Role |
|
| 42 |
+
|------|------|
|
| 43 |
+
| `src/mlx_vision/generate.py` | `ocr()` API |
|
| 44 |
+
| `src/mlx_vision/load.py` | Robust model/processor load |
|
| 45 |
+
| `src/mlx_vision/speculative/` | DSpark/DFlash/confidence/lookup |
|
| 46 |
+
| `src/mlx_vision/models/registry.py` | Presets + drafter registry |
|
| 47 |
+
| `src/mlx_vision/pdf.py` | PDF → PNG |
|
| 48 |
+
| `scripts/train_drafter/` | DeepSpec OCR drafter recipe |
|
| 49 |
+
|
| 50 |
+
## Related
|
| 51 |
+
- GitHub: https://github.com/will702/mlx-vision
|
| 52 |
+
- Hugging Face: https://huggingface.co/will702/mlx-vision
|
| 53 |
+
- Upstream OCR weights: `mlx-community/DeepSeek-OCR-2-*`, `mlx-community/GLM-OCR-*`, `baidu/Unlimited-OCR`
|
CITATION.cff
ADDED
|
@@ -0,0 +1,24 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
cff-version: 1.2.0
|
| 2 |
+
title: "mlx-vision"
|
| 3 |
+
message: "If you use mlx-vision, please cite it as below."
|
| 4 |
+
type: software
|
| 5 |
+
authors:
|
| 6 |
+
- family-names: "Willson"
|
| 7 |
+
given-names: "Gregorius"
|
| 8 |
+
alias: "will702"
|
| 9 |
+
repository-code: "https://github.com/will702/mlx-vision"
|
| 10 |
+
url: "https://huggingface.co/will702/mlx-vision"
|
| 11 |
+
abstract: >-
|
| 12 |
+
DSpark-style speculative decoding for OCR vision-language models on
|
| 13 |
+
Apple Silicon via MLX. Speeds up DeepSeek-OCR-2, GLM-OCR, and Unlimited-OCR
|
| 14 |
+
with lossless target verification when neural drafters are available.
|
| 15 |
+
keywords:
|
| 16 |
+
- mlx
|
| 17 |
+
- ocr
|
| 18 |
+
- speculative-decoding
|
| 19 |
+
- dspark
|
| 20 |
+
- apple-silicon
|
| 21 |
+
- deepseek-ocr
|
| 22 |
+
license: MIT
|
| 23 |
+
version: 0.1.0
|
| 24 |
+
date-released: "2026-07-28"
|
LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
MIT License
|
| 2 |
+
|
| 3 |
+
Copyright (c) 2026 faster-inference-ai
|
| 4 |
+
|
| 5 |
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
| 6 |
+
of this software and associated documentation files (the "Software"), to deal
|
| 7 |
+
in the Software without restriction, including without limitation the rights
|
| 8 |
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
| 9 |
+
copies of the Software, and to permit persons to whom the Software is
|
| 10 |
+
furnished to do so, subject to the following conditions:
|
| 11 |
+
|
| 12 |
+
The above copyright notice and this permission notice shall be included in all
|
| 13 |
+
copies or substantial portions of the Software.
|
| 14 |
+
|
| 15 |
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
| 16 |
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
| 17 |
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
| 18 |
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
| 19 |
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
| 20 |
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
| 21 |
+
SOFTWARE.
|
README.md
ADDED
|
@@ -0,0 +1,71 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: mlx-vision
|
| 3 |
+
emoji: 👁️
|
| 4 |
+
colorFrom: blue
|
| 5 |
+
colorTo: green
|
| 6 |
+
sdk: static
|
| 7 |
+
pinned: false
|
| 8 |
+
license: mit
|
| 9 |
+
tags:
|
| 10 |
+
- mlx
|
| 11 |
+
- ocr
|
| 12 |
+
- apple-silicon
|
| 13 |
+
- speculative-decoding
|
| 14 |
+
- dspark
|
| 15 |
+
- deepseek-ocr
|
| 16 |
+
- glm-ocr
|
| 17 |
+
- vision-language-model
|
| 18 |
+
- region:us
|
| 19 |
+
library_name: mlx-vision
|
| 20 |
+
pipeline_tag: image-to-text
|
| 21 |
+
---
|
| 22 |
+
|
| 23 |
+
# mlx-vision
|
| 24 |
+
|
| 25 |
+
**DSpark-style faster OCR on Apple Silicon (MLX).**
|
| 26 |
+
|
| 27 |
+
Speculative decoding for OCR VLMs — DeepSeek-OCR-2, GLM-OCR, Unlimited-OCR.
|
| 28 |
+
Vision/OCR counterpart to [mlx-dspark](https://github.com/ARahim3/mlx-dspark).
|
| 29 |
+
|
| 30 |
+
| Resource | Link |
|
| 31 |
+
|----------|------|
|
| 32 |
+
| **Code** | [github.com/will702/mlx-vision](https://github.com/will702/mlx-vision) |
|
| 33 |
+
| **Agents** | [AGENTS.md](https://github.com/will702/mlx-vision/blob/main/AGENTS.md) · [llms.txt](https://github.com/will702/mlx-vision/blob/main/llms.txt) |
|
| 34 |
+
| **Weights used** | [`mlx-community/DeepSeek-OCR-2-8bit`](https://huggingface.co/mlx-community/DeepSeek-OCR-2-8bit), [`GLM-OCR`](https://huggingface.co/mlx-community/GLM-OCR-bf16), [`Unlimited-OCR`](https://huggingface.co/baidu/Unlimited-OCR) |
|
| 35 |
+
|
| 36 |
+
## Install
|
| 37 |
+
|
| 38 |
+
```bash
|
| 39 |
+
pip install git+https://github.com/will702/mlx-vision.git
|
| 40 |
+
# with PDF support:
|
| 41 |
+
pip install "mlx-vision[pdf] @ git+https://github.com/will702/mlx-vision.git"
|
| 42 |
+
```
|
| 43 |
+
|
| 44 |
+
## Quick start
|
| 45 |
+
|
| 46 |
+
```python
|
| 47 |
+
from mlx_vision import ocr
|
| 48 |
+
|
| 49 |
+
result = ocr("document.png", model="deepseek-ocr-2", mode="auto")
|
| 50 |
+
print(result.text)
|
| 51 |
+
```
|
| 52 |
+
|
| 53 |
+
```bash
|
| 54 |
+
mlx-vision -m deepseek-ocr-2 -i page.png --mode dspark -v
|
| 55 |
+
mlx-vision -m deepseek-ocr-2 -i doc.pdf --page 0 --dpi 200
|
| 56 |
+
```
|
| 57 |
+
|
| 58 |
+
## Modes
|
| 59 |
+
|
| 60 |
+
- `baseline` — mlx-vlm generate
|
| 61 |
+
- `dspark` / `lookup` — training-free semi-AR draft + confidence schedule (OCR DSpark weights not public yet)
|
| 62 |
+
- `dflash` / `eagle3` / `mtp` — neural speculation via mlx-vlm (`--drafter`)
|
| 63 |
+
- `auto` — best available path
|
| 64 |
+
|
| 65 |
+
## Cite
|
| 66 |
+
|
| 67 |
+
See [CITATION.cff](https://github.com/will702/mlx-vision/blob/main/CITATION.cff). Related paper: [DSpark](https://arxiv.org/abs/2607.05147).
|
| 68 |
+
|
| 69 |
+
## License
|
| 70 |
+
|
| 71 |
+
MIT
|
llms.txt
ADDED
|
@@ -0,0 +1,30 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# mlx-vision
|
| 2 |
+
|
| 3 |
+
> DSpark-style speculative decoding for OCR vision-language models on Apple Silicon (MLX).
|
| 4 |
+
|
| 5 |
+
## Summary
|
| 6 |
+
mlx-vision accelerates local OCR VLMs (DeepSeek-OCR-2, GLM-OCR, Unlimited-OCR) using speculative decoding inspired by DeepSeek DSpark. It wraps mlx-vlm with confidence-scheduled drafting, vision feature cache, PDF rasterization, CLI, and an OpenAI-compatible server.
|
| 7 |
+
|
| 8 |
+
## Install
|
| 9 |
+
pip install mlx-vision
|
| 10 |
+
|
| 11 |
+
## CLI
|
| 12 |
+
mlx-vision -m deepseek-ocr-2 -i page.png --mode auto -v
|
| 13 |
+
mlx-vision -m deepseek-ocr-2 -i doc.pdf --page 0
|
| 14 |
+
|
| 15 |
+
## Python
|
| 16 |
+
from mlx_vision import ocr
|
| 17 |
+
result = ocr("page.png", model="deepseek-ocr-2", mode="auto")
|
| 18 |
+
|
| 19 |
+
## Modes
|
| 20 |
+
baseline, lookup, dspark (training-free semi-AR until OCR drafters exist), dflash, eagle3, mtp, auto
|
| 21 |
+
|
| 22 |
+
## Links
|
| 23 |
+
- https://github.com/will702/mlx-vision
|
| 24 |
+
- https://huggingface.co/will702/mlx-vision
|
| 25 |
+
- https://github.com/Blaizzy/mlx-vlm
|
| 26 |
+
- https://github.com/ARahim3/mlx-dspark
|
| 27 |
+
- https://arxiv.org/abs/2607.05147 (DSpark)
|
| 28 |
+
|
| 29 |
+
## Keywords
|
| 30 |
+
mlx, apple-silicon, ocr, deepseek-ocr, glm-ocr, unlimited-ocr, speculative-decoding, dspark, dflash, vision-language-model, metal, faster-inference
|