MaLoW (Qwen3.5-4B)
Inference checkpoint for Memory as Weights: Internalizing Long-Term History for Streaming Videos.
Evaluation requires the MaLoW code release. The code has been prepared locally; public GitHub publication is pending. Intended repository: MaLoW (forthcoming).
Download the base model separately: Qwen/Qwen3.5-4B. This repository contains inference modules, not a standalone base model.
Follow docs/video_evaluation.md in the MaLoW code release for OVO-Bench and StreamingBench. The JSON metadata records the inference configuration; load it through the MaLoW runtime.
Reported reference results (percent; not a fresh evaluation of this export):
| OVO-Bench backward | Real-time | Forward | Average |
|---|---|---|---|
| 61.15 | 70.52 | 50.82 | 60.83 |
| StreamingBench average | Real-time | Omni | Proactive | SQA |
|---|---|---|---|---|
| 71.18 | 80.8 | 60.2 | 56.8 | 55.2 |
Overall computed from listed subtask scores using benchmark weights 2500:1500:250:250.
Weights and inference metadata are distributed under Apache-2.0. See LICENSE and NOTICE; base models and benchmark assets retain their respective upstream terms.