MaLoW (Qwen3.5-4B, Last 4 Seconds)
Inference checkpoint for Memory as Weights: Internalizing Long-Term History for Streaming Videos.
Evaluation requires the MaLoW code release. The code has been prepared locally; public GitHub publication is pending. Intended repository: MaLoW (forthcoming).
Download the base model separately: Qwen/Qwen3.5-4B. This repository contains inference modules, not a standalone base model.
Evaluation uses the last 4 seconds as the current observation, represented by 4 frames. Earlier history is sampled at 1 FPS with a maximum of 766 history frames, and MaLoW processes 2 frames per memory block. Download with python -m malow.download --model video-4b-last4s --output checkpoints --with-base. The evaluator selects the four-second window automatically from metadata; --preset last4s is the explicit override.
Training code and training data will be released soon.
Follow docs/video_evaluation.md in the MaLoW code release for OVO-Bench and StreamingBench. The JSON metadata records the inference configuration; load it through the MaLoW runtime.
Reported reference results (percent; not a fresh evaluation of this export):
| OVO-Bench backward | Real-time | Forward | Average |
|---|---|---|---|
| 62.54 | 74.67 | 49.32 | 62.17 |
| StreamingBench average | Real-time | Omni | Proactive | SQA |
|---|---|---|---|---|
| 73.51 | 82.28 | 63.7 | 68 | 50 |
Rounding note: the listed rounded StreamingBench subtasks yield 73.50 with weights 2500:1500:250:250, while the supplied reported average is 73.51. Reproduction should use the scorer's unrounded results.
Weights and inference metadata are distributed under Apache-2.0. See LICENSE and NOTICE; base models and benchmark assets retain their respective upstream terms.