Instructions to use MATLOWAI/MiniMax-H3-ORB360-CardSpin with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Inference
- Notebooks
- Google Colab
- Kaggle
Add ORB360 step 1500 orbit LoRA and CardSpin v2; new cat videos first; four-way 360 sharpness comparison
Browse filesminimax_h3_orb360_step1500 (high-resolution orbit LoRA, heights, multi-reference) and minimax_h3_orb360_cardspin_v2_step50 (step 1500 + 50 card-spin steps). README: new videos at the top, table and detail crops showing the CardSpin variants are much softer for plain 360s, v1 videos kept below. Example prompts in the step-1500 caption format; NOTICE updated.
- .gitattributes +3 -0
- NOTICE +4 -1
- README.md +142 -67
- assets/whitecat_cardspin_v2_step50.mp4 +3 -0
- assets/whitecat_cardspin_v2_step50_poster.jpg +0 -0
- assets/whitecat_orbit_detail_4way.jpg +3 -0
- assets/whitecat_orbit_step1500.mp4 +3 -0
- assets/whitecat_orbit_step1500_poster.jpg +0 -0
- minimax_h3_orb360_cardspin_v2_step50.safetensors +3 -0
- minimax_h3_orb360_step1500.safetensors +3 -0
- prompts/orbit_crossheight_front_back_example.txt +31 -0
- prompts/orbit_underneath_3ref_example.txt +31 -0
.gitattributes
CHANGED
|
@@ -38,3 +38,6 @@ assets/whitecat_cardspin_step50.mp4 filter=lfs diff=lfs merge=lfs -text
|
|
| 38 |
assets/whitecat_orbit_step50.mp4 filter=lfs diff=lfs merge=lfs -text
|
| 39 |
examples/herschel_1867_cameron_met263166.png filter=lfs diff=lfs merge=lfs -text
|
| 40 |
examples/whitecat.png filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
assets/whitecat_orbit_step50.mp4 filter=lfs diff=lfs merge=lfs -text
|
| 39 |
examples/herschel_1867_cameron_met263166.png filter=lfs diff=lfs merge=lfs -text
|
| 40 |
examples/whitecat.png filter=lfs diff=lfs merge=lfs -text
|
| 41 |
+
assets/whitecat_cardspin_v2_step50.mp4 filter=lfs diff=lfs merge=lfs -text
|
| 42 |
+
assets/whitecat_orbit_detail_4way.jpg filter=lfs diff=lfs merge=lfs -text
|
| 43 |
+
assets/whitecat_orbit_step1500.mp4 filter=lfs diff=lfs merge=lfs -text
|
NOTICE
CHANGED
|
@@ -1,3 +1,6 @@
|
|
| 1 |
MiniMax H3 is licensed under the MiniMax H3 Community License Agreement, Copyright © 2026 MiniMax. All Rights Reserved.
|
| 2 |
|
| 3 |
-
Modified by MatlowAI
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
MiniMax H3 is licensed under the MiniMax H3 Community License Agreement, Copyright © 2026 MiniMax. All Rights Reserved.
|
| 2 |
|
| 3 |
+
Modified by MatlowAI (September 2026). This repository contains adapter (LoRA) weights, not base-model weights:
|
| 4 |
+
- minimax_h3_orb360_cardspin_step50.safetensors: ORB360 orbit LoRA fine-tuning (750 updates) followed by CardSpin fine-tuning (+50 updates).
|
| 5 |
+
- minimax_h3_orb360_step1500.safetensors: ORB360 orbit LoRA fine-tuning (1500 updates) on a larger, higher-resolution rendered orbit set.
|
| 6 |
+
- minimax_h3_orb360_cardspin_v2_step50.safetensors: the step-1500 orbit LoRA followed by CardSpin fine-tuning (+50 updates).
|
README.md
CHANGED
|
@@ -11,46 +11,92 @@ tags:
|
|
| 11 |
- ref2va
|
| 12 |
- orbit
|
| 13 |
- camera-control
|
|
|
|
| 14 |
- musubi-tuner
|
| 15 |
- comfyui
|
| 16 |
pipeline_tag: image-to-video
|
| 17 |
---
|
| 18 |
|
| 19 |
-
# MiniMax-H3 ORB360
|
| 20 |
|
| 21 |
-
|
|
|
|
|
|
|
| 22 |
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
|
|
|
|
|
|
|
|
|
| 28 |
|
| 29 |
Powered by MiniMax H3.
|
| 30 |
|
| 31 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 32 |
|
| 33 |
<video controls muted playsinline loop preload="metadata" width="100%"
|
| 34 |
poster="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_step50_poster.jpg">
|
| 35 |
<source src="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_step50.mp4" type="video/mp4">
|
| 36 |
-
<a href="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_step50.mp4">Download the card-spin clip (mp4)</a>
|
| 37 |
</video>
|
| 38 |
|
| 39 |
-
|
| 40 |
-
20 steps, seed 20260926.
|
| 41 |
-
|
| 42 |
-
## Same file, normal orbit
|
| 43 |
|
| 44 |
<video controls muted playsinline loop preload="metadata" width="100%"
|
| 45 |
poster="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step50_poster.jpg">
|
| 46 |
<source src="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step50.mp4" type="video/mp4">
|
| 47 |
-
<a href="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step50.mp4">Download the orbit clip (mp4)</a>
|
| 48 |
</video>
|
| 49 |
|
| 50 |
-
|
| 51 |
-
card-spin steps did not break the orbit.
|
| 52 |
-
|
| 53 |
-
## How it happened
|
| 54 |
|
| 55 |
We were testing an orbit LoRA (the ORB360 project) on real photos instead of the grey Blender renders it was trained
|
| 56 |
on, and fed it an 1867 portrait of Sir John Herschel by Julia Margaret Cameron. Instead of orbiting a man, it decided
|
|
@@ -63,17 +109,15 @@ the *photograph* was the object. This is the clip that started it all:
|
|
| 63 |
</video>
|
| 64 |
|
| 65 |
It was too good to leave as a one-off. So we took that single generated video, wrote a caption describing what
|
| 66 |
-
happens in it with timestamps, and trained **50 more steps** on top of the
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
slightly crisper card edges, but 50 has the most charm and is the lightest touch on the orbit behaviour, so that is
|
| 70 |
-
the one released here.
|
| 71 |
|
| 72 |
## Using it
|
| 73 |
|
| 74 |
-
**ComfyUI:** load
|
| 75 |
-
|
| 76 |
-
MiniMax-H3 weights. Give it
|
| 77 |
|
| 78 |
**musubi-tuner** ([kohya-ss/musubi-tuner](https://github.com/kohya-ss/musubi-tuner), v0.3.5 or later):
|
| 79 |
|
|
@@ -82,27 +126,52 @@ python src/musubi_tuner/minimax_h3_generate_video.py --task ref2va \
|
|
| 82 |
--dit minimax_h3_ref2va_bf16.safetensors --prune_adaln \
|
| 83 |
--video_vae minimax_h3_video_vae_fp16.safetensors --audio_vae minimax_h3_audio_vae_fp32.safetensors \
|
| 84 |
--text_encoder qwen3vl_32b_minimax_h3_int8_convrot.safetensors --text_encoder_attn_mode sdpa \
|
| 85 |
-
--lora_weight
|
| 86 |
-
--video_size
|
| 87 |
-
--prompt "$(cat prompts/
|
| 88 |
```
|
| 89 |
|
| 90 |
-
|
| 91 |
-
|
| 92 |
-
|
| 93 |
-
-
|
| 94 |
-
|
| 95 |
-
-
|
| 96 |
-
|
| 97 |
-
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 98 |
- The prompts follow MiniMax's official Ref2VA prompt layout (`subject_definitions`, `summary`,
|
| 99 |
`retention_analysis`, `detailed_description`, `overall_soundscape`, `non_diegetic_music`). Both audio sections are
|
| 100 |
`N/A`, which asks for silence; without them H3 tends to invent a soundtrack.
|
| 101 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 102 |
## Training
|
| 103 |
|
| 104 |
-
|
| 105 |
-
`--resume`) with the same settings:
|
| 106 |
|
| 107 |
| | |
|
| 108 |
|---|---|
|
|
@@ -114,45 +183,51 @@ Both stages used musubi-tuner (v0.3.5, `--task ref2va`; one local patch that onl
|
|
| 114 |
| Batch | 1, no accumulation, video only (no audio loss), seed 20260926 |
|
| 115 |
| Hardware | one NVIDIA RTX PRO 6000 (96 GB) |
|
| 116 |
|
| 117 |
-
**
|
| 118 |
-
|
| 119 |
-
|
| 120 |
-
|
| 121 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 122 |
|
| 123 |
-
**
|
| 124 |
-
|
| 125 |
-
|
| 126 |
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
- Learned from a single example. The card always turns the same way with roughly the same timing, and some subjects
|
| 130 |
-
or seeds may commit to the card less than others.
|
| 131 |
-
- The first stage saw only four rendered assets at 512 x 512, so orbit quality on complex real scenes varies.
|
| 132 |
-
- Evaluated by eye on a handful of images and seeds, not with a benchmark.
|
| 133 |
|
| 134 |
-
##
|
| 135 |
|
| 136 |
-
|
| 137 |
-
|
| 138 |
-
|
| 139 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 140 |
|
| 141 |
## Files
|
| 142 |
|
| 143 |
-
- `
|
| 144 |
-
- `
|
|
|
|
|
|
|
| 145 |
- `examples/whitecat.png`: reference image for the cat clips
|
| 146 |
- `examples/herschel_1867_cameron_met263166.png`: the Herschel portrait (Julia Margaret Cameron, 1867; The
|
| 147 |
Metropolitan Museum of Art, Open Access, CC0), resized
|
| 148 |
-
- `assets/`: the
|
| 149 |
- `LICENSE` (MiniMax H3 Community License Agreement), `NOTICE`
|
| 150 |
|
| 151 |
## Base model and licence
|
| 152 |
|
| 153 |
-
|
| 154 |
-
their base weights, and
|
| 155 |
-
[MiniMax H3 Community License Agreement](LICENSE), including its acceptable-use, distribution, commercial and
|
| 156 |
-
territorial provisions. **
|
| 157 |
-
`NOTICE` for attribution and the modification notice. Thanks to MiniMax for releasing H3, to Ostris for the
|
| 158 |
-
adapter, and to kohya-ss for musubi-tuner.
|
|
|
|
| 11 |
- ref2va
|
| 12 |
- orbit
|
| 13 |
- camera-control
|
| 14 |
+
- novel-view
|
| 15 |
- musubi-tuner
|
| 16 |
- comfyui
|
| 17 |
pipeline_tag: image-to-video
|
| 18 |
---
|
| 19 |
|
| 20 |
+
# MiniMax-H3 ORB360: 360 orbit + CardSpin
|
| 21 |
|
| 22 |
+
Rank-32 LoRAs for MiniMax-H3 **Ref2VA**. Give it a photo, get a smooth clockwise 360-degree camera orbit around the
|
| 23 |
+
frozen subject (`ORB360_CW`), or, with the CardSpin files, turn the photo in space like a thin physical card
|
| 24 |
+
(`ORB360_CARDSPIN`).
|
| 25 |
|
| 26 |
+
| File | What it is | Use it for |
|
| 27 |
+
|---|---|---|
|
| 28 |
+
| `minimax_h3_orb360_step1500.safetensors` | **New.** ORB360 orbit LoRA, 1500 steps on high-resolution Blender orbits, camera heights from below to above, 1-4 labelled reference pictures | **Clean 360 orbits** (sharpest; our pick for multi-view use) |
|
| 29 |
+
| `minimax_h3_orb360_cardspin_v2_step50.safetensors` | **New.** CardSpin v2 = step 1500 + 50 card-spin steps | The card spin |
|
| 30 |
+
| `minimax_h3_orb360_cardspin_step50.safetensors` | CardSpin v1 (the original release) = the old 512 x 512 orbit LoRA + 50 card-spin steps | The original card spin |
|
| 31 |
+
|
| 32 |
+
Both CardSpin files can still orbit with the orbit prompt, but they are noticeably softer; see
|
| 33 |
+
[Which file for a clean 360?](#which-file-for-a-clean-360).
|
| 34 |
|
| 35 |
Powered by MiniMax H3.
|
| 36 |
|
| 37 |
+
## New: CardSpin v2
|
| 38 |
+
|
| 39 |
+
<video controls muted playsinline loop preload="metadata" width="100%"
|
| 40 |
+
poster="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_v2_step50_poster.jpg">
|
| 41 |
+
<source src="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_v2_step50.mp4" type="video/mp4">
|
| 42 |
+
<a href="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_v2_step50.mp4">Download the CardSpin v2 clip (mp4)</a>
|
| 43 |
+
</video>
|
| 44 |
+
|
| 45 |
+
`minimax_h3_orb360_cardspin_v2_step50.safetensors`, one reference image (`examples/whitecat.png`), the prompt in
|
| 46 |
+
`prompts/cardspin_caption.txt`, 1024 x 768, 124 frames, 20 steps, seed 20260926.
|
| 47 |
+
|
| 48 |
+
## New: the 360 from ORB360 step 1500
|
| 49 |
+
|
| 50 |
+
<video controls muted playsinline loop preload="metadata" width="100%"
|
| 51 |
+
poster="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step1500_poster.jpg">
|
| 52 |
+
<source src="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step1500.mp4" type="video/mp4">
|
| 53 |
+
<a href="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step1500.mp4">Download the step-1500 orbit (mp4)</a>
|
| 54 |
+
</video>
|
| 55 |
+
|
| 56 |
+
`minimax_h3_orb360_step1500.safetensors`, same image and seed, prompt `prompts/orbit_realscene_caption.txt`. It was
|
| 57 |
+
trained only on grey-backdrop Blender renders, and it keeps the real snowy scene around the cat.
|
| 58 |
+
|
| 59 |
+
## Which file for a clean 360?
|
| 60 |
+
|
| 61 |
+
The same orbit (same photo, prompt, seed, 1024 x 768, 124 frames) from four LoRAs. Numbers are per-frame image
|
| 62 |
+
statistics averaged over the clip; "detail vs step 1500" is the median per-frame ratio of the Laplacian variance (fine
|
| 63 |
+
detail) to the step-1500 clip.
|
| 64 |
+
|
| 65 |
+
| White cat 360 | Orbit step 750 (old, not in this repo) | CardSpin v1 (750 + 50) | **ORB360 step 1500** | CardSpin v2 (1500 + 50) |
|
| 66 |
+
|---|---|---|---|---|
|
| 67 |
+
| Fine detail (Laplacian variance) | 30.2 | 19.8 | **56.1** | 20.7 |
|
| 68 |
+
| Detail vs step 1500 | 0.54x | 0.34x | **1.00x** | 0.35x |
|
| 69 |
+
| Edge energy (Tenengrad) | 1098 | 809 | **1343** | 841 |
|
| 70 |
+
| High-frequency share | 0.0035 | 0.0028 | **0.0050** | 0.0029 |
|
| 71 |
+
| First frame vs the photo, PSNR / SSIM | 28.0 dB / 0.87 | 24.5 dB / 0.85 | **28.5 dB / 0.87** | 22.5 dB / 0.77 |
|
| 72 |
+
|
| 73 |
+

|
| 74 |
+
|
| 75 |
+
- The high-resolution training nearly doubles the fine detail (step 750 -> step 1500: 1.85x, higher in all 124 frames).
|
| 76 |
+
- **The 50 card-spin steps soften everything**, whichever orbit LoRA they start from: both CardSpin files keep only
|
| 77 |
+
about a third of step 1500's fine detail and drift further from the photo in the first frame. The single card-spin
|
| 78 |
+
training clip is a soft generated video, and 50 steps on it most likely pull the whole look towards it.
|
| 79 |
+
- So: **step 1500 for orbits, CardSpin only for the card trick.**
|
| 80 |
+
|
| 81 |
+
## The original v1 videos
|
| 82 |
+
|
| 83 |
+
CardSpin v1 (`minimax_h3_orb360_cardspin_step50.safetensors`), card-spin prompt:
|
| 84 |
|
| 85 |
<video controls muted playsinline loop preload="metadata" width="100%"
|
| 86 |
poster="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_step50_poster.jpg">
|
| 87 |
<source src="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_step50.mp4" type="video/mp4">
|
| 88 |
+
<a href="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_step50.mp4">Download the v1 card-spin clip (mp4)</a>
|
| 89 |
</video>
|
| 90 |
|
| 91 |
+
Same file, orbit prompt:
|
|
|
|
|
|
|
|
|
|
| 92 |
|
| 93 |
<video controls muted playsinline loop preload="metadata" width="100%"
|
| 94 |
poster="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step50_poster.jpg">
|
| 95 |
<source src="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step50.mp4" type="video/mp4">
|
| 96 |
+
<a href="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step50.mp4">Download the v1 orbit clip (mp4)</a>
|
| 97 |
</video>
|
| 98 |
|
| 99 |
+
## How the card spin happened
|
|
|
|
|
|
|
|
|
|
| 100 |
|
| 101 |
We were testing an orbit LoRA (the ORB360 project) on real photos instead of the grey Blender renders it was trained
|
| 102 |
on, and fed it an 1867 portrait of Sir John Herschel by Julia Margaret Cameron. Instead of orbiting a man, it decided
|
|
|
|
| 109 |
</video>
|
| 110 |
|
| 111 |
It was too good to leave as a one-off. So we took that single generated video, wrote a caption describing what
|
| 112 |
+
happens in it with timestamps, and trained **50 more steps** on top of the orbit LoRA using only that one clip. After
|
| 113 |
+
50 steps the effect transferred to photos it had never seen: other Cameron portraits, and a colour photo of a cat in a
|
| 114 |
+
hat. CardSpin v1 did this on the old 750-step orbit LoRA; CardSpin v2 repeats the same recipe on top of step 1500.
|
|
|
|
|
|
|
| 115 |
|
| 116 |
## Using it
|
| 117 |
|
| 118 |
+
**ComfyUI:** load one file with the standard **Load LoRA** node (model only, strength 1.0) on a MiniMax-H3 **Ref2VA**
|
| 119 |
+
model. These are kohya-style LoRAs (`lora_unet_*`, `lora_down`/`lora_up`/`alpha`); all 200 modules map onto ComfyUI's
|
| 120 |
+
MiniMax-H3 weights. Give it the reference image(s) and paste one of the prompts from `prompts/` as the text.
|
| 121 |
|
| 122 |
**musubi-tuner** ([kohya-ss/musubi-tuner](https://github.com/kohya-ss/musubi-tuner), v0.3.5 or later):
|
| 123 |
|
|
|
|
| 126 |
--dit minimax_h3_ref2va_bf16.safetensors --prune_adaln \
|
| 127 |
--video_vae minimax_h3_video_vae_fp16.safetensors --audio_vae minimax_h3_audio_vae_fp32.safetensors \
|
| 128 |
--text_encoder qwen3vl_32b_minimax_h3_int8_convrot.safetensors --text_encoder_attn_mode sdpa \
|
| 129 |
+
--lora_weight minimax_h3_orb360_step1500.safetensors --lora_multiplier 1.0 --lora_runtime_attach \
|
| 130 |
+
--video_size 768 1024 --video_length 124 --infer_steps 20 --attn_mode sdpa --seed 20260926 \
|
| 131 |
+
--prompt "$(cat prompts/orbit_realscene_caption.txt)" --ref your_photo.png --save_path out/
|
| 132 |
```
|
| 133 |
|
| 134 |
+
For the card spin, use `minimax_h3_orb360_cardspin_v2_step50.safetensors` and `prompts/cardspin_caption.txt`.
|
| 135 |
+
|
| 136 |
+
- The first `--ref` becomes `<Picture 1>`: the first and last frame of the orbit.
|
| 137 |
+
- Keep **124 frames at 24 fps** for these prompts: their timestamps (for example "At 1.708333 seconds (zero-based
|
| 138 |
+
frame 41) ... 120 degrees") assume that length.
|
| 139 |
+
- Prompts:
|
| 140 |
+
- `orbit_realscene_caption.txt`: one photo, keeps the photo's real surroundings.
|
| 141 |
+
- `orbit_3ref_caption.txt`: three views of the same subject (front, 120 and 240 degrees clockwise), for example
|
| 142 |
+
three renders or a turnaround sheet. With real side and back views the model does not have to invent them.
|
| 143 |
+
- `orbit_underneath_3ref_example.txt`, `orbit_crossheight_front_back_example.txt`: the exact caption format step
|
| 144 |
+
1500 was trained on, taken from two of our held-out tests (an orbit from 30 degrees below with no floor, and front
|
| 145 |
+
and back pictures taken from 45 degrees with the orbit at 5 degrees). The numbers in them (camera distance, how
|
| 146 |
+
much of the frame the subject fills) describe those test objects; edit the heights, times and angles for your own
|
| 147 |
+
use.
|
| 148 |
+
- `cardspin_caption.txt`: the card spin.
|
| 149 |
+
- Resolutions we have seen work: 800 x 800, 1024 x 768 and 1152 x 768 landscapes, 672 x 832 and 832 x 1024
|
| 150 |
+
portraits.
|
| 151 |
- The prompts follow MiniMax's official Ref2VA prompt layout (`subject_definitions`, `summary`,
|
| 152 |
`retention_analysis`, `detailed_description`, `overall_soundscape`, `non_diegetic_music`). Both audio sections are
|
| 153 |
`N/A`, which asks for silence; without them H3 tends to invent a soundtrack.
|
| 154 |
|
| 155 |
+
## How well step 1500 orbits
|
| 156 |
+
|
| 157 |
+
Tested on three objects that were never rendered for training (a vintage film camera, a stylised character and a
|
| 158 |
+
stone cat statue), one seed per test, against Blender renders of the true orbit:
|
| 159 |
+
|
| 160 |
+
| Test | Result |
|
| 161 |
+
|---|---|
|
| 162 |
+
| Orbit at 20 degrees, 3 views (0/120/240) | Within one frame (2.9 degrees) of the true orbit on every frame |
|
| 163 |
+
| Orbits from below (30 and 10 degrees below the centre, no floor) | Right height on 118-124 of 124 frames, one clean turn |
|
| 164 |
+
| Pictures taken from 45 degrees, orbit asked for at 5 degrees | Low orbit on all three objects (most frames within 5-10 degrees of the target), one turn |
|
| 165 |
+
| High orbit (45 degrees) from one front picture | High orbit (30-45 degrees on most frames); the unseen sides are invented |
|
| 166 |
+
| Pictures taken from 45 degrees, orbit asked for at 30 degrees below | **Fails**: drifts back towards the pictures' height |
|
| 167 |
+
| Height changing during the clip (for example 45 above to 30 below) | **Fails**: the object tumbles instead of the camera orbiting |
|
| 168 |
+
|
| 169 |
+
The old 750-step LoRA mostly copies the height of the reference pictures: it gets the underneath orbits right when the
|
| 170 |
+
pictures are taken from below, but none of the "pictures from 45 degrees, orbit lower" cases.
|
| 171 |
+
|
| 172 |
## Training
|
| 173 |
|
| 174 |
+
All stages used musubi-tuner (v0.3.5, `--task ref2va`; small local wrappers for resumable segments) with the same settings:
|
|
|
|
| 175 |
|
| 176 |
| | |
|
| 177 |
|---|---|
|
|
|
|
| 183 |
| Batch | 1, no accumulation, video only (no audio loss), seed 20260926 |
|
| 184 |
| Hardware | one NVIDIA RTX PRO 6000 (96 GB) |
|
| 185 |
|
| 186 |
+
**ORB360 step 1500 (new), from scratch.** 23 assets (Poly Haven CC0 props and a few stylised character models),
|
| 187 |
+
rendered in Blender (Cycles) as 345 one-turn clockwise orbits at three lengths, each at the highest resolution that
|
| 188 |
+
fits in memory: 124 frames at 800 x 800, 73 frames at 1024 x 1024, and 22 frames at 1600 x 1600. Camera heights from
|
| 189 |
+
45 degrees below to 55 degrees above the subject's centre (the ones below with the floor switched off), tight and wide
|
| 190 |
+
framing, varied start angles. Each clip appears in three training rows with different reference pictures (the 73-frame clips lost their four-picture sets to fit in memory): two sets
|
| 191 |
+
taken from the clip itself (front only, front and back, front and side, front, side and back, a four-way turnaround,
|
| 192 |
+
or thirds at 0/120/240 degrees) and one set rendered from a different height than the orbit. The captions label every
|
| 193 |
+
picture with its exact time, frame and angle, and describe the camera's height, distance, framing and lens. 960 rows,
|
| 194 |
+
1500 steps, about 62 s per step (about 26 hours).
|
| 195 |
|
| 196 |
+
**CardSpin v2, +50 steps.** Initialised from the step-1500 weights (fresh optimizer). One training clip: the Herschel
|
| 197 |
+
glitch above (generated at 832 x 1024, trained at 672 x 832, 124 frames), with the original Herschel photo as the
|
| 198 |
+
single reference and the caption in `prompts/cardspin_caption.txt`.
|
| 199 |
|
| 200 |
+
**Original release (CardSpin v1).** Stage 1: four Blender assets, one 124-frame, 512 x 512 orbit each, three
|
| 201 |
+
reference pictures at 0, 120 and 240 degrees, 750 steps. Stage 2: the same 50 card-spin steps as above.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 202 |
|
| 203 |
+
## Limitations
|
| 204 |
|
| 205 |
+
- The card spin was learned from a single example: it always turns the same way with roughly the same timing, and
|
| 206 |
+
some subjects or seeds commit to the card less than others. Both CardSpin files soften the image (table above).
|
| 207 |
+
- Step 1500 was trained on grey-backdrop renders of 23 objects. It orbits real photos well in our tests, but was
|
| 208 |
+
evaluated on a handful of images and three held-out objects, one seed each, not a large benchmark.
|
| 209 |
+
- Big jumps between the pictures' height and the orbit's height, and orbits whose height changes during the clip,
|
| 210 |
+
do not work yet.
|
| 211 |
+
- From one picture, the sides the camera has not seen are invented. More pictures (front and back, a turnaround, or
|
| 212 |
+
three views) fix that.
|
| 213 |
|
| 214 |
## Files
|
| 215 |
|
| 216 |
+
- `minimax_h3_orb360_step1500.safetensors`: ORB360 orbit LoRA, step 1500 (fp32, 597 MB)
|
| 217 |
+
- `minimax_h3_orb360_cardspin_v2_step50.safetensors`: CardSpin v2 (fp32, 597 MB)
|
| 218 |
+
- `minimax_h3_orb360_cardspin_step50.safetensors`: CardSpin v1, the original release (fp32, 597 MB)
|
| 219 |
+
- `prompts/`: the prompts described above
|
| 220 |
- `examples/whitecat.png`: reference image for the cat clips
|
| 221 |
- `examples/herschel_1867_cameron_met263166.png`: the Herschel portrait (Julia Margaret Cameron, 1867; The
|
| 222 |
Metropolitan Museum of Art, Open Access, CC0), resized
|
| 223 |
+
- `assets/`: the videos above, their poster frames, and the four-way detail comparison
|
| 224 |
- `LICENSE` (MiniMax H3 Community License Agreement), `NOTICE`
|
| 225 |
|
| 226 |
## Base model and licence
|
| 227 |
|
| 228 |
+
These are LoRAs for [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3) by MiniMax; they do nothing without
|
| 229 |
+
their base weights, and the card-spin stages were trained on a clip generated with them. They are distributed under
|
| 230 |
+
the [MiniMax H3 Community License Agreement](LICENSE), including its acceptable-use, distribution, commercial and
|
| 231 |
+
territorial provisions. **They are not MIT.** Read the complete upstream terms; this repository does not expand them.
|
| 232 |
+
See `NOTICE` for attribution and the modification notice. Thanks to MiniMax for releasing H3, to Ostris for the
|
| 233 |
+
training adapter, and to kohya-ss for musubi-tuner.
|
assets/whitecat_cardspin_v2_step50.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:dbb7b3a1af06b1346ae22dbbfdcdf6801118ae4cd4d2cf29d845b054b8d05f9a
|
| 3 |
+
size 3268303
|
assets/whitecat_cardspin_v2_step50_poster.jpg
ADDED
|
assets/whitecat_orbit_detail_4way.jpg
ADDED
|
Git LFS Details
|
assets/whitecat_orbit_step1500.mp4
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:13baf0d556bc530715498661bea5319c5ec04698df7d1e1b8736c34757afd2f9
|
| 3 |
+
size 6001372
|
assets/whitecat_orbit_step1500_poster.jpg
ADDED
|
minimax_h3_orb360_cardspin_v2_step50.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0e405410ed07c8c63cfb28fa5de3b1f1447b2165fdf5099542b9ad64194a053b
|
| 3 |
+
size 596449416
|
minimax_h3_orb360_step1500.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e9275b2f456f6bd004c9cb3cf5c636cad87217c5bc96e6a5b5533e917de121da
|
| 3 |
+
size 596450168
|
prompts/orbit_crossheight_front_back_example.txt
ADDED
|
@@ -0,0 +1,31 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
subject_definitions:
|
| 2 |
+
<Subject 1> is the exact same rigid subject shown in <Picture 1> and <Picture 2>. Its geometry, proportions, materials, colors, markings, and surface details remain unchanged.
|
| 3 |
+
<Picture 1> shows <Subject 1> from the direction the orbit starts from, but from a higher camera: 45 degrees above the subject's centre instead of the orbit's 5 degrees above. It is not a frame of the video.
|
| 4 |
+
|
| 5 |
+
summary:
|
| 6 |
+
[reference generation] ORB360_CW. A turntable-style camera orbit of <Subject 1> in front of a plain grey seamless backdrop: one full clockwise turn in 5.125000 seconds, filmed from a low, near eye-level camera 5 degrees above the subject, while the reference pictures were taken from a higher camera at 45 degrees.
|
| 7 |
+
|
| 8 |
+
retention_analysis:
|
| 9 |
+
<Subject 1> (appears in [Shot 1]): fully_preserved - its geometry, proportions, materials, colors, markings and surface details are retained, and it stays completely frozen.
|
| 10 |
+
<Picture 1> (not a frame of [Shot 1]; seen from 0 degrees clockwise of the starting direction at 45 degrees elevation): fully_preserved - its geometry, materials and surface details define <Subject 1>; the orbit passes the same side at 5 degrees elevation.
|
| 11 |
+
<Picture 2> (not a frame of [Shot 1]; seen from 180 degrees clockwise of the starting direction at 45 degrees elevation): fully_preserved - its geometry, materials and surface details define <Subject 1>; the orbit passes the same side at 5 degrees elevation.
|
| 12 |
+
|
| 13 |
+
detailed_description:
|
| 14 |
+
[Shot 1] The shot begins facing <Subject 1> from the same direction as <Picture 1>, but from 5 degrees elevation, lower than <Picture 1>'s 45 degrees. The camera performs exactly one smooth clockwise 360-degree orbit around the completely stationary <Subject 1>, at constant radius, constant elevation, constant focal length, constant speed, and constant framing, ending at the starting viewpoint. The subject remains perfectly rigid and stationary. No articulation, deformation, secondary motion, wind, zoom, dolly, camera roll, cuts, lighting changes, or scene changes.
|
| 15 |
+
Camera: a low, near eye-level camera, 5 degrees above the subject's centre, looking down 5 degrees at it throughout the orbit. The camera is below the top of the subject, at about 68% of its height above the ground, about 3.8 times the subject's bounding radius from its centre. At the widest view the subject spans about 90% of the frame, and about 85% on average. 60 mm lens on a full-frame sensor (33.4-degree field of view), no roll, no zoom.
|
| 16 |
+
The background is a plain, evenly lit grey seamless backdrop with only a soft contact shadow under the subject.
|
| 17 |
+
|
| 18 |
+
All 5 milestones below belong to the same uninterrupted [Shot 1], not separate shots. The video runs at 24 fps and has 124 frames. The full turn takes 5.125000 seconds (frame 0 to frame 123), a constant 70.2 degrees per second (2.9268 degrees per frame). Angles describe camera motion clockwise as viewed from above, relative to the starting camera azimuth; the subject itself does not rotate. The camera moves continuously at constant angular speed between these milestones, without pausing, reversing or completing extra turns.
|
| 19 |
+
At 0.000000 seconds (zero-based frame 0), the camera has advanced 0 degrees clockwise from its starting azimuth; it faces the same side of <Subject 1> as <Picture 1> (0 degrees), from 5 degrees elevation instead of 45.
|
| 20 |
+
At 1.708333 seconds (zero-based frame 41), the camera has advanced 120 degrees clockwise from its starting azimuth.
|
| 21 |
+
At 2.583333 seconds (zero-based frame 62), the camera has advanced 181.5 degrees clockwise from its starting azimuth; it faces the same side of <Subject 1> as <Picture 2> (180 degrees), from 5 degrees elevation instead of 45.
|
| 22 |
+
At 3.416667 seconds (zero-based frame 82), the camera has advanced 240 degrees clockwise from its starting azimuth.
|
| 23 |
+
At 5.125000 seconds (zero-based frame 123), the camera has advanced 360 degrees clockwise from its starting azimuth; it faces the same side of <Subject 1> as <Picture 1> (0 degrees), from 5 degrees elevation instead of 45.
|
| 24 |
+
|
| 25 |
+
The entire subject is a single frozen three-dimensional scene throughout the shot. Every visible eyelid keeps its reference position: no blinking. Each eye keeps a fixed gaze relative to the head and never tracks the moving camera. Facial expression, mouth, head, hands, feet and body pose remain unchanged. Hair strands, clothing, ribbons, accessories and loose parts are rigidly fixed in the same positions. There is no breathing, sway, settling or secondary motion. Only the camera changes position. Camera radius, elevation, focal length and look-at point stay fixed. Lighting stays fixed in world space; normal viewpoint-dependent shading and reflections remain physically consistent.
|
| 26 |
+
|
| 27 |
+
overall_soundscape:
|
| 28 |
+
N/A
|
| 29 |
+
|
| 30 |
+
non_diegetic_music:
|
| 31 |
+
N/A
|
prompts/orbit_underneath_3ref_example.txt
ADDED
|
@@ -0,0 +1,31 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
subject_definitions:
|
| 2 |
+
<Subject 1> is the exact same rigid subject shown in <Picture 1>, <Picture 2>, and <Picture 3>. Its geometry, proportions, materials, colors, markings, and surface details remain unchanged.
|
| 3 |
+
<Picture 1> is the first-frame viewpoint of [Shot 1].
|
| 4 |
+
|
| 5 |
+
summary:
|
| 6 |
+
[reference generation] ORB360_CW. A turntable-style camera orbit of <Subject 1> floating in a plain grey space with no floor: one full clockwise turn in 5.125000 seconds, filmed from a camera far below the subject, looking steeply up 30 degrees below the subject.
|
| 7 |
+
|
| 8 |
+
retention_analysis:
|
| 9 |
+
<Subject 1> (appears in [Shot 1]): fully_preserved - its geometry, proportions, materials, colors, markings and surface details are retained, and it stays completely frozen.
|
| 10 |
+
<Picture 1> ([Shot 1] first frame and last frame): fully_preserved - the orbit starts and ends on this view.
|
| 11 |
+
<Picture 2> ([Shot 1] keyframe at 1.708333 seconds, zero-based frame 41, 120 degrees clockwise from <Picture 1>): fully_preserved - the camera passes through this exact view.
|
| 12 |
+
<Picture 3> ([Shot 1] keyframe at 3.416667 seconds, zero-based frame 82, 240 degrees clockwise from <Picture 1>): fully_preserved - the camera passes through this exact view.
|
| 13 |
+
|
| 14 |
+
detailed_description:
|
| 15 |
+
[Shot 1] The shot begins from <Picture 1>. The camera performs exactly one smooth clockwise 360-degree orbit around the completely stationary <Subject 1>, at constant radius, constant elevation, constant focal length, constant speed, and constant framing, ending at the starting viewpoint. The subject remains perfectly rigid and stationary. No articulation, deformation, secondary motion, wind, zoom, dolly, camera roll, cuts, lighting changes, or scene changes.
|
| 16 |
+
Camera: a camera far below the subject, looking steeply up, 30 degrees below the subject's centre, looking up 30 degrees at it throughout the orbit. The camera is below the level of the subject's base, about 0.7 times the subject's height beneath it, about 3.7 times the subject's bounding radius from its centre. At the widest view the subject spans about 90% of the frame, and about 83% on average. 60 mm lens on a full-frame sensor (33.4-degree field of view), no roll, no zoom.
|
| 17 |
+
There is no floor: <Subject 1> floats in a plain, evenly lit grey space with no ground and no shadow, so it can be seen from below.
|
| 18 |
+
|
| 19 |
+
All 4 milestones below belong to the same uninterrupted [Shot 1], not separate shots. The video runs at 24 fps and has 124 frames. The full turn takes 5.125000 seconds (frame 0 to frame 123), a constant 70.2 degrees per second (2.9268 degrees per frame). Angles describe camera motion clockwise as viewed from above, relative to the starting camera azimuth; the subject itself does not rotate. The camera moves continuously at constant angular speed between these milestones, without pausing, reversing or completing extra turns.
|
| 20 |
+
At 0.000000 seconds (zero-based frame 0), the camera has advanced 0 degrees clockwise from its starting azimuth; the viewpoint corresponds to <Picture 1>.
|
| 21 |
+
At 1.708333 seconds (zero-based frame 41), the camera has advanced 120 degrees clockwise from its starting azimuth; the viewpoint corresponds to <Picture 2>.
|
| 22 |
+
At 3.416667 seconds (zero-based frame 82), the camera has advanced 240 degrees clockwise from its starting azimuth; the viewpoint corresponds to <Picture 3>.
|
| 23 |
+
At 5.125000 seconds (zero-based frame 123), the camera has advanced 360 degrees clockwise from its starting azimuth; the viewpoint corresponds to <Picture 1>.
|
| 24 |
+
|
| 25 |
+
The entire subject is a single frozen three-dimensional scene throughout the shot. Every visible eyelid keeps its reference position: no blinking. Each eye keeps a fixed gaze relative to the head and never tracks the moving camera. Facial expression, mouth, head, hands, feet and body pose remain unchanged. Hair strands, clothing, ribbons, accessories and loose parts are rigidly fixed in the same positions. There is no breathing, sway, settling or secondary motion. Only the camera changes position. Camera radius, elevation, focal length and look-at point stay fixed. Lighting stays fixed in world space; normal viewpoint-dependent shading and reflections remain physically consistent.
|
| 26 |
+
|
| 27 |
+
overall_soundscape:
|
| 28 |
+
N/A
|
| 29 |
+
|
| 30 |
+
non_diegetic_music:
|
| 31 |
+
N/A
|