matlod commited on
Commit
9aa5c29
·
verified ·
1 Parent(s): 616b421

Add ORB360 step 1500 orbit LoRA and CardSpin v2; new cat videos first; four-way 360 sharpness comparison

Browse files

minimax_h3_orb360_step1500 (high-resolution orbit LoRA, heights, multi-reference) and minimax_h3_orb360_cardspin_v2_step50 (step 1500 + 50 card-spin steps). README: new videos at the top, table and detail crops showing the CardSpin variants are much softer for plain 360s, v1 videos kept below. Example prompts in the step-1500 caption format; NOTICE updated.

.gitattributes CHANGED
@@ -38,3 +38,6 @@ assets/whitecat_cardspin_step50.mp4 filter=lfs diff=lfs merge=lfs -text
38
  assets/whitecat_orbit_step50.mp4 filter=lfs diff=lfs merge=lfs -text
39
  examples/herschel_1867_cameron_met263166.png filter=lfs diff=lfs merge=lfs -text
40
  examples/whitecat.png filter=lfs diff=lfs merge=lfs -text
 
 
 
 
38
  assets/whitecat_orbit_step50.mp4 filter=lfs diff=lfs merge=lfs -text
39
  examples/herschel_1867_cameron_met263166.png filter=lfs diff=lfs merge=lfs -text
40
  examples/whitecat.png filter=lfs diff=lfs merge=lfs -text
41
+ assets/whitecat_cardspin_v2_step50.mp4 filter=lfs diff=lfs merge=lfs -text
42
+ assets/whitecat_orbit_detail_4way.jpg filter=lfs diff=lfs merge=lfs -text
43
+ assets/whitecat_orbit_step1500.mp4 filter=lfs diff=lfs merge=lfs -text
NOTICE CHANGED
@@ -1,3 +1,6 @@
1
  MiniMax H3 is licensed under the MiniMax H3 Community License Agreement, Copyright © 2026 MiniMax. All Rights Reserved.
2
 
3
- Modified by MatlowAI: ORB360 orbit LoRA fine-tuning (750 updates) followed by CardSpin fine-tuning (+50 updates), September 2026. This repository contains adapter (LoRA) weights, not base-model weights.
 
 
 
 
1
  MiniMax H3 is licensed under the MiniMax H3 Community License Agreement, Copyright © 2026 MiniMax. All Rights Reserved.
2
 
3
+ Modified by MatlowAI (September 2026). This repository contains adapter (LoRA) weights, not base-model weights:
4
+ - minimax_h3_orb360_cardspin_step50.safetensors: ORB360 orbit LoRA fine-tuning (750 updates) followed by CardSpin fine-tuning (+50 updates).
5
+ - minimax_h3_orb360_step1500.safetensors: ORB360 orbit LoRA fine-tuning (1500 updates) on a larger, higher-resolution rendered orbit set.
6
+ - minimax_h3_orb360_cardspin_v2_step50.safetensors: the step-1500 orbit LoRA followed by CardSpin fine-tuning (+50 updates).
README.md CHANGED
@@ -11,46 +11,92 @@ tags:
11
  - ref2va
12
  - orbit
13
  - camera-control
 
14
  - musubi-tuner
15
  - comfyui
16
  pipeline_tag: image-to-video
17
  ---
18
 
19
- # MiniMax-H3 ORB360 CardSpin (step 50)
20
 
21
- A rank-32 LoRA for MiniMax-H3 **Ref2VA** that does two things from one file, picked by the prompt:
 
 
22
 
23
- - **`ORB360_CARDSPIN`**: your photo turns in space like a thin physical photo card. The person or animal inside it
24
- turns with it: profile on the edge-on sliver (nose sticking out past the edge), the back of their head on the back
25
- of the card, the other profile on the way round, and a clean landing back on your photo after exactly 5.125 s.
26
- - **`ORB360_CW`**: the ordinary thing it was built for, a single smooth clockwise 360-degree camera orbit around a
27
- frozen subject that ends on the starting view.
 
 
 
28
 
29
  Powered by MiniMax H3.
30
 
31
- ## The card spin
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
32
 
33
  <video controls muted playsinline loop preload="metadata" width="100%"
34
  poster="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_step50_poster.jpg">
35
  <source src="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_step50.mp4" type="video/mp4">
36
- <a href="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_step50.mp4">Download the card-spin clip (mp4)</a>
37
  </video>
38
 
39
- One reference image (`examples/whitecat.png`), the prompt in `prompts/cardspin_caption.txt`, 1024 x 768, 124 frames,
40
- 20 steps, seed 20260926.
41
-
42
- ## Same file, normal orbit
43
 
44
  <video controls muted playsinline loop preload="metadata" width="100%"
45
  poster="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step50_poster.jpg">
46
  <source src="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step50.mp4" type="video/mp4">
47
- <a href="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step50.mp4">Download the orbit clip (mp4)</a>
48
  </video>
49
 
50
- Same image, same seed, same LoRA file; only the prompt changed (`prompts/orbit_realscene_caption.txt`). The 50
51
- card-spin steps did not break the orbit.
52
-
53
- ## How it happened
54
 
55
  We were testing an orbit LoRA (the ORB360 project) on real photos instead of the grey Blender renders it was trained
56
  on, and fed it an 1867 portrait of Sir John Herschel by Julia Margaret Cameron. Instead of orbiting a man, it decided
@@ -63,17 +109,15 @@ the *photograph* was the object. This is the clip that started it all:
63
  </video>
64
 
65
  It was too good to leave as a one-off. So we took that single generated video, wrote a caption describing what
66
- happens in it with timestamps, and trained **50 more steps** on top of the 750-step orbit LoRA using only that one
67
- clip. After 50 steps (about half an hour on one GPU) the effect transferred to photos it had never seen: other
68
- Cameron portraits, and a colour photo of a cat in a hat. We also checked 100 and 150 steps; they work too, with
69
- slightly crisper card edges, but 50 has the most charm and is the lightest touch on the orbit behaviour, so that is
70
- the one released here.
71
 
72
  ## Using it
73
 
74
- **ComfyUI:** load it with the standard **Load LoRA** node (model only, strength 1.0) on a MiniMax-H3 **Ref2VA** model.
75
- It is a kohya-style LoRA (`lora_unet_*`, `lora_down`/`lora_up`/`alpha`); all 200 modules map onto ComfyUI's
76
- MiniMax-H3 weights. Give it one reference image and paste one of the prompts from `prompts/` as the text.
77
 
78
  **musubi-tuner** ([kohya-ss/musubi-tuner](https://github.com/kohya-ss/musubi-tuner), v0.3.5 or later):
79
 
@@ -82,27 +126,52 @@ python src/musubi_tuner/minimax_h3_generate_video.py --task ref2va \
82
  --dit minimax_h3_ref2va_bf16.safetensors --prune_adaln \
83
  --video_vae minimax_h3_video_vae_fp16.safetensors --audio_vae minimax_h3_audio_vae_fp32.safetensors \
84
  --text_encoder qwen3vl_32b_minimax_h3_int8_convrot.safetensors --text_encoder_attn_mode sdpa \
85
- --lora_weight minimax_h3_orb360_cardspin_step50.safetensors --lora_multiplier 1.0 --lora_runtime_attach \
86
- --video_size 832 672 --video_length 124 --infer_steps 20 --attn_mode sdpa --seed 20260926 \
87
- --prompt "$(cat prompts/cardspin_caption.txt)" --ref your_photo.png --save_path out/
88
  ```
89
 
90
- - One reference image (`--ref` in musubi). It becomes `<Picture 1>` in the prompt and is the first and last frame.
91
- - Keep **124 frames at 24 fps**: the prompt's timestamps (edge-on at 1.0 s, back of the card from 1.7 s, other
92
- profile at 3.8 s, back on the photo at 5.125 s) assume that length.
93
- - Resolutions we have seen work: 672 x 832 and 832 x 1024 portraits, 1024 x 768 and 1152 x 768 landscapes.
94
- - Portraits of people and animals are the sweet spot. It works on colour photos as well as old sepia prints.
95
- - For the normal orbit, use `prompts/orbit_realscene_caption.txt`. It keeps the photo's real surroundings; earlier
96
- wording like "clean studio" pulled some scenes towards a grey studio backdrop.
97
- - Roll a few seeds. This was taught from one clip, so some seeds commit to the card harder than others.
 
 
 
 
 
 
 
 
 
98
  - The prompts follow MiniMax's official Ref2VA prompt layout (`subject_definitions`, `summary`,
99
  `retention_analysis`, `detailed_description`, `overall_soundscape`, `non_diegetic_music`). Both audio sections are
100
  `N/A`, which asks for silence; without them H3 tends to invent a soundtrack.
101
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
102
  ## Training
103
 
104
- Both stages used musubi-tuner (v0.3.5, `--task ref2va`; one local patch that only affects step counting on
105
- `--resume`) with the same settings:
106
 
107
  | | |
108
  |---|---|
@@ -114,45 +183,51 @@ Both stages used musubi-tuner (v0.3.5, `--task ref2va`; one local patch that onl
114
  | Batch | 1, no accumulation, video only (no audio loss), seed 20260926 |
115
  | Hardware | one NVIDIA RTX PRO 6000 (96 GB) |
116
 
117
- **Stage 1: ORB360 orbit, 750 steps, from scratch.** Four assets rendered in Blender, each as one 124-frame, 24 fps,
118
- 512 x 512 clockwise orbit with exactly one full turn. Three
119
- reference images per clip, rendered with the exact cameras of frames 0, 41 and 82 (0, 120 and 240 degrees). The
120
- captions give timestamped milestones (for example "At 1.708333 seconds (zero-based frame 41), the camera has advanced
121
- 120 degrees clockwise ... corresponds to <Picture 2>") and ask for a completely frozen subject. About 18 s per step.
 
 
 
 
122
 
123
- **Stage 2: CardSpin, +50 steps.** Initialised from the stage-1 weights (fresh optimizer). One training clip: the
124
- Herschel glitch above (generated at 832 x 1024, trained at 672 x 832, 124 frames), with the original Herschel photo as
125
- the single reference and the card-spin caption in `prompts/cardspin_caption.txt`. About 40 s per step, 73 GB peak VRAM.
126
 
127
- ## Limitations
128
-
129
- - Learned from a single example. The card always turns the same way with roughly the same timing, and some subjects
130
- or seeds may commit to the card less than others.
131
- - The first stage saw only four rendered assets at 512 x 512, so orbit quality on complex real scenes varies.
132
- - Evaluated by eye on a handful of images and seeds, not with a benchmark.
133
 
134
- ## What's next
135
 
136
- This is a side quest. The main ORB360 work is training a much steadier 360 LoRA on a larger rendered set: more
137
- assets, several orbit lengths and speeds, higher and lower camera heights, closer and wider framing, and captions that
138
- describe the exact camera for every clip. The aim is to see whether consistent, controllable orbits can lead to some
139
- kind of stable geometry adapter. More on that when it works.
 
 
 
 
140
 
141
  ## Files
142
 
143
- - `minimax_h3_orb360_cardspin_step50.safetensors`: the LoRA (fp32, 597 MB)
144
- - `prompts/cardspin_caption.txt`, `prompts/orbit_realscene_caption.txt`: the two prompts behind the cat videos
 
 
145
  - `examples/whitecat.png`: reference image for the cat clips
146
  - `examples/herschel_1867_cameron_met263166.png`: the Herschel portrait (Julia Margaret Cameron, 1867; The
147
  Metropolitan Museum of Art, Open Access, CC0), resized
148
- - `assets/`: the three videos above and their poster frames
149
  - `LICENSE` (MiniMax H3 Community License Agreement), `NOTICE`
150
 
151
  ## Base model and licence
152
 
153
- This is a LoRA for [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3) by MiniMax; it does nothing without
154
- their base weights, and its second stage was trained on a clip generated with them. It is distributed under the
155
- [MiniMax H3 Community License Agreement](LICENSE), including its acceptable-use, distribution, commercial and
156
- territorial provisions. **It is not MIT.** Read the complete upstream terms; this repository does not expand them. See
157
- `NOTICE` for attribution and the modification notice. Thanks to MiniMax for releasing H3, to Ostris for the training
158
- adapter, and to kohya-ss for musubi-tuner.
 
11
  - ref2va
12
  - orbit
13
  - camera-control
14
+ - novel-view
15
  - musubi-tuner
16
  - comfyui
17
  pipeline_tag: image-to-video
18
  ---
19
 
20
+ # MiniMax-H3 ORB360: 360 orbit + CardSpin
21
 
22
+ Rank-32 LoRAs for MiniMax-H3 **Ref2VA**. Give it a photo, get a smooth clockwise 360-degree camera orbit around the
23
+ frozen subject (`ORB360_CW`), or, with the CardSpin files, turn the photo in space like a thin physical card
24
+ (`ORB360_CARDSPIN`).
25
 
26
+ | File | What it is | Use it for |
27
+ |---|---|---|
28
+ | `minimax_h3_orb360_step1500.safetensors` | **New.** ORB360 orbit LoRA, 1500 steps on high-resolution Blender orbits, camera heights from below to above, 1-4 labelled reference pictures | **Clean 360 orbits** (sharpest; our pick for multi-view use) |
29
+ | `minimax_h3_orb360_cardspin_v2_step50.safetensors` | **New.** CardSpin v2 = step 1500 + 50 card-spin steps | The card spin |
30
+ | `minimax_h3_orb360_cardspin_step50.safetensors` | CardSpin v1 (the original release) = the old 512 x 512 orbit LoRA + 50 card-spin steps | The original card spin |
31
+
32
+ Both CardSpin files can still orbit with the orbit prompt, but they are noticeably softer; see
33
+ [Which file for a clean 360?](#which-file-for-a-clean-360).
34
 
35
  Powered by MiniMax H3.
36
 
37
+ ## New: CardSpin v2
38
+
39
+ <video controls muted playsinline loop preload="metadata" width="100%"
40
+ poster="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_v2_step50_poster.jpg">
41
+ <source src="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_v2_step50.mp4" type="video/mp4">
42
+ <a href="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_v2_step50.mp4">Download the CardSpin v2 clip (mp4)</a>
43
+ </video>
44
+
45
+ `minimax_h3_orb360_cardspin_v2_step50.safetensors`, one reference image (`examples/whitecat.png`), the prompt in
46
+ `prompts/cardspin_caption.txt`, 1024 x 768, 124 frames, 20 steps, seed 20260926.
47
+
48
+ ## New: the 360 from ORB360 step 1500
49
+
50
+ <video controls muted playsinline loop preload="metadata" width="100%"
51
+ poster="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step1500_poster.jpg">
52
+ <source src="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step1500.mp4" type="video/mp4">
53
+ <a href="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step1500.mp4">Download the step-1500 orbit (mp4)</a>
54
+ </video>
55
+
56
+ `minimax_h3_orb360_step1500.safetensors`, same image and seed, prompt `prompts/orbit_realscene_caption.txt`. It was
57
+ trained only on grey-backdrop Blender renders, and it keeps the real snowy scene around the cat.
58
+
59
+ ## Which file for a clean 360?
60
+
61
+ The same orbit (same photo, prompt, seed, 1024 x 768, 124 frames) from four LoRAs. Numbers are per-frame image
62
+ statistics averaged over the clip; "detail vs step 1500" is the median per-frame ratio of the Laplacian variance (fine
63
+ detail) to the step-1500 clip.
64
+
65
+ | White cat 360 | Orbit step 750 (old, not in this repo) | CardSpin v1 (750 + 50) | **ORB360 step 1500** | CardSpin v2 (1500 + 50) |
66
+ |---|---|---|---|---|
67
+ | Fine detail (Laplacian variance) | 30.2 | 19.8 | **56.1** | 20.7 |
68
+ | Detail vs step 1500 | 0.54x | 0.34x | **1.00x** | 0.35x |
69
+ | Edge energy (Tenengrad) | 1098 | 809 | **1343** | 841 |
70
+ | High-frequency share | 0.0035 | 0.0028 | **0.0050** | 0.0029 |
71
+ | First frame vs the photo, PSNR / SSIM | 28.0 dB / 0.87 | 24.5 dB / 0.85 | **28.5 dB / 0.87** | 22.5 dB / 0.77 |
72
+
73
+ ![Same frames from the four LoRAs, native-pixel crops](https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_detail_4way.jpg)
74
+
75
+ - The high-resolution training nearly doubles the fine detail (step 750 -> step 1500: 1.85x, higher in all 124 frames).
76
+ - **The 50 card-spin steps soften everything**, whichever orbit LoRA they start from: both CardSpin files keep only
77
+ about a third of step 1500's fine detail and drift further from the photo in the first frame. The single card-spin
78
+ training clip is a soft generated video, and 50 steps on it most likely pull the whole look towards it.
79
+ - So: **step 1500 for orbits, CardSpin only for the card trick.**
80
+
81
+ ## The original v1 videos
82
+
83
+ CardSpin v1 (`minimax_h3_orb360_cardspin_step50.safetensors`), card-spin prompt:
84
 
85
  <video controls muted playsinline loop preload="metadata" width="100%"
86
  poster="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_step50_poster.jpg">
87
  <source src="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_step50.mp4" type="video/mp4">
88
+ <a href="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_cardspin_step50.mp4">Download the v1 card-spin clip (mp4)</a>
89
  </video>
90
 
91
+ Same file, orbit prompt:
 
 
 
92
 
93
  <video controls muted playsinline loop preload="metadata" width="100%"
94
  poster="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step50_poster.jpg">
95
  <source src="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step50.mp4" type="video/mp4">
96
+ <a href="https://huggingface.co/MATLOWAI/MiniMax-H3-ORB360-CardSpin/resolve/main/assets/whitecat_orbit_step50.mp4">Download the v1 orbit clip (mp4)</a>
97
  </video>
98
 
99
+ ## How the card spin happened
 
 
 
100
 
101
  We were testing an orbit LoRA (the ORB360 project) on real photos instead of the grey Blender renders it was trained
102
  on, and fed it an 1867 portrait of Sir John Herschel by Julia Margaret Cameron. Instead of orbiting a man, it decided
 
109
  </video>
110
 
111
  It was too good to leave as a one-off. So we took that single generated video, wrote a caption describing what
112
+ happens in it with timestamps, and trained **50 more steps** on top of the orbit LoRA using only that one clip. After
113
+ 50 steps the effect transferred to photos it had never seen: other Cameron portraits, and a colour photo of a cat in a
114
+ hat. CardSpin v1 did this on the old 750-step orbit LoRA; CardSpin v2 repeats the same recipe on top of step 1500.
 
 
115
 
116
  ## Using it
117
 
118
+ **ComfyUI:** load one file with the standard **Load LoRA** node (model only, strength 1.0) on a MiniMax-H3 **Ref2VA**
119
+ model. These are kohya-style LoRAs (`lora_unet_*`, `lora_down`/`lora_up`/`alpha`); all 200 modules map onto ComfyUI's
120
+ MiniMax-H3 weights. Give it the reference image(s) and paste one of the prompts from `prompts/` as the text.
121
 
122
  **musubi-tuner** ([kohya-ss/musubi-tuner](https://github.com/kohya-ss/musubi-tuner), v0.3.5 or later):
123
 
 
126
  --dit minimax_h3_ref2va_bf16.safetensors --prune_adaln \
127
  --video_vae minimax_h3_video_vae_fp16.safetensors --audio_vae minimax_h3_audio_vae_fp32.safetensors \
128
  --text_encoder qwen3vl_32b_minimax_h3_int8_convrot.safetensors --text_encoder_attn_mode sdpa \
129
+ --lora_weight minimax_h3_orb360_step1500.safetensors --lora_multiplier 1.0 --lora_runtime_attach \
130
+ --video_size 768 1024 --video_length 124 --infer_steps 20 --attn_mode sdpa --seed 20260926 \
131
+ --prompt "$(cat prompts/orbit_realscene_caption.txt)" --ref your_photo.png --save_path out/
132
  ```
133
 
134
+ For the card spin, use `minimax_h3_orb360_cardspin_v2_step50.safetensors` and `prompts/cardspin_caption.txt`.
135
+
136
+ - The first `--ref` becomes `<Picture 1>`: the first and last frame of the orbit.
137
+ - Keep **124 frames at 24 fps** for these prompts: their timestamps (for example "At 1.708333 seconds (zero-based
138
+ frame 41) ... 120 degrees") assume that length.
139
+ - Prompts:
140
+ - `orbit_realscene_caption.txt`: one photo, keeps the photo's real surroundings.
141
+ - `orbit_3ref_caption.txt`: three views of the same subject (front, 120 and 240 degrees clockwise), for example
142
+ three renders or a turnaround sheet. With real side and back views the model does not have to invent them.
143
+ - `orbit_underneath_3ref_example.txt`, `orbit_crossheight_front_back_example.txt`: the exact caption format step
144
+ 1500 was trained on, taken from two of our held-out tests (an orbit from 30 degrees below with no floor, and front
145
+ and back pictures taken from 45 degrees with the orbit at 5 degrees). The numbers in them (camera distance, how
146
+ much of the frame the subject fills) describe those test objects; edit the heights, times and angles for your own
147
+ use.
148
+ - `cardspin_caption.txt`: the card spin.
149
+ - Resolutions we have seen work: 800 x 800, 1024 x 768 and 1152 x 768 landscapes, 672 x 832 and 832 x 1024
150
+ portraits.
151
  - The prompts follow MiniMax's official Ref2VA prompt layout (`subject_definitions`, `summary`,
152
  `retention_analysis`, `detailed_description`, `overall_soundscape`, `non_diegetic_music`). Both audio sections are
153
  `N/A`, which asks for silence; without them H3 tends to invent a soundtrack.
154
 
155
+ ## How well step 1500 orbits
156
+
157
+ Tested on three objects that were never rendered for training (a vintage film camera, a stylised character and a
158
+ stone cat statue), one seed per test, against Blender renders of the true orbit:
159
+
160
+ | Test | Result |
161
+ |---|---|
162
+ | Orbit at 20 degrees, 3 views (0/120/240) | Within one frame (2.9 degrees) of the true orbit on every frame |
163
+ | Orbits from below (30 and 10 degrees below the centre, no floor) | Right height on 118-124 of 124 frames, one clean turn |
164
+ | Pictures taken from 45 degrees, orbit asked for at 5 degrees | Low orbit on all three objects (most frames within 5-10 degrees of the target), one turn |
165
+ | High orbit (45 degrees) from one front picture | High orbit (30-45 degrees on most frames); the unseen sides are invented |
166
+ | Pictures taken from 45 degrees, orbit asked for at 30 degrees below | **Fails**: drifts back towards the pictures' height |
167
+ | Height changing during the clip (for example 45 above to 30 below) | **Fails**: the object tumbles instead of the camera orbiting |
168
+
169
+ The old 750-step LoRA mostly copies the height of the reference pictures: it gets the underneath orbits right when the
170
+ pictures are taken from below, but none of the "pictures from 45 degrees, orbit lower" cases.
171
+
172
  ## Training
173
 
174
+ All stages used musubi-tuner (v0.3.5, `--task ref2va`; small local wrappers for resumable segments) with the same settings:
 
175
 
176
  | | |
177
  |---|---|
 
183
  | Batch | 1, no accumulation, video only (no audio loss), seed 20260926 |
184
  | Hardware | one NVIDIA RTX PRO 6000 (96 GB) |
185
 
186
+ **ORB360 step 1500 (new), from scratch.** 23 assets (Poly Haven CC0 props and a few stylised character models),
187
+ rendered in Blender (Cycles) as 345 one-turn clockwise orbits at three lengths, each at the highest resolution that
188
+ fits in memory: 124 frames at 800 x 800, 73 frames at 1024 x 1024, and 22 frames at 1600 x 1600. Camera heights from
189
+ 45 degrees below to 55 degrees above the subject's centre (the ones below with the floor switched off), tight and wide
190
+ framing, varied start angles. Each clip appears in three training rows with different reference pictures (the 73-frame clips lost their four-picture sets to fit in memory): two sets
191
+ taken from the clip itself (front only, front and back, front and side, front, side and back, a four-way turnaround,
192
+ or thirds at 0/120/240 degrees) and one set rendered from a different height than the orbit. The captions label every
193
+ picture with its exact time, frame and angle, and describe the camera's height, distance, framing and lens. 960 rows,
194
+ 1500 steps, about 62 s per step (about 26 hours).
195
 
196
+ **CardSpin v2, +50 steps.** Initialised from the step-1500 weights (fresh optimizer). One training clip: the Herschel
197
+ glitch above (generated at 832 x 1024, trained at 672 x 832, 124 frames), with the original Herschel photo as the
198
+ single reference and the caption in `prompts/cardspin_caption.txt`.
199
 
200
+ **Original release (CardSpin v1).** Stage 1: four Blender assets, one 124-frame, 512 x 512 orbit each, three
201
+ reference pictures at 0, 120 and 240 degrees, 750 steps. Stage 2: the same 50 card-spin steps as above.
 
 
 
 
202
 
203
+ ## Limitations
204
 
205
+ - The card spin was learned from a single example: it always turns the same way with roughly the same timing, and
206
+ some subjects or seeds commit to the card less than others. Both CardSpin files soften the image (table above).
207
+ - Step 1500 was trained on grey-backdrop renders of 23 objects. It orbits real photos well in our tests, but was
208
+ evaluated on a handful of images and three held-out objects, one seed each, not a large benchmark.
209
+ - Big jumps between the pictures' height and the orbit's height, and orbits whose height changes during the clip,
210
+ do not work yet.
211
+ - From one picture, the sides the camera has not seen are invented. More pictures (front and back, a turnaround, or
212
+ three views) fix that.
213
 
214
  ## Files
215
 
216
+ - `minimax_h3_orb360_step1500.safetensors`: ORB360 orbit LoRA, step 1500 (fp32, 597 MB)
217
+ - `minimax_h3_orb360_cardspin_v2_step50.safetensors`: CardSpin v2 (fp32, 597 MB)
218
+ - `minimax_h3_orb360_cardspin_step50.safetensors`: CardSpin v1, the original release (fp32, 597 MB)
219
+ - `prompts/`: the prompts described above
220
  - `examples/whitecat.png`: reference image for the cat clips
221
  - `examples/herschel_1867_cameron_met263166.png`: the Herschel portrait (Julia Margaret Cameron, 1867; The
222
  Metropolitan Museum of Art, Open Access, CC0), resized
223
+ - `assets/`: the videos above, their poster frames, and the four-way detail comparison
224
  - `LICENSE` (MiniMax H3 Community License Agreement), `NOTICE`
225
 
226
  ## Base model and licence
227
 
228
+ These are LoRAs for [MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3) by MiniMax; they do nothing without
229
+ their base weights, and the card-spin stages were trained on a clip generated with them. They are distributed under
230
+ the [MiniMax H3 Community License Agreement](LICENSE), including its acceptable-use, distribution, commercial and
231
+ territorial provisions. **They are not MIT.** Read the complete upstream terms; this repository does not expand them.
232
+ See `NOTICE` for attribution and the modification notice. Thanks to MiniMax for releasing H3, to Ostris for the
233
+ training adapter, and to kohya-ss for musubi-tuner.
assets/whitecat_cardspin_v2_step50.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:dbb7b3a1af06b1346ae22dbbfdcdf6801118ae4cd4d2cf29d845b054b8d05f9a
3
+ size 3268303
assets/whitecat_cardspin_v2_step50_poster.jpg ADDED
assets/whitecat_orbit_detail_4way.jpg ADDED

Git LFS Details

  • SHA256: 2f1e1199d87979e829259c1edadc2149529da79e60afe00a1f529772e708e8ec
  • Pointer size: 131 Bytes
  • Size of remote file: 242 kB
assets/whitecat_orbit_step1500.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:13baf0d556bc530715498661bea5319c5ec04698df7d1e1b8736c34757afd2f9
3
+ size 6001372
assets/whitecat_orbit_step1500_poster.jpg ADDED
minimax_h3_orb360_cardspin_v2_step50.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0e405410ed07c8c63cfb28fa5de3b1f1447b2165fdf5099542b9ad64194a053b
3
+ size 596449416
minimax_h3_orb360_step1500.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e9275b2f456f6bd004c9cb3cf5c636cad87217c5bc96e6a5b5533e917de121da
3
+ size 596450168
prompts/orbit_crossheight_front_back_example.txt ADDED
@@ -0,0 +1,31 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ subject_definitions:
2
+ <Subject 1> is the exact same rigid subject shown in <Picture 1> and <Picture 2>. Its geometry, proportions, materials, colors, markings, and surface details remain unchanged.
3
+ <Picture 1> shows <Subject 1> from the direction the orbit starts from, but from a higher camera: 45 degrees above the subject's centre instead of the orbit's 5 degrees above. It is not a frame of the video.
4
+
5
+ summary:
6
+ [reference generation] ORB360_CW. A turntable-style camera orbit of <Subject 1> in front of a plain grey seamless backdrop: one full clockwise turn in 5.125000 seconds, filmed from a low, near eye-level camera 5 degrees above the subject, while the reference pictures were taken from a higher camera at 45 degrees.
7
+
8
+ retention_analysis:
9
+ <Subject 1> (appears in [Shot 1]): fully_preserved - its geometry, proportions, materials, colors, markings and surface details are retained, and it stays completely frozen.
10
+ <Picture 1> (not a frame of [Shot 1]; seen from 0 degrees clockwise of the starting direction at 45 degrees elevation): fully_preserved - its geometry, materials and surface details define <Subject 1>; the orbit passes the same side at 5 degrees elevation.
11
+ <Picture 2> (not a frame of [Shot 1]; seen from 180 degrees clockwise of the starting direction at 45 degrees elevation): fully_preserved - its geometry, materials and surface details define <Subject 1>; the orbit passes the same side at 5 degrees elevation.
12
+
13
+ detailed_description:
14
+ [Shot 1] The shot begins facing <Subject 1> from the same direction as <Picture 1>, but from 5 degrees elevation, lower than <Picture 1>'s 45 degrees. The camera performs exactly one smooth clockwise 360-degree orbit around the completely stationary <Subject 1>, at constant radius, constant elevation, constant focal length, constant speed, and constant framing, ending at the starting viewpoint. The subject remains perfectly rigid and stationary. No articulation, deformation, secondary motion, wind, zoom, dolly, camera roll, cuts, lighting changes, or scene changes.
15
+ Camera: a low, near eye-level camera, 5 degrees above the subject's centre, looking down 5 degrees at it throughout the orbit. The camera is below the top of the subject, at about 68% of its height above the ground, about 3.8 times the subject's bounding radius from its centre. At the widest view the subject spans about 90% of the frame, and about 85% on average. 60 mm lens on a full-frame sensor (33.4-degree field of view), no roll, no zoom.
16
+ The background is a plain, evenly lit grey seamless backdrop with only a soft contact shadow under the subject.
17
+
18
+ All 5 milestones below belong to the same uninterrupted [Shot 1], not separate shots. The video runs at 24 fps and has 124 frames. The full turn takes 5.125000 seconds (frame 0 to frame 123), a constant 70.2 degrees per second (2.9268 degrees per frame). Angles describe camera motion clockwise as viewed from above, relative to the starting camera azimuth; the subject itself does not rotate. The camera moves continuously at constant angular speed between these milestones, without pausing, reversing or completing extra turns.
19
+ At 0.000000 seconds (zero-based frame 0), the camera has advanced 0 degrees clockwise from its starting azimuth; it faces the same side of <Subject 1> as <Picture 1> (0 degrees), from 5 degrees elevation instead of 45.
20
+ At 1.708333 seconds (zero-based frame 41), the camera has advanced 120 degrees clockwise from its starting azimuth.
21
+ At 2.583333 seconds (zero-based frame 62), the camera has advanced 181.5 degrees clockwise from its starting azimuth; it faces the same side of <Subject 1> as <Picture 2> (180 degrees), from 5 degrees elevation instead of 45.
22
+ At 3.416667 seconds (zero-based frame 82), the camera has advanced 240 degrees clockwise from its starting azimuth.
23
+ At 5.125000 seconds (zero-based frame 123), the camera has advanced 360 degrees clockwise from its starting azimuth; it faces the same side of <Subject 1> as <Picture 1> (0 degrees), from 5 degrees elevation instead of 45.
24
+
25
+ The entire subject is a single frozen three-dimensional scene throughout the shot. Every visible eyelid keeps its reference position: no blinking. Each eye keeps a fixed gaze relative to the head and never tracks the moving camera. Facial expression, mouth, head, hands, feet and body pose remain unchanged. Hair strands, clothing, ribbons, accessories and loose parts are rigidly fixed in the same positions. There is no breathing, sway, settling or secondary motion. Only the camera changes position. Camera radius, elevation, focal length and look-at point stay fixed. Lighting stays fixed in world space; normal viewpoint-dependent shading and reflections remain physically consistent.
26
+
27
+ overall_soundscape:
28
+ N/A
29
+
30
+ non_diegetic_music:
31
+ N/A
prompts/orbit_underneath_3ref_example.txt ADDED
@@ -0,0 +1,31 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ subject_definitions:
2
+ <Subject 1> is the exact same rigid subject shown in <Picture 1>, <Picture 2>, and <Picture 3>. Its geometry, proportions, materials, colors, markings, and surface details remain unchanged.
3
+ <Picture 1> is the first-frame viewpoint of [Shot 1].
4
+
5
+ summary:
6
+ [reference generation] ORB360_CW. A turntable-style camera orbit of <Subject 1> floating in a plain grey space with no floor: one full clockwise turn in 5.125000 seconds, filmed from a camera far below the subject, looking steeply up 30 degrees below the subject.
7
+
8
+ retention_analysis:
9
+ <Subject 1> (appears in [Shot 1]): fully_preserved - its geometry, proportions, materials, colors, markings and surface details are retained, and it stays completely frozen.
10
+ <Picture 1> ([Shot 1] first frame and last frame): fully_preserved - the orbit starts and ends on this view.
11
+ <Picture 2> ([Shot 1] keyframe at 1.708333 seconds, zero-based frame 41, 120 degrees clockwise from <Picture 1>): fully_preserved - the camera passes through this exact view.
12
+ <Picture 3> ([Shot 1] keyframe at 3.416667 seconds, zero-based frame 82, 240 degrees clockwise from <Picture 1>): fully_preserved - the camera passes through this exact view.
13
+
14
+ detailed_description:
15
+ [Shot 1] The shot begins from <Picture 1>. The camera performs exactly one smooth clockwise 360-degree orbit around the completely stationary <Subject 1>, at constant radius, constant elevation, constant focal length, constant speed, and constant framing, ending at the starting viewpoint. The subject remains perfectly rigid and stationary. No articulation, deformation, secondary motion, wind, zoom, dolly, camera roll, cuts, lighting changes, or scene changes.
16
+ Camera: a camera far below the subject, looking steeply up, 30 degrees below the subject's centre, looking up 30 degrees at it throughout the orbit. The camera is below the level of the subject's base, about 0.7 times the subject's height beneath it, about 3.7 times the subject's bounding radius from its centre. At the widest view the subject spans about 90% of the frame, and about 83% on average. 60 mm lens on a full-frame sensor (33.4-degree field of view), no roll, no zoom.
17
+ There is no floor: <Subject 1> floats in a plain, evenly lit grey space with no ground and no shadow, so it can be seen from below.
18
+
19
+ All 4 milestones below belong to the same uninterrupted [Shot 1], not separate shots. The video runs at 24 fps and has 124 frames. The full turn takes 5.125000 seconds (frame 0 to frame 123), a constant 70.2 degrees per second (2.9268 degrees per frame). Angles describe camera motion clockwise as viewed from above, relative to the starting camera azimuth; the subject itself does not rotate. The camera moves continuously at constant angular speed between these milestones, without pausing, reversing or completing extra turns.
20
+ At 0.000000 seconds (zero-based frame 0), the camera has advanced 0 degrees clockwise from its starting azimuth; the viewpoint corresponds to <Picture 1>.
21
+ At 1.708333 seconds (zero-based frame 41), the camera has advanced 120 degrees clockwise from its starting azimuth; the viewpoint corresponds to <Picture 2>.
22
+ At 3.416667 seconds (zero-based frame 82), the camera has advanced 240 degrees clockwise from its starting azimuth; the viewpoint corresponds to <Picture 3>.
23
+ At 5.125000 seconds (zero-based frame 123), the camera has advanced 360 degrees clockwise from its starting azimuth; the viewpoint corresponds to <Picture 1>.
24
+
25
+ The entire subject is a single frozen three-dimensional scene throughout the shot. Every visible eyelid keeps its reference position: no blinking. Each eye keeps a fixed gaze relative to the head and never tracks the moving camera. Facial expression, mouth, head, hands, feet and body pose remain unchanged. Hair strands, clothing, ribbons, accessories and loose parts are rigidly fixed in the same positions. There is no breathing, sway, settling or secondary motion. Only the camera changes position. Camera radius, elevation, focal length and look-at point stay fixed. Lighting stays fixed in world space; normal viewpoint-dependent shading and reflections remain physically consistent.
26
+
27
+ overall_soundscape:
28
+ N/A
29
+
30
+ non_diegetic_music:
31
+ N/A