Aether Phase-1 โ€” assembled + wired multimodal omni-model

Qwen3-VL-8B base grafted with InternViT-6B (vision), MiMo-Audio (audio), and TRELLIS-SLAT (3D geometry) via 5 projectors, expanded vocab (+4627 tokens), and a factorized / 3D-spatial M-RoPE. Verified end-to-end on ROCm (MI300).

15.16B params - 32GB VRAM (bf16) - full multimodal forward verified (seq 5542, logits (1,5542,156296) finite).

This repo is a resumable assembly package, not a full weight dump - it pins the base + encoders in MANIFEST.json so the exact model re-assembles on any box (6900XT / Kaggle / Colab / rental) without duplicating the 16GB base.

Contents

  • MANIFEST.json - base + 3 encoder repos/revisions/dims, freeze scheme, M-RoPE routing, verification log
  • tokenizer/ - expanded tokenizer (151669 to 156296)
  • new_modules.safetensors - 5 projectors + cam_pose (98.6M params, fresh init pre-alignment)
  • scripts/ - assemble.py, forward.py (routing + M-RoPE), surgery.py, curate_slivers.py, save_package.py
  • BUILD.md - full build spec, graft verifications, ROCm recipes, model-building knowledge

Resume

Clone, pull base+encoders per MANIFEST, run assemble.py then forward.py, load new_modules.safetensors, begin sliver alignment training.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support