File size: 6,081 Bytes
ac46da3
c25c678
 
 
ac46da3
5ac8f75
c25c678
5ac8f75
 
 
c25c678
 
 
5ac8f75
 
 
ac46da3
 
 
 
 
c25c678
 
 
 
 
 
 
 
 
 
 
6d378d8
c25c678
 
 
 
 
 
 
6d378d8
 
c25c678
 
ac46da3
 
 
c25c678
 
 
ac46da3
cf24148
 
 
c25c678
 
cf24148
 
 
c25c678
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
cf24148
 
 
c25c678
 
 
 
 
fb4d1c4
 
 
 
 
 
 
 
 
 
 
cf24148
 
 
c25c678
 
 
 
 
 
 
 
 
 
cf24148
 
 
 
 
c25c678
 
cf24148
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
---
license: other
license_name: per-model
license_link: https://github.com/fernandotonon/QtMeshEditor/blob/master/THIRD_PARTY_AI_MODELS.md
tags:
- onnx
- gguf
- pbr
- texture
- normal-map
- 3d
- rigging
- image-to-3d
- qtmesheditor
- qtmesh
- qtmesh-cloud
library_name: onnx
---

# QtMeshEditor β€” AI models

The models used by [QtMeshEditor](https://github.com/fernandotonon/QtMeshEditor)'s
AI-assisted authoring features. **This repo is what the app downloads from at
runtime** (each model on first use, then it runs locally/offline).

**Licenses are per model** β€” see the table and each dedicated repo. The
dedicated repos carry the full model cards (I/O contracts, provenance,
reproduction scripts) for anyone who wants the converted weights standalone.

| folder / files | feature | dedicated repo (full card) | license |
|---|---|---|---|
| `1x-PBRify_*.onnx` | PBR maps from albedo | [QtMeshEditor-pbrify-onnx](https://huggingface.co/fernandotonon/QtMeshEditor-pbrify-onnx) | CC0-1.0 |
| `RealESRGAN_x{2,4}plus.onnx` | texture upscaling | [QtMeshEditor-realesrgan-onnx](https://huggingface.co/fernandotonon/QtMeshEditor-realesrgan-onnx) | BSD-3-Clause |
| `unirig/` | auto-rig skeleton prediction | [QtMeshEditor-unirig-onnx](https://huggingface.co/fernandotonon/QtMeshEditor-unirig-onnx) | MIT |
| `skintokens/` | ML skin-weight prediction | [QtMeshEditor-skintokens-onnx](https://huggingface.co/fernandotonon/QtMeshEditor-skintokens-onnx) | MIT |
| `triposr/` | image β†’ 3D (triplane) | [QtMeshEditor-triposr-onnx](https://huggingface.co/fernandotonon/QtMeshEditor-triposr-onnx) | MIT |
| `triposg/` | image β†’ 3D (rectified-flow DiT) | [QtMeshEditor-triposg-onnx](https://huggingface.co/fernandotonon/QtMeshEditor-triposg-onnx) | MIT |
| `inbetween/rmib.onnx` | animation in-betweening (ours) | [QtMeshEditor-rmib-inbetween](https://huggingface.co/fernandotonon/QtMeshEditor-rmib-inbetween) | CC-BY-4.0 |
| `motion/` | text-to-motion + clip library (ours) | [QtMeshEditor-t2m](https://huggingface.co/fernandotonon/QtMeshEditor-t2m) | CC0-1.0 |
| `segment/meshseg.onnx` | mesh part segmentation (ours) | [QtMeshEditor-mesh-segmentation](https://huggingface.co/fernandotonon/QtMeshEditor-mesh-segmentation) | CC-BY-4.0 |
| `rembg/u2net.onnx` | background removal | [QtMeshEditor-u2net-onnx](https://huggingface.co/fernandotonon/QtMeshEditor-u2net-onnx) | Apache-2.0 |
| `caption/SmolVLM-500M-*.gguf` | image captioning | [QtMeshEditor-smolvlm-gguf](https://huggingface.co/fernandotonon/QtMeshEditor-smolvlm-gguf) | Apache-2.0 |

## PBR map synthesis

`1x-PBRify_NormalV3.onnx`, `1x-PBRify_RoughnessV2.onnx`, `1x-PBRify_Height.onnx`
generate tangent-space normal / roughness / height maps from a single albedo
texture. ONNX re-exports of the CC0 SPAN models from
**[Kim2091/PBRify_Remix](https://github.com/Kim2091/PBRify_Remix)** β€” all
credit to Kim2091. I/O: `1Γ—3Γ—HΓ—W` float `[0,1]` β†’ `1Γ—3Γ—HΓ—W`, dynamic H/W.

## Texture upscaling

`RealESRGAN_x2plus.onnx`, `RealESRGAN_x4plus.onnx` β€” 2Γ—/4Γ— super-resolution.
ONNX re-exports of **Real-ESRGAN**
([xinntao](https://github.com/xinntao/Real-ESRGAN), BSD-3-Clause). Credit: xinntao.

## Auto-rig skeleton prediction (UniRig)

`unirig/{encoder,decoder,embed}.onnx` β€” autoregressive skeleton prediction for
unrigged meshes. ONNX re-export of the skeleton stage of
**[VAST-AI/UniRig](https://huggingface.co/VAST-AI/UniRig)** (SIGGRAPH 2025,
MIT code + weights). Credit: VAST-AI-Research.

## ML skin weights (SkinTokens / TokenRig)

`skintokens/` β€” five ONNX graphs + manifest; QtMeshEditor's **default
skinner**. ONNX re-export of **VAST-AI SkinTokens/TokenRig** (MIT code +
weights, Qwen3-0.6B backbone). `decoder.onnx.data` holds the LM weights as
external data (ORT can't parse the >1.6 GB single-file proto). Credit:
VAST-AI-Research.

## Image β†’ 3D

- `triposr/` β€” **TripoSR** (Tripo AI + Stability AI, MIT): triplane encoder
  (fp32 + int8 tiers) + per-point density/colour decoder.
- `triposg/` β€” **TripoSG** (VAST-AI, SIGGRAPH 2025, MIT): DINOv2 image
  encoder, rectified-flow DiT step graph (fp32 external weights; the int8
  tier here is deprecated β€” it degrades to blobs over the CFG flow loop),
  VAE latent + field-decoder graphs. Geometry-only; colour comes from
  TripoSR's colour field.
- `rembg/u2net.onnx` β€” **UΒ²-Net** saliency for background removal
  (Apache-2.0, the rembg model).

## Animation in-betweening (RMIB) β€” trained by us

`inbetween/rmib.onnx` β€” fills the gap between two keyframes. **Trained from
scratch** on the permissive
[CMU MoCap database](http://mocap.cs.cmu.edu); beats slerp by >2Γ— on held-out
CMU motion. License: CC-BY-4.0.

## Text-to-motion β€” trained by us

`motion/t2m.onnx` + `motion/t2m-vocab.json` β€” **v8.0**, a flow-matching DiT
(21.7M params) mapping a text keyword to a 22-joint world-frame clip, over a
30-action vocab. Trained on the curated template clips in
`motion/motion-library-v2.json` (which is also the fallback for prompts outside
the vocab). v8.0 fixed the backwards-facing problem inherited from the CMU
corpus and a walk defect where the bend sat in the ankle rather than the knee;
`motion/t2m-v61.onnx` is kept for rollback. Full details and metrics:
[QtMeshEditor-t2m](https://huggingface.co/fernandotonon/QtMeshEditor-t2m).
License: CC0-1.0.

## Mesh part segmentation β€” trained by us

`segment/meshseg.onnx` β€” per-point head/torso/arm/leg labels
(PointNet++-style). Trained on synthetic bodies we own + CC0 rigged
characters (Quaternius). 94.7% per-vertex accuracy on rig-truth eval.
License: CC-BY-4.0.

## Image captioning

`caption/SmolVLM-500M-Instruct-Q8_0.gguf` + `mmproj` β€” quantized
**SmolVLM-500M-Instruct** (HuggingFaceTB, Apache-2.0) for llama.cpp-based
captioning. Credit: Hugging Face TB.

---

These models power the AI-assisted authoring features in
**[QtMeshEditor](https://github.com/fernandotonon/QtMeshEditor)** and its
companion **QtMesh Cloud** ([qtmesh.dev](https://qtmesh.dev)). Provenance and
licensing decisions are documented in the QtMeshEditor repo's
`THIRD_PARTY_AI_MODELS.md`.