File size: 9,319 Bytes
67ff78d
 
5c331a4
 
67ff78d
 
5c331a4
 
67ff78d
 
 
 
5c331a4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
---
title: Component Studio Reference
emoji: 🧩
colorFrom: gray
colorTo: blue
sdk: gradio
sdk_version: 6.27.0
python_version: "3.12"
app_file: app.py
pinned: false
---

# Component Studio

A new Gradio + uv pipeline targeting Hugging Face ZeroGPU. It accepts an image,
a prompt, or both and produces a textured GLB with individually editable components.
Add an animation prompt to run learned humanoid rigging, NVIDIA Kimodo motion generation,
retargeting, video export, full-frame collision checks and Astra visual review.

The animated workflow is under production verification. A completed export is not a quality
pass: rejected models or animations retain their artifacts with `needs_review` status.

## Inputs

| Input | Example | Behavior |
| --- | --- | --- |
| Image only | Upload a stool photograph | Reconstruct the supplied image; skip image generation |
| Prompt only | “A brass desk lamp with a green glass shade” | Generate a reference with FLUX.2 Klein, then reconstruct it |
| Both | Stool image + “Walnut seat and black metal legs; keep the shape” | Edit the image with FLUX.2 Klein, then reconstruct the edited reference |

## Pipeline stages

1. Validate inputs and prepare a reference. Images keep EXIF orientation and alpha.
2. TRELLIS.2 generates geometry at its 1024 cascade setting and a 1024px PBR atlas. Texture sampling
   uses twelve steps before component upscaling, giving refinement more source detail.
3. PartField predicts learned 3D features. Cluster those features into the requested number
   of parts, partition the original faces, crop component texture bounds and remap UVs.
   This is learned part segmentation, not merely splitting disconnected islands. The
   target count controls granularity; semantic correctness is not guaranteed.
4. Real-ESRGAN x4plus upscales each component's base color, with tiled overlap. Profile a
   maximal tile, inspect free RAM, CUDA memory and cgroup CPU limits, and admit the maximum
   workers fitting 90% of those resources. Recalculate each wave. Minimum concurrency is
   one; if one cannot fit safely, shrink tiles or fail instead of overriding the reserve.
   CUDA OOM halves concurrency and then tile size. GPU leases and requests are serialized;
   component workers run concurrently within the lease. Alpha is interpolated separately.
   Normal, metallic, roughness, occlusion and emission maps preserve their data and UVs.
5. Astra (`openai/gpt-6-astra`) reviews the reference and rendered views. An embedded fx
   agent inspects per-run evidence inside a fresh just-bash virtual filesystem and returns
   a validated repair plan. Its tools cannot access the host filesystem, shell or network.
   The host applies bounded repairs, renders again and preserves the original on regression.
   Quality acceptance is a separate decision from non-regression.
6. For an animation prompt, learn the humanoid skeleton and weights, transfer them onto
   the original textured components, and generate the requested motion with pinned native
   NVIDIA Kimodo. Retarget it in Blender and export the actual animated GLB and MP4.
7. Measure every exported frame for new triangle intersections and ground penetration;
   render chronological front/side evidence and ask Astra to judge the requested action,
   deformation, texture fidelity and contacts. A failing review remains `needs_review`.
8. Package the original inputs, candidate meshes, textures, rig, source motion, GLB,
   video, audit evidence and provenance. Downloads belong to one unique run directory.

## Local development

```sh
uv sync --locked
uv run pytest
uv run app.py
```

On a Mac this starts the interface without loading CUDA models. Generation reports that
the GPU runtime is required; it does not substitute sample meshes. `PORT=7861 uv run app.py`
starts another local instance. No public sharing is enabled.

## Hugging Face ZeroGPU deployment

Create a **Gradio Space**, select **ZeroGPU** hardware, and upload this checkout including
`requirements.txt`, `requirements-hf.lock`, `requirements-gpu.txt`, `packages.txt`, `models.lock.json`, `uv.lock`, `vendor/fx/`,
`studio/`, `scripts/`, and `assets/`. Do not upload `.runtime/`, credentials, or old outputs.
Add `OPENROUTER_API_KEY` as a Space secret. Add `HF_TOKEN` from an account with access to TRELLIS’s gated DINOv3 and BRIA
RMBG-2.0 dependencies. Obtain that access on the model pages before startup; setup
does not accept model terms for you. The exact `openai/gpt-6-astra` route is used for review and fx repair planning.
OpenRouter billing and availability apply; no substitute review model is selected.
Native Kimodo also requires access to `meta-llama/Meta-Llama-3-8B-Instruct`.

HF installs the CUDA dependencies and Blender at build time. On startup `app.py` detects
`SPACE_ID`, downloads pinned source/models and installs the pinned embedded fx/Node runtime, then loads models before handling
requests. GPU stages use `@spaces.GPU`; CPU exports, rendering and fx run outside leases.
Only CPU inputs/results cross worker queues. Upscaled textures are returned explicitly
to the parent process; models and UI callbacks are never serialized as worker arguments.
The supported wheel target is **Linux x86_64, Python 3.12, torch 2.11, CUDA 13** and follows
Microsoft's current TRELLIS.2 ZeroGPU Space. This is not a Docker Space.

For a matching local NVIDIA machine:

```sh
uv sync --locked
uv pip install -r requirements-gpu.txt
uv run --no-sync scripts/prepare_runtime.py
STUDIO_NATIVE=1 uv run --no-sync app.py
```

Use `--no-sync` after adding the GPU requirements; normal uv sync intentionally manages
only the portable development environment. GPU requirements are kept separate because
the upstream CUDA wheels cannot install on macOS. Startup requires substantial model
storage and RAM; downloads are cached. ZeroGPU quota expiry, unavailable models and memory
failures are surfaced with retained intermediate files, never reported as completed runs.

Refinement uses `libfx` with an explicit host-owned OpenRouter transport and a fresh
in-memory just-bash filesystem for each attempt. Only the run's review evidence is mounted;
credentials remain outside tools. Tool calls, model calls, memory, output size and elapsed
time are bounded. Runtime dependencies and the upstream cleanup skill are pinned.


## Model choices and known limits

See [model decisions](docs/models.md) for sources, revision pins and the quality/latency
tradeoffs. “Newest” is not a claim of superior texture reconstruction. Neural upscaling
cannot recover details that the source model never generated.

The pinned fx source's custom OpenRouter adapter does not accept native image inputs. A separate
Ling vision call supplies visual evidence to Ling inside fx. The fx stage is constrained
to producing a validated plan; it does not execute arbitrary model-authored scripts in a
public Space. Technical validation and a VLM verdict are not a guarantee of visual quality.

The original code is in `stash@{0}` (`Archive pre-rebuild forge3d and Mercy experiments`).
Ignored old outputs, weights and local secret pointers were moved to the sibling directory
`../.3dgen-legacy-20260916/`. They are not runtime dependencies of this implementation.
Use `git stash show --stat stash@{0}` to inspect the archived source without restoring it.

The released fx v0.0.10 binary predates custom OpenRouter support. The Linux deployment uses a checksummed binary built from
commit `b51be034fea6a26cf9ac7aaa647fbe4bf089914b` with Zig 0.16; source and rebuild details
are in `vendor/fx/BUILD.md`. Other platforms build the same source during setup.

See [verification status](docs/verification.md) for executed checks and the remaining
ZeroGPU smoke test. Regenerate the Linux lock with:

```sh
uv pip compile requirements-hf.in --python-version 3.12 --python-platform x86_64-manylinux_2_28 --emit-index-url -o requirements-hf.lock
sed -e '/^-e \.$/d' -e 's/^torch==2.11.0+cu130$/torch==2.11.0/' -e 's/^torchvision==0.26.0+cu130$/torchvision==0.26.0/' requirements-hf.lock > requirements.txt
```

Archived source stash object: `0899c5725eddd965045cf645af45fbf9a60e10c2`.

## Continue an interrupted generation

If a later GPU lease fails, retain `source.glb` and `reference.png` from that run.
The **Continue an existing textured mesh** panel accepts those files and resumes
PartField, texture upscaling, review and packaging in a fresh request. Its API is
`/continue_asset(image, source_mesh, prompt, parts, seed)`. The prompt supplies the
review brief; it does not regenerate the reference. The continued run records
`continued_from_source` and preserves the supplied GLB bytes.

The September 18 astronaut experiment is in `outputs/animated-astronaut/`.
`scripts/animate_kimodo_astronaut.py` fits that specific character to an exported
Kimodo NPZ/BVH pair in Blender; it is not a generic auto-rigger. Its delivered
GLB retains eight components and a 21-bone skin. Revision 2 projects the reference
into a 4096px front texture and constrains the Kimodo motion to a smaller elbow-led
gesture. All 180 exported frames have zero new triangle intersections relative
to rest; seven static-mesh intersections remain. See `docs/verification.md` and
`outputs/animated-astronaut/final/collision-comparison.json` for the exact scope.