mantrakp's picture
Isolate reference inference in a dedicated ZeroGPU worker
5c331a4 verified
|
Raw History Blame Contribute Delete
9.32 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade
metadata
title: Component Studio Reference
emoji: 🧩
colorFrom: gray
colorTo: blue
sdk: gradio
sdk_version: 6.27.0
python_version: '3.12'
app_file: app.py
pinned: false

Component Studio

A new Gradio + uv pipeline targeting Hugging Face ZeroGPU. It accepts an image, a prompt, or both and produces a textured GLB with individually editable components. Add an animation prompt to run learned humanoid rigging, NVIDIA Kimodo motion generation, retargeting, video export, full-frame collision checks and Astra visual review.

The animated workflow is under production verification. A completed export is not a quality pass: rejected models or animations retain their artifacts with needs_review status.

Inputs

Input Example Behavior
Image only Upload a stool photograph Reconstruct the supplied image; skip image generation
Prompt only “A brass desk lamp with a green glass shade” Generate a reference with FLUX.2 Klein, then reconstruct it
Both Stool image + “Walnut seat and black metal legs; keep the shape” Edit the image with FLUX.2 Klein, then reconstruct the edited reference

Pipeline stages

  1. Validate inputs and prepare a reference. Images keep EXIF orientation and alpha.
  2. TRELLIS.2 generates geometry at its 1024 cascade setting and a 1024px PBR atlas. Texture sampling uses twelve steps before component upscaling, giving refinement more source detail.
  3. PartField predicts learned 3D features. Cluster those features into the requested number of parts, partition the original faces, crop component texture bounds and remap UVs. This is learned part segmentation, not merely splitting disconnected islands. The target count controls granularity; semantic correctness is not guaranteed.
  4. Real-ESRGAN x4plus upscales each component's base color, with tiled overlap. Profile a maximal tile, inspect free RAM, CUDA memory and cgroup CPU limits, and admit the maximum workers fitting 90% of those resources. Recalculate each wave. Minimum concurrency is one; if one cannot fit safely, shrink tiles or fail instead of overriding the reserve. CUDA OOM halves concurrency and then tile size. GPU leases and requests are serialized; component workers run concurrently within the lease. Alpha is interpolated separately. Normal, metallic, roughness, occlusion and emission maps preserve their data and UVs.
  5. Astra (openai/gpt-6-astra) reviews the reference and rendered views. An embedded fx agent inspects per-run evidence inside a fresh just-bash virtual filesystem and returns a validated repair plan. Its tools cannot access the host filesystem, shell or network. The host applies bounded repairs, renders again and preserves the original on regression. Quality acceptance is a separate decision from non-regression.
  6. For an animation prompt, learn the humanoid skeleton and weights, transfer them onto the original textured components, and generate the requested motion with pinned native NVIDIA Kimodo. Retarget it in Blender and export the actual animated GLB and MP4.
  7. Measure every exported frame for new triangle intersections and ground penetration; render chronological front/side evidence and ask Astra to judge the requested action, deformation, texture fidelity and contacts. A failing review remains needs_review.
  8. Package the original inputs, candidate meshes, textures, rig, source motion, GLB, video, audit evidence and provenance. Downloads belong to one unique run directory.

Local development

uv sync --locked
uv run pytest
uv run app.py

On a Mac this starts the interface without loading CUDA models. Generation reports that the GPU runtime is required; it does not substitute sample meshes. PORT=7861 uv run app.py starts another local instance. No public sharing is enabled.

Hugging Face ZeroGPU deployment

Create a Gradio Space, select ZeroGPU hardware, and upload this checkout including requirements.txt, requirements-hf.lock, requirements-gpu.txt, packages.txt, models.lock.json, uv.lock, vendor/fx/, studio/, scripts/, and assets/. Do not upload .runtime/, credentials, or old outputs. Add OPENROUTER_API_KEY as a Space secret. Add HF_TOKEN from an account with access to TRELLIS’s gated DINOv3 and BRIA RMBG-2.0 dependencies. Obtain that access on the model pages before startup; setup does not accept model terms for you. The exact openai/gpt-6-astra route is used for review and fx repair planning. OpenRouter billing and availability apply; no substitute review model is selected. Native Kimodo also requires access to meta-llama/Meta-Llama-3-8B-Instruct.

HF installs the CUDA dependencies and Blender at build time. On startup app.py detects SPACE_ID, downloads pinned source/models and installs the pinned embedded fx/Node runtime, then loads models before handling requests. GPU stages use @spaces.GPU; CPU exports, rendering and fx run outside leases. Only CPU inputs/results cross worker queues. Upscaled textures are returned explicitly to the parent process; models and UI callbacks are never serialized as worker arguments. The supported wheel target is Linux x86_64, Python 3.12, torch 2.11, CUDA 13 and follows Microsoft's current TRELLIS.2 ZeroGPU Space. This is not a Docker Space.

For a matching local NVIDIA machine:

uv sync --locked
uv pip install -r requirements-gpu.txt
uv run --no-sync scripts/prepare_runtime.py
STUDIO_NATIVE=1 uv run --no-sync app.py

Use --no-sync after adding the GPU requirements; normal uv sync intentionally manages only the portable development environment. GPU requirements are kept separate because the upstream CUDA wheels cannot install on macOS. Startup requires substantial model storage and RAM; downloads are cached. ZeroGPU quota expiry, unavailable models and memory failures are surfaced with retained intermediate files, never reported as completed runs.

Refinement uses libfx with an explicit host-owned OpenRouter transport and a fresh in-memory just-bash filesystem for each attempt. Only the run's review evidence is mounted; credentials remain outside tools. Tool calls, model calls, memory, output size and elapsed time are bounded. Runtime dependencies and the upstream cleanup skill are pinned.

Model choices and known limits

See model decisions for sources, revision pins and the quality/latency tradeoffs. “Newest” is not a claim of superior texture reconstruction. Neural upscaling cannot recover details that the source model never generated.

The pinned fx source's custom OpenRouter adapter does not accept native image inputs. A separate Ling vision call supplies visual evidence to Ling inside fx. The fx stage is constrained to producing a validated plan; it does not execute arbitrary model-authored scripts in a public Space. Technical validation and a VLM verdict are not a guarantee of visual quality.

The original code is in stash@{0} (Archive pre-rebuild forge3d and Mercy experiments). Ignored old outputs, weights and local secret pointers were moved to the sibling directory ../.3dgen-legacy-20260916/. They are not runtime dependencies of this implementation. Use git stash show --stat stash@{0} to inspect the archived source without restoring it.

The released fx v0.0.10 binary predates custom OpenRouter support. The Linux deployment uses a checksummed binary built from commit b51be034fea6a26cf9ac7aaa647fbe4bf089914b with Zig 0.16; source and rebuild details are in vendor/fx/BUILD.md. Other platforms build the same source during setup.

See verification status for executed checks and the remaining ZeroGPU smoke test. Regenerate the Linux lock with:

uv pip compile requirements-hf.in --python-version 3.12 --python-platform x86_64-manylinux_2_28 --emit-index-url -o requirements-hf.lock
sed -e '/^-e \.$/d' -e 's/^torch==2.11.0+cu130$/torch==2.11.0/' -e 's/^torchvision==0.26.0+cu130$/torchvision==0.26.0/' requirements-hf.lock > requirements.txt

Archived source stash object: 0899c5725eddd965045cf645af45fbf9a60e10c2.

Continue an interrupted generation

If a later GPU lease fails, retain source.glb and reference.png from that run. The Continue an existing textured mesh panel accepts those files and resumes PartField, texture upscaling, review and packaging in a fresh request. Its API is /continue_asset(image, source_mesh, prompt, parts, seed). The prompt supplies the review brief; it does not regenerate the reference. The continued run records continued_from_source and preserves the supplied GLB bytes.

The September 18 astronaut experiment is in outputs/animated-astronaut/. scripts/animate_kimodo_astronaut.py fits that specific character to an exported Kimodo NPZ/BVH pair in Blender; it is not a generic auto-rigger. Its delivered GLB retains eight components and a 21-bone skin. Revision 2 projects the reference into a 4096px front texture and constrains the Kimodo motion to a smaller elbow-led gesture. All 180 exported frames have zero new triangle intersections relative to rest; seven static-mesh intersections remain. See docs/verification.md and outputs/animated-astronaut/final/collision-comparison.json for the exact scope.