OPEN SOURCE DEEP DIVE
Modly: Local, Open-Source Image-to-3D Mesh Generation Desktop App
Modly is an open-source desktop app from Lightning Pixel (Windows, Linux, Apple Silicon macOS; v0.4.3; MIT plus an attribution clause) that turns a photo into a usable 3D mesh with open models running on your own GPU, so inference never leaves the machine. An Electron shell supervises a Python FastAPI backend bound to 127.0.0.1:8765. The default model is TripoSR (about 2.4 GB); Hunyuan3D 2 Mini, TripoSG and Trellis2 GGUF arrive as extensions, where an extension is a GitHub repo with a manifest.json that can declare multi-repo model_sources, shared weight_groups and separately installable weight_variants carrying size_gb and vram_gb. Generation is expressed as a node-graph workflow (Image to Generate Mesh to Add to Scene), scenes follow the modly.scene-manifest.v1 contract, and export covers glb/stl/obj/ply with a dedicated stl/obj printing path that scales to millimetres for OrcaSlicer. Three agent surfaces ship in the repo: a stdlib-only JSON-first CLI with its own SKILL.md, an MCP server exposing nine modly_* tools, and an in-app tool-calling agent that can run on a local llama.cpp GGUF pool or an external model endpoint. AMD ROCm auto-detection, memory-budgeted Apple Silicon workflows and headless Jetson runs are documented with measured numbers.
What it is
Modly is an open-source desktop application from Lightning Pixel that bundles three commitments into one installer: local, open source, and image-to-3D mesh generation. You hand it a photo (or a prompt), it runs open models on your own GPU, and it hands back a usable 3D mesh without inference ever leaving the machine. The repository was created in March 2026 and is at version 0.4.3, written mainly in TypeScript, with roughly 8,100 GitHub stars and 760 forks; the project site is modly3d.app. It ships for Windows, Linux and Apple Silicon macOS, and the architecture decision record (ADR) explicitly puts Intel Macs, universal binaries and Rosetta fallback out of scope.
The license is MIT plus an attribution clause: redistributions must keep the Lightning Pixel credit visible in the app UI or documentation. Alongside platform installers, you can clone the repo and run it without installing anything via launch.bat or ./launch.sh.
Process model: an Electron shell supervising a local FastAPI
Modly is not a front end calling a cloud service; it is a desktop shell that owns an inference service. The Electron main process spawns and manages a Python FastAPI (uvicorn) bound to 127.0.0.1:8765. The renderer is React with @react-three/fiber and drei for the 3D view, @xyflow/react for node graphs, zustand for state, three-mesh-bvh for geometry acceleration, and gaussian-splats-3d for splat assets. The routers registered in api/main.py lay the capability surface out plainly: status, settings, model, generate, optimize, extensions, export, workflow-runs, agent and llm, plus a /workspace/{full_path} endpoint that serves generated files straight out of the workspace.
Two small details are worth stealing. ensure_utf8_stdio() must run before any print or logging touches the pipe, or Windows subprocess pipes mangle UTF-8 output. And CORS sets expose_headers=["Content-Length"] because, as the code comment says, drei's SplatLoader sizes its buffers from that header and cross-origin JS cannot see headers the server does not expose. This architecture is full of constraints that fail silently unless written down.
Extension system: models are installed from GitHub, not built in
Modly ships no large model weights and never installs PyTorch on an extension's behalf. The default model is TripoSR (about 2.4 GB), downloaded into ~/.modly/models/TripoSR/. Everything else is an extension, and an extension is simply a GitHub repository containing a manifest.json: open the Models page, click Install from GitHub, paste the HTTPS URL. Five official extensions exist today: Hunyuan3D 2 Mini with its Turbo and Fast variants, TripoSG, and Trellis2 GGUF.
The manifest fields solve the unglamorous parts of real distribution. model_sources lets one model node pull weights split across several Hugging Face repositories (each validated, downloaded sequentially, and mutually exclusive with weight_variants). weight_groups lets nodes inside one extension share weights. weight_variants exposes separately installable quantizations, each declaring size_gb (download size) and vram_gb (approximate VRAM), so the user knows the disk and memory cost before clicking download. params_schema declares node parameters and defaults. Downloads are observable by design: byte-level progress, .part resume issued against the resolved final URL so Range still works when upstream redirects to a CDN, and completion verified by the declared download_check rather than by "the directory exists".
Node-graph workflows and the scene contract
Generation is expressed as a node graph. The documented starter graph is Image → Generate Mesh → Add to Scene: wire it on the Workflows tab, select it on the Generate tab, click Generate 3D Model, and read failures in the Settings/Logs/Errors panel. Graphs go through preflight validation before execution, and an invalid graph leaves the current mesh view in place and surfaces inline warnings and toasts instead of replacing the viewport with a terminal error state. That choice is written into the Apple Silicon ADR as well: a long-running tool must not wipe a user's existing result because of one bad wire.
The harder edge is the scene contract. A scene is a workspace directory containing scene-manifest.json with schema modly.scene-manifest.v1, not an arbitrary JSON file, and the Load Scene node selects and validates it. Scene-capable generators implement generate_artifact(input_kind, artifact_path, ...); the generic POST /generate/from-artifact boundary currently accepts only scene, and in this first contract scene is model-only and must be the single input value (not inside inputs) — process nodes and mixed-input scene nodes are rejected. Legacy image generators and POST /generate/from-image are untouched. Mesh optimization (smooth and decimate) accepts both workspace-relative meshes and imported absolute-path meshes, and writes results back into the workspace so they stay visible and reusable.
Three agent surfaces: CLI, MCP, in-app agent
Modly treats "driven by an AI" as a first-class mode and offers three non-interchangeable paths. The first is tools/modly-cli/agent.py, a stdlib-only CLI whose canonical commands are health, model, workflow-run, capability and process-run. Final machine-readable JSON goes to stdout, progress JSON lines to stderr. The friendly generate command wraps POST /workflow-runs/from-image plus polling plus export, returns recovery metadata such as status_command and cancel_command, and fails structurally (for example {"ok": false, "code": "API_UNAVAILABLE"}). The repo even ships a SKILL.md stating the prerequisite that the desktop app must be running first, which is clearly aimed at Claude Code / Codex-style skill systems. Compatibility surfaces are deliberately separated: legacy wraps the old /generate/* job endpoints, dev serve-api starts only FastAPI and therefore does not prove the Electron bridge is ready, and experimental comfy-* are external ComfyUI orchestration helpers rather than the canonical contract.
The second is api/mcp_server.py, which exposes nine modly_* tools to any MCP client: list_models, switch_model, generate_from_image, get_generation_status, decimate_mesh, smooth_mesh, import_mesh, unload_models and get_settings. The third is the in-app agent: api/routers/agent.py carries its own nine tools (list_models, unload_models, get_mesh_info, decimate_mesh, smooth_mesh, get_generation_status, list_workflows, run_workflow, create_workflow), builds workflow graphs on demand, executes tool calls, and can run against local GGUF models or an external model endpoint with an API key.
Backing the in-app agent is a local LLM engine. api/services/llm_server.py manages a pool of llama.cpp llama-server subprocesses (default port 8791, one process and one port per model, concurrency configurable with "auto" sized from VRAM). The binary variant is chosen per platform: CUDA when an NVIDIA driver is present, otherwise Vulkan, otherwise CPU. Models come from a catalog or from any .gguf the user drops into ~/.modly/llm/models/. The comment states the trade-off bluntly: Modly is first and foremost a 3D-generation app, so an idle model is evicted after IDLE_TTL_SECONDS (300 by default) and never sits on VRAM the generator needs. The agent also unloads the local LLM once a workflow finishes.
Export: from glTF to the slicer
Export supports glb, stl, obj and ply, with a separate path aimed at 3D printing. Slicer formats are restricted to stl and obj, and the code comment explains why: OrcaSlicer cannot import glTF/GLB, so .glb is deliberately excluded there. The print path also scales the mesh to a real-world size via longest_mm and rotates it 90 degrees about X to match print orientation. For an image-to-3D app, that block of code is the last mile from generated geometry to a physical object.
Platform engineering: ROCm, Apple Silicon, Jetson
Three documents expose the most solid part of this project: it treats "what about non-NVIDIA desktops" as an engineering problem rather than a disclaimer. For AMD, detection is centralised in electron/main/gpu-detect.ts; NVIDIA keeps priority so dual-vendor machines behave as before; AMD machines always report gpu_sm = 0 and cuda_version = 0 and never a synthesised compute capability. Extensions learn they need ROCm wheels through a torch_flavor setup argument, and third-party extensions that hardcode --index-url .../whl/cu124 and cannot be edited are corrected by a rewrite shim that intercepts their pip calls in electron/main/setup-launcher.ts. The ADR explains that gpu_sm = 0 is load-bearing: older extensions branch on that number to pick their most conservative path, which also keeps them off onnxruntime-gpu (CUDA-only). PyTorch's HIP build answers the whole torch.cuda API, so get_device_capability() reports (12, 0) for a gfx1200 Radeon, indistinguishable from an sm_120 Blackwell — which is why GPU queries go to nvidia-smi rather than torch. Compute-target discovery is platform-specific: Linux reads gfx_target_version from the kernel KFD topology with no ROCm install and no external binary, while Windows maps PCI device ids from Win32_VideoController through a table keyed by silicon; an unmapped AMD card falls back to CPU with an actionable message instead of guessing a wheel. The verification record is equally explicit: on a Radeon RX 9060 XT (gfx1200), torch 2.13.0+rocm7.2 loads, 14 GB of the card's 16 GB allocates and reads back cleanly, and a full image-to-3D generation through the normal ExtensionProcess path takes 221 s. Windows had its wheel URLs, cp311 availability and index layout checked but no end-to-end run. The same work also fixed an unrelated AppImage bug: ensureStableEmbeddedPython() copied the bundled runtime with fs.cp, which rewrites relative symlinks into absolute paths pointing at the ephemeral /tmp/.mount_Modly-XXXXXX/ mount, so the "stable" copy was not stable and every extension venv built from it died on the next launch with a misleading No module named 'PIL'.
On Apple Silicon the binding constraint is unified memory: overlapping heavy GPU stages can destabilise a 16 GB machine, so workflows are sequential and memory-budgeted, with one heavy generative stage resident at a time and handoffs through files. Python-side cleanup does not reliably return MPS memory, so the dependable release boundary is terminating the owning subprocess: the Python bridge runs as its own process-group leader on Unix, the group is killed on quit, and cancel sends a cooperative request before escalating to a kill after a short grace period. For Jetson there is an honest unofficial guide. Modly does not target Jetson, but because the FastAPI backend is fully standalone and HTTP-driven you can skip Electron and a display entirely, run the backend headless on a Jetson AGX Orin (64 GB, JetPack 6.2, sm_87) and drive it with curl. The guide lists the three patches you must apply: replace torch with the jetson-ai-lab Jetson build (the pinned torch 2.5.1+cu124 SBSA wheel has no sm_87 kernels — it loads, then every kernel fails with "no kernel image is available"), pin numpy < 2, and bypass rembg / ONNX Runtime, which aborts hard on the Tegra CPU. The verified model is Hunyuan3D-2 Mini, mesh-only, in the ~6 GB VRAM class.
Who it is for, and what to take from it
As a tool, Modly's position is clear: people who need to turn reference images into mesh assets locally and in volume, without uploading those images to any cloud service — individual creators, the asset-prep stage of game and visualization pipelines, and research setups that need offline reproducibility. The repo also contains api/texture_baker and api/uv_unwrapper (native modules for texture baking and UV unwrapping), but the AMD ADR states plainly that these native extensions are not built as part of standard extension setup, so texture generation is not covered on the ROCm path; the Jetson run was mesh-only too. Geometry is the stable part today, while texturing still depends on the platform.
As an engineering sample, three things are worth copying. First, outsourcing model distribution to "GitHub repo plus manifest.json" keeps the core free of weights and of torch installation, so the extension ecosystem can evolve independently. Second, treating the agent surface as a contract — JSON-first CLI, MCP server, structured error codes, a SKILL.md shipped in-repo — instead of leaving only a UI. Third, writing non-NVIDIA platform support as dated, status-tagged ADRs with measured numbers that separate "verified on Linux" from "not verified on Windows". That level of honesty is not common in open-source AI tooling.