Skip to content
←Back to Open Source

OPEN SOURCE DEEP DIVE

3DGenerationWorldModelGaussianSplatting

image-blaster: blast one image into a collidable 3D world

A set of Claude Code skills that turns one image into an explorable Gaussian-splat environment with a collider mesh and metric scale, plus interactable object meshes and sound effects, in under five minutes.

neilsonnn/image-blaster5.1kTypeScriptMIT10 min read

What it is

image-blaster (neilsonnn/image-blaster, 5,076 stars / 521 forks, TypeScript, MIT) is not another image-to-3D tool. It is a set of skills that ride on Claude Code: hand it one image and it splits that image into three things — an explorable static environment (a Gaussian splat .spz, a collider .glb, and a panorama), dynamic object meshes you can lift back out (.glb / .obj), and ambient plus per-object sound effects (.mp3). The README puts the budget at under five minutes from a single image to a fully meshed 3D environment.

It trains no models and invents no generative algorithm. Four external models do the work: World Labs marble-1.1 builds the environment, hunyuan3d-v3 on FAL builds the objects, nano-banana or gpt-image-2 handles image editing, and ElevenLabs via FAL produces the audio. The entire value of image-blaster is that it orchestrates that multi-model pipeline into a stateful, resumable, disk-first piece of engineering.

The numbers

FieldValue
Reponeilsonnn/image-blaster
Stars / Forks5,076 / 521
Language / LicenseTypeScript / MIT
Created / last commit2026-04-21 / 2026-05-14
Tracked files140 (the .claude/ skills and scripts, plus the React viewer under app/)
External modelsWorld Labs marble-1.1 (world), FAL hunyuan3d-v3 (objects), nano-banana / gpt-image-2 (image edit), FAL elevenlabs-sfx (audio)
Keys requiredWORLD_LABS_API_KEY and FAL_KEY in .env; a SessionStart hook validates them
Artifacts.spz splat, .glb collider, .glb/.obj object meshes, .png panorama and thumbnail, .mp3 audio
How you run itRun claude inside the repo, drop an image into input/, say blast it

Design decision one: the clean plate

This is the move worth stealing, and the dividing line between this project and "just hand the raw photo to a world model."

Feed the original image straight into Marble and the chairs, mugs and statues get baked into the splat together with the floor and walls. They look fine and they are physically dead: you can never pick one up or push it. image-blaster instead runs image-blast-plate first, which uses an image-edit model to erase every confirmed object from the source image, producing a clean plate — an empty scene — and only then sends that plate to Marble to generate the static environment.

The rules around it are unusually precise. The plate prompt must be removal-only: name what to delete, never add fill-in instructions like "repair the background" or "keep the lighting consistent," and never list the objects that should stay. Every removal has to happen in a single edit pass — no splitting across agents or one edit per object. The plate is also a new source artifact and must take the next free visible index in source/ (if the source is 0-room.png, the first plate is 1-room-plate.png, never 0-room-plate.png), because later steps default to "the highest-index visible source image."

Design decision two: the world prompt is subtraction, not copy

The prompt that generates the environment is not the image caption verbatim. image-blast-world reads worlds/<slug>/image.json for the original scene description and then subtracts every confirmed, removed object from it, synthesizing a text clean plate: keep the setting, materials, lighting, atmosphere, camera feel and spatial layout, but describe the scene as empty. Reusing imageJson.short_caption directly is explicitly forbidden, as is naming, implying or reintroducing removed objects.

Why it matters: the world model receives a scene prior that contains no objects, so the splat it produces carries no ghosts of them. The objects are generated separately as 3D meshes and dropped back in as interactable rigid bodies. Environment and objects are decoupled at generation time rather than being separated after the fact — the single most important engineering judgement in the project.

Design decision three: disk-first, provider URLs are provenance only

One rule repeats throughout the project docs: provider URLs are provenance and resume metadata; the frontend loads only local /worlds/... files. After Marble responds, generate-world.mjs downloads every .spz, collider .glb, panorama and thumbnail to its matching local filename, and strips base64 out of the request JSON before writing it. ensure-local-assets.mjs backfills missing files from recorded metadata, but it never creates a new generation — it only repairs local state.

Everything follows one indexed naming convention:

ConventionMeaning
N-slug.extVisible artifact. N is the generation index; 0 is the source, higher numbers are derived generations
.N-slug-request.jsonHidden request metadata sitting beside the file it produced
Shared indexOne world generation produces N-world.json, N-world-plate.png, N-world.glb, N-world-pano.png, N-world-thumbnail.webp and N-world-full_res.spz
.N-world-request.jsonAn unfinished request is resumed by the next run instead of restarting

project-state.mjs is the read-side entry point for this convention; every skill calls it before and after, deriving state from ls -a and JSON sidecars rather than from conversation memory. There is also a disciplined instruction worth quoting: do not Read generated PNG/JPG files just to QC them — if the user wants to look, open the folder instead of pulling images into context.

The nine-step IMAGE-BLAST

.claude/rules/project.md fixes the one-shot order. You can checkpoint with the user after each step or let it run end to end:

  1. Inspect project state and input/ (a UserPromptSubmit hook injects the staged file list into context);
  2. Initialize the project by slug and stage inputs into worlds/<slug>/source/;
  3. Check port 5173 with lsof -i :5173; if free, run bun install && bun run dev and report the viewer URL via show-url.mjs;
  4. image-blast-uncover runs multimodal image analysis and extracts separable object candidates;
  5. Confirm objects, write one object.json each, and decide on the clean plate;
  6. image-blast-world builds the static environment from the newest source image (possibly the plate just generated);
  7. Launch one image-blast-3d per confirmed object;
  8. Launch audio jobs: one ambient loop for the world, one impact set per object;
  9. Report project state and all artifact paths. Done.

The extraction criterion is deliberately conservative: only single, cleanly segmentable items a human could lift or push. Rugs, flooring, walls and fixed architectural features are excluded, and compound assets are forbidden — no table-with-chairs, no table-including-what-is-on-it. That rule decides whether the 3D step succeeds, because a compound object fed to Hunyuan comes back as mush.

The viewer: splats plus trimesh colliders plus rigid bodies

The bundled app/ is a Vite + React 19 browser viewer, and its dependency list tells you the ambition: three@0.180 and @react-three/fiber@9 for rendering, @sparkjsdev/spark@2 for splats, @react-three/rapier@2 for physics, plus drei and postprocessing.

The interesting part is WorldCollider.tsx. It loads the .glb collider Marble produced into a type="fixed", colliders="trimesh" rigid body and aligns it using three semantic fields from the world metadata: metric_scale_factor scales it to real-world size, ground_plane_offset lifts the floor to the origin, and flip_y rotates it by pi about X. In other words, the environment that merely looks like a pretty splat also carries trimesh collision geometry at true scale — the character controller can walk on it and dropped objects land instead of falling through. Placement lives in scene.json: every instance records position, rotation and scale, plus sun intensity and orientation, and each object can be switched between rigidbody, static and ghost.

The splat itself is tiered. Marble returns 500k, 150k, 100k and full_res .spz files; the script downloads all of them and the frontend picks by quality level.

Tunable knobs for 3D objects

The default provider is hunyuan3d-v3/image-to-3d on FAL, with meshy as an alternative. The generator takes four parameters:

FlagDefaultEffect
--face-count50,000 (range 40,000–1,500,000; the Hunyuan API default is 500,000)Target face count
--enable-pbrtrueGenerate PBR materials
--generate-typeNormalNormal textured, LowPoly decimated, Geometry white geometry only
--polygon-typetriangleTriangles or quads under LowPoly

The first run generates a reference image with the edit model: white background, centered, tight crop, studio lighting, the target object isolated and reproduced exactly, with explicit exclusion of anything clustered, adjacent, overlapping or resting on it. Later runs reuse that reference; only --regenerate-reference re-extracts from the source. The prompt insists on "one single object that is true to the source image" — not a pair, a set, or a category example.

Audio has three modes. Ambience uses --loop --count 2 --kind world-ambience --prefix ambient-loop --duration-seconds 10. Object impacts use --count 4 --kind object-impact --prefix impact-<id> --duration-seconds 1 and must not pass --loop. Non-loop output is post-processed with ffprobe / ffmpeg: leading and trailing silence or noise is trimmed, loudness is normalized, and the analysis is stored in the hidden request JSON. Loop output is left as raw provider audio so the seam stays intact.

Why this matters for embodied AI

The README example list includes "Need an environment for a robot? IMAGE-BLAST it." That line is not filler. Assembled from the sections above, the output is an explorable environment with real-world scale, trimesh collision, and objects separated from the scene with their own physics — exactly the three things simulation assets usually lack. Getting from a photograph to something a robot can run in normally means modeling, UV work, collider simplification and unit calibration by hand; image-blaster compresses that into one sentence and five minutes.

The contrast makes it clearer. Single-object image-to-3D lines (Trellis and its kin) hand you a pretty mesh that is an isolated asset — no scale, no ground, nothing to walk on. Pure world-model / video generation gives you continuous scenes but the output is pixels, with no geometry to grab. image-blaster sits between them. It does not chase physical accuracy; it makes a layered compromise — splat for looking, collider mesh for touching, object meshes for moving — and gets you a usable intermediate state you can drop into an engine today.

The cost should be stated plainly. This is not a physically accurate reconstruction pipeline. metric_scale_factor is estimated by the world model, the collider is an approximate trimesh Marble emits as a byproduct rather than CAD-grade geometry, and object meshes come from image-to-3D with uncontrolled topology and face counts. The README's own framing is jumpstarting 3D work: a rough block you can run with, refined afterwards by a human in Blender or the engine.

Caveats

  • Fully hosted, no offline path. Without WORLD_LABS_API_KEY and FAL_KEY nothing runs, and every generation is a paid API call.
  • The repo has gone quiet. The last commit is 2026-05-14 while provider APIs keep moving; hardcoded endpoints like marble-1.1 and hunyuan3d-v3 are future maintenance points.
  • Artifacts are heavy. A full-res splat plus panorama plus several meshes plus audio runs to tens of megabytes per project; the repo itself ships a 43 MB sample butterfly scene.
  • Object extraction sets the ceiling. Anything uncover misses or wrongly merges propagates through the 3D step and the clean plate, which is why the flow deliberately puts user confirmation of objects before any generation.

Bottom line

The real contribution is not that it wired up a few generative models. It is that image-blaster turned "one image into a world you can walk into" into an engineered flow with explicit artifact contracts, resumable requests, and environment/object decoupling at generation time — and then shipped the viewer, the physics and the placement editor along with it. If you need a rough simulation scene for a robot or a game, it is currently the shortest path there.

As an Amazon Associate, we earn from qualifying purchases.