Five product photographs of the S.H.MonsterArts figure went in: four for the shape, one close-up for the colour. This textured, glowing, AR-ready model came out about ten minutes later, on one RTX 3080, with nothing sent to the cloud.
The inputs are four full-figure photographs (front, two three-quarter views, and a back view of the same mould in another colourway, used for shape only) plus a close-up of the back for colour. A multi-view model builds the shape from the four; the front photograph gives the colour; every colour photograph is then aligned to the model by silhouette and painted on where the surface faces its camera. 865,210 raw triangles come back; the shipped model is 59,112 triangles wearing the figure's own colour, its own relief, and its glow.
The front photograph on the left; on the right, the model rendered from the camera the alignment found for that photograph. Nothing here was posed by hand.
Things to look for as it turns: the red lines down the flanks continue onto the back, where no single photograph could have put them; the tail tip and jaw glow from any angle because they are emission, not paint; the feet keep their claws.
Every stage is local. The photograph never leaves the machine, and there is no per-model licence.
A trained segmentation model (BiRefNet) separates the figure from the studio gradient. A colour flood-fill cannot; it kept 77% of this frame as "subject".
Hunyuan3D, fed the four full-figure views, returns the shape: the jagged dorsal plates come out as plates, because a back view saw them. TripoSplat, fed the front, returns the figure's own colour, red glow included, on a mesh of its own. The colour mesh is registered onto the shape mesh by silhouette and its colour transferred vertex by vertex. One photograph alone had turned the plates into a row of knobs.
Pieces that came back detached from the figure (a caption strip, a patch of backdrop) are dropped; every photograph is segmented with a trained model first, so there is little to drop. When the colour engine works alone its mesh is a double wall with ghost sheets inside, and a separate stage carves one clean skin from 64 depth views; the hybrid's shape mesh does not need it.
Colour is sampled per texel from the dense reconstruction, not from the shipped mesh's vertices, so the atlas carries the eye, the teeth and the skin folds rather than a blur between thirty thousand points. The raw mesh's scales and fin edges, 6° of mean relief, go into a normal map on the same atlas, so a 59,112-triangle model shades like the dense one.
Each extra photograph is segmented, aligned to the model by silhouette (a search over camera angle, then scale and offset), and projected onto the atlas where the surface faces its camera and is not hidden. Exposure is matched to the existing colour first, so a darker shot cannot paint a darker stripe, and a photograph whose colours refuse to correspond is refused. The extra views here painted 746,087 texels the front view could not see.
Saturated colour several times brighter than the body's median is emission, not paint: 0.9% of this surface. It ships as an emissive map, so the gill lines and the tail tip light up in a dark room.
Creases in the surface divide the figure into 18 parts. Any part can be repainted or given a different material in the studio, or left as photographed.
One bake writes a glTF for Android's Scene Viewer and a USDZ for iOS AR Quick Look, 18 cm tall, with pinch-to-resize disabled on purpose: the question AR answers is whether it fits.
One photograph is enough for a recognisable, textured model: the reconstruction invents the sides it never saw, and on a roughly symmetric product that guess is good. Here the side view keeps the whole tail and its ridges, at the cost of a thinner body. The four-photograph build above differs in the colour of the flanks and back, which the extra views painted from the real thing instead of from a guess, and in taking three minutes instead of one. Pose matters more than count: the two photographs where the figure's tail was posed differently aligned under 60% and were refused.
The model above builds its shape from the front photograph alone, and the front photograph shows the dorsal plates only as a jagged outline. So it invents them: a neat row of rounded knobs. A second engine, fed all four photographs for the shape and the front one for colour, gets the plates.
The price of the second engine, stated plainly: it fits the front photograph at 76% instead of 81% (it bent the tail forward, because the only back photograph in the set is a close-up and the engine had to reconcile it with three full-figure views), its chest picked up misplaced colour from the three-quarter photographs, and it takes about ten minutes instead of three, most of that loading a second model onto the same 10 GB card. With a full-figure back photograph in the same pose it would have had a consistent set to work from. The viewer at the top shows the first engine's model; the plates are the second's.
Same pipeline, a different kind of product: a Nike Dunk Low from four of its store photographs (lateral, medial, back, three-quarter). Silhouette match of each: 95%, 95%, 90%, 92%. With four photographs that well placed the second pass carved the reconstruction back to their outlines.
The Swoosh, the panel colours and the white midsole land where they are in the photographs; the laces and the lumpy surface do not, because the reconstruction's geometry is the weak half for a shoe. The gallery had no true front view, and the toe shows it. Glow detection correctly found nothing: yellow is bright, but nothing on the shoe is brighter than its own white midsole.
Surfaces no photograph saw are still invented from the model's prior, and a photograph only paints where the figure's pose matches the model's: a shot with the tail posed differently is refused rather than smeared on. Detail finer than the 2.5 mm voxel grid softens. The dorsal plates on the main model are the reconstruction's invention; the section above shows what the all-photo engine does about that, and what it costs. Small square flecks along the flanks are seams of the six-tile atlas where neighbouring faces were painted from different views.
AR itself runs from the app on the local network, where a phone opens the same glTF or USDZ in its own AR viewer. This page embeds the identical file in a web viewer so it can be shared without the app.