A 3D person built from a single photograph, driven by a pose-estimation model running inside this page. It's playing a sample clip now. Turn on your camera and it copies you instead.
Loading the avatar and the pose model…
Private by design. Your camera feed is processed on your own device and never uploaded. There is no server behind this page.
How we built it
One T-pose photo, cut out with Grounding DINO and SAM 2.1, then turned into a textured 31k-triangle mesh by Meshy image-to-3D in under 3 minutes. The first attempt fused a front view and a back view into one mesh. A clean white background fixed it.
Our first rig (the AniGen skeleton plus Kimodo motion, retargeted by us) walked like a stiff robot. We swapped in Meshy's auto-rigger, which built the 24-bone skeleton in 37 seconds, and its motion-capture library: six clips retargeted server-side in 48 seconds, for 23 credits in total. Textures were shrunk from 7.9 MB to 3 MB for the web.
Google's MediaPipe Pose Landmarker (Apache-2.0) finds 33 body landmarks in 3D on every video frame, in WebAssembly, on your GPU where available.
For each limb we measure the bone's direction in the avatar's rest pose and rotate it onto the matching landmark direction. We solve parents before children, so the elbow inherits the shoulder. A torso frame from the shoulders and hips turns the spine, and the ears and nose turn the head. The idle clip keeps playing underneath, and each tracked limb fades in over it only while the camera can see it, so a half-visible person never snaps to a stiff default pose.
three.js r186 with physically based materials, in a plain static page. No build step, no backend, no per-user cost.
What it doesn't do yet