EXP-050 · live demo

Move, and the avatar moves with you.

A 3D person built from a single photograph, driven by a pose-estimation model running inside this page. It's playing a sample clip now. Turn on your camera and it copies you instead.

Input · sample performer
Output · 3D avatar Drag to orbit
Pose model
–
Tracking
–
Landmarks
–
Runs on
–
Or play a motion-capture clip:

Loading the avatar and the pose model…

Private by design. Your camera feed is processed on your own device and never uploaded. There is no server behind this page.

How we built it

From one photo to a live puppet

  1. Photo → 3D person

    One T-pose photo, cut out with Grounding DINO and SAM 2.1, then turned into a textured 31k-triangle mesh by Meshy image-to-3D in under 3 minutes. The first attempt fused a front view and a back view into one mesh. A clean white background fixed it.

  2. Skeleton and motion library

    Our first rig (the AniGen skeleton plus Kimodo motion, retargeted by us) walked like a stiff robot. We swapped in Meshy's auto-rigger, which built the 24-bone skeleton in 37 seconds, and its motion-capture library: six clips retargeted server-side in 48 seconds, for 23 credits in total. Textures were shrunk from 7.9 MB to 3 MB for the web.

  3. Pose estimation in the browser

    Google's MediaPipe Pose Landmarker (Apache-2.0) finds 33 body landmarks in 3D on every video frame, in WebAssembly, on your GPU where available.

  4. Retargeting and blending, our part

    For each limb we measure the bone's direction in the avatar's rest pose and rotate it onto the matching landmark direction. We solve parents before children, so the elbow inherits the shoulder. A torso frame from the shoulders and hips turns the spine, and the ears and nose turn the head. The idle clip keeps playing underneath, and each tracked limb fades in over it only while the camera can see it, so a half-visible person never snaps to a stiff default pose.

  5. Rendering

    three.js r186 with physically based materials, in a plain static page. No build step, no backend, no per-user cost.

What it doesn't do yet

  • No fingers or facial expressions: the pose model gives wrists, not hands. Hand and face landmarkers are the next step.
  • One camera can't measure depth precisely, so arms reaching straight at the lens are the least accurate.
  • The legs follow only when your whole body is in frame. Otherwise the idle clip drives them.
  • The avatar stays in place: it copies your pose, not your position in the room.