← Lab bench
EXP-050Live

One photo to an avatar that copies you through your webcam

26ms

per video frame to find 33 body landmarks, in the browser, with no server

Full-body photo of a man in a navy suit standing in a T-pose, cut out on white
1/4One T-pose photo, cut out
  1. 1 · Photo
  2. 2 · 3D model
  3. 3 · Talking
  4. 4 · Waving

A single photo becomes a textured, rigged 3D person with a motion-capture library. A pose model running inside the web page then reads your body from your webcam, and our retargeting turns those landmarks into the avatar's bone rotations, live.

What we tried

  • Grounding DINO and SAM 2.1 to cut the person out onto white, then Meshy image-to-3D for the mesh.
  • Our own rig first: an AniGen skeleton with Kimodo text-to-motion, retargeted in Blender.
  • Meshy's auto-rigger and its motion-capture library instead, six clips retargeted server-side.
  • MediaPipe Pose Landmarker in the browser, with our own bone-direction retargeting blended over the idle clip.

What we measured

MeasureFirst attemptShippedNote
Photo to 3D meshFused front, back and side copies165 s, one clean figureFixed by cutting the person out onto white first
Skeleton and motionAniGen + Kimodo: moved like a robot85 s: rig plus 6 mocap clips
Avatar download11.2 MB3.4 MB
Pose model per frame—26 ms on GPU, 33 landmarks

What went wrong

  • Meshy's first attempt from the raw photo, a circular crop against a grey wall, returned one mesh containing three fused copies of the person.
  • Our first skeleton and motion retarget worked, but the result walked stiffly, so we replaced both with Meshy's rig and library.
  • The texture atlas has hundreds of small islands, and mipmapping bled the white shirt into the navy suit as speckles. We turned mipmaps off for that texture.

What happens next

  • Hand and face landmarkers, for fingers and expressions.
  • A voice: speech in, a grounded answer out, lip-synced on the avatar.
  • Keyframe in-betweening: set three poses and let a motion model fill in the rest.

Built with

  • Meshy image-to-3D, rigging and animation Paid API, commercial rights
  • MediaPipe Pose Landmarker Apache-2.0
  • Grounding DINO + SAM 2.1 Apache-2.0
  • three.js r186 MIT
Next experiment
Reflective car capture in a quarter of the time →
Want this for your data?