HumanCLAW: Can Vision-Language Models Act Through a Body?
Separating the decision from the motor execution that carries it out makes embodied failure attributable: the model chose wrong, or the body could not do it.
I’m a research scientist at Meta Reality Labs Research, where I work on building robotic foundation models and scaling up egocentric data for human-level dexterity. I obtained my Ph.D. in Computer Engineering from University of Wisconsin-Madison in 2023, advised by Prof. Umit Y. Ogras and I also worked closely with Prof. Yin Li. Prior to that, I received my B.S. from UESTC in 2018.
Separating the decision from the motor execution that carries it out makes embodied failure attributable: the model chose wrong, or the body could not do it.
Dexterous manipulation needs contact-rich data from real environments, not a capture dome — this is the rig and the pipeline that produce it.
Streaming is what separates a video model from a world model an agent can act inside — it has to keep up with the agent, in real time.
One pretrained motion prior, prompted rather than retrained, covers behaviors and constraints that normally need a model each.
Body morphology changes how a motion actually looks — generating shape and motion together keeps the two from contradicting each other.
Continuous tokens avoid the quantization jitter of discrete motion codes while leaving the pretrained language model intact.
Knowing whose body it is turns fitting from a guess about shape into a pose problem, which is where the accuracy comes from.
The first 3D-aware GAN to synthesize a full head rather than a frontal face — back of the skull and hair included.
Follow-up to PanoHead: the geometry representation, not the generator, was what capped full-head quality.
mRI: Multi-modal 3D Human Pose Estimation Dataset using mmWave, RGB-D, and Inertial Sensors
PAniC-3D: Stylized Single-view 3D Reconstruction from Portraits of Anime Characters