Generalist AI Robots Learn New Tasks From Short Videos, Reaching 59% Success
Updated
Updated · WIRED · Aug 19
Generalist AI Robots Learn New Tasks From Short Videos, Reaching 59% Success
2 articles · Updated · WIRED · Aug 19
Summary
Generalist AI showed robot arms learning chores such as stacking cups, sweeping blocks and unzipping purses after watching short instructional videos, then adapting on the fly when objects or conditions changed.
One robot used a dustpan as a substitute brush when the brush was removed, and a two-armed system switched grippers to pull cash from a different purse—behaviors engineers said had not been specifically trained.
The startup says its edge comes from large-scale physical training data gathered with camera-equipped glove grippers, aiming to teach transferable physics-based skills rather than task-by-task routines that often need thousands of examples.
Generalist built its models from scratch, and outside roboticists at Georgia Tech and Stanford said the company appears unusually close to commercial deployment because its data strategy is not tightly tied to one robot.
Reliability remains the main hurdle: the robots complete shown tasks only about 59% of the time on average, well short of the 99%+ level likely needed for broad real-world use.
Will Generalist AI's human-data strategy beat out massive synthetic simulation engines in the high-stakes race to deploy commercial general-purpose robots?
If Generalist AI robots fail 41% of the time on familiar tasks, how can they safely navigate the unpredictable chaos of a real home?
Can a machine truly achieve physical AGI from human videos alone, or is the missing link active, multi-sensory exploration of the real world?