Updated
Updated · MIT Technology Review · Oct 8
Physical Intelligence's π0.7 Shows First Compositional Generalization on 1 Unseen Robot Task
Updated
Updated · MIT Technology Review · Oct 8

Physical Intelligence's π0.7 Shows First Compositional Generalization on 1 Unseen Robot Task

1 articles · Updated · MIT Technology Review · Oct 8

Summary

  • April 2026 tests showed Physical Intelligence’s π0.7 making a passable attempt at an unseen command—loading a sweet potato into an air fryer—marking what the company says are its first signs of compositional generalization.
  • π0.7 pairs a lightweight world model with prior robot training so it can generate intermediate visual steps, letting the machine recombine learned skills instead of only repeating tasks explicitly seen in training.
  • Researchers later found just two air-fryer-related teleoperation examples in the training data, leaving open how far the model truly generalized beyond sparse prior exposure.
  • That result stands out because today’s vision-language-action robots still usually fail outside their training sets, and even 70% task success is widely viewed as too unreliable for real-world deployment.
  • The advance lands amid a humanoid-robot boom—Morgan Stanley sees nearly 1 billion humanoids by 2050—but many researchers argue useful home or factory generalists remain years away.

Insights

Will hidden human operators continue pulling the strings, or can world models finally grant humanoid robots true autonomy?
Are tech leaders selling an illusion of autonomous humanoids while relying on teleoperation to hide fatal flaws in real-world environments?
If physical data is the ultimate bottleneck, which tech giant will crack the simulation code to bring humanoids into our homes?