Robot in-context learning from 30-second video clip has recently been achieved in August 2026. e.g. https://www.youtube.com/watch?v=hr39FlEiCcQ Paired with how much progress has been made thus far in video generation, e.g. "A humanoid robot uses its grippers to clear 2 empty cups in a coffeeshop": https://imgur.com/ziQg165 Is it likely that text-to-action (or voice-to-action) robotic moment is near just around the corner (1 or 2 years away)?

Full article content could not be extracted automatically. Read the original below.