Robotics: The 'Next Frontier' Stuck for 70 Years
Deepak Pathak notes that the excitement surrounding current robotics progress is misplaced, challenging audiences to differentiate between robotic demonstrations from 30 years ago and today.
Deepak Pathak, co-founder and CEO of Skild AI, argues that the robotics field has been stuck due to a hardware-centric approach and a lack of scalable data.
Deepak Pathak notes that the excitement surrounding current robotics progress is misplaced, challenging audiences to differentiate between robotic demonstrations from 30 years ago and today.
The 'MIT copy demo' from the 1960s, predating modern computers, featured a robot arranging blocks from an image, a task mechanically similar to those performed today.
Teleoperation systems, which are still widely used for data collection in labs, trace their origins back to a 1957 Nuclear Congress exhibit, highlighting the lack of evolution in basic robotic control.
Historically, robotics focused on hardware rather than the creation of a general brain, a prioritization Deepak Pathak argues exposes the Moravec paradox and reaffirms robotics as the core intent of AI.
This aligns with the 'Moravec paradox,' where tasks humans find difficult (e.g., complex calculations) are easy for computers, but simple physical tasks (e.g., climbing stairs) are extremely hard.
Deepak Pathak argues that robotics is not merely an application of AI, but the original intent of AI, necessitating a fundamental, first-principles approach.
Robotics lacks a massive, internet-scale dataset comparable to those that propelled large language models like GPT-3.
Manual data collection through teleoperation, a 68 to 70-year-old method, is inherently slow and expensive, hindering progress.
Deepak Pathak calculates that even if the entire US population were to collect data via teleoperation, at one minute per example, it would take over a century to reach the token scale of GPT-3.
This approach aims for greater generality than human biology, which is limited to controlling specific bodies.
A unified brain is crucial because robot data scarcity necessitates learning from diverse hardware, tasks, and scenarios to achieve true scale and robust intelligence.
Curiosity-driven exploration, where robots learn through play, is limited by physical-world constraints and struggles to scale effectively without human input.
Teleoperation, while yielding high-quality data, is extremely expensive and lacks diversity due to fixed setups and the high cost of human operators.
Simulation offers scalability but demands tedious, manual engineering of every scene by humans, restricting environmental diversity, while video learning provides diversity but poor data quality for motor control.
The ideal data combines all three, enabling robots to generalize across tasks and environments efficiently and robustly.
This dual approach mirrors the successful training paradigms employed in large language models, with deployment data expected to become the most significant category as systems scale.
Omni-bodied brains are crucial because robotics cannot rely on a single hardware version indefinitely, requiring adaptability across diverse form factors.
This hardware diversity contrasts with traditional chip manufacturing, preventing the robotics market from being locked into a single form factor.
Deepak Pathak points out that tasks like folding laundry are often overestimated in difficulty, as humans perform them with high tolerance for error, not requiring millimetric precision.
Deep learning-based robotics simplifies fabric manipulation compared to classical physics-based models, while precision tasks, such as correctly inserting an AirPod into its case, demand higher intelligence due to strict orientation requirements.
The system trains robots using egocentric human videos and aims to expand to third-person perspectives, which could enable learning from open-source data like YouTube.
This approach requires less than one hour of robot-specific data for successful skill transfer to humanoid forms, highlighting that intelligence, rather than hardware, is the primary hurdle in robotics.
A demonstration featured a low-cost $4,000 hardware setup successfully preparing omelettes, operating solely vision-based without complex force sensors or state machines.
The robot exhibited robustness to unseen objects and changing environments, having been trained on less than 10 hours of data, showcasing its ability to generalize.
The only intermediate step is a standard PID controller to adjust the signal frequency, enabling tasks that prioritize environmental understanding over discrete movement planning and allowing long-term operation without manual resets.
Stunt-like behaviors such as dancing or backflips are easier for robots because they only require self-awareness of their own body, without needing to perceive or interact with the environment.
Conversely, climbing stairs demands vision, height estimation, and constant adjustment to environmental disturbances, highlighting the Moravec paradox's relevance to humanoid robotics and criticizing the industry's focus on flashy 'checkbox' demos over practical utility.
It operates entirely from a torso-mounted camera, demonstrating robustness to external disturbances, such as being pulled while moving, and utilizes the same core model across diverse scenarios.
While complex parkour movements require extensive training, simple robot tasks, such as the highlighted demo, can be trained in as little as half an hour.
Deepak Pathak emphasizes the urgency of starting deployments now to gather crucial data for the flywheel, citing Nvidia factory automation as a key case study for real-world integration.
The robot system handles the randomization and noise prevalent in real factory floors, demonstrating high robustness to disturbances.
This level of robustness allows robots to operate safely near humans without the need for traditional safety cages, with the same model applicable to both precise factory tasks and package delivery scenarios.
Answers come from the transcript, with the exact spot cited.
Want the next article from AI Engineer?
When AI Engineer publishes, we'll write it up like the one you just read and email it to you.
AI Engineer published 82 in the last 7 days.
Akamai Functions Achieve Zero Cold Starts for AIAI Engineer23 hours ago · 22:37 · 205 views · Created 21 hours ago
From Laptop to Pipeline: Scaling AI AgentsAI Engineeryesterday · 19:39 · 69 views · Created yesterday
Stop Prompting AI: Codify Rules InsteadAI Engineer2 days ago · 15:55 · 77 views · Created 2 days ago