From Algorithm to Data: The Evolving Challenge in AI Development
This paradigm shift underscores that data quality and relevance are paramount for AI to function effectively and for continued innovation in the field.
A Bright Data executive explains why collecting the right real-world video data is crucial for advancing physical AI.
This paradigm shift underscores that data quality and relevance are paramount for AI to function effectively and for continued innovation in the field.
In 2022, Google initiated the training of robots using real-world actions and images, marking a significant step in the field. A year later, in 2023, there was a leap in capabilities, with AI systems demonstrating control over multiple robots simultaneously.
The landscape for robotics training dramatically changed in 2024 with the emergence of open-source frameworks, democratizing access for various laboratories. Waymo's autonomous driving system serves as a prime example of data-driven learning, heavily relying on dashcam footage for its development.
While large language models (LLMs) are trained on trillions of words and image generation models benefit from billions of labeled images, robotics faces a significant data deficit. The field is currently limited to approximately one million videos, a dataset considered extremely small and restrictive.
Traditional methods of recording "instructed" actions further compound this issue by producing unnatural and biased data, hindering robots from learning real-world behaviors effectively.
When individuals are instructed to perform actions for a recording, their movements and behaviors are often unnatural, diverging from spontaneous, real-life interactions. Rafael Levi cited the example of opening a door: a staged recording of the action differs significantly from someone naturally entering their home.
This "instructed" data lacks the inherent intuition and natural flow of real activity, a problem exacerbated by camera shyness. Consequently, essential elements like natural physics, gravity, and authentic cause-and-effect relationships are missing, leading to ineffective training for robots.
Existing data sourcing strategies, such as simulations and human-controlled robotics (teleoperation), present significant limitations. While virtual simulations are cost-effective, their physics models, though improving, are not yet robust enough to accurately train robots for real-world scenarios.
Human-controlled methods are labor-intensive, offering limited daily recording hours and requiring extensive human input, which makes them unsuitable for generating the vast datasets needed for scalable robotic AI development. Furthermore, readily available pre-built datasets are often insufficient in scale and diversity.
The web presents a compelling alternative to the restrictive and manually recorded datasets currently used in robotics. Platforms like YouTube host billions of videos containing countless hours of real-world interactions and actions.
This abundance of organic content offers a unique opportunity to capture natural behaviors that robots can learn to replicate, providing a scalable and authentic source of training data.
Web videos offer an immense resource for AI training, naturally capturing fundamental physics such as gravity, motion, and intricate cause-and-effect relationships, including accidents and object manipulation. These billions of hours of footage represent an unparalleled training material for robots.
Meta demonstrated this potential by training an AI model on approximately one million hours of public video data. This model, with only 62 hours of specific robotics data, successfully controlled a real robot without needing simulations, highlighting the efficiency of leveraging web content.
Advanced AI models can analyze video content frame by frame, measuring subtle differences in movement to calculate precise angles, distances, and trajectories. This capability allows robots to learn complex physical interactions through observation, eliminating the need for raw sensor data.
While video data offers a rich source, post-processing remains essential to convert this observational information into a format that is 100% ready for robot ingestion and action, ensuring accuracy and utility.
Robotics training currently suffers from extreme inefficiency, as evidenced by Nvidia's Cosmos project, which discards 96% of the video data it downloads. Similarly, Stable Video Diffusion discards 74% of the videos it downloads for training purposes.
These high discard rates result in substantial waste of computational resources, storage capacity, and financial investment, indicating a critical need for more targeted and efficient data collection strategies.
This method allows users to query for precise activities, such as "person washing dishes," and receive pre-trimmed video snippets.
This targeted indexing significantly reduces data waste, optimizing collection, storage, and bandwidth usage. By expanding its index to 1.1 billion videos, the platform aims to provide ready-to-use clips for immediate robot ingestion.
The system's ability to pinpoint exact moments within videos minimizes the need to download and process irrelevant footage, streamlining the data pipeline for AI training.
The system allows for highly precise video searches using detailed descriptions, which significantly increases the accuracy and relevance of results. Users can specify complex activities like "human folding clothes" or "person putting on makeup" to retrieve exact matches.
This granular search capability supports automated triggers via APIs, removing the need for manual intervention in data collection. Instead of entire video files, users receive direct snippets, URLs, and timestamps, ensuring maximum efficiency.
Furthermore, the technology can assist brands in tracking specific product demonstrations within videos, even if the video's title is unrelated to the product, providing valuable insights and targeted content.
The system provides users with precise metrics, including specific timestamps, matching scores indicating query proximity, and frame counts for retrieved video segments. This detailed output enhances the utility of the data for AI training and various applications.
By delivering highly relevant data and filtering out noise, the retrieval process significantly reduces the occurrence of "hallucinations" in AI models. Additionally, brands can effectively track specific product demonstrations across diverse video content, irrespective of video titles.
The current reliance on human-recorded data, often involving millions of people recording mundane actions like opening doors, is highly inefficient. Utilizing existing global video databases online presents a far more effective alternative for training these sophisticated AI systems.
Autonomous driving systems, such as Waymo, have already demonstrated the power of leveraging millions of hours of dashcam footage for training. The internet, particularly platforms like YouTube, contains a wealth of public videos capturing diverse real-world behaviors, from traffic light interactions to various turns.
This existing web content represents a massive, largely underutilized resource for developing sophisticated 'world models' in AI. However, current search limitations make it challenging to effectively harvest specific, actionable data from these vast online repositories.
Many firms still resort to paying individuals to record specific actions, like opening doors, which invariably yields unnatural and biased results. This method is increasingly inefficient given the vast amount of authentic real-world data available online.
Physical AI requires genuine physics, gravity, and natural movement, which are best captured in real-world scenarios rather than staged environments. Processing billions of hours of available web video content proves to be a more effective and scalable solution than organizing individual tasks.
Bright Data is actively developing a tool to address this by indexing video content based on specific actions, moving beyond simple keyword tags to unlock the full potential of online video as a training resource.
Beyond its role in AI training, this video indexing technology enables users to isolate specific actions, such as a gamer searching for a precise solution to a level. For example, a gamer could search for videos showing how to beat a particular level in a game, even if no exact video exists.
This versatility means the technology can cater to diverse user needs, from detailed research to practical problem-solving in everyday scenarios, making the world's vast video content more accessible and actionable for everyone.
Answers come from the transcript, with the exact spot cited.
Want the next article from AI Engineer?
When AI Engineer publishes, we'll write it up like the one you just read and email it to you.
AI Engineer published 82 in the last 7 days.
Akamai Functions Achieve Zero Cold Starts for AIAI Engineer23 hours ago · 22:37 · 205 views · Created 21 hours ago
From Laptop to Pipeline: Scaling AI AgentsAI Engineeryesterday · 19:39 · 69 views · Created yesterday
Stop Prompting AI: Codify Rules InsteadAI Engineer2 days ago · 15:55 · 77 views · Created 2 days ago