Training vision-language-action (VLA) models for robotics requires massive amounts of data, but current methods face a scalability crisis. Teleoperation, while precise, is too slow and expensive to scale. This article explores why web videos present a viable alternative, offering diverse and abundant data. It discusses the technical challenges of using web data, such as domain gap and quality control, and potential solutions. For researchers and engineers in embodied AI, understanding this shift is essential for building the next generation of robots.
Explore the data bottleneck in VLA models and why web videos are the key to scaling robotics AI.