Does Scaling Web-Video Pre-training Help Real Robots Do Real Work?
We investigate scaling model size and pre-training compute for video-based robot policies, and benchmark on a real industrial manipulation task.
Read article
We investigate scaling model size and pre-training compute for video-based robot policies, and benchmark on a real industrial manipulation task.
Read articleOur team focuses on deploying robotic systems into the real world. We build general purpose foundation models that can adapt to the variability of commercial and industrial environments.
FutureVision brings the capability to handle real world industrial tasks autonomously. This unlocks generalization that conventional VLA pipelines struggle to achieve. We call this the Direct Video Action model. And it’s why Rhoda works not just in labs but in factories, warehouses, and real production environments.
We first pre-train our model with over a million videos, giving it a strong prior on motion, physics, and dynamics. We then post-train on action data collected on the robot to teach the model to learn specific tasks. This combination of pre- and post-training creates generalist robot policies that can work and adapt in dynamic environments.
We explore a new paradigm for scaling robot intelligence with web-scale video pre-training. We’ve published an in-depth research blog to discuss the novel architectural designs and the benefits of such an approach.
We work with a variety of customers across verticals in automotive, manufacturing, logistics, and ecommerce. If you’re interested in working with us at your facility, reach out here.
Strength: Custom actuators enable 25kg rated, 40kg peak
Safety: Wheel-base, brakes in every actuator, safety-rated vision
Reliability: 3 years of continuous operation at rated payload
AI Control: High component stiffness, linear response