Robots have never learned to use their hands
Three months ago, EgoVista was an idea in my head. Today we are a team of three, we have delivered six proofs of concept, and contributors are recording with us across 32 countries with a strong focus in Europe.
This is what we are building, and why.

The bottleneck is not the model
Robotics has spent the last two years absorbing the lessons of large language models. Vision-Language-Action architectures have arrived, transfer learning works, and the gap between a research demo and a deployable manipulation policy is narrowing faster than most people expected.
What has not scaled is the data.
A language model can be trained on text that humanity has already written. There is no equivalent corpus for manipulation. Nobody has been quietly recording, for the last thirty years, how a human hand approaches a door handle, adjusts grip mid-motion when a glass turns out to be heavier than expected, or reorients a screwdriver without looking at it. That data does not exist as a byproduct of anything. It has to be produced deliberately.
Teleoperation produces it, at a cost per hour that makes scale painful. Simulation produces something adjacent to it, and the gap between synthetic contact dynamics and real ones is precisely where manipulation policies tend to fail. Web video contains fragments of it, buried in footage that was framed for an audience rather than for a robot, from a camera position no robot will ever occupy.
Egocentric video is the closest thing to a native format for this problem. First-person, hands in frame, the same viewpoint the policy will eventually have to act from. It is also the format that is hardest to collect at scale legally, because it is recorded in people's homes and workplaces, by people, about people.
That difficulty is the reason we exist.
What we do
We collect and annotate egocentric video of human manipulation tasks, and we turn it into training-ready datasets.
Collection runs through a distributed contributor network. Contributors record ordinary manual tasks in real environments using a phone or an action camera with a simple head mount. Not staged demonstrations in a lab, not actors performing a script, but the actual variability of how a task gets done when nobody is optimising for the camera.
Annotation is where the raw footage becomes usable. Video alone does not train a manipulation policy. The pipeline produces layered annotation across pose and hand tracking, depth, hand-object segmentation, contact timing, and action labels, and exports to the formats labs are already working in, LeRobot and RLDS compatible, along with HuggingFace Datasets.
We work three ways. Annotation on video a client already holds. Custom collection to a specification, where the environments, tasks and capture protocol are defined by the buyer. And raw footage with faces blurred, for teams who prefer to run their own annotation.
Three bets
Europe first, then wider. We built the collection network in Europe deliberately, and we are extending it worldwide, because manipulation policies that only ever see one region's kitchens, tools and workflows will fail in every other region's. Diversity in this data is not a values statement, it is a generalisation requirement. What stays constant is where the data is processed and under which framework.
Annotation that goes deep. The market for raw footage is commoditising, and it will keep commoditising. The value sits in what makes footage trainable, which is precise multi-layer annotation with real QA on every delivery. This is where most of our engineering effort goes, and it is the part of the roadmap we are still building out.
Compliance as product, not paperwork. Faces are blurred before any external processing. Data is processed in the EU. Every dataset ships with a provenance record: where the footage came from, under what legal basis, with what consent, documented per contributor. Contributors are adults, they consent explicitly and separately, and they can withdraw.
This last point is usually treated as a cost centre. We think it is the product. If you are training a model that will end up in a commercial system, the AI Act makes you a deployer with obligations about the provenance of your training data, and the honest position nobody in this market states plainly is that there is no zero-risk dataset. What there is, is a documented and defensible one. That is what a general counsel actually needs. Not a promise that risk is zero, but a chain of title they can show someone.
Where we are
Early, and honest about it. Six proofs of concept delivered. A contributor network across 32 countries as of July 2026, weighted toward Europe. Three founders: myself on the commercial and legal side, and two technical co-founders, an engineer and a robotics researcher.
What we are building next is more capacity, faster processing, deeper annotation layers, and a second vertical in industrial environments where manual work is the whole job.
If you are building manipulation policies
We would like to hear what data is actually blocking you. Not to pitch, at this stage the useful conversation is the other direction: which tasks, which environments, which capture setup, which annotation layers your training pipeline needs and is not getting.
If that is a conversation worth having, get in touch at contact@egovista.app.
If you would like to get paid to record yourself doing everyday manual tasks, we are onboarding contributors. Fully remote, open to adults anywhere, a few minutes to join at egovista.app/signup.
