Buyer's guide
Buying egocentric video data: 10 questions to ask a vendor
Robot foundation models, VLA models and world models are hungry for first-person video of real work. The market has filled with vendors, and most look alike on a landing page. These ten questions separate a dataset you can train on — and defend in front of legal — from one that becomes a liability.
Short answer
Ask for consent and provenance per clip, a real distribution across countries, scenes and trades, device metadata, a human QA gate that rejects footage before you pay for it, and pricing per approved hour. Then ask for a free sample pack carrying the same metadata a full delivery would.
1. Who owns the footage, and does the licence explicitly cover training?
The contributor agreement should assign the footage to the vendor and name AI training explicitly. Anything vaguer ("content may be used to improve services") will not survive diligence. You should receive a clear licence for training and evaluation.
2. Can every clip be traced to a person, a place and a payment?
Egocentric footage is filmed inside workplaces, often around coworkers. A defensible dataset lets the vendor produce, for any clip: who recorded it, where, when, on which device, under which agreement, and what they were paid. If the vendor cannot answer that for a random clip ID, assume nobody can.
3. Did the business agree in writing?
Contributor consent is not enough when the camera is inside a shop, garage or kitchen. Look for written permission from each business, with the owner able to stop recording at any time.
4. What is never filmed?
Minors, documents, screens, internal pricing and identifiable customers should be excluded by contract and enforced at review — not just mentioned in a training deck. Ask how face and object redaction is handled and whether it follows your rules.
5. How diverse is the distribution, really?
Many public egocentric datasets are dominated by two or three countries and clean, well-lit environments. Models trained that way break in cramped, cluttered, badly lit real workplaces. Ask for the split by country, scene type and domain, and whether you can set the mix. Different regions mean different tools, layouts and working practices — exactly the distribution shift that hurts deployment.
6. Is it real work or staged work?
Scripted demonstrations by paid actors look tidy and generalise poorly. Unscripted footage of people doing their actual shift — trimmed, not graded or stabilised — carries the recoveries, hesitations and mistakes your policy needs to learn from. Ask how the vendor detects repetitive, purposeless or artificially slow "task padding".
7. What metadata ships with each clip?
At minimum: domain and task from a fixed taxonomy, scene type, country, device model and OS, duration, timestamps, session ID and a quality rating. The exact device model matters if you want to filter or stratify by sensor.
8. What does the QA gate reject — and do you pay for rejects?
A serious pipeline rates every session before it counts and never charges for unusable footage. Ask for rejection rates and the top rejection reasons. (In Dovra's network the dominant reason is clarity: light and phone position.)
9. How fast can they stand up a new domain or country?
Standing capacity matters more than a library you have already seen. Ask how long from brief to first clips in a covered domain (days, not weeks, is realistic) and in a new one, and how they scale — by adding cities, or by squeezing the same contributors.
10. How is it priced and delivered?
Per approved hour is the cleanest unit. Annotation (step segmentation, hand-object interaction, your taxonomy) should be priced separately. Delivery should be an encrypted transfer to your bucket with a manifest and dataset card, in formats that match your ingestion.
How Dovra answers these
- Human-layer capture only: first-person video of real work, head-mounted, on the contributor's own phone, device model logged per clip.
- Signed agreement per contributor, written permission per business, payment record per approved hour, training licence and provenance doc.
- Standing capacity of about 5,000 hours per 4-week cycle across 32 countries, with the deepest bench in Latin America.
- 15 work domains and 84 catalogued task types; new domains in two to three weeks.
- Priced per approved hour; free sample pack with the same metadata and consent record as a full delivery.
Dovra is backed by Simera, a remote talent company for Latin America. Read why Simera backs Dovra.
See it before you commit
Tell us the domains you are training on and roughly how many hours you need. We reply within one business day with a sample pack and a price.
Request a sample pack See dataset coverageFAQ
What is egocentric video data?
First-person video recorded from a head-mounted camera at eye level, showing what a person sees and does with their hands while performing a real task. It is used to train robot foundation models, vision-language-action (VLA) models, world models and video generation.
How is egocentric video data priced?
Usually per delivered or approved hour. Price depends on the domain, how hard contributors are to recruit, exclusivity, and whether annotation is included. Dovra prices per approved hour.
What documentation should come with a licensed dataset?
A signed agreement per contributor, written permission from each business where filming happened, a payment record per approved hour, a training licence and a provenance document that ties every clip to who recorded it, where, when and on which device.