- Published on
Figure AI Unveils Index: Crowdsourcing Real-World Human Video to Train Helix

- Figure AI has publicly launched Index, a consumer-facing app and data collection platform previously developed under the stealth initiative Project Go-Big.
- The platform has amassed 16 million video uploads across 108 countries, processing 30 minutes of footage per second from more than 44,000 weekly active users.
- Figure has distributed $15 million in payouts to contributors and committed over $1 billion to data collection and compute over the next 12 months.
- Submissions pass through a five-stage pipeline—filtering, fraud review, deduplication, rebalancing, and hierarchical annotation—to train the company’s Helix AI architecture.
Nearly a year after quietly initiating its effort to turn human video into robot behaviors under Project Go-Big, Figure AI has formally emerged from stealth with Index. The platform is designed to bypass the traditional data bottlenecks of physical robotics by crowdsourcing real-world physical interaction data directly from smartphone users across the globe.
According to Figure founder and CEO Brett Adcock, Index serves as a dedicated pipeline to feed the company’s Helix AI architecture with the rich, diverse physical interactions required for zero-shot generalization. Alongside the public rollout on iOS and Android, Figure revealed that it has already paid out $15 million to contributors and committed over $1 billion toward data acquisition and compute over the next 12 months.

Industrializing Human Telemetry
While language models scale on the vast corpus of the public internet, robotics has long struggled with a severe lack of embodied interaction data. Standard methods—such as direct robot teleoperation, kinesthetic teaching, or synthetic simulation—remain notoriously slow, labor-intensive, and difficult to scale across varied environments.
Figure’s answer is to crowdsource human physical labor. Index operates as a two-sided platform: individual "Creators" can record themselves executing everyday household or workplace tasks for cash bounties, or users can book gig workers through the app to complete chores on-site while capturing first-person footage.
During its four months in stealth, the platform recorded substantial traction:
- 264,000 app downloads across 108 countries.
- 44,000+ weekly active users contributing data.
- 16 million video uploads, translating to 30 minutes of physical video ingested every second (roughly 4.9 years of human work uploaded per day).
- High environmental diversity: Figure reports that for every 1,000 hours logged, the dataset captures 373 unique tasks, 1,146 distinct manipulated objects, and 116 unique environments.
Captured workflows span domestic chores—such as folding laundry, making beds, and cleaning—to commercial operations in restaurants, retail stockrooms, and logistics facilities.

The Five-Stage Ingestion Pipeline
Ingesting continuous consumer-generated video at high volume introduces major quality-control hurdles. To convert raw smartphone video into structured training data for the Helix model, Figure built an automated, five-stage ingestion pipeline:
- Filtering: Automated semantic and technical vision filters immediately screen uploads for resolution, frame rate, lighting conditions, and task relevance.
- Fraud Review: Dedicated human auditing teams monitor user accounts to detect deliberate attempts to spoof tasks or game payout incentives.
- Deduplication: High-dimensional vector embeddings are generated for each video segment to identify and discard redundant footage that falls above strict similarity thresholds.
- Rebalancing: The dataset is dynamically balanced using task quotas and embedding clusters to prevent over-representation of simple tasks while prioritizing rare edge cases.
- Hierarchical Annotation: Accepted episodes receive structured hierarchical text captions, aligning high-level task descriptions with low-level physical manipulations for multimodal model training.
Introducing Index Today we're coming out of stealth with Index, the largest & most diverse robot dataset in the world → 30min of video uploads/sec → 16M video uploads → Paid $15M to date → 264k downloads We're committed to spending $1B the next 12 months on data & compute
Bridging Physical AI and Commercial Fleets
The launch of Index marks a critical shift in Figure’s broader commercial strategy. The company has already begun manufacturing hardware at scale, recently celebrating its 1,000th Figure 03 build at its BotQ facility and deploying units into real-world commercial pilots with BMW at Plant Spartanburg and Catalyst Brands in Reno.
However, as Adcock recently argued, hardware scalability is no longer the primary hurdle—the true bottleneck is onboard intelligence and general-purpose reasoning.
By scaling Index into a multi-million-dollar data collection engine, Figure is attempting to build the equivalent of ImageNet for embodied AI. Whether human video alone can provide the tactile feedback and force dynamics needed to master intricate physical manipulation remains an ongoing research question, but Figure is putting substantial capital behind the bet.
Share this article
Stay Ahead in Humanoid Robotics
Get the latest developments, breakthroughs, and insights in humanoid robotics — delivered straight to your inbox.




