- Published on
Skild AI’s S1 Learns Soccer Through More Than 140 Years of Simulated Self-Play
One tap, and our reporting appears more often in your Google Top Stories. It makes a real difference to a small newsroom.

Supported by RoboStrategySupport for Humanoids Daily comes from RoboStrategy
SponsoredInvesting involves risk, including possible loss of principal. Read the prospectus before investing.

- Skild AI says S1 accumulated more than 140 years of simulated self-play in NVIDIA Isaac Sim, competing against recent versions of itself.
- Before self-play, S1 learned dribbling and kicking drills, each with its own reward and a human reference. The subsequent self-play stage used goal scoring as its objective without additional demonstrations.
- Skild reports that the resulting policy can play against humans and other robots in the real world, with early passing and coordination emerging in four-agent games.
- Detailed training and simulation-to-hardware transfer methods are still due in a later installment; the announcement does not establish improvements on commercial tasks.
Skild AI is putting its S1 robotics foundation model through a training exercise with a moving target: beat an increasingly capable version of itself at soccer.
In its published research announcement, Skild says the experiment accumulated more than 140 years of simulated play, developing skills including dribbling past defenders, shielding the ball, tackling and recovering from falls. The company also reports that the resulting policy can play against a human or another robot in the real world.
The preview builds on Skild's introduction of S1 in August, when the company demonstrated a model that could execute unfamiliar manipulation tasks from a single video prompt. The new work explores how that existing foundation of physical capabilities can be refined through repeated simulated competition.
The published account describes a preparatory stage before self-play: S1 first learns drills such as dribbling at different speeds and in different directions, and kicking. Each drill has its own reward and uses a human reference. Skild says the model outputs angles for each humanoid joint 50 times a second.
The weekly humanoid robotics briefing
One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.
Read recent issuesThen comes self-play, with the objective of scoring goals. S1 plays against recent versions of itself, combining and adapting its existing skills. As the model improves, the opposition becomes more capable too, giving successful strategies an increasingly difficult test.
Skild says the self-play stage produced increasingly complex behavior without separate rewards for those emerging tactics: goal scoring supplied the signal. That claim applies to self-play, while the preceding drill training used skill-specific rewards and human references.
According to NVIDIA, the policy trained in Isaac Sim, with NVIDIA Omniverse libraries used to create the virtual soccer training ground. The reported total of more than 140 years represents accumulated simulated experience; the materials do not specify the elapsed training time, compute budget or number of parallel matches.
The distinction between this experiment and S1's earlier announcement matters. In the original work, a video demonstration supplied the task at inference time, without task-specific updates to the model's weights. Here, reinforcement learning provides an additional training stage in which S1 can refine and combine its existing skills through experience.
The absence of demonstrations therefore applies specifically to the self-play stage. The full training sequence includes both S1's pretrained capabilities and the human-referenced soccer drills, so it should not be described as learning soccer without any human guidance.
Soccer provides a demanding setting for that approach. Staying upright, moving toward the ball, controlling it and responding to an opponent all have to work together. An opponent also changes the conditions continuously, creating a training challenge that evolves with the model.
There is precedent for using simulated soccer to develop robot skills. Google DeepMind's research published in 2024 trained miniature humanoids through reinforcement learning to play one-on-one soccer, including recovery from falls and tactical responses to opponents, and transferred those policies to physical robots. Skild's preview centers on applying self-play as a post-training strategy for its S1 foundation model.
The commercial context makes the research worth watching. Earlier this month, Skild reported crossing $100 million in annual recurring revenue, ten months after its first commercial deployment. That announcement emphasized the difficulty of moving from compelling demonstrations to robots that can reliably keep pace with customer operations.
Self-play could offer a way to build additional experience without collecting a fresh human demonstration for every behavior. Skild reports early passing and coordination in games involving four agents, and says preliminary results also show promise for social navigation. It proposes extending the approach to collaborative manipulation and other multi-agent tasks where an objective can be defined and a simulator built.
These early results broaden the announcement beyond individual soccer skills, but the company does not provide quantitative evaluations of the teamwork or social-navigation claims.
Skild co-founder and CEO Deepak Pathak framed the work in more ambitious terms in a post on X, likening the emergence of human intelligence to billions of years of physical self-play and arguing that robots would follow a similar path. “This method scales, and we will scale it,” he wrote.
Broader applications in factories, construction sites and homes remain a research direction. The preview does not establish that improvements in simulated soccer translate into better performance on commercial tasks, and it provides no quantitative benchmark showing how much self-play improves S1 over its starting point or alternative training methods.
Skild says a later installment will detail the self-play training recipe, transfer from simulation to a physical humanoid, and social behavior at larger team sizes. Those details will help establish how reliably the approach transfers and whether it can produce useful skills beyond the soccer field.
Updated September 23, 2026, with details from Skild’s published blog, including human-referenced drill training before self-play, reported real-world soccer play and early four-agent teamwork, and with CEO Deepak Pathak’s comments on scaling.
Share this article
Read next
- Published on
- Reading time
- 4 min read
The Human Scale: NVIDIA’s EgoScale Unlocks High-Dexterity Robotics via 20,000 Hours of Human Video
- Published on
- Reading time
- 3 min read
NVIDIA "Supersizes" Humanoid Control: SONIC Open-Sources Whole-Body Tracking at Scale
- Published on
- Reading time
- 4 min read
NVIDIA DreamZero explained: how its robot AI differs from a VLA
- Published on
- Reading time
- 4 min read
NVIDIA’s "DoorMan" Teaches Humanoids to Open Doors Faster Than Humans Can
- Published on
- Reading time
- 5 min read
From Representation to Reality: AGIBOT Unveils Genie Envisioner 2.0 as a Scalable “World Simulator”
- Published on
- Reading time
- 3 min read
MIT’s SoftMimic Framework Teaches Humanoids to Be Compliant, Not Just Capable
The weekly humanoid robotics briefing
One email a week: the launches, funding and research that mattered, with context from Humanoids Daily.
Read recent issues
















