Skild AI has launched S1, a robotic foundation model that learns previously unseen, long-horizon tasks from a single video demonstration without retraining its weights. The model, announced last week, uses in-context learning to interpret a video prompt and map the demonstrated intent, objects, and sequence into actions for the robot in front of it. Skild built S1 and conducted the research on NVIDIA AI infrastructure, part of a collaboration spanning synthetic data generation, model training, simulation, and real-world physical AI deployment.

The company says S1 can perform unfamiliar tasks lasting up to 10 minutes, including plant potting, pancake making, pour-over coffee brewing, and kit assembly. In one plant-potting test, the team moved from recording the demonstration to autonomous execution on hardware in 11 minutes. Skild reports that S1 succeeded about 66% of the time at each step on new multistep tasks, compared with 9% for a similar AI system — a more than sevenfold improvement. The company also estimates that one short video example can be as useful as roughly 380 hands-on training examples, which would take a person 50–100 hours to collect manually.

Confirmed

  • S1 is a robotic foundation model built for in-context learning from a single video demonstration.
  • No weight updates or task-specific post-training are required at inference time.
  • Tasks demonstrated: plant potting, pancake making, pour-over coffee brewing, kit assembly — up to 10 minutes long.
  • Plant-potting demo-to-execution time: 11 minutes.
  • Step success rate on new multistep tasks: ~66% for S1 vs ~9% for a comparable system (per Skild).
  • One video example ≈ 380 hands-on training examples (per Skild estimate).
  • Trained on NVIDIA AI infrastructure using Isaac Lab, Isaac Sim, Omniverse, Cosmos, and Newton physics engine.
  • Skild, NVIDIA, and Foxconn are deploying Skild Brain on dual-arm manipulators for high-precision assembly of NVIDIA Blackwell systems.
  • Skild reached a $100 million annual revenue run rate 10 months after first commercial deployment.
  • More than 60 deployment partnerships across manufacturing, logistics, inspection, security, food preparation, and other applications.

Unknown

  • Independent replication of the 66% step success rate and the 380-example equivalence claim.
  • Generalization beyond the demonstrated task categories and lab conditions.
  • Pricing, licensing, and availability of S1 for customers outside Skild's existing deployment partnerships.
  • Hardware requirements for inference and whether S1 runs on Jetson Thor/Orin or requires data-center GPUs.
  • Long-term sim-to-real gap reduction from the jointly developed GPU-accelerated simulation solvers.
  • Whether the Skild Brain deployed with Foxconn is the same S1 model or a fine-tuned variant.

Our take

S1 shows in-context learning crossing into robotics with measurable results on long-horizon, out-of-distribution tasks. The real signal is the deployment flywheel: 60+ partnerships and a Foxconn line assembling Blackwell systems mean the model is already seeing production contact-rich data that will compound its advantage. The open question is whether the video-prompt interface scales to unstructured environments without a fine-tuning safety net, and whether inference cost per robot-hour stays viable for the SMBs that Skild's OEM partners target.

Sources