Axis Robotics Open Sources One of the largest Franka Arm simulation datasets for physical AI

Axis Robotics Open Sources One of the largest Franka Arm simulation datasets for physical AI

Axis Robotics released Sim Axis Dataset V1one of the largest open source simulation datasets for Franca arm processing, with the complete dataset, training code and benchmarks publicly available. The V1 was created from over 50,000 simulated human-controlled trajectories across 207 manipulation tasks and over 60,000+ different scenes on the Franka Research 3 simulated arm.

This dataset has attracted over 160,000 downloads, making it the most downloaded open source Franca processing dataset on Hugging Face. In benchmarks, continuous pretraining on V1 raised π0.5 and exceeded the magnitude-matched RoboCasa baseline, with each outcome open and verifiable.

Axis Robotics is building the ultimate composite data engine for physical AI, a vertically integrated system that spans large-scale simulation, egocentric real-world capture, human positional manipulation, and subsequent human-enclosed DAgger training. The company has raised $12 million in seed funding led by Hack VC, with participation from Nomad Capital, Pi Network Ventures, 10K Ventures, and angel investors.

Bet against “clean data only”

It’s a common assumption in robotics that demos have to be near-perfect to begin with — filter down to expert paths, standardize setup, and eliminate anything noisy before it’s safe to replicate. Access’s thesis goes in the other direction: data quality lives at the distribution level, not the individual path. When a sufficiently large and diverse crowd produces suboptimal noisy paths and their errors are uncorrelated, the average noise policy continues to work during training.

Axis Sim Dataset V1 puts this thesis to the public test. Its paths span picking and placing, stacking, pouring, manipulating articulated objects, and using tools, all brought together through Axis’ browser-based teleoperation platform, Axis Hub, by a distributed crowd rather than a single team of experts. The dataset was created in collaboration with researchers from UC Berkeley, Johns Hopkins, the University of Michigan, and other institutions.

The results are this big

In LIBERO-Plus, continuous pretraining on V1 raises π0.5 from 83.9% to 88.8% success and outperforms the size-matched RoboCasa365 baseline by 37.3%. Performance continually improves as pre-training data increases from 25% to 100% of the dataset, with no saturation in sight, evidence that gains come from diversity and coverage rather than from a one-time increase. The biggest improvements appear under camera, sensor noise, and layout disturbances, which are the exact axes that the gimbal randomly distributes during construction.

The team says V2 is already underway, scaling to 1.2 million paths across 1,200 tasks, with cross-incarnation generalization and results across multiple VLA models showing that suboptimal simulation data trains robust policies.

The engine behind the data set

A dataset is one output of a larger, more complex data engine. Where a traditional data vendor collects data according to fixed specifications and stops, Axis uses model performance and failures to determine what to collect next, so each training round informs the next. This engine operates on a hybrid strategy across four data pipelines, all of which now operate at scale:

  • simulation: Over 200,000 contributors distributed on Axis Hub, a top 3 dApp on Base, generating over 4.7 million pipelines across 13 models.
  • selfish: Managed network of over 1,000 full-time collectors and QC-trained, capturing first-person activity in real homes and businesses across 14 industries: Over 200,000 watches already stored and growing by over 4,000 watches every day, with hand placement verified by Vicon.
  • Positional manipulation: Over 500 hours of combined mobility and dexterity on real humanoid robots (Unitree G1 and Booster T2) through hardware-independent remote operation.
  • Human DAgger gates after training: Over 500 hours of human-directed debugging for advanced deployment situations.

Every task and path on the chain is logged in Base to identify its provenance, and contributors are rewarded for verified quality of work.

From open data to commercial publishing

Beyond open source simulation data, Axis works directly with… Robot avatar companies To build custom and specific data pipelines for rendering and previous models.

As Booster Robotics’ first data sim partner, Axis has rebuilt Booster’s real workspace as a mission-compliant digital twin, and distributed contributors have collected over 42,000 simulation loops on it, distilling them into a previously Booster-specific model. With only 30 real robot demos, this previously achieved a success rate of 87.5% versus 37.5% for the π0.5 off-the-shelf robot, which matches the π0.5 using half the real-world demos.

Other partners extend Embodiment companies (Vision Robots), Model companies (Manicure Tech, Dexmal) and Industrial automation (Lotus Cars, Geely Auto). The axis is also supplied On-chain botnets: BitRobot on Solana and OpenRoboto on Bittensor.

Redefining the physical AI data basis

“The future of physical AI is not a static data set that you download once,” said Chris Feng, founder of Axis Robotics. “It’s an engine that keeps producing the data that the model then needs. Range gives you broad coverage. Diversity keeps the noise unbiased. Closed-loop turns every failure into progress. That’s what compounds.”

Axis was founded by researchers from UC Berkeley, CMU, Georgia Tech, and SJTU, along with series founders who have scaled consumer platforms to over 30 million users. Her research is supervised by Jiachen Li, an assistant professor at Georgia Tech.

Paper link: https://arxiv.org/abs/2607.21588

Project page: https://axisaiorg.github.io/AXIS-V1/

Dataset link: https://huggingface.co/datasets/axisrobotics/Franka-Dataset

Github database: https://github.com/AxisAIOrg/Axis-V1-Training

Leave a Reply

Your email address will not be published. Required fields are marked *