Axis Robotics has officially released the Axis Sim Dataset V1, marking a significant milestone in the development of open-source resources for robotic manipulation. As one of the largest simulation datasets ever created for the Franka robotic arm, V1 provides researchers and developers with a comprehensive suite that includes the full dataset, training code, and rigorous benchmarks, all made publicly available. Built using a simulated Franka Research 3 arm, the initial version of the dataset comprises more than 50,000 human-teleoperated simulation trajectories spanning 207 distinct manipulation tasks and over 60,000 unique scene variants.
The release has already captured substantial attention within the global robotics and artificial intelligence communities. Accumulating over 160,000 downloads, Axis Sim Dataset V1 has quickly become the most downloaded open-source simulation Franka manipulation dataset on Hugging Face. Benchmark evaluations demonstrate that continual pretraining on the V1 dataset significantly boosts performance on foundational models like $pi_0.5$, outperforming volume-matched RoboCasa baselines while maintaining fully open and verifiable results.
Backed by a $12 million seed funding round led by Hack VC—with participation from Nomad Capital, Pi Network Ventures, 10K Ventures, and a network of angel investors—Axis Robotics is positioning itself as a pioneer in the Physical AI landscape. The company is actively building a vertically integrated, compounding data engine designed to tackle the industry’s most persistent bottlenecks. This infrastructure spans large-scale simulation, egocentric real-world capture, humanoid loco-manipulation, and human-gated DAgger post-training, creating a continuous feedback loop that accelerates autonomous capability development.
A Bet Against "Clean Data Only"
For years, a prevailing assumption across the robotics industry has been that training data must consist of near-optimal demonstrations to be effective. Conventional pipelines typically filter down to expert trajectories, heavily standardize experimental setups, and discard any noisy or suboptimal data out of concern that it might compromise the safety and reliability of imitation learning policies.
Axis Robotics operates on a fundamentally different thesis: true data quality lives at the distribution level rather than within any single isolated trajectory. According to the company’s research approach, when a large and diverse enough crowd generates noisy, suboptimal trajectories, and their underlying errors remain uncorrelated, the noise naturally averages out during training. This allows a robust and functional policy to survive and excel.
Axis Sim Dataset V1 was designed to test this hypothesis in public view. The trajectories encompass a wide variety of essential robotic skills, including pick-and-place operations, object stacking, liquid pouring, articulated-object manipulation, and complex tool use. Rather than relying on a small, isolated team of expert roboticists, these demonstrations were collected through Axis’s browser-based teleoperation platform, Axis Hub, powered by a distributed crowd. The dataset itself was developed in collaboration with researchers from prestigious academic institutions, including the University of California, Berkeley, Johns Hopkins University, and the University of Michigan.

Results That Scale
The empirical results derived from the V1 dataset highlight the power of data diversity and scale. On the LIBERO-Plus benchmark, continual pretraining utilizing the V1 dataset successfully lifts the performance of $pi_0.5$ from an 83.9% success rate up to 88.8%. Furthermore, it outperforms a volume-matched RoboCasa365 baseline by an impressive 37.3%.
Experimental evaluations show that performance improves consistently as pretraining data scales from 25% up to 100% of the complete dataset, with no visible saturation point. This steady upward trajectory serves as strong evidence that the performance gains stem directly from broad diversity and environmental coverage rather than a temporary, one-off statistical bump. Notably, the most pronounced performance improvements manifest under challenging conditions involving camera adjustments, sensor noise, and layout perturbations—the exact experimental axes that Axis deliberately randomizes during the data generation process.
Building on this momentum, the Axis Robotics team has confirmed that V2 is already actively underway. The upcoming release is slated to scale up significantly to 1.2 million trajectories distributed across 1,200 tasks. It will focus heavily on cross-embodiment generalization, with preliminary results across multiple Vision-Language-Action (VLA) models demonstrating that suboptimal simulation data can reliably train highly robust policies.
The Engine Behind the Dataset
The publicly available dataset represents just one visible output of a larger, continuously compounding data engine. While traditional data collection vendors typically gather information according to a fixed specification and subsequently halt operations, Axis utilizes real-time model performance metrics and specific failure cases to dynamically determine what data needs to be collected next. Consequently, every single round of training directly informs and shapes the subsequent data collection phase.
This operational engine relies on a hybrid strategy spanning four distinct data lines, all of which are currently running at commercial scale. To ensure absolute transparency and traceability, every individual task and trajectory is recorded on-chain using Base for provenance, ensuring that contributors are fairly rewarded for verified work quality and adherence to performance standards.
From Open Data to Commercial Deployment
Beyond open-sourcing foundational simulation data, Axis Robotics collaborates directly with commercial robot embodiment companies to engineer customized, embodiment-specific data pipelines and model priors tailored to unique hardware platforms.

As the first simulation-data partner for Booster Robotics, Axis successfully reconstructed Booster’s physical workspace as a task-aligned digital twin. Distributed contributors then collected over 42,000 simulation episodes within this virtual environment, which Axis subsequently distilled into a specialized model prior for Booster. When tested with a mere 30 real-world robot demonstrations, this customized prior achieved an 87.5% success rate, compared to just 37.5% for an out-of-the-box $pi_0.5$ model. This achievement effectively matched the performance of $pi_0.5$ while utilizing only half the required real-world demonstration data.
Axis’s partner ecosystem extends across multiple sectors of the robotics and automation industries. Collaborators include embodiment companies such as Feagine Robotics, artificial intelligence model developers like Manycore Tech and Dexmal, and major industrial automation leaders including Lotus Cars and Geely Auto. Additionally, Axis supplies data infrastructure to decentralized on-chain robotics networks, including BitRobot on the Solana network and OpenRoboto on Bittensor.
Redefining Physical AI’s Data Foundation
The philosophy driving the company’s rapid expansion is rooted in the belief that static datasets are no longer sufficient for the demands of modern robotics.
"The future of Physical AI isn’t a static dataset you download once," said Chris Feng, founder of Axis Robotics. "It’s an engine that keeps producing the data the model needs next. Scale gets you broad coverage. Diversity keeps the noise unbiased. The closed loop turns every failure into progress. That’s what compounds."
Axis Robotics was founded by a multidisciplinary team of researchers hailing from UC Berkeley, Carnegie Mellon University, the Georgia Institute of Technology, and Shanghai Jiao Tong University, alongside serial technology founders who have previously scaled consumer platforms to more than 30 million users. The company’s ongoing research initiatives are advised by Jiachen Li, an Assistant Professor at the Georgia Institute of Technology.
Leave a Reply