Axis Robotics Releases Massive Franka Arm Simulation Dataset for Physical AI
Axis Robotics has released the Axis Sim Dataset V1, featuring over 50,000 human-teleoperated simulation trajectories for Franka arm manipulation. The open-source dataset aims to advance Physical AI research and model training.
Axis Robotics has launched the Axis Sim Dataset V1, presenting one of the largest open-source simulation datasets designed for Franka arm manipulation. The complete dataset, training code, and benchmarks are now available to the public. V1 comprises over 50,000 human-teleoperated simulation trajectories encompassing 207 manipulation tasks and more than 60,000 scene variants using a simulated Franka Research 3 arm.
Achieving over 160,000 downloads, this release stands as the most downloaded open-source simulation Franka manipulation dataset on Hugging Face. Benchmark evaluations show that continual pretraining on V1 enhances π0.5 and outperforms a volume-matched RoboCasa baseline, with all outcomes fully open and verifiable.

Axis Robotics is developing a compounding data engine tailored for Physical AI, structured as a vertically integrated platform that combines large-scale simulation, egocentric real-world data collection, humanoid loco-manipulation, and human-gated DAgger post-training. The enterprise secured $12 million in seed financing led by Hack VC, with participation from Nomad Capital, Pi Network Ventures, 10K Ventures, and several angel investors.
A Bet Against “Clean Data Only”
A prevailing convention in robotics dictates that training demonstrations must start near-optimal—filtering exclusively for expert trajectories, standardizing environments, and discarding noisy inputs before imitation training takes place. Axis takes the opposite approach, arguing that data quality exists at the distribution level rather than within individual trajectories. When a large, diverse crowd generates noisy, suboptimal trajectories with uncorrelated errors, the noise balances out, leaving a functional policy during the training phase.
The Axis Sim Dataset V1 evaluates this theory publicly. Its trajectories cover pick-and-place tasks, stacking, pouring, articulated-object manipulation, and tool utilization. Rather than originating from a single expert team, these were gathered through Axis Hub—a browser-based teleoperation platform—by a distributed crowd. The dataset was created alongside researchers from institutions including UC Berkeley, Johns Hopkins, and the University of Michigan.
Results That Scale
Tested on LIBERO-Plus, continual pretraining using V1 elevates the success rate of π0.5 from 83.9% to 88.8% and exceeds a volume-matched RoboCasa365 baseline by 37.3%. Performance scales steadily as pretraining data increases from 25% to 100% of the dataset without hitting a saturation point, demonstrating that performance boosts stem from broad diversity and coverage rather than an isolated spike. The most pronounced gains manifest during camera, sensor-noise, and layout perturbations, which are the exact variables Axis randomizes during creation.

Development on V2 is already underway, targeting 1.2 million trajectories across 1,200 tasks. This next iteration focuses on cross-embodiment generalization and will demonstrate how suboptimal simulation data helps train robust policies across multiple VLA models.
The Engine Behind the Dataset
This dataset functions as one component of a broader, compounding data engine. While conventional data vendors collect information against a static specification and stop, Axis utilizes model performance and failure analysis to steer future collections, ensuring every training cycle feeds into the next. This architecture operates across four distinct data streams:
- Simulation: Over 200,000 distributed contributors operating on Axis Hub, noted as a top-3 dApp on Base, have generated upwards of 4.7 million trajectories across 13 embodiments.
- Egocentric: A managed network consisting of more than 1,000 full-time, QC-trained collectors record first-person actions within actual homes and businesses across 14 industries, accumulating over 200,000 hours with daily increases exceeding 4,000 hours alongside Vicon-verified hand tracking.
- Loco-manipulation: More than 500 hours fusing mobility and dexterity on physical humanoids like the Unitree G1 and Booster T2 through hardware-agnostic teleoperation.
- Human-gated DAgger post-training: Over 500 hours dedicated to human-in-the-loop corrections focused on deployment edge cases.
Every specific task and trajectory is logged on-chain via Base to establish clear provenance, and contributors receive rewards for verified work quality.
From Open Data to Commercial Deployment
Beyond publishing open-source simulation data, Axis collaborates directly with robot embodiment companies to engineer tailored, embodiment-specific data pipelines alongside model priors.
Serving as Booster Robotics’ primary sim-data partner, Axis replicated Booster’s physical workspace inside a task-aligned digital twin. Distributed contributors then gathered over 42,000 simulation episodes within this environment to distill a custom model prior for Booster. Utilizing a mere 30 real-robot demos, this prior attained an 87.5% success rate—compared to 37.5% for a standard out-of-the-box π0.5—effectively matching π0.5 performance while using half the real-world demonstrations.
Additional partners include embodiment companies such as Feagine Robotics, model companies like Manycore Tech and Dexmal, and industrial automation entities including Lotus Cars and Geely Auto. Axis also provides data for on-chain robotics networks, specifically BitRobot on Solana and OpenRoboto on Bittensor.
Redefining Physical AI’s Data Foundation
“The future of Physical AI isn’t a static dataset you download once,” stated Chris Feng, founder of Axis Robotics. “It’s an engine that keeps producing the data the model needs next. Scale gets you broad coverage. Diversity keeps the noise unbiased. The closed loop turns every failure into progress. That’s what compounds.”
Axis was established by researchers coming from UC Berkeley, CMU, Georgia Tech, and SJTU, alongside serial founders who have previously scaled consumer platforms to more than 30 million users. The company’s research activities receive guidance from Jiachen Li, an Assistant Professor at Georgia Tech.
Paper Link: https://arxiv.org/abs/2607.21588
Project Page: https://axisaiorg.github.io/AXIS-V1/
Dataset Link: https://huggingface.co/datasets/axisrobotics/Franka-Dataset
Github Codebase: https://github.com/AxisAIOrg/Axis-V1-Training
Short FAQs
- What is the Axis Sim Dataset V1? It is one of the largest open-source simulation datasets for Franka arm manipulation, containing over 50,000 human-teleoperated simulation trajectories across 207 tasks and 60,000+ scene variants.
- How much funding did Axis Robotics raise? Axis Robotics raised $12 million in seed funding led by Hack VC, with participation from Nomad Capital, Pi Network Ventures, 10K Ventures, and angel investors.
- What are the key performance improvements of using V1? Continual pretraining on V1 lifts the success rate of π0.5 from 83.9% to 88.8% on LIBERO-Plus and outperforms a volume-matched RoboCasa365 baseline by 37.3%.
- How does Axis collect its data? Axis runs a hybrid strategy across four data lines: distributed simulation via Axis Hub, egocentric real-world video capture, humanoid loco-manipulation, and human-gated DAgger post-training.
