Embodied AI Data Management

EmbodiFlow

Organize multi-robot, multi-sensor captures into datasets suitable for training, and complete model training on the same platform. Available as a cloud instance, or as a workstation or cluster on premises.

Overview

The platform integrates with TeleXperience and SenseXperience. On-premises deployment is available as a workstation edition or a cluster edition. Data remains in the customer’s storage, and the system can run offline. Hardware and site requirements are described in the Deployment Requirements Checklist.

Platform

Collection, annotation, and training on one platform

Data can be ingested from devices or uploaded locally, then annotated, checked, and exported. Training can also be started on the platform.

01

Data Collection

Device ingest and upload

A robot-side agent packages and checks data locally before upload. Local file upload is supported, as is conversion of video and audio to MCAP. Collection can be connected to TeleXperience and SenseXperience.

02

Annotation & QA

Annotation, review, and export

Supports natural-language semantic annotation and object annotation in camera views, with review and rework. Quality control covers frame rate, topic continuity, and time synchronization. Common training formats such as LeRobot, HDF5, and MCAP can be exported.

03

Training

In-platform model training

Includes policy models such as ACT, π₀, SmolVLA, and GR00T. Training jobs can be created in the interface by selecting data and GPUs. GPU hardware is not required if training is not used on the platform.

Core Features

Standardization, quality control, training, and deployment

Unified data format

Captures from different sources are converted to MCAP, with ROS topics aligned in time. Common dual-arm, humanoid, collaborative-arm, and mobile platforms can be replayed in the platform.

Annotation workflow and quality control

Supports multi-user annotation, review, and rework. Quality control checks frame rate, dropped frames, and time synchronization. Datasets that fail quality control are blocked from export by default.

Training export and in-platform training

Exports include LeRobot, HDF5, MCAP, and other formats. ACT, π₀, SmolVLA, GR00T, and other models can also be trained on the platform.

Cloud instance and on-premises deployment

A cloud instance is available, as are workstation and cluster editions for on-premises rooms. Data is stored on the customer’s disks or object storage, and offline operation is supported.

Documentation & Guides

Documentation

More resources

Deployment sizing

Estimate a workstation, cluster, or cloud instance from your capture scale.

Frequently asked questions

Which organizations is EmbodiFlow intended for?
Intended for robot OEMs, data-service organizations, universities and research groups, and system integrators. Typical uses include device ingest and collection, annotation collaboration and delivery, model training, and on-premises deployment.
Which training data formats are supported for export?
Twelve formats are supported, including LeRobot, HDF5, RLDS, MCAP, JSON, and CSV. LeRobot can be exported as v3.0 or v2.1. Datasets that fail quality control are blocked from export by default.
How is on-premises deployment structured, and how do the workstation and cluster editions differ?
On-premises deployment can run fully offline, with no data sent outside the site. The workstation edition is a single tower server. The cluster edition is based on Kubernetes and remains available if one node fails. Hardware and site requirements are in the Deployment Requirements Checklist.
Which models can be trained on the platform?
Supported policies include ACT, Diffusion, π₀ / π₀.5, SmolVLA, GR00T, and Spirit-v1.5, on PyTorch and JAX. Data, model, and GPUs are selected in the interface. GPU hardware is not required if training is not used on the platform.