Overall benchmark results
Final task success rate (%) over 100 rollouts. Click a column to sort.
A Simulation Benchmark for
Multi-Humanoid Collaboration
10 tasks · 2–3 humanoids · 600 demos
Under review
Many physical tasks in human environments require collaboration, from assisting a partner to jointly manipulating an object. Yet, existing humanoid benchmarks largely focus on single-humanoid skills and lack evaluation of multi-humanoid collaboration under egocentric visual observations. We introduce CoHuB (Collaborative Multi-Humanoid Benchmark), a simulation benchmark for multi-humanoid collaboration under egocentric visual observations. CoHuB provides 10 tasks—eight with two humanoids and two with three humanoids—spanning diverse collaboration patterns. We also provide synchronized demonstrations collected through a multi-operator VR teleoperation pipeline, in which each operator controls one humanoid from its egocentric view. Experiments with representative visuomotor policies reveal substantial challenges across different forms of coordinated perception and control. CoHuB provides a foundation for developing and evaluating multi-humanoid collaboration policies.
Task suite & taxonomy
8 two-humanoid tasks in 4 categories, plus 2 three-humanoid tasks.
Click a task for its phases, language instructions and results.
Egocentric observations
By default, each robot observes only its own egocentric RGB image and proprioception.
Multi-operator VR teleoperation & dataset
Each operator controls one humanoid from its egocentric view within a shared, synchronized simulation.
Evaluation protocol
Each task is divided into phases p1 → p2 → p3. In two-humanoid tasks, each phase is non-collaborative or collaborative.
Results
Current policies remain limited in multi-humanoid settings, with lower performance observed during collaborative phases.
Final task success rate (%) over 100 rollouts. Click a column to sort.
Cumulative success rate (%) per phase, averaged over tasks or for a single task. p3 is full-task success.
Average success rate on non-collaborative and collaborative phases over the 8 two-humanoid tasks. Each phase is evaluated from its beginning state in each of the same 50 demonstrations.
Average task success (%) over the 8 two-humanoid tasks. Local: own egocentric view. Global: egocentric views of all humanoids. Separate / shared: policy parameters per humanoid.
Agentic robot policy · exploratory
10%average task success over the 8 two-humanoid tasks, zero-shot without training on CoHuB data
One centralized policy controls both humanoids from their egocentric RGB images and proprioception, with depth available only through queries at selected pixels, and outputs base, hand and wrist commands.
Its profile is opposite to the trained baselines. It succeeds on CoPouring (70%) and Pouring (10%), where no trained policy exceeds 11%, but records no successes on the tasks where they do best, such as Handover, FrameHang, CoCarry and TableAlign. This suggests different strengths and limitations from learned visuomotor policies.
Exploratory evaluation through Codex, 10 episodes per task with a larger time budget. Its separate observation, control and time-budget settings limit direct comparison with the trained baselines.
Failure cases
One GR00T N1.7 rollout per task, replayed from its logged simulator state.
Citation