A Simulation Benchmark for
Multi-Humanoid Collaboration

10 tasks · 2–3 humanoids · 600 demos

Under review

CoHuB: A Simulation Benchmark for
Multi-Humanoid Collaboration

  • 10collaborative tasks
  • 2–3humanoids per task
  • 600human-teleop demos
  • 8baseline policies

Abstract

Many physical tasks in human environments require collaboration, from assisting a partner to jointly manipulating an object. Yet, existing humanoid benchmarks largely focus on single-humanoid skills and lack evaluation of multi-humanoid collaboration under egocentric visual observations. We introduce CoHuB (Collaborative Multi-Humanoid Benchmark), a simulation benchmark for multi-humanoid collaboration under egocentric visual observations. CoHuB provides 10 tasks—eight with two humanoids and two with three humanoids—spanning diverse collaboration patterns. We also provide synchronized demonstrations collected through a multi-operator VR teleoperation pipeline, in which each operator controls one humanoid from its egocentric view. Experiments with representative visuomotor policies reveal substantial challenges across different forms of coordinated perception and control. CoHuB provides a foundation for developing and evaluating multi-humanoid collaboration policies.

Task suite & taxonomy

10 tasks with diverse collaboration patterns

8 two-humanoid tasks in 4 categories, plus 2 three-humanoid tasks.

Base movement
How much the humanoids need to move from their initial positions
Physical coupling
How much the humanoids must physically interact with the same object during a collaborative task

Tasks with three humanoids

Click a task for its phases, language instructions and results.

Egocentric observations

Each humanoid sees only its own view

By default, each robot observes only its own egocentric RGB image and proprioception.

Multi-operator VR teleoperation & dataset

Synchronized demonstrations

Each operator controls one humanoid from its egocentric view within a shared, synchronized simulation.

CoHuB pipeline. Step 1: multi-operator VR teleoperation, where mirror sessions send 50 Hz actions to a shared state-only recording session in Isaac Sim/Lab. Step 2: the state dataset is rendered and replayed into the full dataset with RGB-D, state and instructions. Step 3: policy evaluation with phase-wise success rates on collaborative tasks.
The multi-operator teleoperation system records synchronized state trajectories, which are replayed offline to render RGB-D observations and construct the full demonstration dataset for policy training and evaluation.
  • 50 + 10train + val demos per task
  • 1,320per-humanoid trajectories
  • RGB-Degocentric view per humanoid
  • 35-Daction per humanoid
  • 50 Hzsynchronized recording
  • Isaac Sim + Isaac Lab
  • Unitree G1 + Dex3-1 hands
  • 43 DoF per humanoid
  • Meta Quest 3 teleoperation

Evaluation protocol

Phase-wise evaluation

Each task is divided into phases p1 → p2 → p3. In two-humanoid tasks, each phase is non-collaborative or collaborative.

100rollouts per policy and taskfinal task success and phase completion
50trials per phasecontrolled phase-level evaluation from demonstration states
Standard ILACT · DP
Multi-agent ILLatentToM · GauDP
VLAπ0.5 · GR00T N1.7 · Ψ0
WAMFast-WAM

Results

Collaboration remains challenging

Current policies remain limited in multi-humanoid settings, with lower performance observed during collaborative phases.

Overall benchmark results

Final task success rate (%) over 100 rollouts. Click a column to sort.

Phase-wise results

Cumulative success rate (%) per phase, averaged over tasks or for a single task. p3 is full-task success.

How does performance differ between non-collaborative and collaborative phases?

Average success rate on non-collaborative and collaborative phases over the 8 two-humanoid tasks. Each phase is evaluated from its beginning state in each of the same 50 demonstrations.

How do multi-agent design choices affect performance?

Average task success (%) over the 8 two-humanoid tasks. Local: own egocentric view. Global: egocentric views of all humanoids. Separate / shared: policy parameters per humanoid.

Agentic robot policy · exploratory

GPT-6 Astra

10%average task success over the 8 two-humanoid tasks, zero-shot without training on CoHuB data

One centralized policy controls both humanoids from their egocentric RGB images and proprioception, with depth available only through queries at selected pixels, and outputs base, hand and wrist commands.

Its profile is opposite to the trained baselines. It succeeds on CoPouring (70%) and Pouring (10%), where no trained policy exceeds 11%, but records no successes on the tasks where they do best, such as Handover, FrameHang, CoCarry and TableAlign. This suggests different strengths and limitations from learned visuomotor policies.

Exploratory evaluation through Codex, 10 episodes per task with a larger time budget. Its separate observation, control and time-budget settings limit direct comparison with the trained baselines.

Failure cases

Collaboration failure cases

One GR00T N1.7 rollout per task, replayed from its logged simulator state.

Citation

BibTeX