Understanding the Limits of Open-Source General Robotics Policies

We evaluate zero-shot performance of five open-source robot policies through 6,000 episodes on RoboLab-120. Current policies show generalizable intelligence - but are not ready for drop-in deployment.

Comparing policies head-to-head reveals common behaviors and individual failure modes

“Put the banana, then the cube, in the bowl.”

G0.5 Episode failure Task SR 0/10
RoboLab-120, BananaThenRubiksCubeTask, matched episode 5. Playback is accelerated 1.89× and excludes model inference time; Cosmos is shown from a different viewport, but policy inputs are consistent.

All surveyed open-source policies show common zero-shot capabilities, but also distinct behaviors and failure modes. Failures include grasping the wrong object (GR00T N1.7), struggling to grasp non-standard objects like Bananas (MolmoAct 2 and G0.5), or sometimes even bad luck (π0.5). Even Cosmos 3 Nano Policy succeeds in only 7 of the 10 recorded episodes for this task. Policies are already capable, but reliability is a key concern.

This one scene previews the questions behind the study: how reliably can these policies finish the task, where do their attempts break down, and what would help them do better? The comparison below tests those questions across 120 tasks.

Demo videos on social media provide a biased view of the actual capabilities of robotic policies. To get an intuition for the actual current state of fully-autonomous robotics, we want to understand the zero-shot capabilities of current open-source robotics policies in a best-case scenario: in simulation, using the simple DROID embodiment, and a single stable environment inspired by common academic benchmarks.

Progress in hardware and the adaptation of models and techniques from language and vision models have paved the way for a new wave of excitement in robotics. Right now, we see the emergence of “general-purpose robotics policies”, like Generalist GEN-1.5, Gemini Robotics 2, π0.7, or NVIDIA GR00T. Physical Intelligence co-founder Chelsea Finn has recently spoken about the prospect of a “GPT moment for robotics.” And given the impressive demos, who would dare to disagree?

However, so far these demos are just that. Publicly available showcases remain few and far between. For many people, the most they have seen outside a lab is probably a remotely operated Unitree robot. When we speak about the GPT moment, what we think of is wide generalization, transfer learning, and emerge of behavior from large-scale training. Importantly, the GPT moment happened mostly out-in-the-open. Models were made publically available (not necessairly open-source), for the community to inspect, test, and build upon. In contrast, the current state of frontier robotics is mostly closed-source, with only a few open-source policies available for inspection and testing.

So, what can these openly available policies teach us today about the state of general-purpose robotic policies? To answer this question, we look at publicly available models for robotic manipulation. Open models certainly trail the frontier by several months, but they should still give us an idea of the basic level of capability, as well as the most obvious common challenges.

Specifically, we are driven by three main research questions:

  1. How good are current open-source policies in a zero-shot setting?
  2. What are common and specific failure modes of models?
  3. What can we learn from this for the next steps of policy training and evaluation?

We start our investigation in simulation. This allows for a level of accessibility and reproducibility that is not possible with real robots. Importantly, it allows us to compare the behavior of policies across a broad selection of scenarios and tasks. At the same time, benchmarking in simulation is generally considered easier, which matches our goal of first understanding robotic policies in a best-case scenario.

Methodology

We use NVIDIA RoboLab to compare five released policies on a common DROID embodiment, without task-specific training for this evaluation.

5policies 120tasks 10episodes 6,000policy episodes

We run all five policies on the same 120 tasks and the same ten episode slots per task. Alongside RoboLab's success and partial-credit scores, we analyze where failures happen, wrong-object interactions, motion, and whether different policies succeed on the same episodes.

Success rate measures completed tasks; Score also credits partial progress. We combine these outcomes with recorded trajectories, event logs, and a manual review of 1,150 episodes to understand the behavior behind the scores.

We run our Isaac Sim 6 port of RoboLab. In the evaluation loop, simulation does not advance while the policy computes its actions: inference adds wall-clock time, but does not consume the simulated task budget or appear as waiting time in the videos.

How RoboLab, policy execution, and the analysis work

Why we use DROID

DROID is one of the best-known open manipulation platforms, and DROID-adapted checkpoints were available for all five policies in this comparison. RoboLab’s built-in DROID setup gives each model the same fixed-base, single-arm, parallel-jaw embodiment. That makes the comparison practical. It offers a controlled starting point for comparing manipulation policies. These results describe this model–embodiment pairing; they do not establish an upper bound for other robots or for real-world performance.

Running the policies

RoboLab runs each policy through its model-specific adapter: the adapter prepares observations and the task instruction, queries the policy, and passes its actions back to the simulator. We retain each policy's native preprocessing and action execution rather than treating the five models as interchangeable controllers. The comparison therefore evaluates the released checkpoint together with its adapter.

This means the success scores do not test whether a policy can keep up with a real-time control deadline; we report inference cost separately. RoboLab also measures policy-query time separately from simulator stepping and video recording.

Simulator port and compatibility

The upstream repository documents Isaac Sim 5.0 and 5.1; our runs use the Isaac Sim 6 port with Isaac Lab 3. The port preserves RoboLab's observation, action, and recording conventions while adapting the simulator interface. We checked compatibility with targeted regression tests and matched-run comparisons covering initial states, actions, trajectories, and outcomes; the fork also includes checks for coordinate frames, scene identity, and the DROID robot's structure.

Compatibility does not imply identical trajectories across simulator versions: changes in rendering and contact dynamics can alter closed-loop behavior. All five policies in the reported comparison therefore use the same Isaac Sim 6 stack, and our initial-state audit confirms that their 1,200 non-camera physical starting states match. Overall success rates for π0.5 and Cosmos 3 Nano Policy closely match their published results on the official RoboLab-120 leaderboard. This provides an additional aggregate sanity check for the port, rather than proof of one-to-one episode equivalence.

Matched initial physical states do not imply identical camera observations or control cadences. The compute comparison reports recorded policy-query time under these execution setups. Event and motion measures are descriptive; the manually reviewed subset supports behavioral hypotheses rather than population-wide failure rates.

RoboLab-120 is broad, but many tasks are close relatives

Color denotes task family; shape denotes difficulty. Nearby points share instruction language and annotations.

Hover or select a point to inspect its scene, instruction, tags, and five-policy result. This exploratory projection uses TF-IDF features of the task instructions as a proxy for similarity. Difficulty reproduces RoboLab’s task labels; the six task families are our editorial grouping of RoboLab annotations, not an official benchmark taxonomy.

How good are current open-source policies in a zero-shot setting?

The policies complete a meaningful share of tasks, but none is broadly reliable in this controlled setting. We start with overall performance, then look at task differences and the compute needed to achieve it.

Cosmos performs best overall. That changes when we look at individual tasks.

Success depends on whether the final state holds, while Score also credits completed subtasks. In aggregate, Cosmos is in the lead, although it isn't the best policy for every task. Below, intervals are shown for the aggregate success rate.

Loading the comparative readout…

In our runs, we closely reproduce the aggregate success rates reported on the official RoboLab-120 leaderboard for π0.5 and Cosmos 3 Nano Policy: 27.5% versus 28.0% for π0.5, and 35.1% versus 36.8% for Cosmos 3 Nano Policy.

How to read the metrics
Success rate (SR)
The share of episodes that satisfy the complete task at the end. Each policy has 1,200 episodes: 120 tasks × 10 matched seeds. Aggregate SR whiskers are 95% exact binomial intervals.
Score
Mean credited task progress on a 0–100 scale. It awards fractional credit for completed subtasks, so a policy can have a higher Score without more full successes.
Failure progress
Mean Score among failed episodes only. It separates failures that made useful partial progress from episodes that never completed a credited step.
EE speed
Mean end-effector speed over the rollout, reported in cm/s. Faster is not necessarily better; it describes the policy’s motion profile.
EE SPARC
Spectral arc length of the end-effector speed profile. Values closer to zero are smoother; more-negative values indicate less-smooth motion.
Wrong-object rate
The share of episodes with at least one wrong-object grasp annotation. It records interaction incidence, not proof that the event caused the final failure.

Motion metrics should be compared within similar tasks and outcomes: an immobile failure can appear smoother than a successful, long-horizon trajectory.

The five-policy field

All 1,200 episodes per policy.

Per-task breakdown

The aggregate distribution may not resemble yours. Find a breakdown of different scores for each task below.

120 tasks

The ranking changes depending on what a task requires. A policy that does well on spatial tasks may do poorly on multi-step tasks.

Capability slice comparison

Axes overlap (not to be read as additive categories).

Outcome overlap

Pairwise success Jaccard asks how often two models solve the same matched episodes. Low overlap plus sizeable one-sided wins signals useful complementarity.

Capability is only useful if latency fits the control loop

RoboLab records the time spent in policy queries at every simulator step. That makes policy-query latency, together with the policy time accumulated over a full sweep, the most useful compute comparison available from these runs. π0.5 is fastest here, while Cosmos 3 Nano Policy is the slowest. The ordering does not track model size. G0.5, the smallest policy in the field, is the second slowest.

Measured policy inference cost

Step-weighted across all 120 task batches.

Parameter counts cover the policy model, including its perception and action components, but exclude the simulator and serving environment. G0.5 uses a Qwen3.5-2B policy backbone plus a separate ActionCodec.

Current open-source policies show useful zero-shot capability, but even the strongest succeeds in only about a third of these episodes. Cosmos 3 Nano Policy performs best overall, while π0.5 offers a strong trade-off between task performance and inference speed.

What are common and specific failure modes of models?

Across policies, failures extend from selecting and acquiring an object to transporting it, releasing it, and preserving the final state. Within that shared sequence, each policy has distinctive weaknesses.

Similar scores can mean very different behavior

Dense event annotations capture wrong-object interactions, contact, and final-state failures; HDF5 control channels quantify motion and action variation.

Compare behavior

Values are episode-level means or incidences.

These metrics are descriptive. Early termination and different control cadences can affect motion aggregates, and they don't tell us what caused the behavior.

Failures occur at different stages of manipulation

These signals overlap: one episode may select the wrong object, collide with the table, and later drop the target. They help us find interesting failures. They do not by themselves prove a single causal failure mode.

  1. Select the targetNo progress or wrong-object interaction
  2. ApproachTable or object collision
  3. AcquireClosed gripper without stable acquisition
  4. TransportTarget dropped or object leaves the scene
  5. Place and releasePartial progress stalls before completion
  6. Preserve final stateFull progress, but failed terminal state
Show mutually exclusive outcome-stage counts

What we learned from watching 1,150 episodes

We watched and narrated 1,150 of the 6,000 episodes, and checked every observation back against the recorded outcome. For each policy we watched 50 tasks, with an overlap of 42 directly comparable tasks.

MolmoAct 2, GR00T N1.7 and G0.5 sit within 4.6 points of one another on every cut of this benchmark. But watching them, they behave completely differently. Each has a recurring failure that barely appears in the other four. In two cases, we noticed the behavior before measuring it.

π0.5 Smooth, good at gripping, yet stopped by the rim. Bumps the carried object into the container wall

We saw this in48episodes, across 16 tasks

“it does go for the right target object, bumps into the wall, classic pi move.”

“It is grabbing nicely, but it's not super careful while carrying, so it's bumping into the pink crate like three times.”

Measured π0.5 is the only policy of the five for which lift height does not separate success from failure: −0.8 cm, against +5.9 to +13.3 cm for the others. It gets the object off the table just as reliably in its failures as in its successes, and loses it at the rim.

Cosmos 3 Nano Policy Competent, opportunistic and jittery. Abandons the last centimetre

Careful placements in28episodes, across 19 tasks

“once an object is touching the bin, or is leaning against it, or is on top of it, it deprioritises it a lot.”

“This is the shakiest policy that I've observed so far. There's always this micro-shaking in the policy.”

Measured A jitter index, mean change in velocity divided by mean speed, puts Cosmos 3 Nano Policy at 12.07 against 4.05–8.27 for the rest, with interquartile ranges that do not overlap π0.5’s. RoboLab’s own recorded SPARC metric ranks it the smoothest of the five, which is the opposite of what the videos show.

MolmoAct 2 Sees it, holds onto it, cannot turn its wrist. One approach angle, repeated until the clock runs out

We saw this in32episodes, across 20 tasks

“This policy is just incredibly bad at tilting the angle of the gripper. It always tries to grab it from its same angle, which is often very, very bad.”

“it’s really hard to grip the coffee pot by itself without also touching the bin, and then it’s basically stuck. It doesn’t try another angle for gripping, and the episode just ends there.”

Measured This becomes especially costly in clutter. After dropping an object, MolmoAct often cannot approach it from the remaining angle. It acquires an object in 38% of episodes, against 57% for Cosmos 3 Nano Policy and π0.5.

GR00T N1.7 Commits in the first second, disoriented by the ninetieth. Takes whatever is centred in frame, then loses coherence

We saw this in29episodes, across 18 tasks

“In all of the runs, it just goes quickly to the object that is most in the center, and then, if it's not the right object, it sometimes just stays stuck there and bugs around.”

“seems to lose coherence for the last 40 seconds and just stares at a random angle, making infrequent grips into the air.”

Measured Activity in the last third of an episode, as a fraction of the first third: GR00T 0.68, against 0.94 for Cosmos 3 Nano Policy and 0.87 for π0.5. GR00T is the only policy with a large drop. We had already noticed the behavior during review, before measuring it.

G0.5 Finds it, stands over it, decides it is finished. The test grip: touch, lift a centimetre, release, repeat

We saw this in49episodes, across 21 tasks

“It would have the perfect grip, but it never actually closes its grippers, so it's just hovering up and down on the hand without ever doing anything.”

“The policy straightforwardly goes for the onion and does the onion loop. So it doesn't grip it. It does test grips. … For the entire episode it's just stuck in the onion loop without ever picking it up anywhere.”

Measured This was the most common policy-specific failure we found. Its gripper closes at all in 54% of episodes, against Cosmos 3 Nano Policy’s 82%. We checked the same task with Cosmos 3 Nano Policy and saw no loop.

Different policies are good at different things

We also grouped episodes by what the task required. No policy leads on more than three of these capabilities.

Target selection

Picking the object the instruction names. π0.5 identifies it correctly more often than anyone and then takes everything else on the table too; GR00T does not select at all, it goes to whatever is centred in frame.

Strongest MolmoAct 2 Weakest GR00T N1.7

Getting a grip

Closing the hand on the object and lifting it clear. Only two policies do this reliably: Cosmos 3 Nano Policy and π0.5 acquire an object in 57% of episodes, against 44%, 38% and 30% for the rest.

Strongest Cosmos 3 Nano Policy, π0.5 Weakest GR00T N1.7

Approach angle

Rotating the wrist to suit the object. Cosmos 3 Nano Policy and π0.5 adapt; MolmoAct 2 effectively cannot, which matters most when the remaining approach requires a different wrist angle.

Strongest Cosmos 3 Nano Policy Weakest MolmoAct 2

Transport & release

Carrying the object and letting go in the right place. Cosmos 3 Nano Policy is the only policy that reliably clears a tall container; π0.5 loses objects on the rim; G0.5 releases at effectively random locations; GR00T leaves objects half-on rather than in.

Strongest Cosmos 3 Nano Policy Weakest G0.5

Search

Zooming out to find objects outside the current view. Cosmos 3 Nano Policy and π0.5 both do this successfully; G0.5 performs the motion without acting on what it sees; GR00T never surveys the table at all.

Strongest Cosmos 3 Nano Policy, π0.5 Weakest GR00T N1.7

Endurance

Still doing useful work late in the episode. Measured as motion in the final third relative to the first: 0.94, 0.87, 0.83, 0.75, 0.68. Only GR00T shows a real collapse; π0.5 is slow throughout but never degrades.

Strongest Cosmos 3 Nano Policy Weakest GR00T N1.7

We ran the same experiment twice and got different scores

Due to a score-independent rendering issue, we had to run the π0.5 policy twice on identical seeds. If episode-level success were stable, the two runs should largely agree.

They do not. An episode that succeeded the first time succeeded again only 64% of the time, and per-task success rate moved by 20 points or more on 28 of the 120 tasks. One task fell from 8/10 to 2/10; another rose from 4/10 to 10/10.

80.0% of episodes had the same outcome in both runs (baseline 81.3%).

Gripper-to-object distances correlated at r = 0.69 across 5,052 environment–object pairs. The median difference was 2.8 cm on a median distance of 22.2 cm.

The policy's behavior is much more reproducible than its success rate.

See for yourself

Similar scores conceal distinct weaknesses in grasping, transport, release, and recovery, which become visible when we combine outcomes with trajectories and video review.

What can we learn from this for the next steps of policy training and evaluation?

The findings point toward training on specific execution failures, testing improvements across benchmarks, and measuring what specialized policies contribute as general-purpose models become more capable.

The most useful training signal is often more specific than success or failure. π0.5's collisions with container rims suggest a transport-clearance problem; MolmoAct 2's repeated approach angles point toward grasp adaptation and recovery; G0.5's repeated test grips suggest difficulty committing to stable acquisition. Cosmos's unfinished placements and GR00T's loss of coherent activity add different priorities: completing the last step and sustaining useful behavior over time. These observations motivate targeted training experiments, not conclusions about which part of a model's training recipe caused the behavior.

Evaluation needs the same granularity. The π0.5 rerun shows how much an individual task score can move with only ten episodes, even on identical seeds. Meanwhile, Cosmos's favorable SPARC ranking sits uneasily with the visible shaking and our velocity-based jitter measure. Repeated runs, stage-specific outcomes, and video checks help distinguish a robust improvement from score variation or a misleading proxy. Neither a single outcome nor a single motion metric is enough.

Could the benchmark favor a model used while building it?

Policy rankings are properties of a model–benchmark pairing, not universal model attributes. That makes benchmark construction part of the result.

RoboLab measures the performance of each policy on a specific DROID embodiment. Although we use DROID-specific fine-tuned checkpoints for all policies, the quality of these checkpoints may still vary independently of the underlying general-purpose model. A natural next step is therefore to extend the current comparison into a cross-embodiment investigation.

The authors of Cosmos 3 Nano Policy also benchmarked on RoboLab-120 during its development, which could have impacted its performance. However, we should note that they show strong results on other benchmarks such as RoboArena as well.

More generally, the models available while a benchmark is being built shape its design. Scene choices, cameras, task definitions, and debugging decisions tend to become especially compatible with whichever policies the authors run during development. This is not specific to robotics: LLM benchmarks are routinely developed against GPT or Llama models and inherit the same bias. It is one route by which benchmark overfitting can arise without any intentional tuning.

The cross-benchmark reversal makes this worth flagging. On the official RoboDojo simulation leaderboard, G0.5 scores 20.23 with 14.88% success, compared with 11.41 and 6.91% for Pi-05. Here, π0.5 reaches 27.5% success while G0.5 reaches 10.5%. This does not prove benchmark overfitting. It does show that policy rankings are benchmark-contingent and should be triangulated across suites designed by different groups.

Focus on the manipulation problems that foundation models still leave open

Our working hypothesis is that general-purpose VLMs will increasingly supply scene understanding, task decomposition, and planning. The useful research question is therefore where specialized robot policies add capability beyond an off-the-shelf foundation model equipped with a suitable control interface.

Recent results show that for a fixed arm and parallel-jaw gripper such as DROID, general-purpose models already can propose a target position, orientation, and gripper command, with inverse kinematics (IK) translating the desired pose into joint configurations. That can simplify reaching, but a reachable pose alone does not solve collision-free transport, stable contact, controlled release, or recovery after a slip. Those are precisely the kinds of execution problems that recur in our videos.

The distinction is sharper for dexterous hands and humanoids. IK remains useful for these embodiments, but a sequence of end-effector targets alone is insufficient for coordinating contacts, forces, balance, and whole-body motion. We expect tasks that couple manipulation to these demands to be especially informative about the value of specialized policies. This is a direction to test, rather than evidence that foundation models cannot eventually learn such control.

Training and evaluation should target the execution problems that remain difficult, while testing where specialized policies add value beyond general-purpose models with suitable control interfaces.

Where to go from here

Four directions would make the next round of evaluation more informative:

  • Measure behavior at finer granularity. Track acquisition, transport, release, recovery, and final-state stability alongside success and time; automate trajectory and video analysis, and validate those annotations against human review so we can scale beyond watching each episode by hand.
  • Extend the benchmarks and reserve hold-out tasks. Compare independently designed suites, repeat runs, and keep tasks, objects, and layouts out of development and tuning to test whether improvements generalize beyond familiar evaluation cases.
  • Compare multiple embodiments. Extend beyond a fixed arm and gripper to bimanual systems, dexterous hands, and humanoids, including foundation-model baselines with appropriate control interfaces and explicit latency budgets.
  • Look inside the policy's predictions. Inspect predicted action sequences alongside executed motion to locate where plans and behavior diverge; use controlled changes to observations and instructions to test interpretations of what the policy responds to, rather than inferring its reasoning from video alone.

Related work

No single benchmark can answer all of these questions. This study builds directly on RoboLab and complements efforts that emphasize sim-to-real correspondence, long-horizon reasoning, knowledge transfer, or human-grounded household activity. Our narrower contribution is matched, cross-policy behavioral analysis: not only whether a policy succeeds, but where trajectories diverge and which failures matter.

References

Model links point to the exact evaluated DROID checkpoint where one is available on Hugging Face. π0.5 points to RoboLab’s instructions for the evaluated pi05_droid_jointpos checkpoint, which is distributed from Google Cloud Storage rather than Hugging Face. Contextual links point to first-party project pages, documentation, or original papers.

Acknowledgements

We especially thank the authors of NVIDIA RoboLab for creating and releasing the benchmark, task suite, evaluation harness, dashboard, and initial analysis.

We also thank the developers of π0.5-DROID, Cosmos 3 Nano Policy, MolmoAct 2, GR00T N1.7, and G0.5 for releasing models, checkpoints, and supporting code for open research. We are likewise grateful to the DROID team for the shared embodiment and dataset underpinning this comparison.

Finally, thank you to Jenai Xuning Yang, Haoquan Fang, Zhaodong Loke, Jonas Pai, and Benjamin Alt for reading drafts of this post and giving feedback.