NeurIPS 2026 Workshop on Self-Evolving Diversity-Driven Search for Robust AI Systems (EvoRobust) · Oral

Measuring Failure Yield and Reflection Cost in a Frozen Robot Policy

Qinzhen Ma1Shichen Tang1Qingxi Meng1Jialin Wu2

1Rice University2UC San Diego

TL;DR

When searching for failures of a frozen ACT robot policy, structured reflection finds slightly more failures than matched language-guided selection, but a cheap nearest-neighbor selector finds more, in about half the time and with no model calls.

Abstract

We measure failed rollouts and the resource cost of structured reflection for a frozen Action Chunking with Transformers (ACT) policy in MuJoCo bimanual cube transfer. Five primary methods and three post-primary controls each use five search seeds and 64 rollouts per run. In a retrospective analysis, reflection averages 20.4 failures versus 17.6 for otherwise matched language-guided selection: a paired difference of 2.8 +/- 1.92 (sample standard deviation), with 25.6% more measured time and 6.85% more tokens. The comparison with inexpensive search is less favorable. Sequential QD averages 18.8 failures; a subsequently frozen nearest-neighbor selector using past failure labels averages 29.0 +/- 12.49, with 166.1 s/run and no model calls, versus 315.8 s/run for reflection. All 140 repeat executions of 70 unique primary archive scenes remain failures. These observations show that an increment over language selection need not justify reflection's cost against inexpensive alternatives. They do not establish statistical superiority, cross-task generalization, or improved policy robustness.

Results

Figure 1
Figure 1. Paired failure-count differences at 64 rollouts. Each point is one search seed; all panels use the same horizontal scale. The kNN comparison uses a separately executed post-primary control.

Paper

Page 1Page 2Page 3Page 4

Download PDF (12 pages)

BibTeX

@inproceedings{ma2026measuring,
  title  = {Measuring Failure Yield and Reflection Cost in a Frozen Robot Policy},
  author = {Qinzhen Ma and Shichen Tang and Qingxi Meng and Jialin Wu},
  booktitle = {NeurIPS 2026 Workshop on Self-Evolving Diversity-Driven Search for Robust AI Systems},
  year   = {2026}
}