AI 日报hiw3c.com

NVIDIA CLAROPD教多轮人工智能代理从重大错误中恢复过来

原文标题 · NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes
MarkTechPost www.marktechpost.com RSS 全文
正文为英文,可一键机器翻译(仅首次需要等待)

NVIDIA researchers, with Princeton University and the University of Maryland, have introduced PivotOPD, an on-policy distillation method for multi-turn LLM agents. PivotOPD on-policy distillation trains an agent to avoid its most damaging early mistake, and to recover when it happens anyway. Against 13 baselines, it posts the best average on ALFWorld, WebShop and Search-based QA for Qwen3-1.7B and Qwen3-8B students. The takeaway: recovery is learnable, and standard OPD rarely teaches it.

TL;DR

  • Size: A training method, not a model. Tested on Qwen3-1.7B and Qwen3-8B students, plus a Nemotron-3.5-SFT student on SWE-Bench Verified.
  • Runs on: Trained on NVIDIA H100 nodes. Adds 0 inference cost, so the trained agent runs wherever its base model runs.
  • Performance: First on all 8 per-benchmark averages against 13 baselines, across 3 seeds.
  • Best: Recovers from 72.7% of replayed pivotal mistakes, vs 20.3% for standard OPD.
  • Worst: 55.9% on ALFWorld “Look” tasks with the 1.7B student, vs 83.9% for SOD.
  • Bottom line: Best: teaches recovery that outcome-only RL cannot reach. Worst: depends on replayable environments and a teacher whose pivots match the oracle in 77.8% of failed rollouts.

What is a pivotal mistake in a multi-turn agent?

A pivotal mistake is an action that lengthens the shortest remaining path to finishing a task, or makes it unsolvable. ALFWorld’s symbolic oracle measures this at every turn.

Across Qwen3-8B, Qwen3-30B-A3B and Qwen3-235B-A22B, 59% of failed rollouts (155 of 262) contained one. The first pivotal turn arrived early, at a median of turn 8 to 12 out of 30. The agents then wasted 18 to 21 more turns without recovering.

In replays of Qwen3-8B failures, correcting the pivotal turn raised success from 8% to 59%. Leaving the mistake in place and forcing the right action for the next 2 turns still reached 58%.

Why does standard on-policy distillation miss it?

Standard OPD lowered the held-out failure rate from 79% to 56%. Failures after a pivotal turn only fell from 51% to 49%. The correct action stayed below 1% probability at every pivotal turn, so 8 rollouts rarely sample it.

Outcome-based RL shares the blind spot: if every rollout fails, the group-relative advantage is 0.

How does PivotOPD work?

PivotOPD adds 3 components to group-based RL, combined in a single PPO update.

  1. Pivot detection: A larger teacher model reads each rollout and its outcome in hindsight. It picks candidate turns and names a gold action at each. A turn counts as pivotal when the student’s action differs from the gold action. On ALFWorld, detected pivots land within 1 turn of the oracle’s pivot in 77.8% of failed rollouts on average.
  2. Preventive distillation (reverse KL): A frozen copy of the student, hinted with the gold action, re-scores the student’s own response. This pushes the student away from the committed mistake.
  3. Recovery distillation (forward KL): After each pivot, the teacher names a recovery action for up to K turns. The hinted self-teacher writes recovery responses, and the unhinted student trains on them. Forward KL is mass-covering, so it lifts actions the student almost never samples.

The teacher only names actions. Token-level targets come from the student’s own hinted distribution.

How does PivotOPD perform on agent benchmarks?

With the 1.7B student, PivotOPD averages 73.7% on ALFWorld, 5.5 points above SDAR. It averages 44.5% on Search-based QA, 5.9 points above RLSD. On WebShop it beats RLSD by 1.2 in score but by 14.1 in success rate (76.6%).

With the 8B student, it reaches 93.0% on ALFWorld, 47.4% on Search-based QA and 81.9% WebShop success. Margins are smaller, at least 1.8 points.

With Qwen3-8B as its own teacher, PivotOPD still wins all 3 benchmarks by at least 1.5 points, 3.9 on average.

On SWE-Bench Verified, a Nemotron-3.5-SFT student taught by Nemotron-3-Super went from 62.8% to 66.0%. Standard OPD reached 63.0%, and the teacher scores 73.0%.

Recovery is the standout. Across 72 replayed pivotal mistakes, PivotOPD recovered 72.7% of the time, vs 8.3% for the base model, 20.3% for standard OPD and 45.8% for preventive-only. It averaged 9.7 turns to recover, against an optimal 6.2.

Prev</button><button class="btn" id="mtp-next">Next ▶</button></div> </div> <div class="panel" data-p="2"> <div class="tog" id="mtp-tog"><button class="on" data-m="s">Qwen3-1.7B</button><button data-m="l">Qwen3-8B</button></div> <div id="mtp-bench"></div> <p class="muted">Averages over 3 seeds from Table 1 of the paper. Compared with the strongest baselines. PivotOPD ranks first on all 8 per-benchmark averages against 13 baselines.</p> <div class="bench"><h6>SWE-Bench Verified resolve rate (Nemotron-3.5-SFT student)</h6><div id="mtp-swe"></div></div> </div> <div class="panel" data-p="3"> <p>72 oracle-labeled pivotal mistakes, replayed 8 times per policy. How often does each policy finish the task anyway?</p> <div id="mtp-rec"></div> <p class="muted">Average turns to recover: PivotOPD 9.7, standard OPD 12.3, base 13.4, optimal 6.2.</p> <button class="btn" id="mtp-case">Replay case study</button> <div class="case"> <div class="col"><h6>Base model</h6><div id="mtp-ca"></div></div> <div class="col"><h6 style="color:#76B900">PivotOPD model</h6><div id="mtp-cb"></div></div> </div> </div> <div class="ft"><span>Source: <a href="https://arxiv.org/abs/2609.40285" target="_blank" rel="noopener">arXiv 2609.40285</a></span><b>© Marktechpost</b></div> </div> <script> (function(){ var R=document.getElementById('mtp-pivotopd'); function $(s){return R.querySelector(s);} function $$(s){return R.querySelectorAll(s);} function resize(){try{parent.postMessage({mtpPivotOPDHeight:R.offsetHeight+40},'*');}catch(e){}} var shown={}; $$('.tab').forEach(function(b){b.onclick=function(){ $$('.tab').forEach(function(x){x.classList.remove('on')});b.classList.add('on'); var p=b.getAttribute('data-p'); $$('.panel').forEach(function(x){x.classList.toggle('on',x.getAttribute('data-p')===p)}); if(p==='2')bench(mode); if(p==='3')rec(); setTimeout(resize,50);};}); /* panel 1 */ var T=$('#mtp-turns'),cells=[]; for(var i=1;i<=30;i++){var d=document.createElement('div');d.className='t';d.textContent=i;T.ap