Robot Learning · Interactive Imitation Learning

Calibrated intervention for action-chunking policies

Jinhe Tang1 Weiming Zhi1,2,3,*

1 School of Computer Science and 2 Australian Center For Robotics, The University of Sydney, Australia

3 College of Connected Computing, Vanderbilt University, TN, USA

* Corresponding author: Weiming.Zhi@sydney.edu.au

AutoIntervene selectively transfers control between a visuomotor policy and an operator, turning targeted recovery into supervision for the next policy update.

AutoIntervene project page · 2026

Representative real-world bimanual manipulation task evaluated with AutoIntervene
Representative real-world bimanual manipulation.

Overview

Know when to ask for help—and when to give control back.

Action-chunking policies can produce smooth actions even after perception error or execution drift moves the robot outside demonstrated support. AutoIntervene evaluates proposed chunks against a visual-action memory of successful task executions. Phase-local support triggers policy-to-operator transfer; global support enables autonomous operation to resume after recovery.

Method

A bidirectional control loop

Separate switching thresholds are calibrated from held-out expert demonstrations, avoiding direct manual tuning of score cutoffs.

AutoIntervene online control loop and policy adaptation framework
01

Construct

Pair current visual embeddings with the policy's proposed action chunk.

02

Evaluate

Measure visual support and action risk against successful references.

03

Transfer

Select policy or operator control using calibrated acceptance criteria.

04

Adapt

Retain successful intervention segments as targeted corrective data.

Experiment videos

Nine real-world bimanual tasks

Complete rollouts are compressed for web playback while preserving every experiment from the original project.

01

Peg Disassembly

Disassemble a pegged object and place all parts into the box.

02

Potato Transfer

Transfer a wooden potato between hands and place it into the box.

03

Towel Folding

Fold a towel from a fixed initial configuration.

04

Towel Bagging

Fold a towel, pack it and the pegs into a bag, and move the bag to the target side.

05

Lidded Box Packing

Open the box, pack the objects inside, and close the lid.

06

Plant Sorting

Place three wooden plant objects into their matching box slots.

07

Towel Box Packing

Remove pegs, fold the towel, reposition the box, and place the towel inside.

08

Two-Towel Box Packing

Remove pegs, fold two towels, and place both into the box.

09

Towels-and-Cable Bagging

Fold two towels, pick up the cable, and place all three items into the bag.

Experimental results

Targeted intervention improves success with less control data

Success rates are measured over 25 unassisted physical rollouts per task and method unless otherwise noted. Time reports recorded demonstration or operator-control data.

Table 1

Main seven-task benchmark

Success (%) and recorded control-data time. Δt is cumulative additional operator-control time.

TaskInitialHuman R1Human R2AutoIntervene R1AutoIntervene R2Additional Full Data
Succ.TimeSucc.ΔtSucc.ΔtSucc.ΔtSucc.ΔtSucc.Time
Peg Disassembly16%1173.7s52%123.6s56%286.9s64%49.3s72%127.1s60%1853.0s
Potato Transfer44%368.7s52%20.7s80%74.2s48%29.6s76%50.2s64%547.3s
Towel Folding56%766.2s60%36.4s100%144.8s76%61.4s96%140.4s68%1100.0s
Towel Bagging24%1761.1s32%130.2s36%266.5s44%121.6s60%171.0s40%2536.7s
Lidded Box Packing8%989.6s20%107.5s92%246.7s52%63.1s100%130.2s60%1440.0s
Plant Sorting24%450.2s52%88.8s56%157.1s52%76.6s72%128.9s32%693.1s
Towel Box Packing44%1332.4s52%36.4s60%83.3s80%61.4s84%112.3s68%1930.5s
Average30.9%977.4s45.7%77.7s68.6%179.9s59.4%66.1s80.0%122.9s56.0%1442.9s
Table 2

Long-horizon adaptation

Success (%) over three rounds.

Training stageTwo-Towel Box PackingTowels-and-Cable Bagging
Initial Policy28%8%
AutoIntervene R144%20%
AutoIntervene R264%28%
AutoIntervene R388%48%
Additional Full Data52%28%
Table 3

Action-head compatibility

Two-Towel Box Packing success (%).

Training stageDPFMACT
Initial Policy32%32%28%
AutoIntervene R140%36%44%
AutoIntervene R256%48%64%
AutoIntervene R392%80%88%
Additional Full Data44%40%52%
Table 4

Controlled handoff comparison

Lidded Box Packing; 10 perturbed and 10 nominal rollouts per method. ↑ higher is better; ↓ lower is better.

MethodCut-in Recall ↑Cut-in Precision ↑Cut-out Recall ↑Cut-out Precision ↑False-trigger Rate ↓
Prior handoff monitors
LazyDAgger0.401.000.00N/A0.80
RND-DAgger (Wrec=5)0.800.160.000.000.90
RND-DAgger (Wrec=30)0.800.160.750.120.90
Ablations
w/o visual support0.00N/AN/AN/A0.00
w/o action risk0.901.000.560.560.00

Key result

Target intervention where the learner actually needs it.

Across real-world bimanual tasks, targeted intervention trajectories support iterative policy improvement with less operator-control time than collecting additional full demonstrations.

View complete results in the paper →

Citation

BibTeX

@article{autointervene2026,
  author  = {Jinhe Tang and Weiming Zhi},
  title   = {AutoIntervene: Calibrated Intervention for
             Action-Chunking Imitation Learning Policies},
  year    = {2026}
}

Publication venue and identifier will be updated when finalized.