In-context Robot Learning Made Simple:
A Democratized Recipe for Manipulation Tasks

Minxing Li1*, Minghao Han1*, Weizhi Zhao1, Hanwen Wang1, Xiangshuo Liu1, Shuyao Shang1, Jingxiang Zhou1, Mingchao Sun2, Hongyu Pan2, Mu Xu2, Yu Liu2†, Lue Fan1†✎, Zhaoxiang Zhang1✎

1NLPR, Institute of Automation, Chinese Academy of Sciences (CASIA)
2Amap, Alibaba Group

{lue.fan, zhaoxiang.zhang}@ia.ac.cn

*: Equal contribution.  †: Project lead.  ✎: Corresponding authors.

01Task Execution Demo

Execution of the eight real-world evaluation tasks.

02Real-World Experiments

Eight unseen tasks across three difficulty levels.

Evaluation is conducted at three difficulty levels: Easy, where a single executable task is present; Medium, where multiple executable tasks require discrimination; and Hard, where multiple valid execution modes exist.

LevelMethod WeighingWipeTeaDrawer SoapFruitChopsticksShelfAverage
Easy
Fast-WAM0.900.870.830.800.630.900.730.770.80
π0.50.870.970.800.600.730.600.530.670.72
SimpleICL (ours)0.930.760.730.830.730.860.800.830.81
Medium
Fast-WAM0.830.500.630.430.600.830.630.730.65
π0.50.630.460.370.430.570.470.230.570.47
SimpleICL (ours)0.830.660.600.800.670.630.730.630.70
Hard
Fast-WAM0.400.600.270.530.330.570.270.370.42
π0.50.530.330.200.170.470.400.100.300.31
SimpleICL (ours)0.800.670.570.670.630.700.760.600.68

Real-world evaluation on eight unseen tasks. Success rates (%) are reported for each method. Wipe = wiping a plate; Tea = serving tea; Drawer = pulling a drawer; Soap = pressing a soap dispenser; Fruit = placing fruit; Chopsticks = organizing chopsticks and bowls; Shelf = organizing items on a shelf.