WORLD–ACTION MODELS / ACTIVE VISION

ActiveWAM.

Evidence-Aware Active Vision
for World-Action Models

Learn what to keep.
Decide where to look.

A unified world–action model that connects task-guided history learning with physical head and arm control.

REAL ROBOTPAN +36.2° · TILT +40.2°
REAL HEAD MOTION · ORIGINAL SPEED

A changing view.
A continuing task.

↗
50RoboTwin-AV benchmark tasks
16-DBimanual + pan/tilt actions
3Physical kitchen tasks
Raw RGBAt deployment · inversion during training
01 / THE IDEA

Seeing more
is only half the story.

Moving the camera can reveal a useful relation. It can also move a critical cue out of view.

ActiveWAM treats active vision as a retain–acquire problem: learn which evidence should survive visual changes, then jointly decide how to move the head and hands. A recent, view-aware observation history connects the two.

Read the abstract +

Active vision manipulation requires a policy to control both its camera and its end-effectors, yet camera motion determines which evidence remains visible within finite observation windows. Acquiring a new view can displace task-critical cues, while retaining a view forgoes potentially useful observations. We formulate this as an evidence-aware retain–acquire problem and present ActiveWAM, a unified world–action model that learns observation and manipulation jointly. Training-time inversion constrains a frozen video prior by task-bearing source evidence and visible temporal changes. At deployment, the policy generates bimanual and pan/tilt actions from view-aware raw history and updates context from newly measured RGB observations. Future-video prediction serves as a co-training signal. RoboTwin-AV extends 50 manipulation tasks with executable pan/tilt control and automatically generated demonstrations.

01

Retain the evidence.

Task-guided history inversion preserves source evidence and visible changes while transforming appearance during training.

02

Acquire the next view.

Head motion and bimanual manipulation share one generator. Staying still is a valid observation action.

03

Act on what was seen.

Execute a short action prefix, acquire actual RGB, and rebuild context from the available observation window.

02 / IN ACTION

Follow the task.
Watch the view change.

Explore real RGB, source-aligned head poses, and key moments across three domains.

03 / THE METHOD

One model.
Head and hands together.

Learning to preserve task evidence and learning to control observation, within a unified world–action model.

TRAINING ONLY

Transform appearance. Preserve task evidence.

A frozen Wan video prior transforms each camera’s observed history independently. Task-weighted source preservation and temporal-change constraints guide partial inversion. Raw and accepted transformed histories share the same original action and future-video targets.

FIG. 02 / ARCHITECTURE

Inversion is used during training. Deployment conditions on raw history and returns one joint action trajectory; future-video prediction is a co-training signal.

04 / INSIDE THE VISUAL HISTORY

Change the appearance.
Keep the task in sight.

Explore our recent visualization studies on the six key stages of a real cooking sequence.

ILLUSTRATIVE PREVIEWTask: scramble eggs
Original real-robot keyframe
Source RGB Pour Egg
Designed colored-noise visualization of the same keyframe
Appearance study Level 3 / 5
SourceStronger

These previews use designed RGB perturbations to illustrate a visual target. The levels are display settings, not diffusion timesteps or model measurements.

View the full strength sweep ↗Previously created six-level colored-noise visual design sweep

Archived design study. The original “tau” labels denote synthetic noise multipliers, not the paper’s inversion parameter.

View the model-produced TAVIS prototype ↗Existing TAVIS partial-inversion prototype cache, alternating raw and transformed rows

Archived partial-inversion cache. Raw/transformed pairs are shown for head, left wrist and right wrist, with three history timestamps per row. This is a prototype, distinct from the illustrative real-robot previews above.

RoboTwin appearance studies 5 task sequences +

Design studies from the previous visualization iterations; image post-processing, not checkpoint outputs.

05 / EXPERIMENTS

Evidence across
views and environments.

Comparing manipulation success across visual and camera-pose shifts. Select a benchmark and condition to explore the results.

COMPOUND SHIFT
53.3%

Success under appearance + head-pose shifts.

50 tasks · 3 seeds · 100 episodes per task, seed and condition.

+20.0 pp vs. Fast-WAM
SUCCESS RATE (%) ↑
050100
View exact values & evaluation protocol +
Manuscript-reported success rates

HISTORY MATTERS+11.6 pp

Compound success over the raw-history control (53.3% vs. 41.7%).

OBSERVATION CONTROL1.70 rad

Head travel per episode for joint learned control, compared with 4.10 rad for object tracking.

PHYSICAL KITCHEN30 / 60

Full-task successes across three physical kitchen tasks; 20 attempts per task.

Manuscript-reported results; evaluation protocols and exact values are available above. The gallery presents source demonstrations separately from the scored evaluations.

Citation

Anonymous manuscript · citation draft
@misc{activewam2026,
  title = {ActiveWAM: Evidence-Aware Active Vision
           for World-Action Models},
  author = {Anonymous Authors},
  year = {2026},
  note = {Manuscript}
}

Enlarged research figure