More from our group

DY-LUT Depth-Aware YCbCr Lookup Tables for Real-Time Underwater Image Enhancement GitHub → PIC-UIE Predicting Image-Adaptive Corrections for Lightweight Underwater Image Enhancement GitHub →

AquaWAM: A Dynamics-aware World Action Modelfor Underwater Embodied Agents

Cunhao Zhu1, Yifeng Wang2, Dongliang Xu1, Yunzhong Hou2, Yue Yao1, Chi Harold Liu2

1 Shandong University
2 Beijing Institute of Technology

Paper arXiv Code Dataset Env Model

Wreck footage by Jun Ho Lee, Pexels. Not a USIM simulation.

Abstract

World Action Models (WAMs) are becoming increasingly important and useful for embodied intelligence, as they enable robots to anticipate the consequences of candidate actions before interacting with the physical environment. However, underwater robots are usually subject to passive dynamics, such as inertia, buoyancy, hydrodynamic drag, and persistent drift, which can continue to affect the vehicle even after an action is completed. Existing WAMs, which primarily predict action-conditioned visual observations, are not explicitly designed to capture such passive motion dynamics. In this paper, we present AquaWAM, the first World Action Model designed for underwater embodied agents. Instead of predicting future images, AquaWAM models both action-conditioned and passive physical dynamics, including the thruster dead band, the inertial glide that outlasts each command, and ambient currents.

Specifically, it senses through the DVL, IMU, pressure sensor and joint encoders, while cameras supply only semantics for understanding goals and target pose. By modeling compact navigation states rather than high-dimensional visual observations, AquaWAM substantially reduces the model size and computational cost compared with conventional WAMs. Experimentally, AquaWAM achieves a 72.6% task success rate across 20 underwater tasks on the USIM benchmark, outperforming existing methods while making action decisions 2.7× faster than U0 on an NVIDIA Jetson AGX Orin. Our model also remains effective when some onboard sensor measurements are unavailable. For example, without DVL velocity measurements, our method still achieves a 61.6% success rate, compared with 39.4% for U0.

Method

AquaWAM keeps the interface of the VLA baselines: the same two cameras, DVL, IMU, pressure sensor and joint encoders come in, and the same thruster and joint commands go out. What changes is how a command is chosen. A single physical-state predictor forecasts a 19-dimensional vehicle state, or 35 dimensions within reach of an object, two seconds ahead. A disturbance token computed from the last 1.6 s of history captures the ambient current and each thruster’s efficiency. Glide, dead band, buoyancy and drag are learned from demonstrations and task-agnostic play, not hand-coded.

There is no action head. Every half second AquaWAM rolls out 128 candidate commands, scores the predicted futures, and executes only the first 0.5 s of the best. Within reach of an object, the candidates become short thrust pulses followed by a coast, each scored by where its glide ends, so every pulse crosses the dead band. Cameras only supply the goal and the target pose. When the DVL loses bottom lock, the same network estimates velocity from the masked history.

AquaWAM overview: perception, physical state, candidate futures, and the chosen command.

Results

Table 1 of the paper. Succ. is the number of successful trials. SPL and ASD follow the USIM definitions. Green is best, pale green is second-best.

Model Navigation Grasping Transporting Tracking Overall
Succ. ↑SPL ↑ Succ. ↑ASD (s) ↓ Succ. ↑SSR ↑ Succ. ↑ SR ↑
OpenVLA 50/1600.71±0.19 0/480— 0/400.0% 0/20 7.1% (50/700)
π0.5 58/1600.59±0.19 1/48087.3 0/400.0% 0/20 8.4% (59/700)
GR00T N1.5 122/1600.78±0.23 140/480102.2 10/4032.5% 4/20 39.4% (276/700)
X-VLA 85/1600.64±0.24 4/48063.5 0/400.0% 1/20 12.9% (90/700)
SmolVLA 74/1600.64±0.25 7/48094.7 0/405.0% 1/20 11.7% (82/700)
U0 140/1600.75±0.24 195/48099.8 13/4052.5% 14/20 51.7% (362/700)
AquaWAM 151/1600.84±0.21 314/48062.9 23/4065.0% 20/20 72.6% (508/700)

Main results

With full sensing, AquaWAM succeeds in 72.6% of 700 trials, against 51.7% for U0. Grasping rises from 40.6% to 65.4%, transporting from 32.5% to 57.5%, tracking from 70% to 100%, and navigation from 87.5% to 94.4%. Successful grasps take 62.9 s, against 99.8 s for U0. Navigation SPL is 0.84, the highest in the table.

Execution traces for grasping, transport, navigation, and tracking.
Execution traces. Grasping and transport use full sensing. Navigation and tracking lose the DVL where marked and finish on the estimate.

Grasp experiments

This experiment asks where the grasping margin comes from. U0 succeeds in 12 to 20 of 40 trials on each object. AquaWAM succeeds in 20 to 34, and leads on every object. With the benchmark’s own object pose, so that perception plays no part, predicted pulses grasp in 85.2% of 480 trials and a hand-tuned positioning law in 17.9%. Only 24% of that law’s trials enter the closure window. AquaWAM reading the wrist camera, rather than the benchmark pose, grasps in 65.4%.

Computation

On a Jetson AGX Orin, scoring 128 candidates two seconds ahead takes 22 ms, and a pulse decision takes 19 ms. With the perception heads included, a decision costs 79.5 ms in navigation and 95.7 ms in grasping, which is 2.7× faster than U0’s 214 ms. The 2.4 million and 2.6 million parameters are the predictor. Each DINOv2-base perception backbone adds 86 million, and that time is included in the AquaWAM bars.

Successes out of 40 on each grasp task, for the hand-tuned law, U0, AquaWAM, and predicted pulses.
Successes out of 40 on each grasp task. Outer ring: predicted pulses. Then AquaWAM, U0, and the hand-tuned law.
Parameter count and decision time on Jetson AGX Orin.
Computation on Jetson AGX Orin 32 GB. Predictor parameters on the left, decision time on the right.

When the DVL fails

The DVL loses bottom lock partway through every trial and does not recover. U0 falls from 51.7% to 39.4%: navigation from 87.5% to 45.6%, tracking from 70% to 5%, grasping holds at 42.1%. AquaWAM keeps 61.6%: navigation 93.8%, tracking 80%, grasping 54.6%. Transport after a 2.9 m carry is 7.5% against 0%. As a takeover layer for U0, the same estimator lifts U0 from 39.4% to 49.7%.

Success by task family after DVL loss.
Success by family after DVL loss. AquaWAM is marked on navigation, grasping, transport, and tracking.

Extreme cases

The main judges are already near the ceiling, so this experiment tightens the judges and perturbs the world. AquaWAM leads the 1.0 m waypoint judge, 61 of 80 against 53, and ties U0 at 0.5 m, 44 each. With the DVL lost it still leads the 0.5 m gotos, 45 against 28, and the inspections, 39 of 40 against 7. It also leads every perturbed world, including half thrust and turbid water. A thruster that delivers nothing, unseen in play, leaves 4 of 40 gotos against 1 for U0.

Perturbed worldU0AquaWAM Stricter judgeU0AquaWAM
Start heading rotated 90–180°52/8063/80Goto, all waypoints within 1.0 m53/8061/80
Goto endpoint within 0.5 m, not 1 m26/4035/40Goto, all waypoints within 0.5 m44/8044/80
Thruster 3 at 50%, seen in play3/4030/40Goto, 0.5 m waypoints, DVL loss28/8045/80
Thruster 1 at 50%, unseen4/4021/40Goto, all waypoints within 0.25 m25/8024/80
Thruster 1 at 0%, unseen1/404/40Inspect, 1 m ball, 90% coverage36/4039/40
Turbid water, Jerlov 0.501/209/20Inspect, same judge, DVL loss7/4039/40

USIM-Hard. Left: 20 trials per task in a perturbed world. Right: stricter judges re-score the main trials, and the DVL-loss rows re-score the dropout trials.

Underwater Scenarios

Pick a scene below. Cards marked 7 models play all seven policies on the same task and episode index; each panel freezes when its run ends, at 2× until the last success and 16× after (scan uses 8×). Cards marked AquaWAM show AquaWAM alone at 2×, on episodes not recorded for the other models.

Pick the blue block, shallow water AquaWAM closes at 19 s. U0, GR00T N1.5, π0.5, X-VLA, SmolVLA and OpenVLA all reach the 180 s timeout without a grasp.

Selected episodes, not the benchmark average. Success rates are in Tables 1 and 6. A shared episode index is not claimed to be an identical rendered first frame.

Acknowledgement

AquaWAM is evaluated on the USIM benchmark and compared with U0. We thank Junwen Gu, Zhiheng Wu, and coauthors for USIM and U0 and the U0 dataset.

BibTeX

@misc{zhu2026aquawamdynamicsawareworldaction,
      title={AquaWAM: A Dynamics-aware World Action Model for Underwater Embodied Agents},
      author={Cunhao Zhu and Yifeng Wang and Dongliang Xu and Yunzhong Hou and Yue Yao and Chi Harold Liu},
      year={2026},
      eprint={2609.33299},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2609.33299},
}