World-Action Model · UAV Navigation · Predictive Video Representation

DiffWAM: A Fast and Efficient Navigation World Action Model

Mo Zhu1,2,† · Yuze Wu1,2,†,* · Xijie Huang1,2,† · Xiao Cui2 · Fei Gao1,2 · Xin Zhou2,*

1 Zhejiang University · 2 Differential Robotics

† These authors contributed equally to this work.

* Corresponding author. E-mail: wuyuze000@zju.edu.cn (Yuze Wu), iszhouxin@zju.edu.cn (Xin Zhou)

DiffWAM predictive-to-motion overview: synthetic dataset, frozen world model, generalization verification and deployment efficiency

Abstract: pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be recovered directly from the predictive representations of a frozen video model. To this end, we present DiffWAM, a geometry-conditioned navigation world-action model that directly transforms multi-level predictive features into continuous camera trajectories. Its Grid-Motion module preserves spatial-temporal motion associations, while Latent2Pose grounds them with first-frame geometry to recover metrically meaningful 3D motion. Complete video rollouts and geometric reconstruction are required only for offline supervision, eliminating future-video decoding and multi-frame reconstruction during deployment. We further introduce FastDreamer, which overlaps predictive and geometric computation with ongoing flight and performs timestamp-aware asynchronous trajectory handoff for continuous UAV execution. DiffWAM achieves a trajectory RMSE of 0.3492 m and an endpoint success rate of 74.40% on the 1,000-sample DiffWAM-1000 benchmark, while representative real-world experiments demonstrate complex behaviors including constrained traversal, orbiting, S-shaped flight, and multi-stage navigation. An onboard DiffWAM-Flash implementation further reaches 1.08 s model-pipeline latency on NVIDIA Jetson AGX Thor. These results demonstrate that predictive video representations can be efficiently grounded into continuous 3D motion, providing a direct alternative to generate-then-reconstruct navigation pipelines.

Keywords: World-action models · Predictive video representations · Trajectory grounding · Geometric distillation · 3D navigation
74.40%DiffWAM-1000 benchmark success rate (1 m endpoint criterion; 1st among baselines)
5 sPrediction horizon · 38 future poses + initial pose = 39
3.52–4.31×DiffWAM-Flash vs DiffWAM P50 inference speed-up (2× / 8× H20)
>100%Overall inference speed-up of FastDreamer (per the paper's introduction)
01 · OVERVIEW

How much motion actually lives in the representations of a generative video model?

Embodied navigation usually asks “where to go and what to do next”, yet turning complex language instructions into continuous, executable trajectories that respect direction, clearance and spatial-scale constraints has long been a weakness of VLA / VLN systems. DiffWAM narrows this question: before the future video is fully generated, are the generator's intermediate representations already sufficient to recover the underlying navigation motion?

Q1At which denoising stage does decodable motion start to emerge?

Early predictive features are not explicit pose encodings — scene appearance, camera motion, object motion, language conditioning and future uncertainty are entangled. DiffWAM reads H3 layers 15 / 25 / 35 in a single truncated forward, avoiding the full video-generation schedule.

taps 15 · 25 · 35channel 5,376 → 51239 video-time slots

Q2Which depths and representations help most for action generation?

Different network depths encode different types of information with uneven contributions to motion reconstruction. DiffWAM fuses the three layers with softmax weights, followed by frozen factorized spatiotemporal blocks, letting the fusion and the decoder learn each layer's role from data.

softmax fusion4× factorized spatiotemporal blocks

Q3How to map efficiently to trajectories with geometric and metric scale?

Motion inferred purely from implicit representations may lack physical scale. DiffWAM uses frozen MoGe2 to estimate first-frame geometry (8×14 grid, median-depth scale s₀), grounding the generated features into metric poses in the initial-camera frame.

MoGe2 · 8×14 gridscene scale s₀ ∈ [0.1, 100] m

STwo-stage training: read out first, then distill

S1 navigation-aware pretraining establishes a stable representation→motion map; S2 geometry-guided OPD pairs early states with full-video reconstructed trajectories from the same rollout. The teacher (Pi3 reconstruction + MoGe2 scale calibration) sees only complete videos; the student uses only early features and first-frame geometry at deployment — asymmetric supervision keeps reconstruction out of the deployment path.

S0 backbone distillationS1 readout pretrainS2 OPD
02 · CONTRIBUTIONS

Three contributions

Moving from a “generate-then-reconstruct” paradigm to single-forward decoding, with evidence from controlled studies and a latency-aware execution framework.

01
LATENT2POSE

Single-forward, geometry-conditioned Latent2Pose

Grounds camera trajectories directly in the initial predictive state of a frozen video generator: a spatiotemporal decoder preserves motion evidence across locations and time, while first-frame geometry supplies spatial reference and estimated scale. One truncated generator computation predicts a five-second pose sequence — no full video generation and no multi-frame reconstruction at deployment.

38 future poses · 8 transformer blocks
02
CONTROLLED STUDY

A controlled study of geometric supervision and generated-data scaling

On a common 500-request development-set protocol with a fixed update budget, scaling generated supervision from 1,000 to 64,000 pairs reduces position RMSE from 2.1608 to 1.4620 estimated meters; spatial geometry improves positional agreement, while additional teacher constraints show mixed effects across tasks and metrics.

RMSE 2.1608 → 1.4620 m
03
FASTDREAMER

Latency-aware integration for continuous UAV execution

Predictive features and first-frame geometry are processed in parallel, with timestamp-aware asynchronous trajectory handoff: the next motion proposal is prepared while the UAV executes the current committed trajectory, addressing both proposal-computation latency and the reference-frame mismatch caused by motion during inference. Overall inference speed improves by more than 100%.

>100% inference speed-up
03 · METHOD

DiffWAM: grounding predictive representations into motion

Given an initial RGB observation and a language instruction, DiffWAM performs one truncated forward in the generator's initial predictive state and extracts multi-level features; Grid-Motion establishes timestamp- and geometry-conditioned anchor associations before memory pooling; Latent2Pose combines the motion memory with first-frame geometry and outputs a time-indexed continuous camera-pose sequence.

DiffWAM overview: vision-language inputs, predictive-to-motion decoding, continuous 3D motion output, two-stage training

DiffWAM overview. An initial RGB observation and a language instruction condition a frozen video model; one truncated forward yields predictive features, decoded by Grid-Motion and the geometry-conditioned Latent2Pose into a camera-trajectory proposal, with first-frame geometry providing spatial reference and scale. Full video generation and reconstruction are used only by the offline teacher (S0–S2). (Fig. 1 · DiffWAM_Pipeline)

τ̂Problem formulation

The five-second horizon is represented as 38 camera poses (orientation + position) in the initial-camera frame; with the identity initial pose, 39 poses in total. Poses are motion proposals decoded from the generator’s predictive state, not verified collision-free commands — interpolation and executable-trajectory construction are left to downstream execution.

ℐEfficient predictive readout

The generator’s iterative generation schedule is never completed: one truncated call returns multi-level predictive features (layers 15 / 25 / 35, 5,376 → 512 dims) and clean-observation anchors, while a frozen geometry estimator (MoGe2) reads first-frame geometry from the same initial image. Deployment needs no future-video decoding and no multi-frame reconstruction.

MGrid-Motion & Latent2Pose

Grid-Motion builds timestamp- and geometry-conditioned anchor associations before pooling, turning 39 video-time slots into 312 motion-memory tokens. Latent2Pose then combines this memory with first-frame geometry (scene scale s0 from the median valid depth) and decodes a continuous, metrically scaled SE(3) trajectory with 38 learned pose queries.

LGeometry-guided distillation

Training is asymmetric: the offline teacher reconstructs the complete generated video (Pi3) and calibrates scale with MoGe2 depth; the student sees only early features and first-frame geometry. A scale-normalized position loss plus an SO(3) geodesic rotation loss trains the readout — full video generation and reconstruction stay out of the deployment path.

DiffWAM internal architecture: multi-level features, spatiotemporal factorization, Grid-Motion, trajectory decoder and offline teacher

DiffWAM internal architecture. Multi-level predictive features (layers 15 / 25 / 35, 5,376 → 512 dims) are fused and processed by factorized spatiotemporal modules; Grid-Motion builds geometry-conditioned associations before pooling 39 video-time slots into 312 motion-memory tokens; the direct decoder outputs 38 future poses with position / rotation heads. An offline teacher reconstructs complete videos (Pi3) and calibrates scale with MoGe2. (Fig. 2 · DiffWAM_Internal_model_architecture)

04 · FASTDREAMER

Hiding inference latency inside the execution time of committed trajectories

Two problems in closed-loop execution: an extra resident rewrite model strains on-device memory, and the UAV keeps moving while a trajectory is being prepared. FastDreamer answers both with shared-weight rewriting, parallel proposal inference, flight-time budgeting and prospective handoff.

QShared-weight prompt rewriting

Reuses the vision-language parameters Qω already resident in the world model's conditioning branch to rewrite task-level instruction + current observation + task context hn into a local motion condition cn, avoiding a second large resident model for rewriting. Semantic correctness is not guaranteed; progress maintenance is the task manager's job.

REWRITING
$$c_n=\mathcal{Q}_{\omega}(I_n,\mathcal{I},h_n)$$

∥Parallel proposal inference

MoGe2 does not depend on the rewritten condition and can start together with prompt rewriting. The ideal critical path is input processing + the slower of the two branches + trajectory head + output delivery. Concurrent branches may contend for GPU compute and bandwidth; actual latency must be measured in joint running configurations.

CRITICAL PATH
$$L_{\mathrm{proposal}}^{\mathrm{ideal}}=L_{\mathrm{input}}+\max(L_{\mathrm{pred}},L_{\mathrm{geo}})+L_{\mathrm{head}}+L_{\mathrm{output}}$$

BFlight-time budgeting

The remaining execution horizon of the current committed trajectory serves as the preparation budget: estimated preparation latency L̂n, uncertainty margin Mn and fallback margin Rn must sum to at most Bn = en − un; early results wait for their scheduled activation, late results must be revalidated or rejected. gn = 0 only says the original horizon was not exceeded, not that zero-wait succeeded.

ADMISSION
$$\widehat{L}_n+M_n+R_n\le B_n,\qquad g_n=\max(0,\,d_n-B_n)$$

HProspective trajectory handoff

Proposals are anchored in the camera frame at capture time tn, then lifted through the world frame and transformed by the body extrinsics; the committed motion propagates state to the scheduled handoff time, satisfying the minimal position/velocity continuity conditions. Timestamp rewrites alone cannot repair mismatch; invalid, stale or superseded proposals must not overwrite the current valid plan.

CONTINUITY
$$p_n^{+}(0)=p_n^{-}(t_{h,n}),\qquad v_n^{+}(0)=v_n^{-}(t_{h,n})$$
TIMESTAMPED PREPARATION & SCHEDULED HANDOFF Schematic · not measured latency
Execution Prediction Geometry Track committed trajectory Execute accepted update Prompt rewriting cond. + truncated H3 Pose head Plan + validate MoGe2 + geometry transfer Ready · awaiting handoff Time tₙ uₙ rₙ t_{h,n} eₙ Bₙ = eₙ − uₙ (flight-time budget) Rₙ Fallback decision no later than eₙ − Rₙ

The geometry branch starts together with prompt rewriting; the pose head waits for geometry and predictive features to be ready. Updates accepted in time are prepared before the scheduled handoff th,n; the capture time tn always anchors coordinates. Late or invalid results must be realigned and validated before activation. (Reconstructed from the paper's FastDreamer timing diagram; horizontal spacing is schematic.)

05 · EXPERIMENTS

Benchmarks and trajectory quality: first across all three datasets

The DiffWAM-1000 evaluation set contains 1,000 test samples and 18 task types, split scene-isolated; tasks fall into four families: basic motion, object interaction, spatial navigation and scene understanding. All benchmark results use a unified 1-meter endpoint criterion (FDE ≤ 1 m indoors / ≤ 3 m outdoors).

Success rate across three benchmarks (SR %)

Unified 1 m endpoint criterion; DiffWAM ranks first on IndoorUAV-VLA, UAV-FLOW-Sim and DiffWAM-1000

DiffWAM-1000 across four families vs. the strongest baseline

BM basic motion · OI object interaction · SN spatial navigation · SU scene understanding (baseline = strongest competing method per family)

MethodIndoorUAV-​VLA
avg
UAV-FLOW-​Sim
avg
BMOISNSU · avg
WorldVLN39.3280.2486.0056.6250.7754.00 · 58.40
ImagineUAV33.7869.6568.0039.8532.6756.00 · 43.20
Fast-WAM-UAV35.6971.2773.0052.3438.6751.00 · 52.20
WorldFly25.7153.9864.0024.0828.6241.00 · 30.50
DiffWAM (ours)56.7791.4292.0070.3574.0384.00 · 74.40

Absolute gains over the strongest baseline per family on DiffWAM-1000: +6.00 / +13.73 / +23.26 / +28.00 pp; overall absolute gains across the three datasets: +17.45 / +11.18 / +16.00 pp. (Table 1)

3D navigation task distribution: object interaction 65%, spatial navigation 15%, basic motion 10%, scene understanding 10%

Evaluation task distribution. 18 task types across four families: object interaction 65% (object navigation 30%, orbiting 10%, etc.), spatial navigation 15%, basic motion 10%, scene understanding 10%. Object interaction dominates because language-conditioned target selection and spatial-relation reasoning are at the core of vision-language UAV navigation. (Fig. 3 · uav_flight_task_distribution)

DiffWAM representative trajectory generation: navigation approach, gap crossing, orbiting and landing

Representative trajectory generation. Rows 1–4: target-oriented navigation and approach (approach the bicycle from the right, approach the boat and observe its right side, etc.); rows 5–6: constrained gap crossing (turn left and cross the garage passage, curve left and cross the rooftop frame); rows 7–9: full orbits and direction-conditioned half-orbits; row 10: a generated landing sequence. Each visual sequence is accompanied by a top-view trajectory. These examples are used for qualitative analysis, not for estimating overall success rates. (Fig. 4 · Representative cases)

06 · REAL WORLD

Real UAVs: from single-point tasks to multi-stage combinations

Physical flight experiments indoors and outdoors: standalone motion primitives, language-specified target selection, constrained-space crossing, and multi-stage task combinations. Experiments cover three NCS platforms with both cloud-assisted and fully on-device deployments.

Eight representative real-UAV experiments: orbiting, crossing, S-shaped trajectory, landing and multi-stage tasks

Representative real-UAV experiments. Cases I–IV: orbit the dark blue pillar, pass between the two trees on the left, fly an S-shaped trajectory, pass through the circular frame; Cases V–VI: head toward the wall with the “Differential Robotics” logo, head to the red fire hydrant on the far left; Cases VII–VIII: go around the left side of the tree on the right then land on the yellow mat, pass the left side of the electric fan then reach in front of the rock formation (sequential multi-stage goals). These eight illustrated cases are selected qualitative demonstrations rather than a complete record of evaluation trials; they are not used to infer the number of evaluated episodes, aggregate physical-flight success rates or per-task completion times. (Fig. 5 · real_world)

NCS-α-pro
Cloud-assisted WAM inference
  • NVIDIA Jetson Orin NX · 16 GB memory
  • Language-specified target selection and indoor-outdoor navigation
NCS-β
Cloud-assisted WAM inference
  • Livox Mid-360 LiDAR + Intel RealSense D435 camera
  • Physical flight with a multi-sensor configuration
NCS-Thor-preview
Fully on-device
  • NVIDIA Jetson AGX Thor · 128 GB memory
  • Runs the complete DiffWAM inference pipeline on board
CASE VII
Around tree left side → land on yellow mat

Target disambiguation + specified-side bypass + landing-area approach + descent

CASE VIII
Around fan left side → reach rock formation

Switching target reference while preserving unmet task requirements

CASE V–VI
Outdoor target-oriented navigation

Relative-position constraints for the wall logo and the “far-left” hydrant

CASE I–IV
Indoor structured motion

Orbiting, two-tree crossing, S-shaped flight, circular-frame crossing

Head toward the wall logo

Language-specified target selection: the UAV heads toward the wall with the “Differential Robotics” logo and hovers in front of it.

Pass between the two trees

Constrained traversal: the UAV flies through the narrow space between the two trees on the left.

Fly toward the rock formation

Target-oriented navigation: the UAV flies toward the rock formation in the outdoor scene.

Approach the tallest blue pillar

Object selection: the UAV flies to the front of the tallest blue pillar and stops before it.

Left side of the fire extinguisher

Spatial grounding: the UAV moves to the left side of the fire extinguisher as instructed.

Pass through the circular frame

Precise pass-through: the UAV crosses the circular frame ahead without contact.

Orbit the pillar with avoidance

Orbit with obstacle avoidance: the UAV completes a full loop around the pillar while avoiding obstacles.

S-shaped trajectory flight

Trajectory following: the UAV flies forward along an S-shaped trajectory.

07 · ONBOARD PERFORMANCE

On-device inference: measured latency and memory

Historical FastDreamer pipeline variants measured on 2× / 8× NVIDIA H20 and NVIDIA Jetson AGX Thor (BF16 inference, 15 requests × 3 rounds after warm-up, conditional caching off). Latency is measured from a prepared image and instruction to a usable trajectory, excluding rewriting, communication and file export.

Pipeline latency P50 / P95 (s)

DiffWAM-Flash vs DiffWAM: P50 −71.58% / −76.83% / −39.87% (≈3.52× / 4.31× / 1.66× speed-up)

End-to-end pipeline totals (8× H20, ms)

vs NavDreamer: DiffWAM ≈2.80×, DiffWAM-Flash ≈12.07×; last row Flash @ AGX Thor

Trajectory headPlatform / PrecisionModel P50 / P95 (s)Peak memory (GiB)
DiffWAM2×H20 / BF1611.199 / 11.27376.29 / 149.30
DiffWAM-Flash2×H20 / BF163.182 / 3.24276.29 / 149.30
DiffWAM8×H20 / BF163.605 / 3.82060.33 / 457.67
DiffWAM-Flash8×H20 / BF160.835 / 0.91260.33 / 457.67
DiffWAMAGX-Thor / BF161.796 / 1.85197.60 / 97.60
DiffWAM-FlashAGX-Thor / BF161.080 / 1.10197.60 / 97.60

Flash's advantage is latency, not memory: both variants report identical peak memory on the same platform. Memory is reported in GiB as per-card max / multi-card concurrent total. (Table 3)

Inference pipeline stage-wise latency breakdown: NavDreamer, pure video generation, DiffWAM, DiffWAM-Flash

Stage-wise latency breakdown. On 8× H20: NavDreamer 10,080.2 ms, pure video generation 8,957.8 ms, DiffWAM 3,605.5 ms, DiffWAM-Flash 835.2 ms; Flash on AGX Thor 1,080.4 ms. Components include vision encoding, LLM prefill, input / VAE encode, video predictive compute, VAE decode, overhead and Pi3 reconstruction. NavDreamer keeps future-video decoding and Pi3 reconstruction; DiffWAM variants drop both. (Fig. 7 · h3_latency_breakdown)

08 · ABLATIONS

Ablations: readout architecture, feature depth, compute budget, data scale, objective and geometry

Matched training ablations fix dataset splits, initialization, optimizer, training budget, geometry inputs and the model-selection criterion, changing only the factor under study; the world model and geometry backbone stay frozen. The studies cover the trajectory-readout architecture, world-model feature layers, predictive-computation budget, training-data volume, geometric conditioning and training objectives.

Trajectory-readout architecture (six historical cases)

RMSE / ADE / FDE (est. m) on six historical cases. DiffWAM vs MLP: −74.84% / −83.36%; vs Faster-WAM: −50.19% / −59.63%. RoE: 2.9179° vs Fast-WAM 6.1389° / Faster-WAM 4.5870° / MLP 98.8160°. Trajectory-head P50/P95: 37.35/38.84 ms vs Fast-WAM 59.20/60.97, Faster-WAM 60.37/61.77, WorldVLN 3.14/3.34, MLP 0.22/0.25

World-model feature-layer selection

Same native 15×26 grid, 72M pose-head params. RMSE: mixed 0.3492 / early 0.7250 / middle 0.5397 / deep 0.3790 m. RoE: mixed 2.9179° / early 9.2979° / middle 2.4314° (lowest) / deep 2.5763°

World-model evaluations (1 / 2 / 3 / 4)

RMSE 0.5012 → 0.3492 (−30.33%), ADE −28.76%, FDE −36.53%. RoE (non-monotonic): 3.7897° / 3.1867° / 3.7730° / 2.9179° at 1 / 2 / 3 / 4 evaluations

Generated-supervision scale (DiffWAM-1000)

1k → 64k pairs: RMSE 0.8203 → 0.4873 (−40.59%), FDE 1.4011 → 0.6884 (−50.87%), RoE 9.9409° → 2.7894° (−71.94%)

Compact-loss recipes (1,024 pairs / seed 0 / 600 updates)

Neither simplification nor depth normalization alone helps; selected depth-normalized recipe RMSE −5.69% / FDE −7.78% vs legacy. RoE: legacy 3.5997° / unnorm 4.2803° / depth-norm 3.9662° / selected 3.8131°

LFeature depth matters most for position; orientation ranks differently

Mixed 15/25/35 achieves the lowest positional errors (RMSE 0.3492 m, FDE 0.3151 m); early 3/5/7, middle 23/25/27 and deep 33/35/37 give RMSE 0.7250 / 0.5397 / 0.3790 m — mixing separated depths reduces RMSE by 51.83% / 35.30% / 7.86%. Orientation differs: middle layers reach the lowest RoE 2.4314°, followed by deep 2.5763°, while the mixed configuration is 2.9179°. Multi-depth fusion helps position, but no single combination is optimal for every pose metric.

mixed: 0.3492 / 0.3151early: 0.7250 / 1.0036middle: 0.5397 / 0.7641deep: 0.3790 / 0.5562

SMore generated supervision helps, under a fixed update budget

On DiffWAM-1000 (random init, 20,000 updates, batch 32): 1k → 64k pairs reduce RMSE 0.8203 → 0.4873 m (−40.59%), FDE 1.4011 → 0.6884 m (−50.87%) and RoE 9.9409° → 2.7894° (−71.94%), monotonically. On the common 500-request development set (abstract protocol), the same scaling reduces trajectory RMSE from 2.1608 to 1.4620 m, endpoint error (FDE) from 3.3032 to 2.0652 m and orientation error (RoE) from 11.8227° to 8.3682° under a fixed update budget.

DiffWAM-1000: 0.8203 → 0.4873 m500-request dev set: RMSE 2.1608 → 1.4620 m · FDE 3.3032 → 2.0652 m · RoE 11.82° → 8.37°

GSpatial geometry matters; auxiliary teacher constraints do not

Full-parameter DiffWAM-1000 controls (64k training records): removing auxiliary teacher constraints changes RMSE 1.4620 → 1.4405 m and FDE 2.0652 → 2.0480 m — no measurable benefit in this configuration. Replacing full spatial geometry with scale-only conditioning instead increases RMSE to 1.5105 m and FDE to 2.2030 m, supporting the value of spatial geometric information for trajectory localization.

full: 1.4620 / 2.0652no-aux: 1.4405 / 2.0480scale-only: 1.5105 / 2.2030

DHistorical DEV1000 controls delimit the supported claims

Seed-0 controls on the historical development set (250 object-navigation, 150 fly-through and 100 orbit cases; Pi3 references in MoGe2-estimated meters) show that removing auxiliary losses improves mean reconstruction error under the historical objective protocol, temporal shuffling has little effect on the archived checkpoint, whereas cross-case replacement substantially increases error.

250 obj-nav · 150 fly-through · 100 orbitseed-0 · 64k pairs · 20k updates

Conclusion and boundaries

DiffWAM studies recovering camera motion from the very first predictive computation of a frozen video generator: multi-depth features, Grid-Motion associations and geometry-conditioned Latent2Pose decoding connect predictive motion information with an observed spatial reference and estimated metric scale, while FastDreamer complements it with parallel processing, flight-time computation scheduling and timestamp-aware prospective handoff. Mixed-depth features yield the lowest positional errors, more world-model evaluations improve positional accuracy, and larger generated-supervision sets consistently reduce trajectory errors under a fixed update budget; a tuned two-term objective further improves position in the compact-loss comparison, although the legacy composite objective retains the lowest orientation error.

Stated limitations

Current results measure reconstruction agreement on a development set; independent navigation success and on-device execution still need to be established. The one-meter endpoint criterion does not by itself verify traversal direction, orbit coverage or landing completion; reliable full orbits and physical-scale accuracy are not yet established, and per-task closed-loop completion and complete navigation-update latency remain to be quantified.

10 · BIBTEX

Cite this work

Author list and affiliations are taken from VLA-AN-style.sty (6 authors, Zhejiang University, Differential Robotics); fields that depend on the published record are marked below.

diffwam.bib
@misc{zhu2026diffwam,
  title        = {DiffWAM: A Fast and Efficient Navigation World Action Model},
  author       = {Zhu, Mo and Wu, Yuze and Huang, Xijie and Cui, Xiao and Gao, Fei and Zhou, Xin},
  year         = {2026},
  note         = {Preprint}
}
% arXiv ID and venue to be filled in from the published record

Title, full author list and year are taken from the paper source; add the arXiv ID and venue once the paper is published.

BibTeX copied to clipboard