๐ Highlights
ASENA reframes embodied navigation as programmable robot interaction: coding agents write programs, call sensing and motion tools, inspect execution records, repair failures, and retain reusable notes and skills while keeping model weights fixed.
The optional ASENA-VLN-4B policy gives the agent a learned navigation tool that predicts body-frame waypoints from language and recent monocular observations, complementing geometry and programmatic route planning.
The strongest reported results are broad rather than single-metric: 68.7% SR on R2R and 70.2% SR on RxR for standalone ASENA-VLN-4B, state-of-the-art object-goal navigation rows in HM3D/HM3D-OVON, and EQA accuracy up to 81.2% on HM-EQA in the reported evolved high-effort setting.
- Core claim: fixed-weight embodied agents can improve by evolving persistent workspaces, not only by updating model weights.
- Strongest builder takeaway: learned navigation is most useful as a callable tool inside a broader coding-agent system.
- Best-supported efficiency signal: ASENA-VLN reduces mean tool calls for every coding backend and benchmark reported in Table 3.
- Main caveat: real-world tasks still take 10-20 minutes and require supervised execution with safeguards and operator approval.
๐ฏ Introduction
The paper starts from a practical robotics gap: useful navigation robots must do more than reach a destination. The main-body example asks a robot to check whether a vending machine contains an item and return with an answer, which requires search, recognition, memory, communication, and contingent next-step decisions.
Prior learned navigation policies provide movement capability, but the paper argues that many embodied tasks require inspecting new observations, computing over geometry, deciding what to communicate, and revising procedures beyond predefined skills. Coding agents are positioned as a way to synthesize these procedures during a mission if they have programmatic access to robot sensing, computation, simulation, and action.
ASENA's objective is to connect these pieces into an inspectable, supervised, persistent embodied-agent system. The paper contributes the ASENA system, the ASENA-VLN 4B monocular policy, and evaluations showing that fixed-weight workspace evolution improves navigation and EQA performance in simulation, plus supervised real-world Unitree G1 demonstrations without a pre-built map.
๐ฌ Methodology
ASENA gives a coding agent a unified write-run-inspect workflow over robot sensing, computation, visualization, simulation, and action. On G1, the exposed signals include a 360-degree camera, LiDAR, odometry, body IMU, and joint states. Generated programs can process and visualize these inputs, estimate geometry, inspect objects, and request robot actions through ASENA's execution interface.
The system treats physical safety as independent of program generation. Gestures are validated online in MuJoCo for pose continuity, joint and motion limits, and sampled self-collision. Navigation uses LiDAR-based occupancy maps for route planning and obstacle-clearance checks, with additional safeguards for motion limits, stale odometry, low battery, heartbeat-triggered stops, manual emergency halts, and operator approval before physical execution.
Retained experience is formalized as a workspace revision process. The workspace contains notes and executable skills, and a fixed-weight improver updates it from execution traces and feedback. In simulation, private workers may edit workspaces between passes, writes are disabled during frozen evaluation, and all model, policy, and controller weights remain fixed.
ASENA-VLN is the learned navigation component. It extends Qwen3-VL-4B-Instruct with trajectory and pixel-goal token vocabularies, consumes up to eight front-view frames plus a structured JSON instruction, predicts up to eight cumulative body-frame waypoints, and is trained on a 4.3M-sample mixture with a 3:1 balance of trajectory prediction and spatial VQA. The training mix includes atomic tasks, short-horizon ObjectNav, and VLN data from R2R, RxR, ScaleVLN, and SRDF.
workspace representation
Meaning: This defines the persistent workspace at revision k as notes N_k and executable skills S_k, separating textual memory from reusable code.
equation 1
Meaning: This is the workspace revision rule: a fixed-weight improver updates the workspace from the current workspace, execution traces, and feedback while model and controller weights remain fixed.
trajectory waypoint representation
Meaning: ASENA-VLN represents each predicted waypoint as a body-frame motion target with planar displacement and heading.
pixel goal representation
Meaning: The policy's auxiliary pixel-goal tokens encode image-space goal location, heading, and whether the goal is terminal; the paper notes evaluations use trajectory-token predictions rather than pixel goals.
๐ Experiments
The experiments cover standalone VLN policy evaluation, object-goal navigation, agentic navigation before self-evolution, retained-experience navigation, embodied question answering, and real-world Unitree G1 missions. The central datasets and benchmarks named in the main body are R2R, RxR, HM3D-v2, HM3D-OVON, HM-EQA, MT-HM3D, MP3D, ScaleVLN, and SRDF. For R2R/RxR standalone policy evaluation, NE is final goal distance in meters, OSR measures entering the goal region, SR requires stopping there, SPL measures path efficiency, and nDTW measures reference-route similarity.
As a standalone policy on R2R and RxR val_unseen, ASENA-VLN-4B reports R2R NE 3.52, OSR 72.6, SR 68.7, SPL 64.2 and RxR NE 3.90, nDTW 73.1, SR 70.2, SPL 59.7. In object-goal navigation, ASENA reports 80.8/38.1 SR/SPL on HM3D v2 val and 64.9/38.4 on HM3D-OVON unseen. The paper attributes this setting to pairing ASENA-VLN's local atomic navigation with a separate VLM that plans search locations and updates the plan from evidence.
Table 3 is the main agent-policy complementarity result. On R2R Agentic Split, ASENA with Claude Sonnet 5 improves from SR 61.0 and SPL 34.2 without the VLN tool to SR 72.0 and SPL 54.9 with it, while calls per episode fall from 83.32 to 33.96. On RxR Agentic Split, the same backend improves from SR 50.0 and SPL 31.9 to SR 61.0 and SPL 45.5, with calls per episode falling from 103.83 to 38.91. GPT-6 Astra is stronger without the policy, reaching R2R SR 87.0 and RxR SR 78.0, and the paper notes that enabling the policy slightly reduces its RxR success rate while still reducing calls.
Workspace evolution uses recurring tasks, scenes, instructions, and initial poses. Sixteen Claude Sonnet 5 workers execute tasks and may retain programs, notes, and prior action sequences in private workspaces; an improver updates the canonical workspace between passes. Figure 4 reports that SR, OSR, and SPL improve across evolution passes while cost generally decreases, and that VLN-equipped workspaces maintain higher SPL. EQA evaluations on HM-EQA and MT-HM3D report answer accuracy and normalized exploration steps; evolved pass 17 high reaches 81.2% HM-EQA accuracy, while pass 4 low reaches 66.5% MT-HM3D accuracy.
tab:vln-policy
Caption: Standalone monocular VLN-CE policy results on R2R and RxR val\_unseen. ASENA-VLN-4B achieves a state-of-the-art result compared to prior work, with significantly lower navigation error and higher nDTW.
| l*8Y | R2R | RxR | ||||||
|---|---|---|---|---|---|---|---|---|
| Method | NE $ $ | OSR $ $ | SR $ $ | SPL $ $ | NE $ $ | nDTW $ $ | SR $ $ | SPL $ $ |
| NaVid~zhang2024navid | 5.72 | 49.2 | 41.9 | 36.5 | 5.72 | -- | 45.7 | 38.2 |
| Uni-NaVid~zhang2025uninavid | 5.58 | 53.3 | 47.0 | 42.7 | 6.24 | -- | 48.7 | 40.9 |
| NaVILA~cheng2025navila | 5.22 | 62.5 | 54.0 | 49.0 | 6.77 | 58.8 | 49.3 | 44.0 |
| StreamVLN~wei2025streamvln | 4.98 | 64.2 | 56.9 | 51.9 | 6.22 | 61.9 | 52.9 | 46.0 |
| NaVIDA~zhu2026navida | 4.32 | 69.5 | 61.4 | 54.7 | 5.23 | 67.0 | 57.4 | 49.6 |
| JanusVLN~zeng2026janusvln | 4.78 | 65.2 | 60.5 | 56.8 | 6.06 | 62.1 | 56.2 | 47.5 |
| DecoVLN~xin2026decovln | 5.01 | 63.5 | 56.3 | 50.5 | 5.73 | 63.5 | 54.2 | 46.3 |
| ActiveVLN~zhang2026activevln | 5.31 | 58.3 | 50.1 | 43.7 | 5.84 | 58.1 | 50.7 | 41.2 |
| Qwen-VLA-Instruct~wang2026qwenvla | 5.10 | 69.0 | 57.5 | 51.2 | 5.80 | 57.1 | 59.6 | 47.8 |
| InternVLA-N1 (S2, SPF)~cai2025internvlan1 | 4.25 | 68.3 | 60.9 | 55.2 | 5.71 | 46.8 | 63.5 | 55.0 |
| RynnBrain-Nav-8B (SPF)~dang2026rynnbrain | 4.92 | 71.6 | 58.6 | 49.6 | 6.20 | 59.6 | 56.1 | 49.6 |
| DualVLN~wei2025dualvln | 4.05 | 70.7 | 64.3 | 58.5 | 4.58 | 70.0 | 61.4 | 51.8 |
| InternVLA-N1~cai2025internvlan1 | 4.83 | 63.3 | 58.2 | 54.0 | 5.91 | 65.3 | 53.5 | 46.1 |
| Qwen-RobotNav-4B~qwen2026robotnav | 4.22 | 73.6 | 66.9 | 60.5 | 4.15 | 68.6 | 71.3 | 61.5 |
| asenashade ASENA-VLN-4B | 3.52 | 72.6 | 68.7 | 64.2 | 3.90 | 73.1 | 70.2 | 59.7 |
Why It Matters: This is the core standalone navigation-policy comparison, showing ASENA-VLN-4B's reported R2R and RxR val_unseen performance against prior VLN methods.
tab:objectnav
Caption: Object-goal navigation: paired SR/SPL (\%, both $ $). $^ $ denotes HM3D v1 result, where emphasis excludes these.
| l*4Y | HM3D | HM3D-OVON | ||
|---|---|---|---|---|
| Method | v2 val | seen | syn. | unseen |
| VLFM | 63.6/32.5 | 35.2/18.6 | 32.4/17.3 | 35.2/19.6 |
| OpenFMNav$^ $ | 52.5/24.1 | -- | -- | -- |
| SG-Nav | 49.6/25.5 | -- | -- | -- |
| TriHelper$^ $ | 56.5/25.3 | -- | -- | -- |
| WMNav$^ $ | 58.1/31.2 | -- | -- | -- |
| CogNav$^ $ | 72.5/26.2 | -- | -- | -- |
| Uni-NaVid$^ $ | 73.7/37.1 | 41.3/21.1 | 43.9/21.8 | 39.5/19.8 |
| DAgRL+OD | -- | 38.5/21.1 | 39.0/21.4 | 37.1/19.8 |
| MTU3D | -- | 55.0/23.6 | 45.0/14.7 | 40.8/12.1 |
| NavFoM | -- | 40.1/27.1 | 45.4/32.6 | 45.2/ 31.9 |
| ABot-N0 | -- | 55.3/ 32.1 | 55.4/ 33.2 | 54.0/30.5 |
| ApexNav | 76.2/ 38.0 | -- | -- | -- |
| Qwen-RobotNav-4B | 75.6/30.6 | 57.7/24.4 | 60.1/25.1 | 53.1/20.9 |
| Qwen-RobotNav-8B | 71.2/33.0 | 56.1/28.5 | 57.8/28.8 | 51.2/24.0 |
| asenashade | 80.8/38.1 | 60.5/33.5 | 65.8/37.9 | 64.9/38.4 |
Why It Matters: This table tests whether ASENA's navigation stack can support object-goal search rather than only route following.
tab:navigation
Caption: Navigation before self-evolution on R2R and RxR Agentic Splits. NE is in meters and OSR/SR/SPL are percentages. Calls/ep. is mean tool calls over 100 episodes. Green percentages show reductions from the same backend without VLN.
| lc*4Yrl*3Yrl | R2R Agentic Split | RxR Agentic Split | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method / coding agent | VLN tool | NE $ $ | OSR $ $ | SR $ $ | SPL $ $ | Calls/ep. $ $ | NE $ $ | SR $ $ | SPL $ $ | Calls/ep. $ $ | ||
| InstructNav~long2024instructnav | -- | 6.89 | 47.0 | 31.0 | 24.0 | -- | -- | -- | -- | -- | ||
| Open-Nav~qiao2025opennav | -- | 6.70 | 23.0 | 19.0 | 16.1 | -- | -- | -- | -- | -- | ||
| CA-Nav~chen2024canav | -- | 7.58 | 48.0 | 25.3 | 10.8 | -- | 10.4 | 19.0 | 6.0 | -- | ||
| GC-VLN~yin2025gcvln | -- | 7.30 | 41.8 | 33.6 | 16.3 | -- | 8.80 | 33.8 | 13.8 | -- | ||
| Three-Step Nav~zheng2026threestepnav | -- | 5.87 | 39.0 | 34.0 | 29.1 | -- | 9.21 | 22.0 | 16.1 | -- | ||
| Uni-LaViRA~ding2026unilavira | -- | 3.66 | 73.7 | 60.7 | 47.7 | -- | 6.48 | 51.3 | 34.0 | -- | ||
| Minimal ( , SDK)~zhou2026embodied | -- | 5.80 | 61.3 | 51.3 | 37.8 | -- | -- | -- | -- | -- | ||
| asenashade ASENA-VLN-4B | -- | 3.16 | 82.0 | 78.0 | 70.7 | -- | 4.54 | 62.0 | 52.9 | -- | ||
| 0.3pt0pt0pt asenashade ( ) | $ $ | 5.23 | 70.0 | 61.0 | 34.2 | 83.32 | 5.80 | 50.0 | 31.9 | 103.83 | ||
| asenashade ( ) | $ $ | 3.31 | 83.0 | 72.0 | 54.9 | 33.96 | 59.2 | 4.94 | 61.0 | 45.5 | 38.91 | 62.5 |
| 0.3pt0pt0pt asenashade ( ) | $ $ | 1.84 | 90.0 | 87.0 | 72.4 | 26.37 | 0.82 | 93.0 | 78.0 | 40.80 | ||
| asenashade ( ) | $ $ | 2.18 | 89.0 | 88.0 | 72.8 | 14.00 | 46.9 | 1.65 | 86.0 | 64.0 | 19.98 | 51.0 |
Why It Matters: This table separates the contributions of standalone ASENA-VLN, coding backends, and the optional VLN tool before any self-evolution.
fig:overview

Caption: connects coding agents to physical robots. The system integrates robot sensing, a persistent workspace, supervised execution, and an optional learned navigation policy. Each run is stored as an MCAP record~foxglove2024mcap that aligns sensor and robot states with actions, operator feedback, code, and visualizations for inspection and replay. The agent reflects on these records to update its notes and skills while robot motion is disabled.
Why It Matters: It gives builders the system-level contract: coding agent, robot interfaces, retained workspace, supervised execution, execution records, and optional learned navigation.
fig:policy-architecture

Caption: Navigation policy architecture and atomic training examples. Left: language and monocular history enter a shared autoregressive decoder. Separate token blocks encode trajectories and pixel goals; visual QA uses text tokens. Motion tokens decode to body-frame waypoints. Architecture glyphs are schematic. Right: recorded training observations with instructions and overlaid trajectories. The overlays do not use metric camera projection. More examples appear in fig:atomic-training-gallery.
Why It Matters: It shows how the navigation policy turns recent visual history and language into reusable waypoint-producing navigation behavior.
fig:navevolve

Caption: Navigation workspace self-evolution. Success-related metrics (SR, OSR, and SPL) improve across evolution passes, while the cost per pass generally decreases. Access to the VLN tool consistently improves SPL, enabling the agent to reach destinations more efficiently.
Why It Matters: It is the main visual evidence that retained workspace evolution improves success, path efficiency, and cost over repeated navigation passes.
fig:robot

Caption: Real-world navigation and interaction on the G1 robot. Three supervised tasks pair online maps and recorded trajectories with observations and responses. Blue and dashed orange traces show search and return. Distances are estimated from drive odometry. Prompts and responses are condensed for display.
Why It Matters: It grounds the simulated agentic-navigation claims in real supervised G1 robot tasks involving search, return, inspection, and interaction.
๐ฎ Conclusion
The paper concludes that ASENA unifies coding-agent programming, supervised execution, and persistent skills for embodied navigation and interaction. The main practical finding is that recurring-task performance can improve with fixed model weights through retained workspace evolution.
ASENA-VLN is not presented as a replacement for agentic reasoning; it is a specialized motion tool. The evidence suggests it helps weaker coding backends most directly and consistently reduces repeated agent interaction, while stronger coding agents may already navigate effectively with geometric and programmatic tools.
The real-world G1 demonstrations extend the paper beyond route following: the robot searches for unseen targets, builds online maps, inspects objects, remembers user locations, returns, speaks, and synthesizes gestures after validation.
๐ ๏ธ Future Research Improvements
The paper directly identifies faster reasoning and execution as future work because real-world tasks take 10-20 minutes. For builders, that points to caching, shorter tool loops, better route priors, faster perception, and lower-latency supervision interfaces as practical next targets.
The authors also state that future work will strengthen safeguards for dynamic environments. That is important because the current safeguard stack is well described for validation, limits, odometry, battery, heartbeat stops, emergency halts, and operator approval, but dynamic human-populated deployment remains a harder operating regime.
A third stated direction is extending workspace evolution to learned policy updates. The current paper keeps model, policy, and controller weights fixed, so the next research question is how to preserve inspectability and safety while allowing the learned navigation tool itself to improve.
- Measure whether workspace evolution transfers to new scenes, new instructions, and new task distributions.
- Add stronger latency reporting for coding-agent calls, motion execution, simulator validation, and human approval.
- Evaluate dynamic-obstacle and human-interaction scenarios before treating the system as deployment-ready.
- Study policy-update loops separately from workspace-update loops to avoid mixing memory gains with model-learning gains.
๐ญ Potential Industry Use Scenarios
ASENA's most credible near-term industry use is supervised mobile service robotics where tasks combine navigation, search, inspection, memory, and user response. The G1 examples map naturally to facilities work: finding restrooms, counting machines in a pantry, checking a vending machine against a reference item, returning to the user, and reporting verbally or through gesture.
The architecture also fits robotics R&D and operations teams that need inspectable autonomy rather than opaque end-to-end policies. Execution records, code, visualizations, feedback, and retained skills give engineers artifacts they can audit, replay, repair, and promote into reusable capabilities.
A more speculative use is enterprise robot skill libraries that evolve from repeated supervised tasks. The paper supports the idea that recurring task performance can improve through persistent notes and executable skills, but it does not yet prove broad unsupervised deployment because physical execution remains supervised and real-world runs are slow.
- Facilities inspection and wayfinding robots that search, verify, and return with directions or answers.
- Warehouse or office robots that combine object search with visual confirmation and user communication.
- Robotics development platforms where coding agents generate, test, and retain task programs under safety gates.
- Simulated training and evaluation loops for persistent robot workspaces before promoting skills to hardware.
๐ฌ Critical Analysis
The strongest aspect of ASENA is the system framing. The paper does not overclaim that a single learned policy solves embodied autonomy; instead, it composes coding agents, geometric reasoning, learned navigation, workspace memory, execution records, and independent safety checks. That is a realistic blueprint for builders.
The numeric evidence is broad and useful, but some conclusions are conditional. The optional VLN tool clearly helps Claude Sonnet 5 in Table 3 and reduces calls for both reported coding backends, yet GPT-6 Astra already performs strongly without it and loses RxR success when the policy is enabled. This suggests learned navigation is a backend- and integration-dependent accelerator, not a universal win.
Workspace evolution is compelling but should be read carefully. The navigation evolution protocol repeatedly presents the same tasks, scenes, instructions, and initial poses, so the reported gains are strongest evidence for recurring-task improvement and retained procedural memory. They are weaker evidence for open-ended generalization unless future evaluations separate transfer from repetition.
Reproducibility will depend on access to the exact workspaces, prompts, task splits, feedback signals, safety wrappers, and policy-training data. The paper's main body is unusually concrete about metrics, tables, and system components, but deployment builders still need to verify code availability, runtime cost, latency, and failure handling before treating ASENA as a production recipe.
- Strength: strong integration of coding agents with replayable robot execution evidence.
- Strength: clear separation between fixed-weight workspace evolution and learned policy behavior.
- Weakness: real-world autonomy is supervised and slow in the reported demonstrations.
- Open question: how much of the self-evolution gain transfers beyond recurring tasks.