Accept 7.3/10 cs.RO trendtoknow-paper-summaries codex-pro/gpt-5.5

Benchmarking EMlog Calibration for Autonomous Surface Vehicles

Samuel Cohen-Salmon, Itzik Klein · September 30, 2026 · cs.RO

📌 Highlights

The paper benchmarks EMlog calibration for autonomous surface vehicles using main-body evidence from real MARVEL ASV sea trials, not a simulated-only setup.

The strongest reported model is EM3, a linear bias-and-scale formulation, which outperforms bias-only, scale-only, and second-order polynomial alternatives in the representative comparison.

The central engineering lesson is that maneuver design matters: dynamic figure-8 and circular trajectories improve parameter observability, while straight-line calibration can be rank-deficient.

  • Dataset: 221 minutes of continuous real-world telemetry from MARVEL with two EMlogs and DVL recordings.
  • Best representative result: EMlog1 vs DVL RMSE drops from 0.576 m/s raw to 0.167 m/s with LS-EM3 and 0.170 m/s with KF-EM3.
  • High-excitation calibration maneuvers are reported to improve out-of-sample generalization by 30% compared with straight-line paths.
  • Polynomial EM4 gives no meaningful gain over EM3 in the reported operational speed envelope.
  • No source-backed algorithm block is provided in the main-body extraction.

🎯 Introduction

The task is continuous velocity estimation for ASVs and AUVs when high-precision DVL measurements are unavailable or unreliable. The paper explains that DVL bottom-track availability is bounded by altitude relative to the seafloor: AUVs can lose acoustic bottom-track in deep water, while ASVs in shallow coastal areas can encounter acoustic interference, reverberation, and minimum-altitude violations.

EMlogs provide a robust alternative because they measure speed-through-water using electromagnetic induction, but raw EMlog readings are corrupted by systematic effects such as hydrodynamic flow patterns, boundary-layer disturbances, installation angles, and electronic zero-drift. The paper argues that prior EMlog calibration work is fragmented, with different error models and estimation approaches that are hard to compare under common conditions.

The paper-level objective is therefore a comparative model-based benchmark: evaluate four EMlog error models and two estimation pipelines on a shared real-world dataset, then identify which model-estimator pair gives reliable online calibration under dynamic sea conditions.

marvel_fig

marvel_fig

Caption: The MARVEL platform and its safety boat deployed during dynamic sea trials for sensor data collection.

Why It Matters: It grounds the benchmark in the physical ASV platform used for the sea-trial data collection.

🔬 Methodology

The methodology starts with four physical error models. EM1 models constant additive bias, EM2 models a multiplicative scale factor, EM3 combines bias and scale, and EM4 adds a quadratic hydrodynamic coefficient. The estimation layer then creates four least-squares variants and three Kalman-filter variants, yielding seven total calibration methodologies.

The least-squares pipeline estimates static calibration parameters by minimizing residuals between observations and the model prediction. The Kalman-filter pipeline estimates speed and error states recursively, using an identity transition matrix and observation matrices that encode the selected error model. The paper uses a pseudo-linear operating point where the ground-truth reference speed populates the observation matrix, avoiding an extended Kalman filter.

The identifiability analysis is central: at constant speed, the information matrix for the bias-and-scale model becomes singular, making scale factor and zero-drift bias collinear. Dynamic maneuvers introduce speed variation, satisfy the non-collinearity condition, and make the calibration parameters separable.

The evaluation workflow segments the 221-minute dataset into five trajectories, synchronizes sensors with cross-correlation lag optimization, estimates parameters on one calibration segment, freezes them, and evaluates RMSE improvement on the remaining trajectories.

equation 1

Formula:
\[J(\boldsymbol{\theta}) = \sum_{k=1}^{N} (z_k - \boldsymbol{h}_{k}^{T} \boldsymbol{\theta})^{2} = ||\boldsymbol{z} - \mathbf{H_{LS}}\boldsymbol{\theta}||^{2}\]

Meaning: Defines the least-squares objective minimized to estimate static calibration parameters.

equation 2

Formula:
\[\hat{\boldsymbol{\theta}}_{LS} = (\mathbf{H_{LS}}^T \mathbf{H_{LS}})^{-1} \mathbf{H_{LS}}^T \boldsymbol{Z}\]

Meaning: Gives the closed-form batch solution for the calibration parameter vector when the normal matrix is non-singular.

equation 3

Formula:
\[\boldsymbol{x}_k = \boldsymbol{\Phi}_{k-1} \boldsymbol{x}_{k-1} + \boldsymbol{w}_{k-1}\]

Meaning: Defines the Kalman-filter state transition used to recursively estimate speed and calibration states.

equation 7

Formula:
\[\mathbf{K}_k = \mathbf{P}_{k|k-1} \mathbf{H}_k^T (\mathbf{H}_k \mathbf{P}_{k|k-1} \mathbf{H}_k^T + \mathbf{R})^{-1}\]

Meaning: Computes the Kalman gain that weights the new EMlog measurement innovation against the predicted state.

📊 Experiments

The benchmark uses MARVEL ASV sea trials with two commercial electromagnetic speed logs and a Rowe SeaPILOT DVL. EMlog1 is an airmar DX900+ dual-axis electromagnetic sensor with 1 Hz update rate, while EMlog2 is a ben marine alize sensor with 10 Hz and 5 Hz reported outputs. The DVL operates at 10 Hz and is used as the absolute reference for speed benchmarking.

The dataset is split into five trajectories: transit to experiment location, figure-8 maneuvers, clean or straight lines, circular maneuvers, and transit back. Durations are 125 min, 27 min, 15 min, 24 min, and 30 min respectively, for 221 minutes total. The paper also uses inter-sensor EMlog1-vs-EMlog2 calibration to test relative calibration without the DVL reference.

The main metric is RMSE, where lower is better, and improvement percentage reports relative RMSE reduction against raw EMlog measurements. The consistency metric is NIS, evaluated against the expected scalar level NIS = 1. The main estimator comparison trains on the figure-8 run and evaluates on trajectory 3, while the cross-validation matrix isolates each trajectory as a calibration segment and evaluates the frozen parameters across the remaining trajectories.

The representative results show that LS-EM3 is best for EMlog1 vs DVL with 0.167 m/s RMSE and 71% improvement, while KF-EM3 reaches 0.170 m/s and 70% improvement. EM4 reaches 0.172 m/s and 70%, indicating no useful improvement over the simpler affine EM3 model. Cross-validation shows straight-line calibration performs poorly on circular trajectories, with RMSE 0.507 m/s and -27% improvement, while circles and figure-8 maneuvers generalize more robustly.

rmse_improvement_table_2

Caption: RMSE and improvement - calibration on figure-8 Run (applied to trajectory 3)

Content:
UUT vs GTRAW RMSE [m/s]LS-EM1 RMSE [m/s]LS-EM1 %LS-EM2 RMSE [m/s]LS-EM2 %LS-EM3 RMSE [m/s]LS-EM3 %LS-EM4 RMSE [m/s]LS-EM4 %KF-EM1 RMSE [m/s]KF-EM1 %KF-EM2 RMSE [m/s]KF-EM2 %KF-EM3 RMSE [m/s]KF-EM3 %
EMlog1 vs DVL0.5760.20864%0.27652%0.16771%0.17270%0.20864%0.27552%0.17070%
EMlog2 vs DVL0.3050.11363%0.14253%0.10366%0.10466%0.11363%0.14353%0.10267%
EMlog1 vs EMlog20.3090.19537%0.21929%0.17743%0.17743%0.19935%0.22128%0.18241%

Why It Matters: This table directly compares all seven estimator variants on the representative figure-8-to-straight benchmark.

cv_matrix_table

Caption: Cross-validation :EMlog1 vs DVL calibration and processed trajectories

Content:
Calibration trajectoryDuration [min]Avg. speed [m/s]1 (transit) RMSE [m/s]1 (transit) Imp [%]2 (figure-8) RMSE [m/s]2 (figure-8) Imp [%]3 (straight) RMSE [m/s]3 (straight) Imp [%]4 (circles) RMSE [m/s]4 (circles) Imp [%]5 (return) RMSE [m/s]5 (return) Imp [%]overall RMSE [m/s]overall Imp [%]
1 (transit)1251.950.26454%0.23359%0.19167%0.24040%0.38827%0.27350%
2 (figure-8)271.480.34140%0.18567%0.16771%0.27132%0.32938%0.29146%
3 (straight)151.590.48715%0.21063%0.14774%0.507-27%0.45214%0.42322%
4 (circles)241.850.30946%0.20564%0.20964%0.16559%0.28346%0.26152%
5 (return)301.370.33142%0.22161%0.23160%0.17556%0.28047%0.27749%

Why It Matters: This table shows how calibration path geometry affects generalization across unseen trajectory types.

🔮 Conclusion

The paper concludes that a bias-and-scale EMlog model is the most robust calibration formulation for the examined sensors and operating regime. In the central EMlog1-vs-DVL representative result, EM3 reduces baseline RMSE from 0.576 m/s to 0.167 m/s, yielding a 71% improvement.

The practical conclusion is not just model choice but maneuver choice: high-excitation calibration paths, especially circles and figure-8 paths, are structurally required to raise the information matrix to full rank and separate scale factor from zero-drift bias.

The authors also emphasize operational constraints. Calibration maneuvers impose an energy and path-efficiency trade-off, and EMlog reliability is limited when vehicle speed is close to ambient current speed.

🛠️ Future Research Improvements

A clear next research step is broader platform validation. The main-body evidence comes from MARVEL ASV trials, so the reported calibration behavior should be tested on different hulls, EMlog placements, sea states, and mission speeds before treating the maneuver prescription as universal.

The paper motivates but does not fully exhaust current-aware calibration. Since EMlogs measure water-relative velocity and reliability degrades when ambient current is close to platform speed, future work should evaluate calibration under stronger and more variable currents.

Another useful extension would be adaptive maneuver planning: if KF-EM3 can converge within 15 seconds of dynamic turn entry, an autonomy stack could trigger the minimum excitation needed for identifiable calibration rather than executing full calibration loops.

  • Validate EM3 and KF-EM3 across more ASV and AUV platforms.
  • Quantify performance under stronger ambient currents and lower vehicle speeds.
  • Design calibration maneuvers that minimize energy cost while preserving full-rank observability.
  • Investigate whether higher-order hydrodynamic terms become useful outside the reported operating speed envelope.

🏭 Potential Industry Use Scenarios

The strongest deployment fit is marine autonomy where DVL bottom-track cannot be guaranteed. ASVs operating in shallow coastal environments and AUVs operating beyond bottom-track range could use calibrated EMlogs as continuous velocity aids for dead reckoning and navigation integrity.

The method is also relevant to commercial shipping, offshore inspection, environmental monitoring, and defense operations where EMlogs are already common but calibration is often handled offline or folded into broader navigation filters. The paper’s online KF-EM3 result makes it plausible to integrate calibration directly into a vehicle runtime stack.

For product teams, the key operational feature would be calibration-aware mission planning: schedule a short figure-8 or circular excitation when navigation confidence drops, then freeze or update parameters for subsequent straight-line operation.

  • ASV coastal navigation when DVL readings are disrupted by shallow-water acoustic effects.
  • AUV or ASV dead-reckoning support during DVL outage windows.
  • Fleet calibration workflows for platforms carrying commercial EMlogs.
  • Runtime health monitoring that combines RMSE-style validation with NIS consistency checks.

💬 Critical Analysis

The paper’s strength is that it connects model structure, estimator design, maneuver geometry, and real sensor data. The comparison is builder-facing because it does not simply report that calibration helps; it identifies which error model is worth implementing and when the calibration data are informative enough.

The main technical caveat is external validity. The benchmark is rich for one platform and trial campaign, but the conclusions about EM3 dominance and 15-second convergence should be verified under other hydrodynamic conditions, mounting geometries, speed envelopes, and current profiles.

Reproducibility is partially supported by explicit formulas, trajectory segmentation, metrics, and tables, but the supplied main-body evidence does not include a code link or public dataset path. Builders can reproduce the estimator structure from the paper, but exact result reproduction would require access to the MARVEL telemetry.

The most important reviewer interpretation is that the result is less about AI in the narrow model-learning sense and more about autonomy infrastructure: reliable sensor calibration, observability, and online estimation. That makes it valuable for robotics systems even though the core methods are classical least squares and Kalman filtering.

Original Abstract

Accurate velocity measurement is a fundamental requirement for autonomous surface and underwater vehicles. Commonly, velocity is provided by a Doppler velocity log (DVL) sensor, yet it becomes unavailable due to operational altitude constraints. In such situations, electromagnetic logs (EMLogs) provide a critically robust alternative for continuous velocity estimation. However, raw EMLog measurements are inherently corrupted by systematic errors, which need to be calibrated prior mission begins. Currently, a benchmarking comparative evaluation of how different calibration models perform under rapidly changing dynamic sea conditions is missing in the literature. To bridge this gap, this paper presents a comparative model-based calibration methodology that evaluates four distinct calibration models using two different estimation pipelines. The proposed framework is rigorously validated on a unique 221 minutes of continuous real-world telemetry collected from the MARVEL surface vehicle during dynamic sea trials. The dataset contains two different EMLogs and DVL recordings. Experimental results demonstrate that the bias and scale error model implemented with the Kalman filter improves the speed estimation by 71%. We also demonstrate that dynamical manoeuvres further improve the accuracy compared to standard straight-line paths, ultimately delivering a validated, real-time online calibration EMLog approach for autonomous surface vehicles.