Accept 7.3/10 cs.AI trendtoknow-paper-summaries codex-pro/gpt-5.5

Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation

Zhihan Zhang, Lizi Liao ยท September 30, 2026 ยท cs.AI

๐Ÿ“Œ Highlights

TPAE addresses a sharp weakness in multimodal RLVR: GRPO and DAPO can judge whether a rollout's final answer is correct, but their rollout-level advantage assigns the same credit or blame to every token. The paper argues that this creates gradient noise for multimodal reasoning, where a few visually grounded but wrong intermediate tokens can collapse an otherwise plausible chain.

  • The paper identifies the joint behavior of visual dependency and predictive entropy as a model-intrinsic signal of multimodal reasoning quality.
  • In the preliminary analysis, 12,800 trajectories from Qwen2.5-VL-7B show that correct chains become more certain as visual dependency rises, while incorrect chains exhibit non-resolving grounding.
  • TPAE builds token-specific reference distributions from correct rollouts and penalizes tokens whose vision-entropy states fall outside a Mahalanobis-distance trust region.
  • The strongest Qwen2.5-VL-7B result is TPAE-D-Qwen2.5-7B at 48.71 avg @8 acc %, compared with VPPO-7B at 47.59 and Shuffle-R1-7B at 47.32.
  • The main limitation is that the method depends on reliable correct-rollout reference distributions and is validated in the main body only on two Qwen backbones and exact-match multimodal reasoning benchmarks.

๐ŸŽฏ Introduction

The paper studies multimodal reinforcement learning with verifiable rewards for MLLMs, where the model must synthesize visual perception with multi-step logical deduction. Existing RLVR systems typically use final-answer accuracy or similarly verifiable rollout-level signals. This has been powerful for reasoning models, but in multimodal settings it hides the local structure of failure: a response may fail because of a specific visual mis-grounding token, a reasoning transition, or a calculation step, yet the training algorithm broadcasts the same sequence-level advantage across all tokens.

The authors position process-level reward models as a natural but costly remedy, because they require a separately trained scorer and can introduce reward-hacking risk. TPAE instead asks whether fine-grained perception-grounded advantages can be obtained directly from the policy's behavior. The key empirical setup examines predictive entropy and visual dependency, where visual dependency measures how much a token distribution changes when the image is masked. Correct trajectories exhibit lower entropy as image reliance increases, while incorrect ones can continue to depend on the image without resolving uncertainty. The paper calls this failure pattern non-resolving grounding.

๐Ÿ”ฌ Methodology

The method begins with the baseline GRPO formulation. For a multimodal input consisting of image I and query q, the old policy samples a group of G rollouts. Each rollout receives a binary reward R_i based on final-answer correctness, and the advantage is normalized across the group. The important weakness is visible in Eq. 1: the token-level advantage A_i,j is set equal to the same rollout-level A_i for every token. TPAE keeps the RLVR setting but replaces this undifferentiated token treatment with a token-specific penalty derived from vision-entropy alignment.

For each generated token, TPAE computes two quantities. Predictive entropy H_i,j is the uncertainty of the next-token distribution P_i,j over the vocabulary. Visual dependency S_i,j is the KL divergence between P_i,j with the original image and P'_i,j with a masked image. The paper follows a patch-based masking strategy with a 60% masking ratio, so the visual dependency score estimates how much the model's token prediction relies on available visual cues. Together, z_i,j = [S_i,j, H_i,j]^T is the token's vision-entropy state.

For each question, rollouts are split into correct D_T and incorrect D_F sets. TPAE uses correct rollouts as the reference behavior. For every unique token t observed in D_T, it aggregates that token's vision-entropy states and fits a multivariate Gaussian reference distribution. The covariance estimate uses a shrinkage form to stabilize token-level estimates under limited rollout counts. Token trustworthiness is then posed as a hypothesis test: if the token's z_i,j is likely under the correct-rollout reference distribution, it is trusted; if its Mahalanobis distance exceeds the trust boundary tau, it is treated as a statistically significant outlier.

The final advantage adjustment is deliberately penalty-only. A positive trustworthiness score means the token lies within the 99% confidence trust region and keeps the original advantage. A negative score produces a bounded penalty through sigmoid(min(0, T_i,j)) - 0.5. This design matters because TPAE is not trying to reward every token that looks normal; it is trying to suppress tokens whose visual dependency and entropy jointly resemble the pivotal failure triggers identified in the preliminary analysis.

equation 1

Formula:
\[\hat{A}_{i,j} =\hat{A}_i = \frac{R_i - \text{mean}(\{R_k\}_{k=1}^G)}{\text{std}(\{R_k\}_{k=1}^G)} \label{eq:adv_grpo}\]

Meaning: Shows the coarse rollout-level advantage that TPAE refines into token-level supervision.

equation 3

Formula:
\[\mathcal{H}_{i,j} := -\sum_{v \in \mathcal{V}} P_{i,j}(v) \log P_{i,j}(v)\]

Meaning: Defines token predictive entropy, one coordinate of the vision-entropy state.

equation 4

Formula:
\[\mathcal{S}_{ij} := \mathbb{D}_{\text{KL}} \left( P_{i,j} \parallel P'_{i,j} \right)= \sum_{v \in \mathcal{V}} P_{i,j}(v) \log \frac{P_{i,j}(v)}{P'_{i,j}(v)}\]

Meaning: Defines visual dependency as information gained from the image for a token prediction.

equation 5

Formula:
\[P(\mathtt{z} | \mathcal{\mu}_t, \mathbf{\Sigma}_t) = \frac{1}{2\pi\sqrt{ | \mathbf{\Sigma}_t|}} \exp\left( -\frac{1}{2} (\mathtt{z} - \mathcal{\mu}_t)^\top \mathbf{\Sigma}_t^{-1} (\mathtt{z} - \mathcal{\mu}_t) \right) \label{eq:gaussian}\]

Meaning: Models the correct-rollout reference distribution for each unique token.

๐Ÿ“Š Experiments

The main empirical evaluation trains on ViRL39K, which the paper describes as containing 38.9k diverse multimodal reasoning problems. The evaluated model backbones are Qwen2.5-VL-7B-Instruct and Qwen3-VL-8B-Instruct, each combined with GRPO and DAPO to form TPAE-G-Qwen2.5-7B, TPAE-D-Qwen2.5-7B, TPAE-G-Qwen3-8B, and TPAE-D-Qwen3-8B. The benchmark suite covers seven multimodal reasoning datasets: MathVerse, We-Math, MathVision, DynaMath, Geo3k, LogicVista, and MMMU-Pro. For MathVerse, the authors evaluate on the full set of 3.94k multimodal examples while excluding the text-only subset. Evaluation uses exact-match scoring against ground-truth answers and reports average accuracy@8 at inference temperature 1.0, avoiding LLM-as-a-judge evaluation.

Table 1 reports the main comparison. Among Qwen2.5-VL-7B-based systems, TPAE-D-Qwen2.5-7B reaches 48.71 avg @8 acc %, above VPPO-7B at 47.59 and Shuffle-R1-7B at 47.32. TPAE-G-Qwen2.5-7B reaches 47.63, which is slightly above VPPO-7B's 47.59 and below the DAPO variant. The Qwen3-VL-8B variants are higher overall, with TPAE-G-Qwen3-8B at 60.62 and TPAE-D-Qwen3-8B at 61.35. These Qwen3 rows should be read as scaling evidence for the method across another backbone, not as a same-backbone comparison with the listed 7B baselines, because the paper states the comparative baselines are instantiated from Qwen2.5-VL-7B.

Table 2 is the cleanest ablation for builders because it compares multimodal RL with and without TPAE under the same backbones and base algorithms. On Qwen2.5-VL-7B, GRPO improves from 43.86 without TPAE to 47.63 with TPAE, while DAPO improves from 44.97 to 48.71. On Qwen3-VL-8B, GRPO improves from 58.32 to 60.62, and DAPO improves from 58.85 to 61.35. The paper further reports training dynamics in Figure 4, where TPAE-equipped runs show faster learning on accuracy rewards across GRPO and DAPO for both backbones under identical training configurations.

The trust-boundary ablation in Table 3 tests tau = 2.45, tau = 3.26, and the default tau = 3.03 using TPAE-G-Qwen2.5-7B. The default tau = 3.03 gives the best average score, 47.63, and is highest on six of seven benchmarks. The more permissive tau = 2.45 scores 46.96, while tau = 3.26 scores 47.34. The qualitative analysis in Figure 5 then shows that low-trust tokens concentrate around visually hallucinated geometric reasoning, including the mistaken treatment of angle relationships in a failed trajectory.

Table 1

Caption: Main results (avg 8 acc %) across seven multimodal reasoning benchmarks. All evaluations utilize exact-match scoring on verifiable instances to ensure objective results, avoiding any LLM-as-a-judge. All comparative baselines are instantiated from the Qwen2.5-VL-7B backbone. Best performance of 7B models is marked with bold, second best with underline.

Content:
ModelsMathVerseWe-MathMathVisionDynaMathGeo3kLogicVistaMMMU-ProAvg.
ThinkLite-VL-7B42.5065.1727.3447.4437.7739.1528.5341.13
OpenVLThinker-7B41.3965.6325.9051.4138.6043.8532.3342.73
NoisyRollout-7B39.0663.7624.1553.8642.8046.2535.1343.57
MM-Eureka-7B44.5664.1827.8949.8039.6247.3232.1143.64
Perception-R1-7B42.9969.0525.4252.8445.1944.3234.3644.88
VL-Rethinker-7B44.8067.7029.8652.2639.8945.6436.4445.23
R1-ShareVL-7B44.4969.8127.8753.2442.6246.8134.8145.66
PAPO-G-7B44.8266.7927.7652.8240.2546.0736.6345.02
PAPO-D-7B45.3268.3028.3655.8344.1146.7036.3446.42
Shuffle-R1-7B45.6769.7629.2454.4947.9047.5436.6547.32
VPPO-7B46.4370.0130.4857.0845.6346.2637.2347.59
TPAE-G-Qwen2.5-7B45.5069.0730.0256.5946.7347.2038.2747.63
TPAE-D-Qwen2.5-7B47.1571.7030.5657.3948.1448.5537.4548.71
TPAE-G-Qwen3-8B54.6180.4447.5566.8866.3159.9348.6160.62
TPAE-D-Qwen3-8B55.1481.1048.2866.7068.1861.0748.9561.35

Why It Matters: Reports the main benchmark comparison and the headline accuracy claims.

Table 2

Caption: Performance (avg 8 acc %) comparison of multimodal reinforcement learning with (w/) and without (w/o) TPAE.

Content:
ModelsMathVerseWe-MathMathVisionDynaMathGeo3kLogicVistaMMMU-ProAvg.
Qwen2.5-VL-7B28.5346.1318.5145.2036.2342.5125.6434.68
+ GRPO w/o TPAE42.4265.8827.2751.1242.1243.7134.5343.86
+ GRPO w/ TPAE45.5069.0730.0256.5946.7347.2038.2747.63
+ DAPO w/o TPAE44.2266.1827.1453.1243.1145.8635.1944.97
+ DAPO w/ TPAE47.1571.7030.5657.3948.1448.5537.4548.71
Qwen3-VL-8B40.9865.5827.1560.4053.8351.4833.8647.61
+ GRPO w/o TPAE52.7977.3244.9264.4863.3959.1046.2458.32
+ GRPO w/ TPAE54.6180.4447.5566.8866.3159.9348.6160.62
+ DAPO w/o TPAE53.0878.3045.4465.0664.3558.9846.7658.85
+ DAPO w/ TPAE55.1481.1048.2866.7068.1861.0748.9561.35

Why It Matters: Directly measures the performance difference attributable to adding TPAE.

Table 3

Caption: Ablation study on the trust boundary ๐œbased on TPAE-G-Qwen2.5-7B.

Content:
ConfigurationMathVerseWe-MathMathVisionDynaMathGeo3kLogicVistaMMMU-ProAvg.
ฯ„ = 2.4545.1567.9929.3055.7245.7446.7638.0946.96
ฯ„ = 3.2645.4268.8429.8055.8346.8046.8437.8747.34
ฯ„ = 3.0345.5069.0730.0256.5946.7347.2038.2747.63

Why It Matters: Evaluates whether the trust boundary choice matters for downstream accuracy.

fig:preliminary_charts

fig:preliminary_charts

Caption: Empirical analysis of reasoning trajectories using Qwen2.5-VL-7B. (a) Distribution of average entropy and log-frequency across visual dependency levels (left). (b) Density of token-level vision-entropy misalignment relative to the reference distribution center (upper right). (c) Error analysis of divergent tokens through human annotation (lower right).

Why It Matters: This figure supplies the empirical motivation for TPAE: correct chains reduce entropy as visual dependency rises, while divergent tokens in failed chains concentrate in misaligned regions and are mostly annotated as perception or reasoning errors.

fig:algorithm

fig:algorithm

Caption: Overview of the TPAE algorithm. TPAE utilizes correct rollouts to establish reference distributions of vision-entropy states of unique tokens. From a hypothesis testing perspective, TPAE then applies a trust boundary $ $ to quantify token-level trustworthiness based on distributional alignment. This metric is then integrated with the rollout-level advantage from GRPO.

Why It Matters: This is the central method diagram, showing how image masking, entropy, visual dependency, token reference distributions, Mahalanobis distance, trustworthiness, and rollout-level advantage are connected.

fig:training_dynamics

fig:training_dynamics

Caption: Comparison of RLVR training dynamics on accuracy rewards. Solid lines denote running averages with a window size of 20. TPAE achieves consistently faster learning on both GRPO and DAPO across two model backbones.

Why It Matters: This figure supports the paper's optimization-efficiency claim, showing faster learning dynamics for TPAE-equipped GRPO and DAPO variants on both tested model backbones.

fig:qualitative_analysis

fig:qualitative_analysis

Caption: Qualitative analysis of token-level trustworthiness in a failed reasoning trajectory. Dark red highlights successfully localize the pivotal triggers of multimodal reasoning collapse, demonstrating the effectiveness of our TPAE algorithm.

Why It Matters: This qualitative example shows TPAE's token-level trustworthiness signal localizing visually hallucinated geometric reasoning steps in a failed trajectory.

๐Ÿ”ฎ Conclusion

The paper's conclusion is that token-level vision-entropy dynamics can turn RLVR from a purely outcome-supervised procedure into a perception-aware optimization process. TPAE does not replace verifiable final-answer rewards; it refines their assignment by identifying tokens whose visual dependency and entropy patterns are statistically inconsistent with correct reasoning chains. The reported benchmark and training-dynamics results support the claim that this finer-grained penalty improves both accuracy and learning efficiency across the tested GRPO and DAPO settings.

๐Ÿ› ๏ธ Future Research Improvements

The most important next research question is robustness of the reference distribution. TPAE relies on correct rollouts to estimate token-specific Gaussian distributions, but the main body does not show what happens when correct rollouts are rare, when token counts are extremely sparse, or when the task distribution is far harder than ViRL39K-style training. Future work should stress-test group size, reward sparsity, and fallback behavior for tokens absent from D_T.

A second direction is broadening the visual dependency estimator. The paper uses a 60% patch masking strategy for I', which is simple and model-intrinsic, but it may introduce variance or artifacts depending on the visual encoder and task type. Builders should compare patch masking with structured object masking, saliency-aware masking, or counterfactual image perturbations while preserving the same exact-match evaluation discipline.

  • Verify whether Gaussian token-level reference distributions remain adequate across other MLLM families.
  • Test TPAE on non-exact-match multimodal tasks where final-answer verification is less clean.
  • Measure computational overhead from extra masked-image forward distributions during RL training.

๐Ÿญ Potential Industry Use Scenarios

TPAE is most immediately relevant to teams training multimodal reasoning models for domains where final answers are objectively checkable but intermediate visual grounding errors are costly. Examples supported by the paper's benchmark framing include mathematical, geometric, logical, and multi-discipline visual reasoning. In those settings, a rule-based final-answer reward can be retained while TPAE supplies a denser training signal against visually grounded failure tokens.

The method also suggests a practical debugging layer for RL training runs. Because token trustworthiness is interpretable at the generated-token level, model builders could inspect low-trust spans in failed trajectories to identify recurring visual mis-grounding patterns. The qualitative geometry example shows this kind of diagnostic use: the method highlights the steps where the model confuses angle and arc relationships. Production use would still require validation outside the reported Qwen backbones and benchmark suite, especially where images are noisier or answers are not exact-match verifiable.

  • Training multimodal math and geometry tutors with verifiable final answers.
  • Auditing visual reasoning chains for localized perception errors during RLVR runs.
  • Improving data and reward pipelines for enterprise MLLMs that must reason over diagrams, charts, or technical images.

๐Ÿ’ฌ Critical Analysis

The paper's strongest technical feature is that TPAE connects an interpretable diagnostic to a concrete policy-optimization change. The preliminary analysis is unusually aligned with the proposed method: it first shows a measurable split between correct and incorrect trajectories in the vision-entropy plane, then uses that split to define token-level trustworthiness, and finally converts trustworthiness into a bounded advantage penalty. This gives the method a cleaner motivation than many reward-shaping approaches that introduce dense signals without proving they correspond to actual failure points.

The strongest experimental evidence is Table 2, because it isolates w/ and w/o TPAE under GRPO and DAPO for both tested backbones. The gains are consistent across all four base-algorithm/backbone pairings, and the exact-match protocol avoids evaluator-model ambiguity. Table 1 is useful for context, but the Qwen3 rows should not be overread against Qwen2.5-based baselines. The paper correctly states that all comparative baselines use Qwen2.5-VL-7B, so the fairest state-of-the-art comparison for same-backbone claims is the Qwen2.5 TPAE rows.

The main caveats are around distributional assumptions and training-time practicality. The main body says Gaussian modeling is motivated by observed unimodal centrality, but details comparing alternative parametric models are outside the provided main-body evidence. The approach also requires masked-image token distributions, per-token entropy, and per-question reference statistics during RL. Those costs may be acceptable for research training but need profiling for larger-scale industrial pipelines. Finally, the method is penalty-only and conservative for unseen tokens, which is sensible, but it leaves open how much supervision is lost in domains with many rare tokens or low correct-rollout diversity.

Original Abstract

Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet existing frameworks rely on coarse, sequence-level reward signals that lack the fine-grained supervision over the visually-grounded steps within a multimodal reasoning chain. We investigate this gap through the lens of two token-level metrics: visual dependency (i.e. how much a token's prediction relies on the input image features) and predictive entropy. Our empirical analysis reveals two key findings: (1) correct reasoning chains exhibit a markedly sharper entropy reduction as visual grounding intensifies, compared to incorrect ones; (2) pivotal tokens, those whose misprediction triggers reasoning collapse, are statistical outliers in the joint distribution of visual dependency and predictive entropy derived from correct chains. Motivated by these findings, we propose token-level perception-grounded advantage estimation (TPAE), which estimates token-level advantages by measuring each token's statistical consistency with the vision-entropy patterns of correct rollouts. TPAE leverages this granular score to modulate the sequence-level advantage, producing a fine-grained supervision signal that can be integrated into various RLVR frameworks. Extensive experiments on seven benchmarks show that TPAE consistently outperforms leading strong baselines, yielding more stable and efficient optimization for multimodal reasoning. The code is publicly available at https://github.com/Zhihan72/TPAE.