Accept 7.5/10 cs.LG trendtoknow-paper-summaries codex-pro/gpt-5.6-luna

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin, Wenjing Chen, Ping He, Pan Lu, Alvaro Velasquez · October 01, 2026 · cs.LG

📌 Highlights

The paper's core claim is that mathematical reasoning has a structural stage that is distinct from procedural execution. A model may solve a problem without isolating the load-bearing idea, or identify that idea and still fail in technical execution.

Across the PRIM diagnosis, correct primitives improve Execution by 17.58–29.67 percentage points across 12 models, while 83.6% of Generation failures fall in the failed-Discovery regime. These findings support Discovery as the principal upstream bottleneck.

Post-training analysis finds that D−E+ failures are repaired at 20.4% by SFT and 21.1% by OPSD, compared with 6.8% and 8.2% for D−E− failures.

ABSORB uses primitive-conditioned teacher guidance with bounded override and reports the strongest average performance across Qwen3.5-4B, 9B, and 27B in the supplied comparison table. Its main limitation is that explicit Discovery gains remain smaller than downstream Generation gains.

  • Mathematical Primitives make structural understanding measurable rather than treating final-answer accuracy as a complete reasoning diagnosis.
  • Digestion is often much stronger than Discovery, indicating that models can recognize structure retrospectively after a valid solution exposes it.
  • Teacher Primitives outperform generic Teacher Plans in the reported intervention comparison.
  • Selective privileged distillation is more stable than indiscriminate post-training in the reported experiments.

🎯 Introduction

The paper starts from a gap in contemporary mathematical evaluation: final-answer accuracy and formal proof verifiability establish whether an output is correct, but they do not reveal whether the model discovered the mathematical structure that organizes the solution. A model can produce a valid derivation through opaque or specious reasoning, while another can identify a sound proof strategy and fail only during low-level execution.

To expose this distinction, the authors define a Mathematical Primitive as the concise, load-bearing observation that makes a problem solvable. It must identify the exploited property and explain how that property unlocks the solution, while remaining distinct from routine calculation, generic strategic advice, or a complete proof.

The paper's objective is therefore twofold: diagnose structural mathematical understanding through PRIM's Discovery, Generation, Digestion, and Execution dimensions, and use the diagnosis to design post-training that transfers primitive-guided reasoning without requiring an explicit primitive at inference time.

🔬 Methodology

PRIM is built from Humanity's Last Exam, using its mathematics, text-only, free-form subset with evaluation. The authors randomly sample 200 problems from the Gold and Revision subsets, generate initial primitive annotations with GPT-5.4-High conditioned on the problem, reference answer, and gold rationale, and use three graduate-level mathematics experts to review the problem, answer, rationale, and primitive. Removing 18 unsuitable examples yields a final evaluation set of 182.

Let x be the problem, p the corresponding primitive, y a complete solution, and π an LLM solver. Discovery evaluates π(x)→p; Generation evaluates direct zero-shot solving π(x)→y; Digestion evaluates π(x,y)→p; and Execution evaluates π(x,p)→y. Discovery and Digestion use a primitive-scoring protocol, while Generation and Execution use final-answer accuracy.

For post-training, the authors use 709 Mathematics Ph.D. qualifying-examination problems from 1991–2026, each paired with a human-authored proof with median length 118 words. OPSD compares teacher and student token distributions along the student's on-policy trajectory. ABSORB changes the privileged signal from a full reference solution to a mathematical primitive and adds a bounded reverse-KL mechanism so the teacher can reinforce better alternatives without aggressively overriding student preferences.

The conceptual pipeline and the ABSORB intervention are summarized by the supplied method figures.

Primitive scoring

Formula:
\[\mathrm{Score} = V \cdot \sigma_{\text{gate}} \cdot \big(0.6 + 0.4 \cdot \sigma_{\text{mech}}\big), \quad \mathrm{PrimitiveAcc} = \mathbf{1}\left[\mathrm{Score}\ge\tau\right],\]

Meaning: The scoring rule treats validity and structural identification as necessary conditions and uses mechanism quality to refine the primitive score.

OPSD objective

Formula:
\[\mathcal{L}_{\mathrm{OPSD}} = \frac{1}{T}\sum_{t=1}^{T} D(P_t, Q_t) = \frac{1}{T} \sum_{t=1}^{T} D\left(\pi_S(\cdot \mid \hat{y}_{<t}, x), \pi_T(\cdot \mid \hat{y}_{<t}, x,y) \right),\]

Meaning: The objective supplies token-level teacher supervision along the student's own rollout.

ABSORB loss

Formula:
\[\mathcal{L}_{\mathrm{ABSORB}} = \frac{1}{T}\sum_{t=1}^{T} \sum_{v\in\mathcal{S}_t} \min\left\{ \bar{\pi}_S(v \mid \hat{y}_{<t}, x) \log \frac{\bar{\pi}_S(v \mid \hat{y}_{<t}, x)}{\bar{\pi}_T(v \mid \hat{y}_{<t}, x, p)}, \tau \right\}.\]

Meaning: The clamped divergence bounds negative override from the privileged teacher while retaining positive primitive-guided pressure.

📊 Experiments

The main diagnostic evaluates 12 models from the OpenAI, Qwen, and DeepSeek lineages on the 182-example PRIM set, using high reasoning effort when available and a maximum generation budget of 120k tokens per problem. The post-training experiments use Qwen3.5-4B, Qwen3.5-9B, and Qwen3.5-27B, with bfloat16 precision, seed 42, and NVIDIA RTX 6000 Ada GPUs.

The paper evaluates Generation and Discovery on PRIM, the full mathematics subset of HLE-Verified, HMMT25, and Omni-MATH. HMMT25 and Omni-MATH use pass@4 because of their difficulty, and the reported average is computed over Generation scores across the four benchmarks while excluding Discovery.

The diagnosis finds that similar Generation accuracy can hide different structural profiles. Qwen3.6-27B and gpt-5.4-mini have comparable Generation performance but differ by more than 30 percentage points in Discovery. Digestion is much stronger than Discovery for several models: Qwen3.6-27B rises from 24.73% Discovery to 92.31% Digestion, while R1-Distill-32B rises from 6.04% to 65.38%.

Supplying the correct primitive improves Execution by 17.58–29.67 percentage points across all 12 models. In the intervention comparison, self-generated primitives provide little benefit or hurt performance, Teacher Plans help less than Teacher Primitives, and Gold Primitives reach 71.98% for gpt-5.4-mini, 68.68% for gpt-5.4-nano, and 78.57% for Qwen3.6-27B. The ABSORB table reports its strongest average gains over the listed baselines at all three Qwen3.5 scales.

tab:execution-primitive

Caption: Generation and Execution performance on PRIM under different forms of reasoning support.

Content:
ModelGenerationExecution: Self-Generated PrimitiveExecution: Teacher PlanExecution: Teacher PrimitiveExecution: Gold Primitive
gpt-5.4-mini50.5548.90 (-1.65)60.44 (+9.89)63.74 (+13.19)71.98 (+21.43)
gpt-5.4-nano44.5146.15 (+1.64)54.40 (+9.89)60.99 (+16.48)68.68 (+24.17)
Qwen3.6-27B52.7541.21 (-11.54)56.04 (+3.29)64.84 (+12.09)78.57 (+25.82)

Why It Matters: It tests whether the benefit comes from a true structural primitive rather than merely forcing a two-stage plan-and-solve procedure.

Table 3

Caption: Post-training repair rates of SFT and OPSD across initial PRIM failure quadrants.

Content:
Failure typeSFTOPSD
D−E+20.4%21.1%
D−E−6.8%8.2%

Why It Matters: It quantifies the higher repairability of failures where execution is latent but discovery is missing.

Table 4

Caption: Comparison of SFT, OPSD, and ABSORB across three Qwen3.5 model scales.

Content:
Model / methodDiscoveryGenerationHLE MathHMMT25Omni-MATHAvg.
Qwen3.5-4B6.5926.3726.7076.6778.0051.94
+ SFT5.49 (-1.10)25.27 (-1.10)27.36 (+0.65)83.33 (+6.66)79.33 (+1.33)53.82 (+1.89)
+ OPSD8.79 (+2.20)30.22 (+3.85)29.58 (+2.88)83.33 (+6.66)79.33 (+1.33)55.62 (+3.68)
+ ABSORB9.89 (+3.30)31.32 (+4.95)31.02 (+4.32)83.33 (+6.66)80.00 (+2.00)56.42 (+4.48)
Qwen3.5-9B13.1938.4635.6090.0078.6760.68
+ SFT12.64 (-0.55)32.42 (-6.04)34.55 (-1.05)90.00 (+0.00)80.67 (+2.00)59.41 (-1.27)
+ OPSD10.99 (-2.20)33.52 (-4.95)35.08 (-0.52)90.00 (+0.00)81.33 (+2.67)59.98 (-0.70)
+ ABSORB14.29 (+1.10)43.96 (+5.49)38.35 (+2.75)93.33 (+3.33)82.00 (+3.33)64.41 (+3.73)
Qwen3.5-27B28.5747.8042.6796.6787.3368.62
+ SFT28.02 (-0.55)48.35 (+0.55)43.59 (+0.92)96.67 (+0.00)86.67 (-0.67)68.82 (+0.20)
+ OPSD28.57 (+0.00)44.51 (-3.30)42.80 (+0.13)96.67 (+0.00)88.00 (+0.67)67.99 (-0.62)
+ ABSORB28.02 (-0.55)48.90 (+1.10)46.60 (+3.93)100.00 (+3.33)88.67 (+1.33)71.04 (+2.42)

Why It Matters: It shows that ABSORB improves average benchmark performance at every listed model scale, while the baselines can regress on Generation or average performance.

fig:teaser

fig:teaser

Caption: Diagnosing and internalizing mathematical primitives. We introduce Mathematical Primitives to probe structural mathematical understanding in LLMs across four dimensions: Discovery, Generation, Digestion, and Execution. Our diagnosis further motivates , which selectively transfers primitive-guided reasoning into the student without requiring primitives at inference time.

Why It Matters: This figure gives the paper's central conceptual pipeline: diagnose structural understanding with four dimensions, then use primitive-guided supervision to improve reasoning without exposing primitives at inference time.

fig:s-d-composition

fig:s-d-composition

Caption: Joint composition of Generation and Discovery outcomes across models.

Why It Matters: The joint composition makes the paper's main diagnostic claim visible: similar Generation accuracy can arise from different combinations of successful or failed Discovery.

Figure 3

Figure 3

Caption: Figure 3: Discovery versus Generation across all evaluated models on PRIM.

Why It Matters: The comparison directly exposes the gap between identifying a load-bearing primitive and producing a correct final answer across the evaluated models.

Figure 5

Figure 5

Caption: Figure 5: Illustration of ABSORB paradigm.

Why It Matters: This architecture illustration clarifies how primitive-conditioned teacher guidance is transferred selectively into the student's on-policy reasoning.

🔮 Conclusion

The authors conclude that Mathematical Primitives and PRIM expose a structural layer of mathematical reasoning that is obscured by final-answer evaluation. Similar solution accuracy can arise from different Discovery and Execution profiles, correct primitives can unlock substantial latent execution capacity, and independent Discovery is the dominant bottleneck.

The post-training results support a targeted intervention strategy. Discovery-limited failures are more repairable than failures lacking both Discovery and Execution, but conventional SFT and OPSD can introduce regressions. ABSORB addresses this instability by giving the teacher primitive information and bounding how strongly that privileged signal overrides the student's own trajectory.

For builders, the practical takeaway is to measure and train structural discovery separately from execution, then use privileged structural supervision selectively rather than forcing the student to reproduce full reference proofs.

🛠️ Future Research Improvements

The paper explicitly points toward generalizing the diagnostic framework to structurally rigorous domains such as theoretical physics and computational chemistry. A useful next experiment would test whether domain-specific primitives show the same Discovery–Execution asymmetry outside mathematics.

The observed asymmetry also motivates agent designs that separate structural search from procedural execution and allocate inference-time computation differently between the two stages. This should be evaluated with controlled compute budgets and error categories rather than only aggregate accuracy.

Integration with formal verification systems such as Lean is another concrete direction: primitives could provide informal structural guidance while formal proof generation and checking validate the resulting derivation. The current evidence does not yet establish whether primitive-guided reasoning transfers reliably into formal proof assistants.

🏭 Potential Industry Use Scenarios

Mathematical reasoning assistants could use a primitive-first internal pipeline to decide whether a failure arose from missing structural insight or from execution error. This would support targeted retries, teacher routing, and more interpretable debugging for systems used in education, theorem proving, scientific research, or technical analysis.

For model training teams, PRIM-style evaluations could become a quality gate for reasoning updates: a checkpoint should be assessed not only on answer accuracy but also on Discovery, Digestion, Execution, repair rates, and regressions on previously solved problems.

In scientific AI workflows, a primitive can serve as a compact handoff between problem understanding and downstream derivation. However, deployment would require validation on domain-specific data, calibrated uncertainty, and independent verification because the paper's direct evidence is limited to the reported mathematics benchmarks and Qwen3.5 post-training setup.

💬 Critical Analysis

The strongest aspect of the work is its decomposition of a single final-answer metric into interpretable capability dimensions. The intervention comparison is especially persuasive because self-generated primitives provide little benefit, Teacher Plans help less than Teacher Primitives, and Gold Primitives produce large Execution gains. This directly supports the claim that the useful signal is structural rather than merely procedural.

The main empirical caution is that the paper's most complete model-comparison table is not numerically recoverable in the supplied main-body evidence, so the review should not reconstruct those missing cells. The reported analyses still provide strong directional evidence, but the magnitude and consistency of every model-level PRIM result require verification against the full paper.

Reproducibility also depends on choices that deserve stress testing: the definition and scoring of a primitive, the expert-review process, the HLE-derived sample, the 709-example post-training corpus, and the privileged-teacher configuration. The ABSORB results are promising, but the relatively modest explicit Discovery changes suggest that its gains may reflect implicit guidance into downstream reasoning rather than a fully learned ability to articulate primitives.

Overall, the paper offers a compelling builder-facing hypothesis and a useful evaluation vocabulary, but broader validation is needed before treating Mathematical Primitives as a general solution to reasoning failures.

Original Abstract

While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce \abs{}, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that \abs{} consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.