Accept 7.4/10 cs.AI trendtoknow-paper-summaries codex-pro/gpt-5.5

Multi-LLM Collaborative Alignment via Stackelberg Games

Christina Hahn, Shangbin Feng, Dean Light, Swastik Roy, Hila Gonen, Yulia Tsvetkov ยท September 30, 2026 ยท cs.AI, cs.LG

๐Ÿ“Œ Highlights

Stackelberg Alignment turns instruction selection for multi-LLM collaborative training into an adaptive leader-follower game rather than a uniform sampling routine.

The leader is an EXP3 bandit that learns which instructions produce informative duels; the followers are language models that generate responses, judge peers, update reputations, and train with DPO or GRPO.

The strongest headline result is that Stackelberg GRPO achieves the highest macro-average in all three reported model pools: 0.563, 0.544, and 0.609.

The main caveat is that gains are not uniform across every benchmark: some narrow scientific tasks are still led by AggLM or Heterogeneous Swarms, and DPO slightly trails Sparta Alignment in Pool 3 Avg.

  • Core claim: adaptive instruction curricula improve training-time multi-LLM collaboration.
  • Strongest result: Stackelberg GRPO ranks first by macro-average across all three pools.
  • Key mechanism: EXP3 concentrates duels on prompts where model disagreement remains informative.
  • Important limitation: peer-judgment quality and compute cost become central engineering risks.

tab:main

Caption: Performance across three model pools and 12 datasets grouped by domain. Bold: best per column within each pool; underline: second best; --: not applicable. Shaded rows ( gray!12 ) are variants. The Avg column is the macro-average over all available datasets, with AlpacaEval min-max normalized to $[0,1]$.

Content:
PoolSparta Alignment AvgBest static inference AvgStackelberg (DPO) AvgStackelberg (GRPO) Avg
Pool 1: Specialized Expert LLMs0.5240.4510.5480.563
Pool 2: LLMs from Diverse Academic Research0.5330.4820.5400.544
Pool 3: General-Purpose LLMs0.5790.4870.5760.609

Why It Matters: The Avg column is the paper's central aggregate result: Stackelberg GRPO ranks first in all three pools, while Stackelberg DPO is second in Pools 1 and 2 but slightly below Sparta in Pool 3.

๐ŸŽฏ Introduction

The paper targets training-time model collaboration, where multiple LLMs improve by competing, judging, and learning from each other rather than only aggregating answers at inference time. Prior methods such as Sparta Alignment use pairwise duels and peer judgment, but the paper argues that a crucial part of the loop remains fixed: instructions are sampled uniformly throughout training.

The authors' central observation is that instruction informativeness is heterogeneous and non-stationary. Some instructions are too easy and yield no useful preference gap, some are too hard and produce tied or unreliable scores, and the useful frontier shifts as the pool improves. This makes uniform sampling wasteful in exactly the setting where duel budgets are expensive.

Stackelberg Alignment models this as a bilevel game: a leader commits to an instruction distribution, and follower LLMs respond through duels and training. The paper's objective is to jointly learn the instruction curriculum and the follower models so that compute is concentrated on currently informative disagreement.

๐Ÿ”ฌ Methodology

The method starts with a model pool M0 and instruction set X. Each iteration runs D duels. For each duel, the leader samples an instruction, two models generate responses, the remaining models judge both responses on a 1-10 scale, and reputation-weighted scores determine the preference signal. Non-tied comparisons feed DPO, while GRPO uses peer scores online as rewards.

The leader uses EXP3 because instruction rewards change over training. Its sampling distribution mixes learned instruction weights with a uniform exploration floor, so high-reward prompts become more likely without permanently excluding the rest of the pool. The reward combines a difficulty term and a preference-quality term: instructions should be learnable, not mastered, and should produce a score gap near the scheduled target.

Opponent matching is also scheduled. Early training targets larger reputation gaps to create clear preference pairs; later training moves toward near-peer duels for finer distinctions. Peer judgment is reputation-weighted, and reputations are updated with an Elo-style rule that incorporates score gap, stability, and a gap factor.

equation 1

Formula:
\[p_k^t = (1 - \gamma)\,\frac{w_k^t}{\|\mathbf{w}^t\|_1} + \frac{\gamma}{K}, \label{eq:exp3-dist}\]

Meaning: Defines the EXP3 leader's instruction sampling probability, mixing exploitation of high-weight instructions with uniform exploration.

equation 2

Formula:
\[r_{\mathrm{diff}}(s) = \begin{cases} \dfrac{s_{\max} - s}{s_{\max} - \tau_{\mathrm{th}}} & s \geq \tau_{\mathrm{th}} \\[4pt] 0 & s < \tau_{\mathrm{th}} \end{cases} \label{eq:diff-reward}\]

Meaning: Rewards instructions that are hard but still above a learnability threshold, avoiding prompts that are either mastered or too difficult.

equation 3

Formula:
\[g^*_t = g_{\mathrm{end}} + (g_{\mathrm{start}} - g_{\mathrm{end}})\,(1 - \tau), \qquad \tau = \frac{t}{T-1}, \label{eq:gap-schedule}\]

Meaning: Schedules the target score or reputation gap over training, enabling broad early distinctions and finer late-stage comparisons.

equation 4

Formula:
\[r_{\mathrm{pref}}(\bar{s}_i, \bar{s}_{i'}) = \exp\!\left(-\frac{\bigl(|\bar{s}_i - \bar{s}_{i'}|\,/\,(s_{\max} - s_{\min}) - g^*_t\bigr)^2}{2\sigma_r^2}\right). \label{eq:pref-reward}\]

Meaning: Measures how close a duel's preference gap is to the scheduled target, penalizing both ties and trivially large gaps.

Algorithm 1 STACKELBERG ALIGNMENT

Steps: ['Initialize model reputations, reputation deviations, and uniform leader weights over the instruction set.', 'For each training iteration, initialize judged duels and preference pairs.', 'For each duel, sample an instruction from the leader distribution, select an active model, and draw an opponent using reputation-based matching.', 'Generate both model responses, have non-combatant models score the responses, aggregate scores using reputation weights, update reputations, and add non-tied comparisons to preference data.', 'Update leader weights with EXP3 using all judged duels, then train all models with either DPO on preference pairs or GRPO with online peer rewards.']

Why It Matters: This algorithm is the operational recipe for reproducing the framework: it specifies the alternation between adaptive instruction sampling, dueling, judging, reputation updates, and follower fine-tuning.

๐Ÿ“Š Experiments

The evaluation uses three model pools: Pool 1 has 9 specialized expert LLMs, Pool 2 has 8 models from diverse academic research projects, and Pool 3 has 4 general-purpose LLMs. The paper compares Stackelberg DPO and Stackelberg GRPO against Majority Vote, MoA, Multiagent Debate, Heterogeneous Swarms, Sparta Alignment, Multiagent Fine-tuning, Trained Router, and AggLM.

The 12 benchmarks span five domains: BixBench, LabBench, SMDD, and AssayBench for scientific discovery; GPQA-Diamond and MATH for reasoning; HumanEval and MBPP for code; AlpacaEval and IFEval for instruction following; and TruthfulQA plus CulturalBench-Hard for knowledge and truthfulness. The Avg column is the macro-average over available datasets, with AlpacaEval min-max normalized to [0,1].

The main result is that Stackelberg GRPO achieves the highest macro-average across all three pools: 0.563, 0.544, and 0.609. Relative to Sparta Alignment, the strongest training-based baseline, the paper reports macro-average improvements of 7.4%, 2.1%, and 5.2%. Relative to the best static inference baseline, the gap exceeds 20% in Pools 1 and 3 and 12% in Pool 2.

The analysis figures support the mechanism behind the results. Reputation scores diverge and stabilize by iteration 4-5, leader opponent weights become task-specific, and instruction weights on MATH move from near-uniform exploration to an implicit curriculum by iteration 5.

  • Datasets: BixBench, LabBench, SMDD, AssayBench, GPQA-Diamond, MATH, HumanEval, MBPP, AlpacaEval, IFEval, TruthfulQA, CulturalBench-Hard.
  • Metric direction: higher table values indicate better reported performance; the Avg column is a macro-average with AlpacaEval normalized to [0,1].
  • Training setup from main body: T = 8 iterations, D = 64 duels per iteration, EXP3 exploration rate gamma = 0.2, DPO or GRPO follower training.

tab:main

Caption: Performance across three model pools and 12 datasets grouped by domain. Bold: best per column within each pool; underline: second best; --: not applicable. Shaded rows ( gray!12 ) are variants. The Avg column is the macro-average over all available datasets, with AlpacaEval min-max normalized to $[0,1]$.

Content:
PoolSparta Alignment AvgBest static inference AvgStackelberg (DPO) AvgStackelberg (GRPO) Avg
Pool 1: Specialized Expert LLMs0.5240.4510.5480.563
Pool 2: LLMs from Diverse Academic Research0.5330.4820.5400.544
Pool 3: General-Purpose LLMs0.5790.4870.5760.609

Why It Matters: The Avg column is the paper's central aggregate result: Stackelberg GRPO ranks first in all three pools, while Stackelberg DPO is second in Pools 1 and 2 but slightly below Sparta in Pool 3.

tab:main-open-ended

Caption: Selected open-ended and instruction-following results from Table 1.

Content:
PoolMethodSMDDAlpacaEvalIFEvalTruthfulQACulturalBench
Pool 1Sparta Alignment0.1687.5280.7070.6860.656
Pool 1Stackelberg (DPO)0.2048.0810.7490.6820.747
Pool 1Stackelberg (GRPO)0.2377.8190.7300.6840.724
Pool 2Sparta Alignment0.2166.2730.6720.6320.680
Pool 2Stackelberg (DPO)0.1964.9860.6760.6560.696
Pool 2Stackelberg (GRPO)0.2345.2070.6800.6260.680
Pool 3Sparta Alignment0.2947.1090.7750.7700.749
Pool 3Stackelberg (DPO)0.3904.1060.7780.7700.810
Pool 3Stackelberg (GRPO)0.3537.1380.7910.7800.841

Why It Matters: These columns show where peer-judgment-driven preference learning is especially relevant. Stackelberg variants are strongest on SMDD across all pools, while GRPO leads Pool 3 on AlpacaEval, IFEval, TruthfulQA, and CulturalBench.

fig:overview

fig:overview

Caption: Overview of for one iteration. The EXP3 leader maintains learned weights $w^t$ over the instruction pool and selects instruction $x_k$ non-uniformly. Two follower models ($M_3$, $M_4$ in the figure) are matched by reputation and duel on $x_k$, generating responses $y_3$ and $y_4$. Remaining models act as peer judges, scoring both responses; scores are aggregated weighted by judge reputation, determining the winner. The Update step performs three actions: (1) reputation scores are adjusted, (2) leader weights are updated via EXP3 based on duel informativeness, and (3) the winning preference pair is added to the dataset $ P$ for DPO training, or the selected instruction and online judge score...

Why It Matters: This is the clearest architecture figure: it shows the full loop connecting instruction selection, model duels, peer judgment, reputation updates, EXP3 updates, and DPO or GRPO training.

fig:rating_dynamics

fig:rating_dynamics

Caption: Reputation score trajectories over training iterations for TruthfulQA and HumanEval (Pool~2, Stackelberg GRPO). Reputation scores diverge from the same initialization and stabilize by iteration~4--5. The leading model differs across tasks, confirming genuine task-specific competitive structure and collaborative learning landscape.

Why It Matters: It supports the paper's claim that reputation dynamics capture task-specific model strengths rather than collapsing to one universally dominant model.

fig:leader_strategy

fig:leader_strategy

Caption: Normalized opponent weight assigned by the Stackelberg leader over iterations (Pool~2, Stackelberg GRPO). Dashed line = uniform selection baseline. The leader shields the emerging dominant model from being excessively used as an opponent and adapts its strategy per task.

Why It Matters: It provides evidence that the learned strategy is task-adaptive and not merely random or uniform opponent selection.

fig:curriculum

fig:curriculum

Caption: Leader instruction weights over training iterations (the MATH dataset, model pool 1). Left: heatmap of normalized weight deviation from uniform for the top-32 most dynamic instructions. Right: trajectory of top-5 and bottom-5 instructions by final-iteration weight; dashed line = uniform. The leader converges to an implicit difficulty curriculum by iteration~5.

Why It Matters: This is the most direct visual evidence for the paper's curriculum-learning claim: the leader moves from near-uniform exploration to concentrated instruction weights.

๐Ÿ”ฎ Conclusion

The paper concludes that adaptive instruction selection improves collaborative alignment by concentrating pairwise duels on instructions where the model pool still disagrees. This is a meaningful upgrade over uniform sampling because the useful frontier of preference learning changes as models improve.

For builders, the main takeaway is that multi-LLM collaboration is not only about model aggregation or judge prompts. The curriculum policy, opponent policy, and reputation system form a coupled training environment, and this paper provides a concrete EXP3-based implementation that reports consistent macro-average gains.

๐Ÿ› ๏ธ Future Research Improvements

A natural next research direction is stronger validation of the judging system. Because peer scores determine preference pairs, reputations, leader rewards, and GRPO rewards, future work should isolate judge calibration, judge collusion risks, and sensitivity to low-quality or domain-mismatched judges.

The paper's evidence suggests that specialized scientific tasks may need richer instruction pools or task-specific data curation. LabBench in Pool 3 and AssayBench in Pools 1-2 are exceptions where non-Stackelberg baselines lead, so follow-up work should test whether adaptive curricula help more when the underlying instruction pool has better domain coverage.

Another improvement path is cost-aware curriculum learning. The main-body setup uses repeated duels, peer judging, and iterative fine-tuning, so a practical extension would optimize not only informativeness but also judge cost, model latency, and marginal learning gain per GPU-hour.

  • Verify robustness under noisy, biased, or adversarial peer judges.
  • Test larger and more specialized instruction pools for scientific-discovery benchmarks.
  • Add compute-aware rewards that account for training and judging cost.
  • Compare EXP3 against other non-stationary bandit or curriculum-learning policies.

๐Ÿญ Potential Industry Use Scenarios

The most credible near-term use case is internal model-pool improvement: organizations with several domain-tuned models could use Stackelberg-style duels to generate preference data where models disagree, rather than paying for uniform prompt labeling across the whole instruction set.

The method also fits evaluation-heavy environments where multiple LLMs already act as specialists, such as code assistants, scientific copilots, customer-support routing systems, and enterprise knowledge agents. In those settings, adaptive dueling could identify which prompts expose meaningful differences among models and convert those disagreements into training data.

Deployment remains uncertain for safety-critical domains because the paper's main-body evidence does not establish human-verified judge quality, production latency, or governance controls. Builders should treat the framework as a training-data generation and research pipeline before treating it as an autonomous production alignment system.

  • Adaptive preference-data generation for heterogeneous enterprise model pools.
  • Curriculum learning for domain-tuned assistants where uniform prompt sampling wastes fine-tuning budget.
  • Model reputation tracking for routing, evaluation, or internal benchmark triage.
  • GRPO-style online peer-reward training in controlled research environments.

๐Ÿ’ฌ Critical Analysis

The strongest part of the paper is the alignment between problem diagnosis and mechanism. Uniform instruction sampling is plausibly inefficient in multi-model preference training, and the EXP3 leader directly attacks that bottleneck while preserving exploration. The analysis figures add useful evidence that the leader is not just adding complexity: reputations diverge, opponent weights become task-specific, and instruction weights sharpen into a curriculum.

The main weakness is dependence on peer-generated supervision. The same model pool supplies responses, judgments, reputation signals, and training rewards, so systematic blind spots could reinforce themselves. Reputation weighting helps discount weak judges, but the main-body evidence does not fully resolve whether reputation is measuring reliability, task specialization, or compatibility with the pool's own preferences.

The experimental result is persuasive at the aggregate level, especially against Sparta Alignment, but it is not a clean win everywhere. Pool 3 DPO trails Sparta on Avg, and several narrow scientific tasks are led by other methods. The right interpretation is that adaptive instruction selection is a promising control layer for collaborative training, not a universal replacement for better data, stronger domain specialists, or external evaluation.

Reproducibility looks plausible but incomplete from main-body evidence alone. The paper reports many key hyperparameters in the main body and includes a code link on the first page, but some implementation details, confidence intervals, dataset statistics, baseline configurations, and judge prompt details are stated as appendix material and therefore fall outside this summary's evidence boundary.

  • Strength: clear mechanism for non-stationary instruction informativeness.
  • Strength: strong aggregate results across model-pool types and task domains.
  • Risk: peer-judgment feedback loops may amplify shared model biases.
  • Risk: compute and orchestration complexity may limit adoption outside well-resourced training teams.

Original Abstract

A pool of language models can collaborate and improve collectively by learning from one another's responses. These interactions depend on the instructions used during training. Existing methods typically sample instructions uniformly, even though their usefulness may change as the models improve: an instruction on which models' responses once differed in quality may later be answered equally well, while a previously difficult instruction may begin to provide a useful learning signal. We propose Stackelberg Alignment, a game-theory-inspired leader-follower framework that turns instruction selection into an adaptive curriculum. An EXP3 bandit acts as the leader, allocating a fixed sampling budget across instructions and updating its sampling distribution using a reward that combines instruction difficulty and response discriminability. The language models act as followers: they respond to the selected instructions, evaluate one another's responses, and learn from the resulting preference signals through DPO or GRPO. The framework uses Elo-style reputation-weighted peer judgment and reputation-based opponent matching to support reliable and competitive model interactions. Experiments across three heterogeneous model pools and 12 benchmarks spanning scientific discovery, reasoning, code, instruction following, and knowledge show that Stackelberg Alignment achieves the highest macro-average across three diverse model pools, outperforming the strongest training-time baseline by up to 7.4% and the best static inference baseline by 12-25%. Analysis confirms that the adaptive leader concentrates duels on the most informative instructions, and ablations show that both reputation-weighted judgment and reputation-based matching improve the effectiveness of multi-LLM evolution.