Accept 7.4/10 cs.CV trendtoknow-paper-summaries codex-pro/gpt-5.5

CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding

Yiduo Jia, Muzhi Zhu, Jinchuan Shi, Hao Zhong, Yuling Xi, Ke Liu, Hao Chen ยท September 30, 2026 ยท cs.CV

๐Ÿ“Œ Highlights

CoEvoWhen treats ultra-long video temporal grounding as an agentic evidence-acquisition problem: a frozen VLM must decide where to search, what to inspect, and when evidence is sufficient.

The central claim is that policies and tools should coevolve. Policies determine evidence needs and orchestration; tools determine what visual evidence can be sampled and presented.

The strongest main-body result is that the evolved Qwen3.5-27B skill improves VUE-LVTR IoU AUC from 0.4107 to 0.5137 and ExtremeWhenBench Mean IoU from 0.1507 to 0.2636 while reducing VUE-LVTR visual tokens from 202.58k/query to 141.89k/query.

  • Task: localize natural-language event intervals in videos lasting tens of minutes to hours.
  • Mechanism: evolve an external skill containing high-level policies and executable media tools, without changing VLM weights.
  • Benchmarks: VUE-LVTR, ExtremeWhenBench, CoMET-Bench, LVBench, and LSDBench.
  • Main caveat: reproducibility depends on the external skill updater and details deferred outside the main body.

tab:grounding-main

Caption: Gains from policy--tool coevolution on ultra-long video temporal grounding. Bold: best; AcademicBlueblue: gains over the base skill, with AcademicBlue$ $ indicating relative improvements.

Content:
MethodVUE-LVTR Precision AUC โ†‘VUE-LVTR Recall AUC โ†‘VUE-LVTR IoU AUC โ†‘VUE-LVTR IoU@0.5 โ†‘ExtremeWhenBench Mean IoU โ†‘ExtremeWhenBench Recall@0.5 โ†‘CoMET-Bench Mean IoU โ†‘CoMET-Bench Recall@0.5 โ†‘CoMET-Bench F1@0.5 โ†‘
Qwen3.5-27B + Base Skill0.47060.47890.41070.42670.15070.14780.13550.14400.1529
Qwen3.5-27B + Evolved Skill0.59180.57730.51370.56030.26360.27850.16350.17020.1831

Why It Matters: This table contains the core reported grounding comparison between the base skill and evolved skill on the main temporal-grounding benchmarks.

๐ŸŽฏ Introduction

The paper studies ultra-long video temporal grounding: given a long video and a natural-language query, identify the temporal interval or intervals corresponding to the described event. The main-body motivation is operational: agents for footage retrieval, clip extraction, and editing assistance need to map semantic intent onto concrete timeline segments.

The prior limitation is that very long videos create a search-versus-detail tradeoff. Coarse observations can miss brief events or key details, while repeated fine-grained inspection is visually expensive. Existing multi-step and tool-using agents improve evidence acquisition, but their observation mechanisms and media-processing capabilities are largely predefined.

CoEvoWhen's paper-level objective is to leverage task experience to jointly improve both tool capabilities and the policies that guide tool use. The authors argue that image-based observations support broad temporal coverage and distant-candidate comparison, while video-based observations support action order, state changes, continuity, and boundary judgment.

๐Ÿ”ฌ Methodology

The method represents task experience as a reusable external skill. At round k, the skill contains high-level policies for planning and observation orchestration plus executable media tools. Each tool has a description, interface specification, and source code, and the number of tools may change as tools are created, consolidated, or retired.

During skill update, the frozen VLM executes a batch of grounding tasks with the current skill. Each trajectory records the complete execution process and final prediction, and the updater pairs trajectories with ground-truth intervals. The updater then distills reusable policy changes and uses coding capability to upgrade or create media tools.

At inference, the evolved skill is fixed. The VLM starts from an empty history, selects tool invocations conditioned on query and accumulated evidence, receives observations from the selected tool, and appends the interaction to history. It stops when evidence is sufficient or when it reaches the maximum inference-round limit, reported as at most 24 rounds per query.

  • Base skill: basic temporal-grounding protocol plus primitive image, video, and auxiliary operations.
  • Evolution data: 100 challenging VUE-LVTR queries from distinct videos.
  • Updater: Codex (GPT-5.5, xhigh) in the reported setup.
  • Inference: evolved policy in the system prompt, evolved tools registered as callable functions, thinking enabled, greedy decoding.

display equation 2

Formula:
\[\mathcal S^{(k)} = \left(\Pi^{(k)},\mathcal T^{(k)}\right), \qquad \mathcal T^{(k)} = \{\tau_j^{(k)}\}_{j=1}^{J_k}.\]

Meaning: Defines the external skill at evolution round k as the combination of high-level policies and a set of executable media tools.

display equation 3

Formula:
\[\mathcal B_k = \left\{\left(\mathcal R_i(\mathcal S^{(k)}),\mathcal Y_i^\ast\right)\right\}_{i\in\mathcal I_k}, \quad \Delta\mathcal S^{(k)} = \mathcal U\left(\mathcal S^{(k)},\mathcal B_k\right) = \left(\Delta\Pi^{(k)},\Delta\mathcal T^{(k)}\right).\]

Meaning: Defines the feedback batch from execution trajectories and ground-truth intervals, and the updater that produces policy and tool updates.

display equation 4

Formula:
\[\Pi^{(k+1)} = \Pi^{(k)}\oplus\Delta\Pi^{(k)}, \quad \mathcal T^{(k+1)} = \mathcal T^{(k)}\oplus\Delta\mathcal T^{(k)} = \left\{\tau_j^{(k)}\oplus\Delta\tau_j^{(k)}\right\}_{j\in\mathcal J_k^{\mathrm{ret}}} \cup\mathcal T_k^{\mathrm{new}}.\]

Meaning: Formalizes how policy patches and tool patches are applied, including retaining updated tools and adding newly created or consolidated tools.

display equation 5

Formula:
\[a_t=(j_t,\eta_t) \sim M_\theta \left( \cdot\mid q_i,H_{t-1};\Pi^\ast,\mathcal T^\ast \right), \quad H_t = H_{t-1}\mathbin{\Vert} \left( a_t, f_{j_t}^\ast(V_i;\eta_t) \right).\]

Meaning: Describes inference-time tool selection and history accumulation: the VLM chooses a tool and arguments conditioned on the query, prior history, evolved policy, and evolved tools.

๐Ÿ“Š Experiments

The main-body evaluation covers three ultra-long video temporal-grounding benchmarks: VUE-LVTR, ExtremeWhenBench, and CoMET-Bench. VUE-LVTR uses visual queries from VUE-TR and VUE-TR-V2 restricted to videos of at least 30 minutes; ExtremeWhenBench videos average approximately 76 minutes; CoMET-Bench covers multi-event grounding in videos of at least 30 minutes. Transfer is assessed on long-video QA benchmarks LVBench and LSDBench.

Baselines include the base skill and evolved skill under identical task protocols and inference settings, plus open-source VLMs Qwen3.5-27B, InternVL3.5-8B, and TimeLens-7B, closed-source VLMs Gemini 2.5 Flash and GPT-5.6 Luna, and agent-based baselines VideoMind-7B and EvoGround-7B. Performance uses official benchmark metrics; efficiency uses cumulative visual-token cost averaged over queries with separate image/video costs, plus counts of model calls, local tool calls, and visual observations.

The authors report that the evolved skill achieves the best results on all reported grounding metrics in the main comparison. Against the base skill, Qwen3.5-27B improves VUE-LVTR IoU AUC from 0.4107 to 0.5137, ExtremeWhenBench Mean IoU from 0.1507 to 0.2636, and CoMET-Bench Mean IoU from 0.1355 to 0.1635. Efficiency also improves: on VUE-LVTR, visual tokens fall from 202.58k/query to 141.89k/query.

Ablations in the main body state that joint policy-tool evolution beats policy-only and tool-only variants by 0.0243 and 0.0735 IoU AUC while reducing visual-token cost by 40.6% and 25.5%, respectively. The image-video ablation is described qualitatively in the extract: image+video coordination outperforms single-modality skills across grounding metrics while using fewer visual tokens, but the numeric table for that ablation was not included in the supplied extracted tables.

  • Main grounding benchmarks: VUE-LVTR, ExtremeWhenBench, CoMET-Bench.
  • Transfer benchmarks: LVBench and LSDBench.
  • Metric direction: AUC, IoU, Recall@0.5, F1, Rejection-F1, and Overall Accuracy are reported as higher-is-better; visual tokens and call/observation counts are lower-is-better efficiency metrics.
  • Key efficiency mechanism: on ExtremeWhenBench, image tokens increase from 179.3k to 197.5k while video tokens fall from 47.8k to 3.7k, indicating a shift away from expensive video observation.

tab:efficiency

Caption: Inference efficiency with the base and evolved skills. EfficiencyGreenGreen ($ $) and AllocationPurplepurple ($ $) indicate decreases and increases relative to the base skill.

Content:
BenchmarkMethodVisual Tokens โ†“ (k/query)Image Tokens (k/query)Video Tokens (k/query)Model Calls โ†“Local Tool Calls โ†“Visual Obs. โ†“Image Obs.Video Obs.
VUE-LVTRQwen3.5-27B + Base Skill202.58162.0140.5717.038.168.117.350.76
VUE-LVTRQwen3.5-27B + Evolved Skill141.89141.370.529.634.793.863.820.04
VUE-LVTREvolved โˆ’ Base-60.69-20.64-40.05-7.40-3.37-4.25-3.53-0.72
VUE-LVTRRelative change29.96%12.74%98.72%43.5%41.3%52.4%48.0%94.7%
ExtremeWhenBenchQwen3.5-27B + Base Skill227.10179.2647.8416.658.986.935.691.24
ExtremeWhenBenchQwen3.5-27B + Evolved Skill201.17197.493.6811.195.694.564.350.21

Why It Matters: This table shows that the evolved skill is not merely more accurate; it reduces visual-token and interaction overhead, especially video-token usage.

tab:generalization

Caption: Generalization across VLMs and transfer to long-video QA. Bold: results with the evolved skill; AcademicBlueblue: gains over the base skill, with AcademicBlue$ $ indicating relative improvements.

Content:
BackboneSkillVUE-LVTR Precision AUC โ†‘VUE-LVTR Recall AUC โ†‘VUE-LVTR IoU AUC โ†‘ExtremeWhenBench Mean IoU โ†‘ExtremeWhenBench Recall@0.5 โ†‘CoMET-Bench Mean IoU โ†‘CoMET-Bench Rejection F1 โ†‘LVBench Overall Acc. (%) โ†‘LSDBench Overall Acc. (%) โ†‘
Qwen3.5-9B+ Base Skill0.32400.34250.24050.06260.05760.067267.5140.0949.08
Qwen3.5-9B+ Evolved Skill0.42870.47070.32440.12470.11700.085268.8045.8453.68
Qwen3.5-9BEvolution Gain+0.1047+0.1282+0.0839+0.0621+0.0594+0.0180+1.29+5.75+4.60
Qwen3.5-9BRelative gain32.3%37.4%34.9%99.2%103.1%26.8%1.9%14.3%9.4%

Why It Matters: This table supports the claim that the coevolved skill generalizes beyond the main Qwen3.5-27B setup and transfers to long-video QA.

fig:policy_tool_coevolution

fig:policy_tool_coevolution

Caption: Policy--tool coevolution framework for ultra-long video temporal grounding. A frozen VLM executes grounding tasks with an external skill comprising high-level policies and executable media tools. Across evolution rounds, an external skill updater jointly refines the policies and evolves the tools from execution trajectories, synthesizing code to upgrade existing tools or create new ones.

Why It Matters: This is the central architecture figure: it shows the frozen VLM, the external skill, and the feedback loop that jointly updates policies and executable tools.

fig:agentic_inference

fig:agentic_inference

Caption: Policy-guided agentic inference with the evolved skill. Guided by the evolved policy, the VLM autonomously orchestrates media tools, using accumulated evidence to decide on subsequent observations without relying on a separate, stronger planning model. The illustrated trajectory combines image-based global search and candidate refinement with selective video verification to localize the queried event.

Why It Matters: This figure explains the runtime behavior builders would need to implement: iterative evidence acquisition through tool calls, governed by the evolved policy.

fig:evolution-dynamics

fig:evolution-dynamics

Caption: [t] figures/skill_evolution_fig.pdf minipage[t]0.375 0pt >0pt minipage[c][ ][c] figures/evolution_dynamics_fig.pdf minipage figures/evolution_dynamics_fig.pdf figure performance--cost dynamics during skill evolution. DynamicsBlueBlue and EfficiencyGreengreen show changes in IoU and visual tokens relative to the base skill on matched queries. fig:evolution-dynamics minipage minipage[t]0.605 0pt >0pt minipage[c][ ][c] figures/skill_evolution_fig.pdf minipage figures/skill_evolution_fig.pdf figure...

Why It Matters: Despite malformed extraction, the caption verifies that the figure reports performance-cost dynamics during evolution, directly supporting the claim that the skill improves accuracy while reducing visual-token use.

๐Ÿ”ฎ Conclusion

The paper concludes that jointly evolving policies and executable media tools can turn long-video grounding experience into a reusable skill. The evolved skill lets the VLM coordinate image and video observations autonomously, without a stronger external planner and without parameter updates.

The practical takeaway is that long-video agents can become more accurate and cheaper by learning better observation strategies and better observation tools. In the reported main-body evidence, the gains are not limited to the evolution benchmark: improvements appear across grounding benchmarks, multiple VLM backbones, and long-video QA transfer.

๐Ÿ› ๏ธ Future Research Improvements

A natural next step is broader evolution-data coverage. The main-body setup evolves on 100 challenging VUE-LVTR queries, so future work should verify whether the same coevolution procedure remains stable across larger and more diverse video domains, event types, and target-absent distributions.

Reproducibility would benefit from releasing the evolved policies, executable tools, trajectory logs, and updater prompts. The method's key mechanism is external skill evolution, so builders need to inspect exactly how code changes and policy changes are proposed, validated, and retained.

Future evaluations should separate the effects of model backbone, updater capability, tool-library design, and inference-round budget. The current evidence supports policy-tool coevolution, but production adopters will need to know which component is most responsible for the gains under their cost constraints.

  • Test evolution on non-video-editing domains such as industrial monitoring, medical procedure videos, and robotics logs.
  • Add guardrails for tool-code generation, including sandboxing, regression tests, and rollback criteria.
  • Evaluate whether smaller or open-source skill updaters can reproduce the reported gains.

๐Ÿญ Potential Industry Use Scenarios

The most direct use case is long-form media search and editing assistance. A system could use CoEvoWhen-like skills to find moments in recordings, produce candidate intervals, refine boundaries, and verify motion continuity before generating clips or edits.

Another plausible deployment area is enterprise video review: security footage, training recordings, compliance sessions, livestream archives, or customer-support videos. The paper's emphasis on reducing visual-token cost is particularly relevant when many hour-scale videos must be searched under cost or latency limits.

The long-video QA transfer evidence suggests the evolved skill may be useful beyond strict timestamp localization. However, the main-body evidence only reports LVBench and LSDBench accuracy gains, so product teams should treat broader long-video understanding claims as promising but requiring domain-specific validation.

  • Video archive search with sparse target events.
  • Clip extraction and timeline editing assistants.
  • Compliance and safety review over long recordings.
  • Long-video QA systems that need targeted evidence acquisition instead of full-video ingestion.

๐Ÿ’ฌ Critical Analysis

The strongest aspect of the paper is the coupling between policy evolution and tool evolution. The main-body ablation supports the argument: policy-only and tool-only variants improve over the base skill, but joint evolution achieves a better accuracy-cost combination, exceeding policy-only and tool-only by 0.0243 and 0.0735 IoU AUC while reducing visual-token cost by 40.6% and 25.5%.

The main reproducibility concern is that the skill updater is a powerful external coding agent, and the most operationally important details are not fully available in the main-body evidence supplied here. For builders, the central risk is not whether the idea is plausible; it is whether the update loop can be made stable, inspectable, and repeatable outside the authors' exact setup.

The evidence is compelling for the reported benchmarks but still benchmark-centered. Because tools can shape what the VLM sees, failures may come from undersampling rare visual evidence, overconfident termination, or tool-generated presentations that omit context. Any deployment should log observation choices and final answers together so that missed events and boundary errors can be audited.

Overall, CoEvoWhen is a strong builder-facing paper because it offers a concrete path away from brute-force long-context video processing: evolve compact, task-aware evidence acquisition. The remaining open question is whether the coevolution machinery can be productized with enough reliability, safety, and transparency.

Original Abstract

Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters. During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition. Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model. Extensive experiments spanning five benchmarks and three VLMs show that policy-tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.