๐ Highlights
CoEvoWhen treats ultra-long video temporal grounding as an agentic evidence-acquisition problem: a frozen VLM must decide where to search, what to inspect, and when evidence is sufficient.
The central claim is that policies and tools should coevolve. Policies determine evidence needs and orchestration; tools determine what visual evidence can be sampled and presented.
The strongest main-body result is that the evolved Qwen3.5-27B skill improves VUE-LVTR IoU AUC from 0.4107 to 0.5137 and ExtremeWhenBench Mean IoU from 0.1507 to 0.2636 while reducing VUE-LVTR visual tokens from 202.58k/query to 141.89k/query.
- Task: localize natural-language event intervals in videos lasting tens of minutes to hours.
- Mechanism: evolve an external skill containing high-level policies and executable media tools, without changing VLM weights.
- Benchmarks: VUE-LVTR, ExtremeWhenBench, CoMET-Bench, LVBench, and LSDBench.
- Main caveat: reproducibility depends on the external skill updater and details deferred outside the main body.
tab:grounding-main
Caption: Gains from policy--tool coevolution on ultra-long video temporal grounding. Bold: best; AcademicBlueblue: gains over the base skill, with AcademicBlue$ $ indicating relative improvements.
| Method | VUE-LVTR Precision AUC โ | VUE-LVTR Recall AUC โ | VUE-LVTR IoU AUC โ | VUE-LVTR IoU@0.5 โ | ExtremeWhenBench Mean IoU โ | ExtremeWhenBench Recall@0.5 โ | CoMET-Bench Mean IoU โ | CoMET-Bench Recall@0.5 โ | CoMET-Bench F1@0.5 โ |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3.5-27B + Base Skill | 0.4706 | 0.4789 | 0.4107 | 0.4267 | 0.1507 | 0.1478 | 0.1355 | 0.1440 | 0.1529 |
| Qwen3.5-27B + Evolved Skill | 0.5918 | 0.5773 | 0.5137 | 0.5603 | 0.2636 | 0.2785 | 0.1635 | 0.1702 | 0.1831 |
Why It Matters: This table contains the core reported grounding comparison between the base skill and evolved skill on the main temporal-grounding benchmarks.
๐ฏ Introduction
The paper studies ultra-long video temporal grounding: given a long video and a natural-language query, identify the temporal interval or intervals corresponding to the described event. The main-body motivation is operational: agents for footage retrieval, clip extraction, and editing assistance need to map semantic intent onto concrete timeline segments.
The prior limitation is that very long videos create a search-versus-detail tradeoff. Coarse observations can miss brief events or key details, while repeated fine-grained inspection is visually expensive. Existing multi-step and tool-using agents improve evidence acquisition, but their observation mechanisms and media-processing capabilities are largely predefined.
CoEvoWhen's paper-level objective is to leverage task experience to jointly improve both tool capabilities and the policies that guide tool use. The authors argue that image-based observations support broad temporal coverage and distant-candidate comparison, while video-based observations support action order, state changes, continuity, and boundary judgment.
๐ฌ Methodology
The method represents task experience as a reusable external skill. At round k, the skill contains high-level policies for planning and observation orchestration plus executable media tools. Each tool has a description, interface specification, and source code, and the number of tools may change as tools are created, consolidated, or retired.
During skill update, the frozen VLM executes a batch of grounding tasks with the current skill. Each trajectory records the complete execution process and final prediction, and the updater pairs trajectories with ground-truth intervals. The updater then distills reusable policy changes and uses coding capability to upgrade or create media tools.
At inference, the evolved skill is fixed. The VLM starts from an empty history, selects tool invocations conditioned on query and accumulated evidence, receives observations from the selected tool, and appends the interaction to history. It stops when evidence is sufficient or when it reaches the maximum inference-round limit, reported as at most 24 rounds per query.
- Base skill: basic temporal-grounding protocol plus primitive image, video, and auxiliary operations.
- Evolution data: 100 challenging VUE-LVTR queries from distinct videos.
- Updater: Codex (GPT-5.5, xhigh) in the reported setup.
- Inference: evolved policy in the system prompt, evolved tools registered as callable functions, thinking enabled, greedy decoding.
display equation 2
Meaning: Defines the external skill at evolution round k as the combination of high-level policies and a set of executable media tools.
display equation 3
Meaning: Defines the feedback batch from execution trajectories and ground-truth intervals, and the updater that produces policy and tool updates.
display equation 4
Meaning: Formalizes how policy patches and tool patches are applied, including retaining updated tools and adding newly created or consolidated tools.
display equation 5
Meaning: Describes inference-time tool selection and history accumulation: the VLM chooses a tool and arguments conditioned on the query, prior history, evolved policy, and evolved tools.
๐ Experiments
The main-body evaluation covers three ultra-long video temporal-grounding benchmarks: VUE-LVTR, ExtremeWhenBench, and CoMET-Bench. VUE-LVTR uses visual queries from VUE-TR and VUE-TR-V2 restricted to videos of at least 30 minutes; ExtremeWhenBench videos average approximately 76 minutes; CoMET-Bench covers multi-event grounding in videos of at least 30 minutes. Transfer is assessed on long-video QA benchmarks LVBench and LSDBench.
Baselines include the base skill and evolved skill under identical task protocols and inference settings, plus open-source VLMs Qwen3.5-27B, InternVL3.5-8B, and TimeLens-7B, closed-source VLMs Gemini 2.5 Flash and GPT-5.6 Luna, and agent-based baselines VideoMind-7B and EvoGround-7B. Performance uses official benchmark metrics; efficiency uses cumulative visual-token cost averaged over queries with separate image/video costs, plus counts of model calls, local tool calls, and visual observations.
The authors report that the evolved skill achieves the best results on all reported grounding metrics in the main comparison. Against the base skill, Qwen3.5-27B improves VUE-LVTR IoU AUC from 0.4107 to 0.5137, ExtremeWhenBench Mean IoU from 0.1507 to 0.2636, and CoMET-Bench Mean IoU from 0.1355 to 0.1635. Efficiency also improves: on VUE-LVTR, visual tokens fall from 202.58k/query to 141.89k/query.
Ablations in the main body state that joint policy-tool evolution beats policy-only and tool-only variants by 0.0243 and 0.0735 IoU AUC while reducing visual-token cost by 40.6% and 25.5%, respectively. The image-video ablation is described qualitatively in the extract: image+video coordination outperforms single-modality skills across grounding metrics while using fewer visual tokens, but the numeric table for that ablation was not included in the supplied extracted tables.
- Main grounding benchmarks: VUE-LVTR, ExtremeWhenBench, CoMET-Bench.
- Transfer benchmarks: LVBench and LSDBench.
- Metric direction: AUC, IoU, Recall@0.5, F1, Rejection-F1, and Overall Accuracy are reported as higher-is-better; visual tokens and call/observation counts are lower-is-better efficiency metrics.
- Key efficiency mechanism: on ExtremeWhenBench, image tokens increase from 179.3k to 197.5k while video tokens fall from 47.8k to 3.7k, indicating a shift away from expensive video observation.
tab:efficiency
Caption: Inference efficiency with the base and evolved skills. EfficiencyGreenGreen ($ $) and AllocationPurplepurple ($ $) indicate decreases and increases relative to the base skill.
| Benchmark | Method | Visual Tokens โ (k/query) | Image Tokens (k/query) | Video Tokens (k/query) | Model Calls โ | Local Tool Calls โ | Visual Obs. โ | Image Obs. | Video Obs. |
|---|---|---|---|---|---|---|---|---|---|
| VUE-LVTR | Qwen3.5-27B + Base Skill | 202.58 | 162.01 | 40.57 | 17.03 | 8.16 | 8.11 | 7.35 | 0.76 |
| VUE-LVTR | Qwen3.5-27B + Evolved Skill | 141.89 | 141.37 | 0.52 | 9.63 | 4.79 | 3.86 | 3.82 | 0.04 |
| VUE-LVTR | Evolved โ Base | -60.69 | -20.64 | -40.05 | -7.40 | -3.37 | -4.25 | -3.53 | -0.72 |
| VUE-LVTR | Relative change | 29.96% | 12.74% | 98.72% | 43.5% | 41.3% | 52.4% | 48.0% | 94.7% |
| ExtremeWhenBench | Qwen3.5-27B + Base Skill | 227.10 | 179.26 | 47.84 | 16.65 | 8.98 | 6.93 | 5.69 | 1.24 |
| ExtremeWhenBench | Qwen3.5-27B + Evolved Skill | 201.17 | 197.49 | 3.68 | 11.19 | 5.69 | 4.56 | 4.35 | 0.21 |
Why It Matters: This table shows that the evolved skill is not merely more accurate; it reduces visual-token and interaction overhead, especially video-token usage.
tab:generalization
Caption: Generalization across VLMs and transfer to long-video QA. Bold: results with the evolved skill; AcademicBlueblue: gains over the base skill, with AcademicBlue$ $ indicating relative improvements.
| Backbone | Skill | VUE-LVTR Precision AUC โ | VUE-LVTR Recall AUC โ | VUE-LVTR IoU AUC โ | ExtremeWhenBench Mean IoU โ | ExtremeWhenBench Recall@0.5 โ | CoMET-Bench Mean IoU โ | CoMET-Bench Rejection F1 โ | LVBench Overall Acc. (%) โ | LSDBench Overall Acc. (%) โ |
|---|---|---|---|---|---|---|---|---|---|---|
| Qwen3.5-9B | + Base Skill | 0.3240 | 0.3425 | 0.2405 | 0.0626 | 0.0576 | 0.0672 | 67.51 | 40.09 | 49.08 |
| Qwen3.5-9B | + Evolved Skill | 0.4287 | 0.4707 | 0.3244 | 0.1247 | 0.1170 | 0.0852 | 68.80 | 45.84 | 53.68 |
| Qwen3.5-9B | Evolution Gain | +0.1047 | +0.1282 | +0.0839 | +0.0621 | +0.0594 | +0.0180 | +1.29 | +5.75 | +4.60 |
| Qwen3.5-9B | Relative gain | 32.3% | 37.4% | 34.9% | 99.2% | 103.1% | 26.8% | 1.9% | 14.3% | 9.4% |
Why It Matters: This table supports the claim that the coevolved skill generalizes beyond the main Qwen3.5-27B setup and transfers to long-video QA.
fig:policy_tool_coevolution

Caption: Policy--tool coevolution framework for ultra-long video temporal grounding. A frozen VLM executes grounding tasks with an external skill comprising high-level policies and executable media tools. Across evolution rounds, an external skill updater jointly refines the policies and evolves the tools from execution trajectories, synthesizing code to upgrade existing tools or create new ones.
Why It Matters: This is the central architecture figure: it shows the frozen VLM, the external skill, and the feedback loop that jointly updates policies and executable tools.
fig:agentic_inference

Caption: Policy-guided agentic inference with the evolved skill. Guided by the evolved policy, the VLM autonomously orchestrates media tools, using accumulated evidence to decide on subsequent observations without relying on a separate, stronger planning model. The illustrated trajectory combines image-based global search and candidate refinement with selective video verification to localize the queried event.
Why It Matters: This figure explains the runtime behavior builders would need to implement: iterative evidence acquisition through tool calls, governed by the evolved policy.
fig:evolution-dynamics

Caption: [t] figures/skill_evolution_fig.pdf minipage[t]0.375 0pt >0pt minipage[c][ ][c] figures/evolution_dynamics_fig.pdf minipage figures/evolution_dynamics_fig.pdf figure performance--cost dynamics during skill evolution. DynamicsBlueBlue and EfficiencyGreengreen show changes in IoU and visual tokens relative to the base skill on matched queries. fig:evolution-dynamics minipage minipage[t]0.605 0pt >0pt minipage[c][ ][c] figures/skill_evolution_fig.pdf minipage figures/skill_evolution_fig.pdf figure...
Why It Matters: Despite malformed extraction, the caption verifies that the figure reports performance-cost dynamics during evolution, directly supporting the claim that the skill improves accuracy while reducing visual-token use.
๐ฎ Conclusion
The paper concludes that jointly evolving policies and executable media tools can turn long-video grounding experience into a reusable skill. The evolved skill lets the VLM coordinate image and video observations autonomously, without a stronger external planner and without parameter updates.
The practical takeaway is that long-video agents can become more accurate and cheaper by learning better observation strategies and better observation tools. In the reported main-body evidence, the gains are not limited to the evolution benchmark: improvements appear across grounding benchmarks, multiple VLM backbones, and long-video QA transfer.
๐ ๏ธ Future Research Improvements
A natural next step is broader evolution-data coverage. The main-body setup evolves on 100 challenging VUE-LVTR queries, so future work should verify whether the same coevolution procedure remains stable across larger and more diverse video domains, event types, and target-absent distributions.
Reproducibility would benefit from releasing the evolved policies, executable tools, trajectory logs, and updater prompts. The method's key mechanism is external skill evolution, so builders need to inspect exactly how code changes and policy changes are proposed, validated, and retained.
Future evaluations should separate the effects of model backbone, updater capability, tool-library design, and inference-round budget. The current evidence supports policy-tool coevolution, but production adopters will need to know which component is most responsible for the gains under their cost constraints.
- Test evolution on non-video-editing domains such as industrial monitoring, medical procedure videos, and robotics logs.
- Add guardrails for tool-code generation, including sandboxing, regression tests, and rollback criteria.
- Evaluate whether smaller or open-source skill updaters can reproduce the reported gains.
๐ญ Potential Industry Use Scenarios
The most direct use case is long-form media search and editing assistance. A system could use CoEvoWhen-like skills to find moments in recordings, produce candidate intervals, refine boundaries, and verify motion continuity before generating clips or edits.
Another plausible deployment area is enterprise video review: security footage, training recordings, compliance sessions, livestream archives, or customer-support videos. The paper's emphasis on reducing visual-token cost is particularly relevant when many hour-scale videos must be searched under cost or latency limits.
The long-video QA transfer evidence suggests the evolved skill may be useful beyond strict timestamp localization. However, the main-body evidence only reports LVBench and LSDBench accuracy gains, so product teams should treat broader long-video understanding claims as promising but requiring domain-specific validation.
- Video archive search with sparse target events.
- Clip extraction and timeline editing assistants.
- Compliance and safety review over long recordings.
- Long-video QA systems that need targeted evidence acquisition instead of full-video ingestion.
๐ฌ Critical Analysis
The strongest aspect of the paper is the coupling between policy evolution and tool evolution. The main-body ablation supports the argument: policy-only and tool-only variants improve over the base skill, but joint evolution achieves a better accuracy-cost combination, exceeding policy-only and tool-only by 0.0243 and 0.0735 IoU AUC while reducing visual-token cost by 40.6% and 25.5%.
The main reproducibility concern is that the skill updater is a powerful external coding agent, and the most operationally important details are not fully available in the main-body evidence supplied here. For builders, the central risk is not whether the idea is plausible; it is whether the update loop can be made stable, inspectable, and repeatable outside the authors' exact setup.
The evidence is compelling for the reported benchmarks but still benchmark-centered. Because tools can shape what the VLM sees, failures may come from undersampling rare visual evidence, overconfident termination, or tool-generated presentations that omit context. Any deployment should log observation choices and final answers together so that missed events and boundary errors can be audited.
Overall, CoEvoWhen is a strong builder-facing paper because it offers a concrete path away from brute-force long-context video processing: evolve compact, task-aware evidence acquisition. The remaining open question is whether the coevolution machinery can be productized with enough reliability, safety, and transparency.