Apple releases LVSum, a benchmark for timestamp-aware long video summarization with 72 videos across 13 domains
The benchmark includes 72 videos averaging 16 minutes each, with up to 10 human-generated summaries per video containing temporal references. Evaluations of leading MLLMs reveal gaps in temporal grounding and cross-modal coherence.
1 source · cross-referenced
- Apple’s Machine Learning Research team introduced LVSum, a human-annotated benchmark for evaluating long-form video summarization with fine-grained temporal alignment.
- The benchmark consists of 72 diverse videos spanning 13 domains, with an average duration of 16 minutes per video.
- Each video is annotated with up to 10 human-generated summaries that include temporal references.
- Evaluations of proprietary and open-source MLLMs using LLM-based and standard metrics identified systematic weaknesses in temporal grounding, instruction adherence, and cross-modal coherence.
Apple’s Machine Learning Research team introduced LVSum, a human-annotated benchmark designed to evaluate long-form video summarization with fine-grained temporal alignment. The benchmark targets a key challenge for multimodal large language models (MLLMs): maintaining temporal fidelity over extended durations while producing summaries that are both semantically and temporally grounded.
LVSum comprises 72 diverse videos spanning 13 domains, with an average duration of 16 minutes per video. Each video is annotated with up to 10 human-generated summaries that include temporal references, providing a robust dataset for evaluating summarization quality and temporal alignment.
In a comprehensive evaluation of leading proprietary and open-source MLLMs, researchers used newly introduced LLM-based metrics for content relevance and modality coherence, alongside standard automatic metrics. The experiments revealed three key findings: transcripts contribute substantially more to summarization quality than visual frames alone; a significant performance gap persists between model-generated and human-written summaries; and current MLLMs exhibit systematic weaknesses in temporal grounding, instruction adherence, and cross-modal coherence.
The benchmark’s design and findings underscore the limitations of current MLLMs in handling long-form video summarization tasks, particularly in maintaining temporal awareness and cross-modal coherence. This highlights the need for further research and development to address these gaps, which are critical for applications requiring precise temporal reasoning.
- Jul 14, 2026 · arXiv cs.CL
Researchers release CLIR-Bench to evaluate multimodal QA over irregular clinical time series
Trust79 - Jul 12, 2026 · Hamel Husain — applied AI engineering
Comparison finds automated evals correlate with human annotations in 100 traces
Trust79 - Jul 9, 2026 · OpenAI — News
OpenAI flags reliability issues in SWE-Bench Pro coding benchmark
Trust79