Apple releases LVSum, a benchmark for timestamp-aware long video summarization with 72 videos across 13 domains
The benchmark includes 72 videos averaging 16 minutes each, with up to 10 human-generated summaries per video containing temporal references. Evaluations of leading MLLMs reveal gaps in temporal grounding and cross-modal coherence.
1 source · cross-referenced
- Apple’s Machine Learning Research team introduced LVSum, a human-annotated benchmark for evaluating long-form video summarization with fine-grained temporal alignment.
- The benchmark consists of 72 diverse videos spanning 13 domains, with an average duration of 16 minutes per video.
- Each video is annotated with up to 10 human-generated summaries that include temporal references.
- Evaluations of proprietary and open-source MLLMs using LLM-based and standard metrics identified systematic weaknesses in temporal grounding, instruction adherence, and cross-modal coherence.
Apple’s Machine Learning Research team introduced LVSum, a human-annotated benchmark designed to evaluate long-form video summarization with fine-grained temporal alignment. The benchmark targets a key challenge for multimodal large language models (MLLMs): maintaining temporal fidelity over extended durations while producing summaries that are both semantically and temporally grounded.
LVSum comprises 72 diverse videos spanning 13 domains, with an average duration of 16 minutes per video. Each video is annotated with up to 10 human-generated summaries that include temporal references, providing a robust dataset for evaluating summarization quality and temporal alignment.
In a comprehensive evaluation of leading proprietary and open-source MLLMs, researchers used newly introduced LLM-based metrics for content relevance and modality coherence, alongside standard automatic metrics. The experiments revealed three key findings: transcripts contribute substantially more to summarization quality than visual frames alone; a significant performance gap persists between model-generated and human-written summaries; and current MLLMs exhibit systematic weaknesses in temporal grounding, instruction adherence, and cross-modal coherence.
The benchmark’s design and findings underscore the limitations of current MLLMs in handling long-form video summarization tasks, particularly in maintaining temporal awareness and cross-modal coherence. This highlights the need for further research and development to address these gaps, which are critical for applications requiring precise temporal reasoning.
- Aug 28, 2026 · Google DeepMind — Blog
Google DeepMind pilots first double-blind AI evaluations with cryptographic safeguards
Trust79 - Aug 26, 2026 · arXiv cs.AI
New NL2SQL benchmark shows enterprise database complexity degrades model performance
Trust79 - Aug 26, 2026 · arXiv cs.AI
New benchmark shows how reader-facing memory formats affect LLM evaluation scores
Trust79