Researchers introduce benchmark testing multimodal models’ abstract perceptual reasoning
Acousto-kinematic word inference task reveals human–model performance gap and paradoxical fusion effect in leading systems like GPT-4o and Gemini 2.5-Pro.
1 source · cross-referenced
- A new benchmark called The Unwritten Benchmark evaluates multimodal models on abstract perceptual reasoning via acousto-kinematic word inference.
Researchers propose The Unwritten Benchmark to probe abstract perceptual reasoning in multimodal models, focusing on acousto-kinematic word inference. In this task, models must infer words being written across three writing styles using only the audio of pen scratches and video of hand movements, without visible ink.
Evaluation results show a stark contrast between human and machine performance: human participants achieve over 80% ordered letter accuracy, while leading multimodal models, including GPT-4o and Gemini 2.5-Pro, fail to surpass 10% accuracy.
The study also identifies a paradoxical fusion effect, where providing both audio and video modalities often degrades model performance instead of improving it, indicating a breakdown in the ability to synthesize complementary perceptual cues for this cognitive task.
The findings underscore significant limitations in current models’ cross-modal causal reasoning and their understanding of micro-kinematics essential for intuitive perceptual reasoning.
- Aug 18, 2026 · arXiv cs.AI
Replication study finds FLOPs-based efficiency metrics unreliable on modern hardware
Trust79 - Aug 18, 2026 · arXiv cs.AI
Study finds metacognitive sensitivity in medical LLMs but highlights calibration gaps in high-uncertainty cases
Trust79 - Aug 17, 2026 · arXiv cs.CL
Researchers propose BCMT architecture to reduce attention’s quadratic cost for long-context modeling
Trust79