This is a driving video. Please determine whether the vehicle in the green box shows subtle movement during the video. Answer Yes or No.
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
Overview
ViSTR-Bench is designed to systematically evaluate whether MLLMs can reason from continuous visual cues in dynamic scenes. This project page follows the paper through four parts: Task Definition, Benchmark statistics, Construction pipeline, and Evaluation Results.
ViSTR-Bench is organized into four complementary dimensions, namely Motion Perception, Spatial Relations, Outcome Prediction, and Physical Dynamics.
Benchmark statisticsThe benchmark comprises 15 distinct subtasks and 1,340 high-quality video question-answer pairs spanning diverse scenarios.
Construction pipelineViSTR-Bench is constructed through data collection, video preprocessing, QA pair generation, and human quality control.
Evaluation ResultsComprehensive evaluations reveal substantial bottlenecks in complex spatial-temporal reasoning.
Task Definition
This dimension evaluates whether MLLMs can perceive and compare object motion from continuous visual observations. The core challenge is to distinguish true object motion from apparent visual changes caused by camera movement, viewpoint variation, or environmental distractors.
This dimension focuses on reasoning about evolving spatial relationships and geometric constraints in dynamic scenes, requiring models to account for camera motion, object positions, scene layout, and spatial clearance over time.
This dimension evaluates whether MLLMs can infer future outcomes from historical motion cues before the final outcomes are directly observed.
This dimension examines whether MLLMs can reason about latent physical properties, dependencies, and stability conditions revealed through dynamic interactions.
Benchmark statistics
ViSTR-Bench consists of 15 subtasks and 1,340 high-quality video question-answer pairs, spanning tabletop, indoor, and outdoor scenes. The data are collected from multiple sources, including public datasets, web videos, and self-collected recordings.
Dimensions and subtasks
Hover over each segment to view detailed benchmark statistics.
Construction pipeline
We construct ViSTR-Bench through a systematic multi-stage pipeline, including data collection, video preprocessing, question-answer (QA) pair generation, and human quality control. This pipeline is designed to ensure that each sample requires temporal evidence, supports high-level reasoning, and has an unambiguous qualitative answer.
Data Collection
To cover diverse spatio-temporal reasoning scenarios across tabletop, indoor, and outdoor environments, we collect videos from established public datasets, curated web videos, and self-collected recordings.
Video Preprocessing
Raw videos inherently contain multiple distinct events, irrelevant temporal context, and explicit outcomes that could trivialize reasoning tasks into simple recognition.
QA Pair Generation
We formulate all instances as binary-choice questions using task-specific templates, and the order of the two candidate options is systematically randomized for each question.
Human Quality Control
Expert annotators manually filter all candidate QA pairs according to four strict criteria: visibility, temporal sufficiency, objectivity, and non-triviality.
Evaluation Results
Existing MLLMs achieve limited performance on ViSTR-Bench, indicating that qualitative spatial-temporal reasoning from continuous visual cues remains a challenging problem for current MLLMs.
The best-performing model, GPT-5.4-thinking, obtains an overall accuracy of 62.0%.
The best-performing model is only 4.1% higher than Chance Level (Frequency).
Human evaluation remains 29.0% above the best-performing model.
Main results
Accuracy (%) summarized across the four reasoning dimensions in ViSTR-Bench. Overall accuracy is averaged over all video question-answer pairs.
| Method | Rank | Average | Motion Perception |
Spatial Relations |
Outcome Prediction |
Physical Dynamics |
|---|---|---|---|---|---|---|
| Baseline | ||||||
| Chance Level (Random) | - | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 |
| Chance Level (Frequency) | - | 57.9 | 58.4 | 63.2 | 54.2 | 60.3 |
| Proprietary General MLLMs | ||||||
| GPT-5.4 | 5 | 56.1 | 63.0 | 67.8 | 49.4 | 48.9 |
| GPT-5.4-thinking | 1 | 62.0 | 69.8 | 74.4 | 51.9 | 60.7 |
| Gemini-3-Flash-Preview-thinking | 15 | 50.7 | 44.0 | 61.2 | 50.9 | 48.9 |
| Gemini-3.1-Flash-Lite-Preview-thinking | 12 | 52.6 | 49.3 | 63.6 | 48.5 | 55.5 |
| Gemini-3.1-Pro-Preview-thinking | 9 | 53.6 | 49.3 | 71.1 | 50.6 | 48.5 |
| Claude-Sonnet-4.6 | 14 | 51.0 | 52.2 | 59.9 | 51.1 | 39.3 |
| Claude-Sonnet-4.6-thinking | 11 | 53.4 | 51.9 | 59.9 | 53.4 | 48.9 |
| Claude-Opus-4.6 | 6 | 55.4 | 52.5 | 64.0 | 51.9 | 59.0 |
| Claude-Opus-4.6-thinking | 7 | 54.7 | 51.3 | 68.6 | 49.2 | 57.6 |
| Seed-2.0-Lite | 8 | 54.4 | 53.7 | 57.4 | 51.3 | 59.4 |
| Seed-2.0-Lite-thinking | 3 | 58.8 | 61.6 | 66.9 | 53.4 | 58.5 |
| Seed-2.0-Pro | 4 | 56.8 | 57.8 | 66.5 | 52.1 | 55.9 |
| Seed-2.0-Pro-thinking | 2 | 60.1 | 65.1 | 71.5 | 55.1 | 52.4 |
| MiMo-V2.5 | 13 | 52.4 | 50.7 | 63.6 | 49.6 | 49.3 |
| MiMo-V2.5-thinking | 10 | 53.6 | 48.4 | 64.5 | 50.4 | 57.2 |
| Open-source General MLLMs | ||||||
| LLaVA-OneVision-1.5-4B | 14 | 47.6 | 46.0 | 58.7 | 47.0 | 39.7 |
| LLaVA-OneVision-1.5-8B | 11 | 49.7 | 54.5 | 57.9 | 47.2 | 39.7 |
| Qwen3.5-27B-thinking | 1 | 55.0 | 54.0 | 66.5 | 53.8 | 47.2 |
| Qwen3.5-35B-A3B-thinking | 13 | 49.0 | 46.3 | 61.6 | 46.4 | 45.4 |
| Qwen3.5-122B-A10B-thinking | 3 | 53.4 | 51.3 | 67.8 | 47.5 | 55.0 |
| Qwen3.5-397B-A17B-thinking | 2 | 54.4 | 54.8 | 68.6 | 49.8 | 49.3 |
| InternVL3.5-8B | 8 | 50.9 | 55.4 | 64.5 | 46.0 | 41.0 |
| InternVL3.5-14B | 7 | 51.3 | 52.2 | 59.5 | 48.9 | 46.7 |
| InternVL3.5-38B | 5 | 52.3 | 42.8 | 64.5 | 49.6 | 59.8 |
| InternVL3.5-30B-A3B | 9 | 50.1 | 51.3 | 62.8 | 45.6 | 45.4 |
| InternVL3.5-241B-A28B | 4 | 53.4 | 51.0 | 69.4 | 47.2 | 54.1 |
| Intern-S1 | 6 | 51.7 | 52.5 | 62.4 | 47.7 | 48.5 |
| Intern-S1-Pro | 12 | 49.7 | 48.1 | 64.0 | 45.3 | 47.2 |
| GLM-4.6V-thinking | 10 | 49.9 | 49.3 | 63.6 | 46.2 | 44.5 |
| Specialized Spatial MLLMs | ||||||
| Spacer-SFT-7B | 2 | 51.6 | 49.6 | 65.7 | 49.8 | 44.1 |
| VG-LLM-4B | 7 | 48.2 | 49.9 | 57.0 | 47.5 | 38.0 |
| VG-LLM-8B | 9 | 47.6 | 46.3 | 52.5 | 48.3 | 42.8 |
| Spatial-MLLM-135K-4B | 10 | 47.3 | 46.0 | 56.2 | 45.8 | 43.2 |
| Spatial-MLLM-820K-4B | 3 | 51.3 | 55.7 | 56.6 | 48.5 | 45.9 |
| SpatialLadder-3B | 8 | 47.8 | 49.9 | 47.9 | 48.7 | 42.8 |
| Spatial-SSRL-7B | 5 | 50.3 | 44.9 | 64.9 | 49.1 | 45.9 |
| VST-7B-RL | 6 | 49.4 | 51.9 | 50.0 | 48.5 | 47.2 |
| GeoThinker-Qwen2.5VL-7B | 1 | 52.8 | 51.0 | 56.6 | 48.5 | 61.6 |
| GeoThinker-Qwen3VL-8B | 4 | 50.5 | 55.4 | 52.1 | 49.1 | 45.0 |
| Human Evaluation | ||||||
| Human | - | 91.0 | 96.2 | 96.5 | 84.0 | 93.9 |
Current MLLMs remain far from human-level spatial-temporal reasoning.
Proprietary general-purpose MLLMs perform best overall.
Spatial MLLMs show limited generalization to dynamic reasoning.
Thinking mode is helpful but not consistently reliable.
Performance varies substantially across reasoning dimensions.
Citation
Please cite ViSTR-Bench if you find the benchmark useful for your research.
@article{vistr-bench,
title={ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?},
author={Li, Han and Liu, Si and Huang, Zehao and Lyu, Dongxin and Xu, Longfei and Fu, Jiahui and Tian, Daxin and Xiu, Yuliang and Wang, Naiyan},
journal={arXiv preprint arXiv:2607.20868},
year={2026}
}