๐Ÿค– AI ่ต„่ฎฏ

ยท ยท โ†—
← ่ฟ”ๅ›žๅˆ—่กจ

SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

arXiv cs.CV2026-09-16 04:00:00็ฎ—ๅŠ›่Šฏ็‰‡,ๆŽจ็†ๆ€่€ƒ,ๅŠžๅ…ฌๆ•ˆ็އ,ๆจกๅž‹่ฏ„ๆต‹,ๆ‹›่˜HR,่ฎบๆ–‡ๅŽŸๆ–‡ โ†—

arXiv:2609.16233v1 Announce Type: new

Abstract: Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clouds that capture geometry but discard rich visual features like texture, text, and materials. Second, annotations treat objects in isolation while ignoring real-world hierarchical organization (scenes, rooms, functional areas, object groups). Third, evaluation tasks focus narrowly on basic recognition rather than multi-step spatial reasoning.

In this context, we introduce SceneBench, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects. These annotations are produced through a human-in-the-loop pipeline combining vision-language models with roughly 1,500 human-hours of iterative refinement and verification, producing over 183K annotated nodes with textual descriptions and 3D bounding boxes. Building on this representation, we define three evaluation tasks: Existence-Based Questions probing object attributes, Spatial Intelligence Questions covering counting, size comparison, distance, and directional relations, and Grounded Question-Reasoning-Answer (QRA) triplets requiring multi-step reasoning across semantic levels. Experiments with state-of-the-art vision-language models show that while models perform well on basic recognition tasks (e.g., up to 85% accuracy for detection), performance drops substantially on hierarchical and compositional reasoning (e.g., down to 60% for counting), revealing limitations not captured by existing benchmarks. SceneBench provides a realistic testbed for developing and evaluating models capable of fine-grained spatial reasoning in photorealistic 3D environments.