PAPER

ARROW: Arbitrary Reconstruction and Tracking of 4D Observations in the Wild

CV VSER Other CV
动态场景可通过移动的相机、多路视频流,或在不同时刻拍摄的图像进行采集。这些观测结果揭示了场景几何结构与运动状态的互补信息;然而,要将它们有效融合,必须在不同视角、不同采集时刻以及可见性变化之间建立准确的对应关系。我们提出了ARROW——一种前馈式模型,能够从任意图像集合中统一完成三维重建与三维点跟踪任务。该模型的核心是一种新颖的、与输入顺序无关的查询机制,使得查询能够灵活地关联到任意输入中的观测数据。我们发现,在训练过程中向模型提供更加多样化的输入数据集,可显著提升其各项任务性能。此外,由此训练得到的模型还具备良好的泛化能力,可拓展应用于更广泛的任务,例如多视角目标跟踪。采用该训练策略后,ARROW在WorldTrack和TAPVid-3D数据集上的三维跟踪任务中均刷新了当前最优性能;在经适配后的纯RGB多视角跟踪基准MVTracker上,其表现亦超越了专为多视角跟踪设计的专用方法;同时,在各类三维重建任务中仍保持极具竞争力的性能。本工作的代码与预训练权重已全部开源。
Dynamic scenes may be captured by a moving camera, multiple video streams, or images taken at different times. These observations reveal complementary aspects of scene geometry and motion, yet bringing them together requires establishing correspondence across viewpoints, capture times, and visibility changes. We introduce ARROW, a feed-forward model that unifies 3D reconstruction and 3D point tracking from arbitrary image sets. At its core is a novel order-invariant querying approach, which allows the association of queries with observations across arbitrary inputs. We show that exposing the model to more diverse sets of inputs during training results in improved task performance. Moreover, the resulting model is capable of generalization to a wider range of tasks including multi-view tracking. Trained with this strategy, ARROW establishes a new state of the art in 3D tracking on WorldTrack and TAPVid-3D and outperforms dedicated multi-view trackers on an adapted RGB-only MVTracker benchmark, while remaining competitive across 3D reconstruction tasks. Code and weights are publicly available.
许愿