PAPER DEEP DIVE
UniQueR: Unified Query-based Feedforward 3D Reconstruction
Reframes feedforward reconstruction from per-pixel 2.5D prediction to inference over a sparse set of 3D queries. Because queries live in global 3D space rather than in any single frame camera space, they can place Gaussians in regions never observed. By default 4096 queries spawn 64 Gaussians each, about 260K primitives, an order of magnitude fewer than per-pixel methods. On Mip-NeRF 360 with three views it reaches 22.70 PSNR (AnySplat 20.08) in a 0.213 s forward pass; at 32 views it uses 15x fewer primitives, 40 percent less memory and improves depth abs-rel from 0.062 to 0.038. In the dense regime feedforward-only quality trails AnySplat, but as initialization for per-scene optimization it overtakes it (25.14 vs 24.99 PSNR). No public code; not reproduced.
Paper
Title: UniQueR: Unified Query-based Feedforward 3D Reconstruction
Authors: Chensheng Peng, Quentin Herau, Jiezhi Yang, Yichen Xie, Yihan Hu, Wenzhao Zheng, Matthew Strong, Masayoshi Tomizuka, Wei Zhan
Affiliations: Applied Intuition, UC Berkeley, Stanford University
Preprint: arXiv:2603.22851v1 [cs.CV], 24 March 2026
Venue: ECCV 2026 (poster linked from the project page)
Project page: uniquer3d.github.io
Code: ❌ No public release at the time of writing. Neither the paper nor the project page links a repository, and no official implementation exists on GitHub. Reproducibility therefore rests entirely on the bases the authors name explicitly: the DINOv2 encoder and point-map head are initialized from Pi3 and frozen, data processing follows MapAnything, and the test-time optimization protocol follows AnySplat.
In one sentence
UniQueR reframes feedforward 3D reconstruction from predicting one depth value or Gaussian per pixel to letting a sparse set of learnable 3D queries infer the scene — because queries live in global 3D space rather than in any single frame's camera space, they can place Gaussians in regions no camera ever saw, completing occluded geometry in a single forward pass.
Background and motivation
Recovering 3D structure from 2D captures is a capability that robotics, autonomous driving and digital content creation all sit on top of. The traditional route runs through Structure-from-Motion, multi-view stereo and SLAM: feature matching and geometric optimization jointly solve for camera poses and dense point clouds. These are reliable under controlled conditions and fall apart when viewpoints are sparse, textures are missing, or the scene is visually ambiguous.
Deep learning brought a second route: neural radiance fields and 3D Gaussian Splatting. They turn reconstruction into a per-scene optimization, with excellent visual fidelity and multi-view consistency, at the cost of minutes to hours per scene — and they cannot transfer geometric priors learned from large-scale data, because every new scene starts fitting from scratch.
The third route is feedforward reconstruction. DUSt3R, VGGT and Pi3 showed that a transformer or cost volume can extract multi-view features and recover 3D geometry in a single forward pass. They typically emit intermediate 2.5D representations — depth maps, normal maps, point maps — and generalize well thanks to depth priors absorbed from massive datasets. The appeal is the prospect of a general-purpose 3D perception backbone for robotics and embodied AI.
A structural flaw hides in that design: these representations are pixel-aligned 2.5D. Outputs are tied to input pixels, so the model can only describe surfaces that were seen. Backsides no camera looked at, regions occluded by foreground — there is no pixel to hang an output on, so the reconstruction simply has holes there. Figure 1 draws that line clearly.
Figure 1: Pixel-aligned versus query-based pipelines. Pixel-aligned representations leave holes where nothing was observed; UniQueR queries cover occluded regions in global space.
Later work pushed feedforward outputs from point clouds toward more expressive representations. MVSplat and NoPoSplat predict pixel-aligned Gaussians directly, improving efficiency and completeness, but the pixel-alignment constraint itself never moves: a Gaussian can only be born at the backprojection of an input pixel. Removing that constraint is where UniQueR starts.
Prerequisites
3D Gaussian Splatting represents a scene as colored, opacity-carrying 3D primitives rendered through differentiable rasterization. Its value here is not rendering quality but that it provides a differentiable bridge from 3D geometry to 2D supervision: with a renderer, abundant RGB and depth signals can supervise 3D structure without costly ground-truth 3D annotations.
The second ingredient is the DETR-style query paradigm: a set of learnable queries aggregates information from image features through attention and decodes predictions directly. UniQueR ports this to 3D reconstruction and immediately hits a problem DETR never had — DETR3D has 3D boxes as supervision, so queries converge from random initialization, whereas 3D reconstruction has no per-query 3D ground truth and random initialization diverges during training. That tension is what produces the hybrid initialization design below.
Method
Overview
The input is $N$ unposed RGB images $\{\mathbf{I}_i\}_{i=1}^{N}$. The model maintains $Q$ learnable 3D queries $\mathcal{Q}=\{\mathbf{q}_i\}_{i=1}^{Q}$, each carrying a 3D position $\mathbf{p}_i$ and spawning $K$ 3D Gaussians. The union over all queries
$$\mathcal{G}=\{\mathbf{g}_{ij}\}_{i=1,j=1}^{Q,K}$$
represents the whole scene and renders differentiably under any camera view $\pi\in\mathbb{SE}(3)$ as $\hat{\mathbf{I}}=\mathrm{Render}(\mathcal{G},\pi)$. Defaults are $Q=4096$ and $K=64$, about 262K Gaussians in total.
In parallel, the image branch follows the VGGT and Pi3 paradigm, mapping inputs to per-frame 3D annotations — camera pose $\mathbf{P}_i$, point map $\mathbf{X}_i$ and confidence map $\mathbf{C}_i$:
$$f\big(\{\mathbf{I}_i\}_{i=1}^{N}\big)=\{(\mathbf{P}_i,\mathbf{X}_i,\mathbf{C}_i)\}_{i=1}^{N}$$
This branch does not produce the final geometry; it supplies priors that guide query updates.
Figure 2: UniQueR pipeline overview. A vision encoder extracts per-frame tokens and decodes poses and point maps; 3D queries are refined through cross-attention with image tokens and self-attention among queries; each query spawns Gaussians that are rendered by differentiable splatting under RGB and depth supervision.
Image tokenization and geometric priors
Each frame first goes through DINOv2, a vision transformer backbone chosen for strong semantic representations:
$$\mathbf{T}_i=\mathrm{DINO}(\mathbf{I}_i)\in\mathbb{R}^{HW/p^{2}\times d}$$
with patch size $p$ and token dimension $d$. An alternating-attention transformer then alternates intra-frame and inter-frame attention, following VGGT and Pi3:
$$\{\mathbf{T}_i^{l+1}\}=\mathrm{AA\text{-}Transformer}(\{\mathbf{T}_i^{l}\})$$
The aggregated tokens serve two purposes: one path decodes camera poses, point maps and confidence maps as geometric priors for observed surfaces; the other feeds the query transformer so queries can absorb image information.
Query initialization: half grounded, half exploring
This is the design most worth reading closely. Every query carries an explicit 3D position and must occupy some spatial region. Following DETR's random learnable initialization makes training diverge for lack of a geometric anchor, since no 3D ground truth exists to assign supervision targets per query.
The fix is a hybrid: half the queries are sampled from the non-metric point maps predicted in the first stage, landing them on coarse structure aligned with observed 2.5D surfaces; the other half are learnable anchors sampled uniformly in 3D space, free to explore and reconstruct under-reconstructed or unobserved regions. The first half buys stability, the second buys completeness. The ablation shows just how irreplaceable the second half is.
Decoupled attention: splitting a quadratic
Queries and image tokens enter the query transformer together:
$$\mathcal{Q}=\mathrm{QueryTransformer}(\mathcal{Q},\mathcal{T})$$
Queries are orders of magnitude fewer than the Gaussians in dense representations, which is where the efficiency comes from. But image tokens scale linearly with the number of input images, so attention cost runs away. The direct approach concatenates queries and image tokens and applies full self-attention:
$$\mathcal{Q},\mathcal{T}=\mathrm{Self\text{-}Attn}([\mathcal{Q},\mathcal{T}])$$
at a complexity of $\mathcal{O}((Q+NHW/p^{2})^{2})$ — prohibitively expensive at high resolution.
Figure 3: Two attention designs. Top: full self-attention over concatenated tokens. Bottom: the decoupled design used here — cross-attention from queries to images, then self-attention among queries.
The authors therefore decouple: cross-attention from queries to image tokens first, then self-attention among queries:
$$\mathcal{Q}=\mathrm{Cross\text{-}Attn}(\mathcal{Q},\mathcal{T}),\qquad \mathcal{Q}=\mathrm{Self\text{-}Attn}(\mathcal{Q})$$
Complexity drops to $\mathcal{O}(QNHW/p^{2}+Q^{2})$. The payoff is real memory savings, which in turn lets the model scale to larger sizes, higher resolutions and more queries.
Positional encoding is handled with equal care: image tokens use Plücker ray embeddings converted from predicted camera poses, while each query uses its own 3D coordinate as its positional embedding. Both sides speak the same 3D coordinate system, so geometric interaction is meaningful.
Gaussian spawning and supervision
With no 3D ground truth available, Gaussian Splatting is the only viable supervision route: render 3D primitives onto the image plane and constrain 3D structure with plentiful 2D signals. Concretely, each query $\mathbf{q}_i$ first predicts a query deformation $\delta\mathbf{q}_i$ that shifts it toward a geometry-consistent position, then an MLP decodes $K$ Gaussian offsets $\delta\mathbf{g}_{ik}$ added residually to the deformed query center:
$$\mathcal{G}=\bigcup_{i=1}^{Q}\bigcup_{k=1}^{K}\big(\mathbf{p}_i+\delta\mathbf{q}_i+\delta\mathbf{g}_{ik},\ \Sigma_{ik},\ \mathbf{c}_{ik},\ \alpha_{ik}\big)$$
where $\Sigma_{ik}$, $\mathbf{c}_{ik}$ and $\alpha_{ik}$ are covariance, color and opacity predicted by appearance decoders. Sparse queries deliver efficiency; the spawned dense Gaussians deliver detail.
Supervision has three terms: an RGB reconstruction loss on rendered color (combining $\ell_1$ and LPIPS), a scale-invariant depth loss on rendered depth, and the camera loss retained from the backbone:
$$\mathcal{L}=\mathcal{L}_{\mathrm{rgb}}+\lambda_d\mathcal{L}_{\mathrm{depth}}+\lambda_c\mathcal{L}_{\mathrm{cam}}$$
What actually makes occlusion completion happen is the choice of supervision views: the supervised views are a superset of the input views. Given 3 input images, the model renders and computes loss on 6 views — the 3 inputs plus 3 held-out. If the model reconstructs only what the inputs show, holes appear in the held-out views and are penalized directly by their ground-truth images. In other words, the model must answer for views it never saw before it will learn to fill in geometry it never saw.
Training is staged: the DINOv2 encoder and point-map head are initialized from Pi3 and kept frozen, while the camera head and query transformer are trained with AdamW (learning rate $1\times10^{-4}$, cosine decay) for 100 epochs at $224\times224$ on 32 A100 GPUs, then fine-tuned for 20 epochs at $448\times448$; each sample randomly uses 2 to 64 input views. At inference, AnySplat's test-time optimization can be applied, initializing from the predicted Gaussians and refining against the input views.
flowchart LR A[Unposed multi-view images] --> B[DINOv2 tokenization] B --> C[Alternating-attention transformer] C --> D[Decode pose point map confidence] C --> E[Image tokens with Plücker encoding] D --> F[Sample half of queries from point map] G[Uniform learnable anchors] --> H[Hybrid query init] F --> H H --> I[Query-to-image cross-attention] E --> I I --> J[Inter-query self-attention] J --> K[Predict query shift and K Gaussian offsets] K --> L[Spawned Gaussian set] L --> M[Differentiable RGB and depth splatting] M --> N[Supervise on input plus held-out views]
Experiments
Sparse views: where the advantage is clearest
Table 2 is the main arena — 3 or 6 input views, with metrics computed on novel views.
| Dataset | Method | Views | PSNR ↑ | SSIM ↑ | LPIPS ↓ | Time (s) ↓ |
|---|---|---|---|---|---|---|
| Mip-NeRF 360 | NoPoSplat | 3 | 18.21 | 0.482 | 0.426 | 0.416 |
| Mip-NeRF 360 | AnySplat | 3 | 20.08 | 0.606 | 0.274 | 0.279 |
| Mip-NeRF 360 | UniQueR | 3 | 22.70 | 0.660 | 0.261 | 0.213 |
| Mip-NeRF 360 | NoPoSplat | 6 | 16.06 | 0.423 | 0.540 | 2.121 |
| Mip-NeRF 360 | AnySplat | 6 | 18.29 | 0.518 | 0.336 | 0.646 |
| Mip-NeRF 360 | UniQueR | 6 | 21.80 | 0.622 | 0.300 | 0.447 |
| VR-NeRF | NoPoSplat | 3 | 21.28 | 0.744 | 0.381 | 0.406 |
| VR-NeRF | AnySplat | 3 | 19.67 | 0.745 | 0.313 | 0.290 |
| VR-NeRF | UniQueR | 3 | 21.99 | 0.708 | 0.446 | 0.198 |
| VR-NeRF | NoPoSplat | 6 | 20.69 | 0.728 | 0.426 | 2.176 |
| VR-NeRF | AnySplat | 6 | 17.25 | 0.683 | 0.411 | 0.656 |
| VR-NeRF | UniQueR | 6 | 20.87 | 0.689 | 0.431 | 0.412 |
On Mip-NeRF 360 at three views, PSNR rises from AnySplat's 20.08 to 22.70 while a single forward pass drops from 0.279 s to 0.213 s — quality and speed improving together, which is uncommon in a comparison table. Equally telling is how the baselines behave as views increase: NoPoSplat falls to 16.06 PSNR and 2.121 s at six views, showing how steeply per-pixel representations degrade with more inputs, while UniQueR holds at 21.80 and 0.447 s.
One caveat on VR-NeRF must be stated: UniQueR wins PSNR (21.99 versus 19.67) but clearly loses LPIPS to AnySplat (0.446 versus 0.313). A pixel-error metric and a perceptual metric disagreeing usually means more complete structure and firmer boundaries but weaker high-frequency texture. The paper does not discuss this.
Efficiency and geometric accuracy
Table 5 best captures the design trade-off: 32 input views, 448×448, a single A100 80GB, batch size 1.
| Method | Gaussians | GPU memory | Time (s) | Depth Abs Rel ↓ |
|---|---|---|---|---|
| AnySplat | 3.85M | 18.42 GB | 4.63 | 0.062 |
| UniQueR | 260K | 11.19 GB | 1.97 | 0.038 |
15× fewer primitives buy 40% less memory and a 2.4× speedup, while depth error improves from 0.062 to 0.038. The geometry result matters most: sparsity here is not trading accuracy for speed. Fewer primitives land in more meaningful places.
Figure 4: Qualitative and geometric comparison. Top rows render RGB on held-out novel views, bottom rows render depth. The pixel-aligned method leaves blank regions in RGB and holes in depth where no input pixel provides coverage; UniQueR fills them through 3D queries.
Dense views: the advantage vanishes, and returns elsewhere
Table 4 covers the dense setting at 32 and 64 views, and it is the easiest table to misread.
| Setting | Method | 32-view PSNR | 32-view LPIPS | 64-view PSNR | 64-view LPIPS |
|---|---|---|---|---|---|
| Feedforward only | AnySplat | 22.32 | 0.258 | 21.26 | 0.303 |
| Feedforward only | UniQueR | 21.51 | 0.325 | 21.58 | 0.335 |
| Feedforward + per-scene opt. | 3DGS + AnySplat | 24.99 | 0.227 | 23.71 | 0.266 |
| Feedforward + per-scene opt. | 3DGS + UniQueR | 25.14 | 0.183 | 26.00 | 0.176 |
In the feedforward-only dense setting, UniQueR loses to AnySplat: 21.51 versus 22.32 PSNR and 0.325 versus 0.258 LPIPS at 32 views. The paper does not hide this and offers an explanation — it has at most $4096\times64=262\text{K}$ Gaussians, whereas per-pixel methods generate $448\times448\times N$ primitives, so the gap widens as inputs grow. The sparse budget is fixed; more inputs do not grant more primitives.
Once you move to the workflow people actually use — feedforward initialization followed by per-scene optimization — the ranking flips. Starting from UniQueR's Gaussians, 3DGS optimization reaches 25.14 at 32 views and 26.00 at 64 views, with LPIPS down to 0.183 and 0.176, clearly ahead of the same pipeline initialized from AnySplat. In dense settings its value is not the rendered image but a much better starting point: more complete initial geometry and more accurate poses keep per-scene optimization out of bad local minima.
Pose estimation and ablations
On camera pose, UniQueR reports RRA@30 of 99.99, RTA@30 of 95.44 and AUC@30 of 83.69 on RealEstate10K, and 99.05, 97.44, 88.52 on Co3Dv2 — essentially level with Pi3 (99.99, 95.62, 85.90 and 99.05, 97.33, 88.41) and marginally ahead on Co3Dv2. That is expected and worth stating: the encoder is initialized from Pi3, so pose ability is largely inherited. The authors attribute the small RealEstate10K gap to Pi3 having used private training data.
Table 6 gives two decisive results. Removing rendered-depth supervision drops PSNR from 20.23 to 19.96, confirming depth cues help stabilize geometry. Replacing hybrid initialization with purely random initialization collapses PSNR to 12.11 and SSIM to 0.259 — far worse than dropping depth. This confirms the authors' starting premise: with no 3D ground truth, some queries must begin on meaningful geometry or training cannot stand up.
Figure 5: Ablation on the number of queries. Increasing query count consistently improves PSNR, showing a clear scaling trend.
Ablations on query count, Gaussians per query (16, 32, 64) and model capacity all trend upward consistently, indicating the representation is far from saturated and still has room to scale.
Limitations
Stated by the authors: the framework does not handle dynamic scenes. The query representation is designed for static scenes; incorporating temporal dynamics into the query formulation is named as future work.
Feedforward quality does not win in the dense regime. The paper admits this, though the headline can obscure it: at 32 and 64 views, feedforward-only PSNR and LPIPS both trail AnySplat. It recovers by serving as a better initialization for per-scene optimization — but if your setting has many inputs and runs a single feedforward pass, UniQueR is not currently the better choice. The fixed primitive budget is by design, so this shortfall will not disappear as inputs increase.
Perceptual metrics disagree on VR-NeRF. In the same experiment PSNR leads while LPIPS trails, and the paper offers no explanation. Applications that care about subjective appearance should verify this themselves.
No public code and no independent reproduction. No official implementation exists at the time of writing, and every figure here is from the authors' own evaluation; we ran no reproduction. Pose ability further depends heavily on the encoder frozen from Pi3, so how much the query mechanism alone contributes cannot be separated until code is released.
Conclusion
UniQueR's contribution is not another incremental bump in metrics but the removal of a default assumption. Feedforward reconstruction had assumed outputs must align to input pixels, a constraint that structurally limited models to surfaces that were seen. Replacing the representation with a sparse set of queries in global 3D space removes it: queries can be placed where nothing was observed, as long as rendering supervision charges a price for geometry missing there.
Three engineering judgments hold it up. Decoupled attention splits a quadratic into a sum of two terms, buying substantial memory headroom. Hybrid initialization — half grounded, half exploring — solves query divergence when no 3D ground truth exists. And making supervision views a superset of input views turns occlusion completion from an aspiration into an optimizable objective. Remove any one and the representation does not stand.
The cost is shown honestly: a sparse budget loses under dense input and recovers only by serving as a better initialization for per-scene optimization. For robotics and embodied AI, where observations are frequently sparse and partial, that trade is favorable; for dense offline reconstruction it is not yet optimal. Making the primitive budget adapt to the number of inputs, or extending queries to dynamic scenes, would widen the reach considerably.
Highlights
"Make the model answer for views it never saw, and it will learn to fill in geometry it never saw." — Setting supervision views as a superset of input views is the step that turns occlusion completion from a slogan into an optimizable objective.
"Half the queries stay grounded, half go exploring." — With no 3D ground truth, anchoring some queries on observed surfaces for stability while letting others roam for completeness is the precondition for training this representation at all.


