PointWeave: Feed-Forward View Synthesis with Explicit Point Interface

1Simon Fraser University   2Alberta Machine Intelligence Institute (Amii)   3Canadian Institute for Advanced Research (CIFAR)

PointWeave uses explicit 3D points to guide learned appearance reconstruction,
matching state-of-the-art geometric-free methods with up to 4× faster rendering and substantially better geometric consistency.

How PointWeave renders a view. Every point, selection, attention weight, feature and image in this animation comes from one forward pass of the PointWeave model on a RealEstate10K test scene. The camera roll and zoom comparison at the end uses PointWeave fine-tuned for 500 steps with camera augmentation.

Abstract

We introduce PointWeave, a feed-forward novel-view renderer that uses explicit 3D points to guide learned appearance reconstruction. Existing 3DGS-based methods render explicit primitives efficiently, but recovering fine detail can require many primitives. Geometric-free methods offer flexible and scalable appearance modeling, but can struggle to maintain geometric consistency under unseen camera transformations. In contrast, PointWeave combines explicit geometry with learned image synthesis. It predicts 3D points with appearance features from input views, and uses learned ray-to-point attention to aggregate their features for each target ray. The resulting feature map is combined with colors retrieved from the input views using the inferred geometry and passed through a decoder to reconstruct the target image. On RealEstate10K, PointWeave achieves 31.38 dB PSNR with 8,192 points, outperforming recent 3DGS-based methods by 1.6 dB while using fewer primitives. It matches the performance of state-of-the-art geometric-free methods with up to 4× faster rendering and substantially better geometric consistency under camera roll, field-of-view, pixel-aspect and world-scale changes. It also outperforms the baselines in zero-shot transfer to unseen datasets, with 0.54–2.02 dB PSNR gains on ACID, DL3DV, ScanNet++, and DTU.

Method Overview

PointWeave framework overview

A reconstruction backbone predicts featured 3D points once per scene. Given a target camera, each ray gathers features from these points via learned ray-to-point attention. The aggregated features are combined with context appearance retrieved through the predicted geometry to decode the target image. Times are measured on an RTX 3090.

Qualitative Results

Comparisons with 3DGS-based and geometric-free methods on in-domain, zero-shot and transformed-camera views.

Quantitative Results

best, second best.

RE10K novel view synthesis

MethodPrimitivesEncode / render (ms)PSNR ↑SSIM ↑LPIPS ↓
3DGS-based
pixelSplat393k139.5 / 3.225.89.858.142
MVSplat131k50.5 / 2.026.39.869.128
DepthSplat131k118.2 / 3.227.47.889.114
TokenGS (1,024 tokens)66k116.8 / 0.628.02.896.147
TokenGS (4,096 tokens)262k249.6 / 1.028.41.903.135
SplatWeaver47k273.4 / 1.329.06.899.102
ReSplat (no refinement)131k105.8 / 1.429.40.909.104
ReSplat (2 refinements)131k393.1 / 1.429.75.912.100
Geometric-free
LVSM (decoder-only)—— / 41.929.68.906.098
CLiFT†—875.7 / 10.227.06.867.132
LagerNVS—191.5 / 30.931.39.926.078
PointWeave, K=208,19265.6 / 10.831.38.927.093
PointWeave, K=58,19267.7 / 7.031.32.927.093

Two context views, 256×256 images. Times are encoding / rendering in ms on an RTX 3090. †: trained with four context views.

Zero-shot transfer from RE10K

MethodACID (1,595)DL3DV short (139)ScanNet++ (50)DTU (64)
PSNR ↑SSIM ↑LPIPS ↓PSNR ↑SSIM ↑LPIPS ↓PSNR ↑SSIM ↑LPIPS ↓PSNR ↑SSIM ↑LPIPS ↓
3DGS-based
pixelSplat27.82.835.15527.20.882.10318.42.720.27811.53.327.633
MVSplat28.16.841.14726.95.873.10117.14.687.29713.94.474.386
DepthSplat28.38.848.14228.14.905.08320.79.761.25414.59.425.437
ReSplat (no refinement)29.69.865.13730.53.927.07522.82.787.23715.33.620.389
Geometric-free
LVSM (decoder-only)30.44.869.12730.03.914.08025.49.785.21216.01.544.343
CLiFT†29.00.840.16426.87.850.13423.09.760.26914.25.492.533
LagerNVS31.70.886.11029.72.907.06625.30.785.19414.75.498.375
PointWeave (ours)32.25.896.12432.28.939.06927.51.829.19917.79.684.315

256×256, two context views; number of test scenes in parentheses. All models are trained on RE10K only. †: trained with four context views.

Geometric consistency under camera and world transformations

MethodRollFOVAnisotropyScale
PSNR ↑SSIM ↑PSNR ↑SSIM ↑PSNR ↑SSIM ↑PSNR ↑SSIM ↑
LagerNVS17.72.48014.05.80116.02.75815.22.512
LVSM (decoder-only)18.03.48816.44.84718.93.81220.63.645
CLiFT‡17.95.49417.07.87217.82.78418.34.548
PVSM†19.96.70920.17.92819.23.84221.70.729
PointWeave (ours)21.85.79021.67.95324.41.93225.20.805

Consistency benchmark of PVSM on 100 RE10K scenes, two context views; scores are averaged over the settings of each transformation and computed over valid pixels. †: released model fine-tuned with additional camera augmentations. ‡: trained with four context views.

BibTeX

@misc{zhang2026pointweave,
  title  = {PointWeave: Feed-Forward View Synthesis with Explicit Point Interface},
  author = {Zhang, Yanshu and Peng, Shichong and Vashist, Chirag and Li, Ke},
  year   = {2026}
}