PointWeave: Feed-Forward View Synthesis with Explicit Point Interface
Abstract
We introduce PointWeave, a feed-forward novel-view renderer that uses explicit 3D points to guide learned appearance reconstruction. Existing 3DGS-based methods render explicit primitives efficiently, but recovering fine detail can require many primitives. Geometric-free methods offer flexible and scalable appearance modeling, but can struggle to maintain geometric consistency under unseen camera transformations. In contrast, PointWeave combines explicit geometry with learned image synthesis. It predicts 3D points with appearance features from input views, and uses learned ray-to-point attention to aggregate their features for each target ray. The resulting feature map is combined with colors retrieved from the input views using the inferred geometry and passed through a decoder to reconstruct the target image. On RealEstate10K, PointWeave achieves 31.38 dB PSNR with 8,192 points, outperforming recent 3DGS-based methods by 1.6 dB while using fewer primitives. It matches the performance of state-of-the-art geometric-free methods with up to 4× faster rendering and substantially better geometric consistency under camera roll, field-of-view, pixel-aspect and world-scale changes. It also outperforms the baselines in zero-shot transfer to unseen datasets, with 0.54–2.02 dB PSNR gains on ACID, DL3DV, ScanNet++, and DTU.
Method Overview
A reconstruction backbone predicts featured 3D points once per scene. Given a target camera, each ray gathers features from these points via learned ray-to-point attention. The aggregated features are combined with context appearance retrieved through the predicted geometry to decode the target image. Times are measured on an RTX 3090.
Qualitative Results
Comparisons with 3DGS-based and geometric-free methods on in-domain, zero-shot and transformed-camera views.
RE10K, two context views. Target, a 4× crop (yellow box) and the crop's error map, where brighter means larger error. 3DGS-based methods blur when the point density is low. Geometric-free methods may hallucinate geometry, such as the extra window frame (bottom).
DL3DV, full-frame protocol, two to six context views. PointWeave retains more detail than the baselines with fewer primitives and better geometric consistency.
PointWeave and LagerNVS, square-input DL3DV, with 2, 4 and 6 input views.
Zero-shot transfer from RE10K to ACID, DL3DV, ScanNet++ and DTU. PointWeave shows better geometric consistency.
Renders under camera transformations. Geometric-free methods fail to render geometrically consistent images under transformations.
Quantitative Results
best, second best.
RE10K novel view synthesis
| Method | Primitives | Encode / render (ms) | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
|---|---|---|---|---|---|
| 3DGS-based | |||||
| pixelSplat | 393k | 139.5 / 3.2 | 25.89 | .858 | .142 |
| MVSplat | 131k | 50.5 / 2.0 | 26.39 | .869 | .128 |
| DepthSplat | 131k | 118.2 / 3.2 | 27.47 | .889 | .114 |
| TokenGS (1,024 tokens) | 66k | 116.8 / 0.6 | 28.02 | .896 | .147 |
| TokenGS (4,096 tokens) | 262k | 249.6 / 1.0 | 28.41 | .903 | .135 |
| SplatWeaver | 47k | 273.4 / 1.3 | 29.06 | .899 | .102 |
| ReSplat (no refinement) | 131k | 105.8 / 1.4 | 29.40 | .909 | .104 |
| ReSplat (2 refinements) | 131k | 393.1 / 1.4 | 29.75 | .912 | .100 |
| Geometric-free | |||||
| LVSM (decoder-only) | — | — / 41.9 | 29.68 | .906 | .098 |
| CLiFT† | — | 875.7 / 10.2 | 27.06 | .867 | .132 |
| LagerNVS | — | 191.5 / 30.9 | 31.39 | .926 | .078 |
| PointWeave, K=20 | 8,192 | 65.6 / 10.8 | 31.38 | .927 | .093 |
| PointWeave, K=5 | 8,192 | 67.7 / 7.0 | 31.32 | .927 | .093 |
Two context views, 256×256 images. Times are encoding / rendering in ms on an RTX 3090. †: trained with four context views.
Zero-shot transfer from RE10K
| Method | ACID (1,595) | DL3DV short (139) | ScanNet++ (50) | DTU (64) | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| PSNR ↑ | SSIM ↑ | LPIPS ↓ | PSNR ↑ | SSIM ↑ | LPIPS ↓ | PSNR ↑ | SSIM ↑ | LPIPS ↓ | PSNR ↑ | SSIM ↑ | LPIPS ↓ | |
| 3DGS-based | ||||||||||||
| pixelSplat | 27.82 | .835 | .155 | 27.20 | .882 | .103 | 18.42 | .720 | .278 | 11.53 | .327 | .633 |
| MVSplat | 28.16 | .841 | .147 | 26.95 | .873 | .101 | 17.14 | .687 | .297 | 13.94 | .474 | .386 |
| DepthSplat | 28.38 | .848 | .142 | 28.14 | .905 | .083 | 20.79 | .761 | .254 | 14.59 | .425 | .437 |
| ReSplat (no refinement) | 29.69 | .865 | .137 | 30.53 | .927 | .075 | 22.82 | .787 | .237 | 15.33 | .620 | .389 |
| Geometric-free | ||||||||||||
| LVSM (decoder-only) | 30.44 | .869 | .127 | 30.03 | .914 | .080 | 25.49 | .785 | .212 | 16.01 | .544 | .343 |
| CLiFT† | 29.00 | .840 | .164 | 26.87 | .850 | .134 | 23.09 | .760 | .269 | 14.25 | .492 | .533 |
| LagerNVS | 31.70 | .886 | .110 | 29.72 | .907 | .066 | 25.30 | .785 | .194 | 14.75 | .498 | .375 |
| PointWeave (ours) | 32.25 | .896 | .124 | 32.28 | .939 | .069 | 27.51 | .829 | .199 | 17.79 | .684 | .315 |
256×256, two context views; number of test scenes in parentheses. All models are trained on RE10K only. †: trained with four context views.
Geometric consistency under camera and world transformations
| Method | Roll | FOV | Anisotropy | Scale | ||||
|---|---|---|---|---|---|---|---|---|
| PSNR ↑ | SSIM ↑ | PSNR ↑ | SSIM ↑ | PSNR ↑ | SSIM ↑ | PSNR ↑ | SSIM ↑ | |
| LagerNVS | 17.72 | .480 | 14.05 | .801 | 16.02 | .758 | 15.22 | .512 |
| LVSM (decoder-only) | 18.03 | .488 | 16.44 | .847 | 18.93 | .812 | 20.63 | .645 |
| CLiFT‡ | 17.95 | .494 | 17.07 | .872 | 17.82 | .784 | 18.34 | .548 |
| PVSM† | 19.96 | .709 | 20.17 | .928 | 19.23 | .842 | 21.70 | .729 |
| PointWeave (ours) | 21.85 | .790 | 21.67 | .953 | 24.41 | .932 | 25.20 | .805 |
Consistency benchmark of PVSM on 100 RE10K scenes, two context views; scores are averaged over the settings of each transformation and computed over valid pixels. †: released model fine-tuned with additional camera augmentations. ‡: trained with four context views.
BibTeX
@misc{zhang2026pointweave,
title = {PointWeave: Feed-Forward View Synthesis with Explicit Point Interface},
author = {Zhang, Yanshu and Peng, Shichong and Vashist, Chirag and Li, Ke},
year = {2026}
}