Physically Grounded Monocular Depth
via Nanophotonic Wavefront Encoding
1 New York University2 Columbia University
* Equal contribution † Corresponding authors
ECCV 2026

Overview
A birefringent metalens records two polarization images with depth-dependent point-spread functions (PSFs). We adapt a pretrained depth model to these image pairs and fine-tune it on simulated data and five real captures for single-shot metric depth estimation.
Video
Simulation Results
NYU Depth V2 · MPI Sintel · Hypersim · MIT-CGH-4K

FlyingThings3D · Depth-prior ablation

Real Scenes




3D View
Prediction preview · approximate projectionReal-World Comparisons
Single- and multi-object scenes captured with the prototype.

Depth Consistency

Method
Our optical simulator maps RGB-D scenes to polarization image pairs. Both simulated and captured pairs are stacked as [Ix, Iy, (Ix + Iy) / 2] and passed to the depth model.

Optical Simulation

Polarization Robustness

Transfer to UniDepth V2
The same three-channel input adaptation with a different depth model.
| Dataset | Original Large | Fine-tuned Small |
|---|---|---|
| NYU Depth V2 | 0.0338 | 0.0217 |
| MIT-CGH-4K | 0.1273 | 0.0708 |
| Real captures | 0.1068 | 0.0348 |
PSF Analysis
A point-spread function (PSF) is the image formed by a point of light. Our design makes it rotate with distance and records opposite lobes in separate polarization channels. These theoretical comparisons use a 1–5 m range and 50 mm focal length, not the near-range prototype.
Point Sources and Photon Budget

Edge Orientation

Extended Edges and Photon Budget

Ambiguity Across Depths

Resources
Code includes the optical simulator, training, and inference. Checkpoints are available in Small, Base, and Large sizes. The dataset contains five training scenes and 42 evaluation scenes. Inputs and depth labels are available under CC BY-NC 4.0 for noncommercial use, including academic research.
Evaluation protocol and limitations
The 42 scenes were also used for validation and model selection, not as an untouched holdout. Their depth labels are manually segmented regions assigned measured object distances, not dense 3D scans. The explorer uses the selected mixed-training Large checkpoint; paper figures retain their original results and are not a fresh benchmark.
The prototype targets well-lit, near-range scenes. Limited light throughput, field of view, and high-frequency polarization imbalance remain limitations. Code is MIT-licensed with retained third-party notices. Small weights follow Apache-2.0; Base and Large follow CC-BY-NC-4.0.
BibTeX
@article{li2026physically,
title={Physically Grounded Monocular Depth via
Nanophotonic Wavefront Encoding},
author={Li, Bingxuan and Wu, Jiahao and Xu, Yuan and
Zhu, Zezheng and Zhang, Yunxiang and Chen, Kenneth and
Liang, Yanqi and Yu, Nanfang and Sun, Qi},
journal={arXiv preprint arXiv:2503.15770},
year={2026}
}