Z3D: Zero-Shot 3D Visual Grounding from Images
| Benchmark | Task | Input | Metric | Result | Evidence |
|---|---|---|---|---|---|
| ScanRefer | Zero-shot 3D visual grounding | Ground-truth point cloud; no bounding-box or text supervision | Overall [email protected] / [email protected] | 54.2 / 46.0 | Z3D paper, Table 1 Reported as state of the art among zero-shot approaches in the associated paper. |
| ScanRefer | Zero-shot 3D visual grounding | Posed multi-view RGB | Overall [email protected] / [email protected] | 42.8 / 24.8 | Z3D paper, Table 1 |
| ScanRefer | Zero-shot 3D visual grounding | Unposed multi-view RGB | Overall [email protected] / [email protected] | 31.2 / 12.9 | Z3D paper, Table 1 |
| Nr3D | Zero-shot 3D visual grounding | Depth-aware setting | Overall top-1 accuracy | 54.8 | Z3D paper, Table 2 The strongest zero-shot result in the paper's main Nr3D table was 54.8 versus 54.3 for SPAZER. |