Abstract
Pretrained 3D encoders are typically developed on globally reconstructed scenes expressed in a consistent world coordinate frame, whereas embodied systems must reason from partial, viewpoint-dependent observations in camera coordinates. We show that this shift from globally learned 3D feature spaces to realistic partial ob- servations exposes a severe representation mismatch, which we find consistently across representative state-of-the-art encoders, including Sonata and Concerto. A frozen Sonata encoder with a global linear probe achieves 72.47 mIoU on full ScanNet scenes, but only 2.57 mIoU on single-frame camera-coordinate inputs. Training-free gravity alignment recovers performance to 41.64 mIoU, showing that coordinate-frame mismatch is a dominant source of degradation but cannot be fully resolved through canonicalization alone. We introduce PAGER, a label- free adaptation method that aligns partial-view features with a frozen global 3D semantic space using only paired partial/global geometry. It learns lightweight adaptation modules while keeping the pretrained encoder and global segmenta- tion probe frozen. Matched-point feature alignment anchors partial features to their global counterparts, while relational supervision preserves their similarity structure with respect to the global representation. Global geometry provides su- pervision only during training. Inference operates directly on the partial obser- vation. Without partial-view labels, PAGER outperforms label-supervised PEFT on both Sonata and Concerto, and in zero-shot ScanNet→ScanNet++ transfer sur- passes fully fine-tuned Sonata (53.93 vs. 48.09 mIoU), suggesting that preserving the frozen global representation can improve cross-dataset transfer
Representation Gap Analysis
We first investigate how representations learned by globally pretrained 3D encoders transfer to increasingly view-constrained observations. Using a frozen Sonata encoder with a global linear probe, we evaluate four increasingly realistic input settings: Global, Room Slice, View Crop, and Single Frame — in both world and camera coordinates.
| Coord. frame | Input type | allAcc | mAcc | mIoU |
|---|---|---|---|---|
| World | Global | 89.73 | 83.27 | 72.47 |
| Room Slice | 88.48 | 81.01 | 69.64 | |
| View Crop | 87.24 | 80.79 | 69.02 | |
| Single Frame | 87.36 | 78.98 | 67.45 | |
| Camera | Global | 26.42 | 9.56 | 5.14 |
| View Crop | 19.92 | 7.52 | 3.00 | |
| Single Frame | 18.30 | 6.94 | 2.57 | |
| + PCA align | 28.39 | 14.70 | 8.38 | |
| + Gravity align (Jin et al., 2023) | 70.85 | 54.74 | 41.64 |
- Missing geometry costs little. In world coordinates, mIoU drops by under 5 points, from 72.47 on Global to 67.45 on Single Frame (the latter uses ground-truth poses).
- The coordinate frame is the dominant factor. Expressing the complete Global scene in camera coordinates, a single rigid transform with no geometry removed, collapses mIoU to 5.14; adding single-frame partiality brings it to 2.57. The encoder relies on world-frame priors such as gravity-aligned axes and floors below.
- Training-free canonicalization is not enough. PCA axis alignment recovers only 8.38 mIoU and gravity alignment (Jin et al., 2023) 41.64.
PAGER therefore treats the remaining failure as a representation mismatch and directly adapts partial-view features toward the global feature space.
Methodology
PAGER adapts a frozen, globally pretrained 3D encoder to partial, camera-frame observations. Only a lightweight adapter on the partial branch is trained; the encoder and linear probe stay frozen, and at inference only the partial branch is used.
- Lightweight geometry-aware adapter. A small adapter inside the frozen backbone refines each point's features using its local 3D neighbourhood. It deliberately has no global attention: scene-level context comes from the training objective, not the architecture.
- Geometric anchor (ℒcos). For points seen in both the partial and the global view, each partial feature is pulled toward the global feature of the same physical point, which is already compatible with the frozen linear probe.
- Relational shaper (ℒKL). Each point's similarity to a random set of scene-wide reference points acts as a compact signature of how it relates to the whole scene. The partial branch learns to reproduce that signature from its limited view.
The two losses are complementary: the anchor fixes where each feature should be, while the relational term preserves how features relate to one another. Together they transfer global scene structure to partial views without any semantic labels.
Results
BibTeX
@misc{adeniranlowe2026pagerpartialtoglobalalignmentgeometric,
title={PAGER: Partial-to-global Alignment via Geometric and Relational Distillation},
author={Akira-Miranda Adeyomi Adeniran-Lowe and Binod Singh and Lars Arnold Dethlefsen and Lazaros Nalpantidis and Theodora Kontogianni},
year={2026},
eprint={2610.01589},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.01589},
}