PAGER: Partial-to-Global Alignment via Geometric and Relational Distillation

1Technical University of Denmark 2Pioneer Center for Artificial Intelligence
Partial point cloud input; a frozen global model mislabels the bed as wall, while PAGER correctly predicts bed, matching the global ground truth.

Given a partial point cloud in a local coordinate system with limited field of view, a frozen global model mislabels the bed as wall. PAGER aligns partial observations with the global model and correctly predicts bed.

Abstract

Pretrained 3D encoders are typically developed on globally reconstructed scenes expressed in a consistent world coordinate frame, whereas embodied systems must reason from partial, viewpoint-dependent observations in camera coordinates. We show that this shift from globally learned 3D feature spaces to realistic partial ob- servations exposes a severe representation mismatch, which we find consistently across representative state-of-the-art encoders, including Sonata and Concerto. A frozen Sonata encoder with a global linear probe achieves 72.47 mIoU on full ScanNet scenes, but only 2.57 mIoU on single-frame camera-coordinate inputs. Training-free gravity alignment recovers performance to 41.64 mIoU, showing that coordinate-frame mismatch is a dominant source of degradation but cannot be fully resolved through canonicalization alone. We introduce PAGER, a label- free adaptation method that aligns partial-view features with a frozen global 3D semantic space using only paired partial/global geometry. It learns lightweight adaptation modules while keeping the pretrained encoder and global segmenta- tion probe frozen. Matched-point feature alignment anchors partial features to their global counterparts, while relational supervision preserves their similarity structure with respect to the global representation. Global geometry provides su- pervision only during training. Inference operates directly on the partial obser- vation. Without partial-view labels, PAGER outperforms label-supervised PEFT on both Sonata and Concerto, and in zero-shot ScanNet→ScanNet++ transfer sur- passes fully fine-tuned Sonata (53.93 vs. 48.09 mIoU), suggesting that preserving the frozen global representation can improve cross-dataset transfer

Representation Gap Analysis

We first investigate how representations learned by globally pretrained 3D encoders transfer to increasingly view-constrained observations. Using a frozen Sonata encoder with a global linear probe, we evaluate four increasingly realistic input settings: Global, Room Slice, View Crop, and Single Frame — in both world and camera coordinates.

Point clouds and PCA of features for global, room slice, view crop and single frame inputs, in world and camera coordinate systems
Two sources of degradation. Left: progressive geometric reduction (Global → Room Slice → View Crop → Single Frame) in world coordinates. Right: the same inputs in camera coordinates cause a much larger drop, visible in both the PCA feature visualizations (bottom row) and segmentation accuracy (table below). Coordinate-system origin shown by black/red arrows.
Coord. frame Input type allAcc mAcc mIoU
World Global 89.7383.2772.47
Room Slice 88.4881.0169.64
View Crop 87.2480.7969.02
Single Frame 87.3678.9867.45
Camera Global 26.429.565.14
View Crop 19.927.523.00
Single Frame 18.306.942.57
+ PCA align 28.3914.708.38
+ Gravity align (Jin et al., 2023) 70.8554.7441.64
  • Missing geometry costs little. In world coordinates, mIoU drops by under 5 points, from 72.47 on Global to 67.45 on Single Frame (the latter uses ground-truth poses).
  • The coordinate frame is the dominant factor. Expressing the complete Global scene in camera coordinates, a single rigid transform with no geometry removed, collapses mIoU to 5.14; adding single-frame partiality brings it to 2.57. The encoder relies on world-frame priors such as gravity-aligned axes and floors below.
  • Training-free canonicalization is not enough. PCA axis alignment recovers only 8.38 mIoU and gravity alignment (Jin et al., 2023) 41.64.

PAGER therefore treats the remaining failure as a representation mismatch and directly adapts partial-view features toward the global feature space.

Methodology

PAGER adapts a frozen, globally pretrained 3D encoder to partial, camera-frame observations. Only a lightweight adapter on the partial branch is trained; the encoder and linear probe stay frozen, and at inference only the partial branch is used.

PAGER architecture: a frozen Sonata encoder on the global view and an adapted Sonata encoder on the partial view, aligned with KL and cosine losses on matched points
PAGER architecture. Only the adaptation layers on the partial branch are trained. For matched points, an anchor ℒcos and a relational loss ℒKL are used. Encoder and linear probe remain frozen. At inference, only the partial branch is used.
  • Lightweight geometry-aware adapter. A small adapter inside the frozen backbone refines each point's features using its local 3D neighbourhood. It deliberately has no global attention: scene-level context comes from the training objective, not the architecture.
  • Geometric anchor (ℒcos). For points seen in both the partial and the global view, each partial feature is pulled toward the global feature of the same physical point, which is already compatible with the frozen linear probe.
  • Relational shaper (ℒKL). Each point's similarity to a random set of scene-wide reference points acts as a compact signature of how it relates to the whole scene. The partial branch learns to reproduce that signature from its limited view.

The two losses are complementary: the anchor fixes where each feature should be, while the relational term preserves how features relate to one another. Together they transfer global scene structure to partial views without any semantic labels.

Results

ScanNet ScanNet++ ScanNet200
Method Family Superv. Params mIoU ↑mAcc ↑allAcc ↑mIoU ↑mAcc ↑allAcc ↑mIoU ↑mAcc ↑allAcc ↑
Sonata (Wu et al., 2025)
Frozenfrozennone– 2.576.9418.300.691.929.780.140.373.11
Linearfine-tunesem.<0.1M 36.1752.4268.6415.7326.1167.668.1312.9861.43
Late blocks + linearfine-tunesem.21.0M 48.2859.1376.8617.7725.3372.9811.4316.5663.55
Full + linear†fine-tunesem.108.5M 69.2677.1187.8835.9944.7585.9524.5832.1377.86
Adapter (Houlsby et al., 2019)adaptsem.6.0M 41.4965.6874.2213.9628.9765.5810.1516.1764.49
GEM (Tang et al., 2025)adaptsem.1.8M 59.8173.9283.9033.2742.5784.2221.9732.5275.45
PAGER (ours)adaptgeom.1.3M 65.4680.5085.9327.3839.1881.4319.2527.7374.19
Concerto (Zhang et al., 2025)
Frozenfrozennone– 4.9710.8321.551.444.2617.460.371.645.20
Linearfine-tunesem.<0.1M 39.5356.5771.1018.1326.3771.6912.6819.9465.89
GEM (Tang et al., 2025)adaptsem.1.8M 69.9080.9688.6432.8645.2884.1225.0635.1077.83
PAGER (ours)adaptgeom.1.3M 71.6787.7789.9633.9746.2884.9226.7350.1976.12
† Fully supervised full-backbone upper bound, excluded from bold/underline ranking.
none: frozen encoder with global linear probe. sem.: partial-view semantic labels. geom.: paired partial/global geometry only, no semantic labels.
Semantic segmentation of partial point clouds: input, Sonata with linear head, GEM, PAGER, and ground truth
wall floor bed chair table door window shelf picture desk toilet
Semantic segmentation of partial inputs. Columns: input point cloud, Sonata (+linear), GEM, PAGER (ours), and ground truth.
In-the-wild results on reconstructed and stereo inputs: input, Sonata with linear head, GEM, and PAGER
Results on in-the-wild captures: inputs reconstructed from ScanNet (top) and stereo inputs on unseen scenes (bottom), compared with Sonata (+linear) and GEM.

BibTeX


      @misc{adeniranlowe2026pagerpartialtoglobalalignmentgeometric,
      title={PAGER: Partial-to-global Alignment via Geometric and Relational Distillation}, 
      author={Akira-Miranda Adeyomi Adeniran-Lowe and Binod Singh and Lars Arnold Dethlefsen and Lazaros Nalpantidis and Theodora Kontogianni},
      year={2026},
      eprint={2610.01589},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.01589}, 
}