Skip to content

Decoder & heads

Source: Model_Traning/v5/models/dpt.py and heads.py.

ReassembleDPT decoderHeadsInput tile512 × 512 RGBat 0.5 m GSDDINOv3 ViT-L/16SAT-493M · 24 blocks · d=1024Patch embed 16 × 16Blocks 1–632 × 32 tokenstap @ 696 ch · ConvT 4× → 1/43×3 conv → 256 chFusion 1RefineNet · GN(8)Blocks 7–1232 × 32 tokenstap @ 12192 ch · ConvT 2× → 1/83×3 conv → 256 chFusion 2RefineNet · GN(8)Blocks 13–1832 × 32 tokenstap @ 18384 ch · identity → 1/163×3 conv → 256 chFusion 3RefineNet · GN(8)Blocks 19–2432 × 32 tokenstap @ 24768 ch · stride-2 → 1/323×3 conv → 256 chFusion 4RefineNet · GN(8)frozen → unfrozen with layer-wise LR decayF · 256 ch@ 1/2 resHead A · regressionsoftplus → aHead B · adaptive binsK=96, 0–120 m → bLearned gate ααa + (1−α)bHeight (m)+ σ from bin spreadHead C · segmentation7 classes + ignorev5 detail branchRGB stem → 64 ch @ 1/2 · inject into F · ConvexUp2x
Decoder and heads. See the Architecture overview for the full walkthrough.

The Dense Prediction Transformer (DPT) decoder turns four same-resolution token maps into a feature pyramid, then fuses it top-down.

Reassemble. Each of the four taps goes through a 1 × 1 conv, is resampled, then projected to 256 channels (decoder_dim) by a 3 × 3 conv:

Tap Channels Resample Scale
Block 6 96 ConvTranspose 4× 1/4 (128²)
Block 12 192 ConvTranspose 2× 1/8 (64²)
Block 18 384 identity 1/16 (32²)
Block 24 768 stride-2 conv 1/32 (16²)

Fusion. Four RefineNet-style blocks, each built from residual conv units with GroupNorm(8), merge the pyramid from coarse to fine. Each block upsamples its input and adds the next finer scale; the deepest block has no skip input. The result, F, has 256 channels at half the input resolution.

Every head starts with the same block: 3×3 conv 256→128 → GroupNorm → ReLU → 1×1 conv → output.

A single-channel regression with a softplus output, so heights are non-negative with a smooth gradient. v2 used ReLU, which can kill gradients. The last bias starts at −2, so initial predictions sit near 0 m, where most pixels are.

Head B follows the AdaBins idea: it predicts a per-image set of bin widths and a per-pixel distribution over those bins.

  1. A pooled, RMS-normalised copy of F feeds a Linear(256 → 96), which gives width logits z (computed in fp32 and clamped to ±15).
  2. The widths are a floored softmax over the height range [hmin, hmax] = [0, 120 m]:
  1. Bin centres ck are the midpoints of the cumulative widths. Each pixel gets probabilities pk from a 1 × 1 conv and a softmax.
  2. The height is the expected value of that distribution, and its spread is the uncertainty:

The width floor (0.05/K) stops any bin from collapsing to zero width, which caused unstable training in early runs.

The gate sees the features plus both candidate heights (256 + 1 + 1 = 258 channels), so it can learn, for example, to trust the bins on tall buildings and the regression on flat ground.

An 8-channel classifier: 7 land-cover classes plus an ignore channel (id 7) for unlabelled pixels.

id Class Counts as flat ground in the losses?
0 other —
1 ground ✓
2 low vegetation ✓
3 building —
4 water ✓
5 road ✓
6 tree —

Its predictions give the ground mask for DTM fitting, colour the Object classes layer in the viewer, and drive the building/tree/water object extraction.

Enabled with detail_branch = true, detail_dim = 64.

  • DetailStem: a small conv stem on the raw RGB, producing 32 channels at full resolution and 64 at half resolution (stride 2).
  • Inject: a 1 × 1 conv (320 → 256) adds these edge features to F before the heads.
  • ConvexUp2x: a learned 2× upsampler, as in RAFT. For each output sub-pixel it predicts softmax weights over the 3 × 3 low-resolution neighbourhood (a 448 → 128 → 36 conv stack) and takes the convex combination.

All new layers are zero-initialised, so at step 0 the network behaves exactly like v4 (ConvexUp2x starts as bilinear upsampling) and v4 checkpoints load without surgery.