Skip to content

Architecture overview

DepthWizard is a dense regression network. It maps a 512 × 512 RGB tile at 0.5 m GSD to a height-above-ground map in metres. It has four parts:

  • a satellite-pretrained Vision Transformer encoder,
  • a DPT decoder that rebuilds a multi-scale feature pyramid,
  • three task heads (direct regression, adaptive bins and segmentation), combined by a learned per-pixel gate,
  • a light detail branch, added in v5 to recover sharp edges.
ReassembleDPT decoderHeadsInput tile512 × 512 RGBat 0.5 m GSDDINOv3 ViT-L/16SAT-493M · 24 blocks · d=1024Patch embed 16 × 16Blocks 1–632 × 32 tokenstap @ 696 ch · ConvT 4× → 1/43×3 conv → 256 chFusion 1RefineNet · GN(8)Blocks 7–1232 × 32 tokenstap @ 12192 ch · ConvT 2× → 1/83×3 conv → 256 chFusion 2RefineNet · GN(8)Blocks 13–1832 × 32 tokenstap @ 18384 ch · identity → 1/163×3 conv → 256 chFusion 3RefineNet · GN(8)Blocks 19–2432 × 32 tokenstap @ 24768 ch · stride-2 → 1/323×3 conv → 256 chFusion 4RefineNet · GN(8)frozen → unfrozen with layer-wise LR decayF · 256 ch@ 1/2 resHead A · regressionsoftplus → aHead B · adaptive binsK=96, 0–120 m → bLearned gate ααa + (1−α)bHeight (m)+ σ from bin spreadHead C · segmentation7 classes + ignorev5 detail branchRGB stem → 64 ch @ 1/2 · inject into F · ConvexUp2x
The v5 network. Tokens from four depths of the encoder are reassembled into a 1/4–1/32 pyramid and fused top-down by RefineNet blocks into a 256-channel feature map F at half resolution. Heads A and B each predict height, and the gate α blends them per pixel. Head C labels land cover; its ground mask is later used to fit terrain. The dashed orange path is the v5 detail branch.
Encoder params
303.1M
all 24 blocks trainable
Decoder + heads
20.6M
logged with encoder frozen
Input
512²px
RGB at 0.5 m GSD
Outputs
3
height · classes · σ
Stage Output Shape (batch 1)
Input tile normalised RGB 3 × 512 × 512
Patch embed (16 × 16) tokens 1024 × 32 × 32
Taps after blocks 6 / 12 / 18 / 24 4 token maps (class and register tokens dropped) 4 × (1024 × 32 × 32)
Reassemble 96 / 192 / 384 / 768 ch at 1/4, 1/8, 1/16, 1/32 96×128², 192×64², 384×32², 768×16²
3 × 3 projection 256 ch at each scale 256 × (128², 64², 32², 16²)
RefineNet fusion ×4 feature map F 256 × 256 × 256 (1/2 res.)
Head A height a (softplus) 1 × 512 × 512
Head B bin probabilities → height b, σ 96 × … → 1 × 512 × 512
Gate α ∈ (0, 1) 1 × 512 × 512
Head C class logits 8 × 512 × 512 (7 classes + ignore)

The final height is the gated blend of the two height estimates:

Here g is a small conv block (258 → 64 → 1 channels). Head B’s bin distribution also gives a per-pixel standard deviation, which is written out as the uncertainty map ndsm_std_m.

  • A satellite-native encoder. DINOv3 pretrained on 493 M satellite images (SAT-493M) already models nadir texture, shadows and roof shapes. Replacing it with Depth Anything V2 (natural images) was tested in an ablation and scored worse on validation: 3.43 m vs 2.61 m, both with TTA (Benchmarks).
  • Two height heads. Direct regression is smooth and accurate on low objects. Adaptive bins handle the long tail of tall structures better. The gate learns where to trust each one.
  • Joint segmentation. Head C costs little and supplies the ground mask that the georeferencing step uses to fit a bare-earth DTM (Absolute DSM).
  • Detail branch (v5). Transformer features at 1/16 resolution give over-smoothed maps: a gradient ratio of 0.23 against LiDAR, where 1.0 would be perfectly sharp. A shallow RGB stem at full and half resolution adds edges back, and a learned convex 2× upsampler replaces bilinear upsampling. Both are zero-initialised, so v4 checkpoints load unchanged.

The trained v5 weights are published on Kaggle: kaggle.com/models/abhaydkale232/depthwizard-v5. Load them with the CLI or FastAPI service. PyTorch checkpoints embed their preprocessing recipe (ONNX exports keep it in the sidecar depthwizard.onnx.json), so no extra configuration is needed.

  • Encoder: DINOv3 SAT-493M, freezing and layer-wise LR decay.
  • Decoder & heads: reassemble, fusion, adaptive bins, gate, segmentation and the detail branch.
  • Loss functions: every term, its weight and the stratum balancer.
  • Inference engine: tiling, blending, TTA and large scenes.