Benchmarks
Headline: held-out and cross-domain
Section titled “Headline: held-out and cross-domain”These are the most trustworthy numbers: v4-modal on data never used for training or checkpoint selection.
| Split (tiles) | RMSE | MAE | r | δ1 | Balanced RMSE |
|---|---|---|---|---|---|
| GAMUS test · plain (2861) | 3.373 | 1.549 | 0.886 | 0.627 | 4.613 |
| GAMUS test · TTA (2861) | 3.321 | 1.511 | 0.890 | 0.629 | 4.556 |
| GAMUS test · sliding+TTA (400) | 3.873 | 2.257 | 0.911 | 0.564 | 4.118 |
| DFC23 val · OOD (246) | 4.968 | 1.789 | 0.881 | 0.776 | 5.544 |
| India val · near-domain (35) | 2.771 | 0.865 | 0.922 | 0.888 | 4.447 |
v4-modal on held-out and cross-domain data. Heights in metres. Source: Model_Traning/V4_modal/Output/metrics.json; Research-Paper/main.tex Table IV
Validation curves
Section titled “Validation curves”src/data/metrics.json ← run metrics.json histories; v1/v2 from Research-Paper/scripts/make_figures.pyFinal validation evaluation
Section titled “Final validation evaluation”src/data/metrics.json ← final_plain / final_tta / final_sliding_tta| Run | RMSE | MAE | r | δ1 | Bal. RMSE | RMSE (TTA) | RMSE (sliding) |
|---|---|---|---|---|---|---|---|
| v3 | 2.715 | 1.312 | 0.917 | 0.672 | 3.584 | 2.605 | 2.723 |
| v4-Kaggle | 3.427 | 1.885 | 0.921 | 0.613 | 3.869 | 3.389 | 3.441 |
| v4-2 | 3.803 | 2.065 | 0.902 | 0.595 | 4.167 | 3.790 | 3.804 |
| v4-modal | 3.754 | 2.001 | 0.905 | 0.601 | 4.060 | 3.631 | 3.646 |
| DAv2 | 3.522 | 1.904 | 0.915 | 0.619 | 3.948 | 3.432 | 3.394 |
GAMUS validation, 400 tiles. RMSE through Bal. RMSE are single-pass. Validation sets differ between v3 and the v4-era runs (see Design findings). Source: Model_Traning/{Logs/v3 (2), V4_Kaggle/outputs/*, V4_modal/Output, DAV2_V1/outputs/dav2-v1}/metrics.json
Same-protocol test comparison
Section titled “Same-protocol test comparison”To remove validation-set differences, v3, v4-modal and the DAv2 ablation were scored with one protocol on the same held-out tiles (eval_test.py --compare).
Model_Traning/v3/outputs/v3_vs_v4_vs_dav2.md| Split · protocol | v3 | v4-modal | DAv2 |
|---|---|---|---|
| GAMUS test · plain | 3.628 | 3.373 | 3.424 |
| GAMUS test · TTA | 3.565 | 3.321 | 3.396 |
| GAMUS test · sliding+TTA | 3.822 | 3.873 | 3.759 |
| India val · plain | 5.456 | 2.780 | 3.027 |
Source: Model_Traning/v3/outputs/v3_vs_v4_vs_dav2.md
v5 runs
Section titled “v5 runs”v5 was trained as a chain of warm-started runs (lineage). None has yet been scored on the held-out test splits. Preliminary · not test-evaluated
| Run | Primary val RMSE | Notes |
|---|---|---|
| with_dfc | 2.988 m (GAMUS) | DFC23 val 5.391 m |
| without_dfc | 2.909 m (GAMUS) | |
| without_dfc / resume | 2.863 m (GAMUS) | |
warm_start / Resume (v5_probe_v4init) |
2.804 m (GAMUS) | DFC23 4.878 m, India 2.767 m; source of the ONNX export |
| resume-v2 | 2.888 m (GAMUS) | adds US3D |
| resume-v3 | 2.865 m (GAMUS) | |
| resume-v4-1.6 | 1.615 m (MVS3DM) | GAMUS 2.864 m, US3D 3.964 m, gradient ratio 0.23 |
| v5_final_forest | 2.481 m selection score | MVS3DM 1.545 m, NEON forest 4.11 m |
Model_Traning/V4_modal/FINAL_RUN_RESULTS.md