# Santiago AI building-outline pipeline

This document describes the building-detection process actually implemented in
this repository. It is intended both as a reproducibility record for the
Santiago, Putumayo pilot and as a diagnostic reference for other projects whose
AI-generated outlines are less accurate.

The most important point is that the visible result is not produced by one
magic model call. Its quality comes from a chain of choices: using appropriate
imagery, matching the inference scale to the model's training scale, preserving
context with overlapping tiles, interpreting the model's classes correctly,
removing tiny components, separating candidate discovery from source
corroboration, and retaining human review as the final decision.

No geometry produced by this pipeline is automatically suitable for upload to
OpenStreetMap. Under the proposed import workflow, a person must compare every
polygon with OAM imagery, keep it only when its shape is accurate, correct it
when necessary, and reject it when uncertain. Community approval and a fresh
OSM conflict check are still required before any upload.

## Pipeline at a glance

```mermaid
flowchart TD
    A["OAM orthomosaic<br/>5.6767 cm native GSD"] --> B["Resample to 0.30 m/pixel<br/>and derive valid-pixel mask"]
    B --> C["Overlapping 256 px tiles<br/>192 px stride"]
    C --> D["ImageNet normalization"]
    D --> E["DINOv3 ViT-S/16 + UperNet<br/>three-class ONNX inference"]
    E --> F["Building softmax probability<br/>midpoint-cropped tile stitching"]
    F --> G["Threshold at 0.4371<br/>remove components below 8 m²"]
    G --> H["Raster-to-polygon conversion<br/>AI screening layer"]
    H --> I["Microsoft corroboration<br/>and unmatched-candidate recovery"]
    I --> J["OAM coverage and live OSM<br/>conflation guards"]
    J --> K["Stratified human-review queue"]
    K --> L["Human validation against OAM<br/>keep, correct, or reject"]
```

The actual implementation is split between:

- [`santiago_ai_pass.py`](../santiago_ai_pass.py), which prepares the imagery,
  runs the neural network, stitches the probability raster, removes tiny
  components, and creates the initial AI polygons; and
- [`build_review_queue.py`](../build_review_queue.py), which compares those AI
  signals with Microsoft footprints, filters against valid imagery and existing
  OSM objects, scores the candidates, and builds the review batch.

## 1. Source imagery

The imagery is the **Santiago Putumayo** OpenAerialMap orthomosaic:

| Property | Value |
|---|---|
| Acquisition date | 2025-11-29 |
| Native ground sample distance | 0.056767 m/pixel |
| Valid image area | 3.391 km² |
| Catalog record | `692bbb0e2183424759916f87` |
| Contributor | Jorge Froilan Mallama Botina |
| Provider | JM |
| License | CC BY 4.0 through OpenAerialMap |

The exact catalog, tile, file, attribution, and hash information is preserved in
[`santiago_results/source_manifest.json`](../santiago_results/source_manifest.json).
OpenAerialMap's [legal page](https://openaerialmap.org/legal/) explicitly states
that its imagery is available to be traced in OpenStreetMap.

### The valid footprint matters

The rectangular image bounds cover roughly 14.64 km², but about 78.4% of those
pixels are black/nodata. The usable flight corridor is irregular and covers only
3.391 km². Counting detections in the rectangular bounds would therefore mix
real imagery with empty space and would substantially overstate coverage.

During inference, a pixel is considered valid when any RGB band is nonzero. A
separate vector footprint in
[`santiago_results/oam_valid_footprint.geojson`](../santiago_results/oam_valid_footprint.geojson)
is then used later for stricter candidate-level coverage checks.

## 2. Model and why its input scale is important

The model is
[`hotosm/dinov3s-buildings`](https://huggingface.co/hotosm/dinov3s-buildings),
at revision:

```text
f7254545dfbbf10a40594d5a5fb62136902000ae
```

It uses a frozen DINOv3 ViT-S/16 visual backbone with a trained UperNet decoder
and segmentation head. It was trained for very-high-resolution building
segmentation using OAM imagery and validated OSM labels.

The source orthomosaic is much sharper than the scale on which the released
model was trained. At Santiago's latitude, a z19 OAM tile is approximately
0.30 m/pixel. The pipeline therefore bilinearly resamples the native 5.68 cm
imagery to **0.30 m/pixel before inference**.

This is probably the single most transferable lesson in this implementation.
A model trained to recognize a 10 m roof as roughly 33 pixels wide may behave
poorly when the same roof is suddenly 176 pixels wide. Feeding the sharpest
available raster at native resolution is not automatically better; matching the
model's learned object scale is usually more important.

The Santiago target grid was:

```text
15,603 × 10,425 pixels at 0.30 m/pixel
```

## 3. Tiling and seam control

The resampled mosaic is processed in:

- 256 × 256 pixel tiles;
- a 192 pixel stride; and
- a 64 pixel, or 25%, overlap between ordinary neighboring tiles.

The final tile along each row or column is anchored to the image edge so the
entire raster is covered even when its dimensions are not exact multiples of
the stride.

### Midpoint stitching

The overlapping predictions are not averaged. Instead, neighboring tiles are
partitioned at the midpoint of their overlap. Only the central ownership region
of each prediction is copied into the full probability raster. Consequently,
almost all pixels near a model tile's outer edge are replaced by predictions
made with more surrounding context in a neighboring tile.

This reduces the square seams and clipped buildings often seen when inference
tiles are simply placed edge-to-edge. It is best described as **overlap plus
midpoint cropping**, not probability blending.

Tiles with fewer than 256 valid pixels are skipped. Because a full tile contains
65,536 pixels, this eliminates essentially empty nodata tiles without rejecting
ordinary tiles that happen to cross the irregular flight boundary.

## 4. Pixel preprocessing

Every RGB tile is converted to floating point in the range 0–1, then normalized
with the standard ImageNet channel statistics:

```text
mean = [0.485, 0.456, 0.406]
std  = [0.229, 0.224, 0.225]
```

The tensor layout is channel-first and batched:

```text
[1, 3, 256, 256]
```

A surprising number of weak segmentation results come from using the correct
weights with the wrong preprocessing: BGR instead of RGB, 0–255 inputs instead
of 0–1, omitted normalization, or an unexpected channel layout.

## 5. Neural-network inference

Inference uses the released ONNX model through ONNX Runtime. The Santiago run
requested Apple's Core ML provider first and CPU as the fallback:

```text
CoreMLExecutionProvider
CPUExecutionProvider
```

The released graph returns per-pixel logits in this class order:

1. building;
2. building boundary; and
3. background.

The pipeline computes a numerically stable softmax over the three classes and
retains the building probability channel. The separate boundary class is
important even though it is not directly vectorized: during training it gives
the network an explicit way to represent the gap between touching roofs rather
than forcing every edge pixel to be either building or background.

## 6. Probability raster and threshold

The building probability is quantized to an unsigned 8-bit raster from 0 to 255.
The model's published default threshold, **0.4371**, is used. In the stored
8-bit representation this becomes an integer cutoff of 111, equivalent to
approximately 0.4353 after quantization.

During a fresh run, the probability raster is written beneath the selected
output directory as:

```text
ai_probability_uint8.npy
```

The historical raster from the recorded run is not included in this lightweight
repository snapshot. Keeping that intermediate in future runs is useful because
thresholds can be retested without paying the inference cost again.

### Threshold sensitivity

The pipeline tested four cutoffs while holding the 8 m² area filter constant:

| Probability threshold | 8-connected components at least 8 m² |
|---:|---:|
| 0.35 | 517 |
| 0.4371 | 529 |
| 0.50 | 539 |
| 0.55 | 537 |

The narrow range is reassuring: the candidate count is not being determined by
one fragile threshold choice. The total detected area does decline steadily as
the threshold rises, which is expected even though the component count does not
move monotonically. A higher threshold can split one merged component into two,
increasing the count while shrinking the total mask.

Threshold selection should still be validated locally. Candidate count alone is
not a quality metric.

## 7. Component cleanup and vectorization

The thresholded mask is analyzed using 8-connected components. Components below
8 m² are removed. At 0.30 m/pixel, one pixel represents 0.09 m², so the cutoff is:

```text
ceil(8 / 0.09) = 89 pixels
```

This suppresses isolated noise while retaining small rural structures. The
cleaned raster is converted to polygons using its georeferenced transform, and
invalid rings are repaired with a zero-width geometry buffer. The polygons are
then transformed from the imagery CRS to WGS84 GeoJSON.

The local script does **not** simplify, orthogonalize, smooth, or square the
polygons. It also does not replace the AI outline with a Microsoft outline when
the two agree. This is important when diagnosing where the apparent quality
comes from: the initial AI shape is a direct polygonization of the cleaned
0.30 m mask.

### Why the run reports both 529 and 537

At the default threshold there are 529 components under OpenCV's 8-neighbor
connectivity. Rasterio's polygon conversion uses 4-neighbor connectivity by
default, so a few components joined only at diagonal pixels become separate
polygons. The final screening GeoJSON therefore contains **537 features**. This
is an implementation detail, not evidence that eight buildings appeared during
vectorization.

Every emitted feature is labeled with:

- `screening_only: true`;
- the model identifier;
- the probability threshold;
- the minimum-area rule; and
- its detected area.

## 8. What Microsoft contributes—and what it does not

Microsoft Global Building Footprints are used as an independent discovery and
corroboration source. They do not geometrically refine, average, or replace an
AI outline when a match exists.

Spatial comparisons are performed in local UTM coordinates. An AI and Microsoft
footprint are considered a screening match if any of these conditions is true:

- intersection-over-union is at least 0.10;
- intersection covers at least 35% of the smaller footprint; or
- their edges are at most 2 m apart **and** their centroids are at most 8 m apart.

Each Microsoft footprint is assigned to its best matching AI candidate. The
candidate construction rules are then:

- matched AI + Microsoft: retain the **AI** geometry and label it consensus;
- AI without Microsoft: retain the AI geometry as an AI-only candidate; and
- Microsoft without AI: retain the Microsoft geometry as a Microsoft-only
  recovery candidate.

When several Microsoft footprints match one AI component, the queue records
multiple provisional roof units and flags a possible merged-roof case. This is
especially useful in dense Santiago blocks, where the AI sometimes joins
attached but structurally separate roofs.

Thus, Microsoft consensus improves **confidence and recall**, not the pixel-level
quality of the AI geometry shown for a consensus candidate.

## 9. Coverage, conflation, and safety guards

Candidate geometry must be at least 95% covered by valid OAM imagery. Candidates
within 8 m of the valid-data boundary receive an edge warning, since partial
roofs and resampling artifacts are more common there.

Existing OSM buildings are matched under a deliberately stricter, overlap-only
rule. Proximity alone can never suppress a candidate as already mapped. Each OSM
feature is assigned to no more than one best candidate, and a whole candidate is
excluded only when the OSM coverage and provisional-unit accounting support that
conclusion. Partial overlaps remain visible for manual conflation and are kept
out of the pilot batch.

IGAC and Google Open Buildings are completely excluded from queue geometry,
matching, scoring, ordering, and pilot selection. They were useful in the
initial source audit, but changing either dataset cannot change the shareable
review queue. Tests enforce this separation.

## 10. Candidate scoring and the review sample

Candidates receive a review priority, not a model confidence score. The starting
scores favor source agreement:

| Category | Base priority |
|---|---:|
| AI + Microsoft consensus | 82 |
| AI only | 55 |
| Microsoft only | 47 |

The score is adjusted for plausible area, geometric compactness, proximity to
the OAM edge, and signs that one AI component contains several roofs. Irregular
or unusually large features are intentionally not hidden; they are demoted and
flagged for review.

The deterministic 200-candidate pilot is stratified as:

- 130 AI + Microsoft consensus candidates;
- 35 AI-only candidates; and
- 35 Microsoft-only candidates.

That mixture prevents the attractive consensus cases from hiding failure modes
in either source. Stable geometry-derived identifiers make the queue and pilot
reproducible even if input feature order changes.

## 11. Human review

The local review application exposes four decisions:

- **Accept:** a real, current building is clearly visible;
- **Needs retrace:** a building is present, but the candidate is merged, offset,
  or malformed;
- **Reject:** the candidate is not a building, is demolished, lies outside valid
  imagery, or is too ambiguous; and
- **Uncertain:** defer the decision.

Even “Accept” is not “upload ready.” The reviewer is deciding whether the region
contains a plausible building, not certifying the candidate vertices. Final OSM
geometry must be manually traced or corrected from OAM, checked against current
OSM, tagged conservatively, and validated.

This human stage is part of the quality system. Removing it changes the nature
and risk of the pipeline.

## 12. Santiago run results

The recorded run processed:

| Metric | Result |
|---|---:|
| Valid pixel fraction | 21.5614% |
| Grid tiles processed | 1,163 |
| Essentially empty tiles skipped | 3,211 |
| Runtime | 129.19 seconds |
| AI screening polygons | 537 |
| Microsoft footprints in valid imagery | 491 |
| Existing OSM buildings in the snapshot | 8 |

After source matching and safety checks, the queue contained:

| Metric | Result |
|---|---:|
| Total queue candidates | 591 |
| Active candidates | 573 |
| AI + Microsoft consensus | 279 active |
| AI only | 252 active |
| Microsoft only | 42 active |
| Provisional remaining roof units | 724 active |
| Upload-ready candidates | 0 |

The provisional roof-unit count is a workload estimate. It is not a confirmed
building count, because attached roofs can be interpreted differently and one
segmentation component can contain several structures.

## 13. Reproducing the inference pass

Install the repository requirements, obtain the exact OAM GeoTIFF and the model's
`model.onnx` at the revision listed above, then run:

```bash
python3 santiago_ai_pass.py \
  --imagery /absolute/path/to/santiago_oam.tif \
  --model /absolute/path/to/model.onnx \
  --output-dir /absolute/path/to/output \
  --gsd 0.30 \
  --tile-size 256 \
  --stride 192 \
  --threshold 0.4371 \
  --minimum-area 8
```

Then build the review queue:

```bash
python3 build_review_queue.py \
  --ai /absolute/path/to/output/ai_buildings_screening.geojson \
  --sources /absolute/path/to/source_comparison.geojson \
  --footprint /absolute/path/to/oam_valid_footprint.geojson \
  --queue-out /absolute/path/to/review_queue.geojson \
  --pilot-out /absolute/path/to/pilot_batch.geojson \
  --summary-out /absolute/path/to/review_queue_summary.json \
  --pilot-size 200
```

Reproduction requires the original source assets, not merely similarly named
datasets. The source manifest records the exact URLs, dates, licenses, model
revision, and hashes used in this run.

## 14. What probably explains the strong outlines

In likely order of importance:

1. **Correct physical scale.** The imagery is resampled to the scale the model
   expects instead of using the much sharper native raster directly.
2. **Very good source imagery.** The OAM flight is recent and has sufficient
   detail for roof boundaries, with fewer compression artifacts than many global
   basemaps.
3. **A strong segmentation architecture.** The DINOv3 representation plus
   UperNet decoder can combine local edges with broader roof context.
4. **A boundary-aware training target.** The three-class output gives the model
   a dedicated representation for roof boundaries.
5. **Overlapping inference tiles.** Midpoint cropping prevents low-context tile
   edges from dominating the stitched result.
6. **Correct normalization and class interpretation.** The ONNX input and output
   contract is implemented explicitly.
7. **An area floor.** Tiny false-positive components never reach the candidate
   layer.
8. **Independent corroboration.** Microsoft improves ranking and recovers some
   omissions, although it does not refine matched AI geometry.
9. **Human triage.** Bad and ambiguous results remain distinguishable instead of
   being silently treated as correct.

There is no learned local Santiago fine-tuning in this repository. The strong
result comes from applying the released general model carefully and surrounding
it with defensible geospatial processing and review.

## 15. Diagnosing a weaker outline pipeline

| Symptom | Most likely checks |
|---|---|
| Roofs are consistently too large or too small | Physical GSD mismatch; wrong probability threshold; unexpected resize behavior |
| Predictions appear as square tile artifacts | No tile overlap; keeping tile-edge predictions; inconsistent padding |
| Neighboring buildings merge together | Wrong scale; insufficient output resolution; boundary class ignored; threshold too low |
| One roof fragments into many pieces | Threshold too high; normalization mismatch; shadows or domain shift |
| Outlines are blocky or staircase-shaped | Coarse inference raster; nearest-neighbor resizing; raw raster polygonization without optional simplification |
| Everything is detected weakly | RGB/BGR error; 0–255 versus 0–1 error; missing ImageNet normalization; incorrect model weights |
| Polygons are spatially offset from roofs | Imagery georeferencing or orthorectification problem; layer offset—not primarily a segmentation problem |
| Many tiny false positives | No minimum-area filter; poor valid-pixel mask; imagery domain differs from training data |
| Good-looking candidates are still missing | Candidate source has limited recall; inspect the full probability raster and add a stratified omission audit |
| Counts change wildly with a small threshold adjustment | Poor probability calibration or strong domain shift; do not select a cutoff by appearance alone |

### Recommended comparison experiment

To isolate the cause in another project, use the same small, manually labeled
area and change only one variable at a time:

1. run at the other project's current GSD;
2. run at 0.30 m/pixel;
3. enable 256/192 overlapping tiles with midpoint cropping;
4. verify RGB and ImageNet normalization;
5. inspect all three output classes, not only one raw channel;
6. compare thresholds 0.35, 0.4371, 0.50, and 0.55;
7. score precision, recall, intersection-over-union, merge rate, and split rate
   against human reference outlines.

Visual attractiveness alone is not enough. A useful evaluation must separately
measure missed buildings, false positives, merged neighbors, fragmented roofs,
and positional error.

## 16. Known limitations

- The AI often merges attached urban roofs.
- Rural vegetation, bare plots, ponds, paved areas, and imagery edges can produce
  false positives.
- Roof outlines are not necessarily wall footprints.
- The 8 m² cutoff can remove genuinely tiny structures.
- Quantizing probabilities to 8 bits slightly changes the exact numerical
  threshold.
- The local polygonizer does not regularize corners or enforce orthogonality.
- Imagery alignment must be checked independently against reliable ground or OSM
  references.
- Source agreement does not prove correctness; correlated datasets can agree on
  the same old or malformed building.
- Review decisions made against a frozen OSM snapshot require a new live OSM
  duplicate check immediately before editing.

## 17. Reproducibility fingerprints

The source manifest records the main provenance fingerprints. At the time of
this document, the relevant local files have these SHA-256 hashes:

```text
santiago_ai_pass.py
dfac112157979e591ffcb86037bc832aa83c14a76c0bf1d895d520cc235616ea

build_review_queue.py
e1068990cc0b26f3a035fafbc1c2d22c563160d3682b3948bc146f2f4849bd0a

santiago_results/ai_buildings_screening.geojson
85892df28b10cd11b2ce7197e3d79ac0cfc76b714e183597c56c1f538a2f16e3

santiago_results/review_queue.geojson
f8db068bad6db376a424d29c85d054824618537d27de0a4d99fd61b67708447b
```

Because this repository does not yet have a Git commit, these fingerprints are
the most precise identifiers for the documented implementation state.
