Reproducible Train/Validation Splits for Spatial Data

A model reports 0.84 IoU on validation and 0.61 in production, and nothing in between changed. The usual cause is not the model: it is a train/validation split drawn by shuffling tiles, which put the tile north of a warehouse in training and the tile south of it in validation. The two share the roof, the shadow, the sensor pass, and often the same building cut in half. The model was evaluated on ground it had already seen.

Fixing it takes three properties that a train_test_split call does not have. The split must be spatially blocked, so held-out ground is genuinely held out. It must be buffered, so no validation tile touches a training tile. And it must be reproducible from a manifest, so an evaluation number from six months ago can be traced to the exact assignment that produced it. This topic builds all three, and the leakage test that proves them.

Prerequisites & Toolchain Alignment

bash
pip install geopandas==0.14.4 shapely==2.0.6 pyproj==3.6.1 \
            numpy==1.26.4 pandas==2.2.2 pyyaml==6.0.1

The split operates on tile footprints, not on imagery, so it is fast and needs no GPU. It does need:

  • Tile footprints in a metric CRS. Block edges and buffer widths are distances. Reproject once, as coordinate reference systems in annotation pipelines describes, and record the CRS in the manifest.
  • A stable tile identifier. The assignment hashes an identifier; if tile ids are regenerated on each tiling pass, the split changes for reasons nobody intended.
  • A scene identifier per tile. Two tiles from the same acquisition share sensor conditions even when they are far apart, and some projects treat that as leakage too.
Random tiles against buffered blocks On the left a random split scatters validation tiles among training tiles, so almost every held-out tile has a training neighbour sharing objects and sensor conditions. On the right the area is divided into blocks, whole blocks are assigned to a split, and a buffer strip along the boundary is dropped from training so nothing in validation touches anything the model saw. random tile split every held-out tile has a training neighbour validation IoU 0.84 production 0.61 buffered spatial blocks train buffer validation validation IoU 0.66 production 0.64 — the number holds

Core Split Workflow

Step 1 — Choose a Block Size From the Data

The block edge should exceed the distance at which two samples stop resembling each other. Guessing works badly; a coarse variogram of the target over the tile centroids gives a defensible number in a few lines.

Reading the block size off a variogram Semivariance of building density rises steeply with distance and flattens at about 400 metres, which is the autocorrelation range: beyond it two tiles no longer resemble each other. The block edge is taken at three times that range, 1200 metres, so two tiles in different blocks are genuinely independent samples rather than neighbours with a boundary drawn between them. 0 400 m 800 m 1 200 m 1 500 m lag distance between tile centroids semivariance sill range ≈ 400 m — beyond here, tiles stop resembling each other block edge = 3 × range a block smaller than the range puts correlated ground on both sides of the split, which is the leak the blocking was for
python
import numpy as np
import geopandas as gpd

def autocorrelation_range(tiles: gpd.GeoDataFrame, value_col: str,
                          max_lag_m: float = 5000.0, n_bins: int = 20) -> float:
    """Rough empirical range: the lag beyond which semivariance stops rising.

    `tiles` must be in a metric CRS; `value_col` is a per-tile target summary such as
    building density or the share of the dominant class.
    """
    pts = np.column_stack([tiles.geometry.centroid.x, tiles.geometry.centroid.y])
    vals = tiles[value_col].to_numpy(dtype=float)
    d = np.linalg.norm(pts[:, None, :] - pts[None, :, :], axis=-1)
    dv = 0.5 * (vals[:, None] - vals[None, :]) ** 2
    iu = np.triu_indices_from(d, k=1)
    d, dv = d[iu], dv[iu]
    edges = np.linspace(0, max_lag_m, n_bins + 1)
    gamma = np.array([dv[(d >= lo) & (d < hi)].mean() if ((d >= lo) & (d < hi)).any() else np.nan
                      for lo, hi in zip(edges[:-1], edges[1:])])
    sill = np.nanmax(gamma)
    reached = np.where(gamma >= 0.95 * sill)[0]
    return float(edges[reached[0] + 1]) if len(reached) else max_lag_m

Take the block edge as roughly three times that range. If the computation is impractical — very large tile counts make the pairwise distance matrix expensive — sample a few thousand tiles; the range estimate is not sensitive to that.

Step 2 — Assign Blocks Deterministically

The assignment must not depend on tile order, dataset size, or a random seed that lives in someone’s notebook. Hashing the block identifier gives a stable, order-independent bucket that new tiles fall into automatically.

python
import hashlib

def block_id(x: float, y: float, block_m: float) -> str:
    """The block a coordinate falls in, as a stable string."""
    return f"{int(np.floor(x / block_m))}_{int(np.floor(y / block_m))}"

def assign_split(block: str, salt: str, val_share: float = 0.2,
                 test_share: float = 0.1) -> str:
    """Deterministic split assignment for a block. Same block + salt → same answer, always."""
    h = hashlib.sha256(f"{salt}:{block}".encode()).digest()
    # first 8 bytes as a fraction in [0, 1)
    frac = int.from_bytes(h[:8], "big") / float(1 << 64)
    if frac < val_share:
        return "val"
    if frac < val_share + test_share:
        return "test"
    return "train"

The salt is the only knob that changes the split, and changing it is a deliberate, versioned act. Recording it in the manifest is what lets a reviewer confirm that last quarter’s numbers and this quarter’s came from the same partition of the world.

Step 3 — Buffer the Boundaries

A tile in a training block that shares an edge with a validation block is still adjacent to held-out ground. Drop it from training — not from validation, which would shrink the evaluation set for no benefit.

python
def apply_buffer(tiles: gpd.GeoDataFrame, buffer_m: float,
                 split_col: str = "split") -> gpd.GeoDataFrame:
    """Drop training tiles within `buffer_m` of any validation or test tile."""
    held = tiles[tiles[split_col].isin(["val", "test"])]
    if held.empty:
        return tiles
    zone = held.geometry.buffer(buffer_m).union_all()
    is_train = tiles[split_col] == "train"
    touching = is_train & tiles.geometry.intersects(zone)
    out = tiles.copy()
    out.loc[touching, split_col] = "buffer"     # kept in the manifest, used by nothing
    return out

Marking them buffer rather than deleting them matters: the manifest then records that these tiles exist and were deliberately excluded, so a later reader does not conclude the tiling was incomplete.

Step 4 — Test for Leakage Before Training

The split is a claim, and the claim deserves an assertion. Two failures matter: a validation tile within the buffer of a training tile, and a scene appearing on both sides.

python
def assert_no_leakage(tiles: gpd.GeoDataFrame, buffer_m: float,
                      split_col: str = "split", scene_col: str = "scene_id") -> None:
    train = tiles[tiles[split_col] == "train"]
    held = tiles[tiles[split_col].isin(["val", "test"])]

    if not train.empty and not held.empty:
        zone = train.geometry.buffer(buffer_m).union_all()
        bleed = held[held.geometry.intersects(zone)]
        if len(bleed):
            raise AssertionError(
                f"{len(bleed)} held-out tile(s) within {buffer_m} m of training data, "
                f"e.g. {bleed.index[0]}")

    shared = set(train[scene_col]) & set(held[scene_col])
    if shared:
        raise AssertionError(f"{len(shared)} scene(s) in both splits, e.g. {sorted(shared)[0]}")

Run it as a step in the CI gate as well as before training. A split is exactly the kind of artifact that is correct when written and quietly wrong three merges later.

The two leaks a split test has to catch The first leak is spatial: a validation tile sits directly across the block boundary from a training tile, sharing objects and shadows. The second is by scene: two tiles far apart in space came from the same acquisition, so they share sensor geometry, atmospheric conditions and processing, and a model can key on that instead of on the ground. leak 1 — adjacency one warehouse, cut by the boundary train validation fixed by a buffer strip leak 2 — shared scene tile 0142 tile 1907 same acquisition kilometres apart, identical sun angle, atmosphere and processing chain fixed by splitting on scene, not only on space a buffer alone does not catch leak 2

Step 5 — Write and Version the Split Manifest

The manifest is the artifact. It is small, it is deterministic, and it is what makes an evaluation number auditable a year later.

python
import json
from pathlib import Path

def write_split_manifest(path: str, tiles: gpd.GeoDataFrame, *, salt: str,
                         block_m: float, buffer_m: float, crs: str) -> str:
    """Serialise the split deterministically and return its digest."""
    payload = {
        "salt": salt,
        "block_m": block_m,
        "buffer_m": buffer_m,
        "crs": crs,
        "counts": tiles["split"].value_counts().sort_index().to_dict(),
        "assignments": {str(t.Index): t.split for t in
                        tiles.sort_index().itertuples()},
    }
    blob = json.dumps(payload, indent=2, sort_keys=True) + "\n"
    Path(path).write_text(blob, encoding="utf-8")
    return hashlib.sha256(blob.encode()).hexdigest()

Track that file with the dataset, not beside it. The pattern is the one in tracking annotation changes with SHA hashing: sorted keys, no timestamp, so two runs that produce the same split produce byte-identical manifests and no spurious version.

Split Parameters & Configuration Reference

Parameter Typical Effect
Block edge 3 × autocorrelation range (0.5 – 5 km) Too small reintroduces leakage; too large limits achievable ratios
buffer_m 1 – 2 tile widths Drops 2 – 8% of tiles from training; below one tile width it does nothing
val_share 0.15 – 0.20 Blocked splits are noisier than random ones, so do not go below 0.15
test_share 0.10, held out entirely Touched once, at the end, or it is a second validation set
salt a project-lifetime constant Changing it invalidates every historical comparison
Scene exclusivity on for single-sensor projects Off only when scenes genuinely do not correlate

Edge Cases & Spatial Gotchas

A rare class living in one block. Hash assignment does not know that every solar farm in the dataset is in block 12_7. If that block lands in validation, the model never trains on solar farms. Check per-class counts per split after assignment and re-salt — not re-shuffle — if a class is absent from training.

Long thin study areas. A corridor along a motorway has almost no interior, so blocks and buffers eat most of it. Use blocks along the corridor axis rather than a square grid, and accept a coarser split ratio.

Multi-temporal stacks. Two dates over the same block are the same ground. They must share a split, which means the block id must be computed from location alone and never include the date.

Tiles that straddle a block boundary. Assign by centroid and be consistent about it. Assigning by intersection puts one tile in two blocks, which quietly makes the assignment order-dependent — the exact property the hashing was there to remove.

Growing the dataset later. New tiles hash into existing blocks and inherit their side automatically, which is the point. What must not happen is re-running block sizing on the larger dataset and getting a different block_m, because that reassigns everything. Pin block_m in the manifest and treat it as fixed for the project’s life.

Integration & Automation Hooks

As a DVC stage. The split is a pure function of the tile footprints, the salt and the two distances, so it belongs in the pipeline as a stage whose outputs are the manifest. When the footprints change, the split regenerates; when they do not, nothing does.

yaml
stages:
  split:
    cmd: python -m pipeline.split --salt ${split.salt} --block-m ${split.block_m} --buffer-m ${split.buffer_m}
    deps:
      - data/tiles/footprints.parquet
      - pipeline/split.py
    params:
      - split.salt
      - split.block_m
      - split.buffer_m
    outs:
      - data/splits/manifest.json

As a CI assertion. assert_no_leakage runs in the same gate as the geometry and schema checks, so a pull request that adds tiles across a boundary fails before it is merged.

In the data loader. The loader reads the manifest and filters, rather than globbing a directory per split. A directory layout encodes the split in the filesystem, which means changing it requires moving terabytes and makes two experiments on different splits impossible to run side by side.

Validation & Testing

python
def test_assignment_is_order_independent() -> None:
    blocks = ["3_4", "1_1", "9_2", "7_7"]
    a = {b: assign_split(b, salt="v1") for b in blocks}
    b = {b: assign_split(b, salt="v1") for b in reversed(blocks)}
    assert a == b

def test_salt_changes_the_split() -> None:
    blocks = [f"{i}_{j}" for i in range(20) for j in range(20)]
    v1 = [assign_split(b, salt="v1") for b in blocks]
    v2 = [assign_split(b, salt="v2") for b in blocks]
    assert v1 != v2                                   # a new salt is a new partition

def test_leakage_detector_rejects_a_random_split(tiles) -> None:
    """Feed the checker the split it exists to refuse."""
    import numpy as np, pytest
    rng = np.random.default_rng(0)
    tiles = tiles.copy()
    tiles["split"] = rng.choice(["train", "val"], size=len(tiles), p=[0.8, 0.2])
    with pytest.raises(AssertionError, match="within"):
        assert_no_leakage(tiles, buffer_m=200.0)

The last test is the one that keeps the rest honest: it constructs the naive random split this whole topic argues against, and asserts that the machinery refuses it.

Frequently Asked Questions

# Can I use k-fold cross-validation with spatial blocks?

Yes — assign blocks to k buckets by the same hash and rotate which bucket is held out. The buffer has to be recomputed per fold, since the boundary moves. It is the right approach when the dataset is small enough that a single validation split is noisy, which is common early in a project.

# What if my validation number gets worse after switching to blocked splits?

That is the expected outcome, and it is good news. The previous number was measuring memorisation of neighbouring ground. Record both for one release so the change in methodology is visible in the history, then keep only the blocked number.

# How does this interact with active learning?

Actively selected tiles enter training, so they must respect the split: a tile in a validation block can never be labelled into training, no matter how uncertain the model is about it. Filter candidates by split before ranking, as part of the selection step described in uncertainty sampling for geospatial active learning.

# Should the test split ever be looked at?

Once, at the end, for the number you report. Every additional look leaks information through the decisions it influences. If you find yourself checking test performance between experiments, what you have is a second validation set and no test set at all.

Splitting is one stage of the broader Dataset Versioning & Spatial Data Sync pipeline, which keeps every artifact — including this one — reproducible from a manifest.