Cold-Start Strategies for New Annotation Projects
Active learning has a chicken-and-egg problem at the start of every project: the loop selects tiles the model is unsure about, and there is no model. The usual response is to label a random few thousand tiles and start the loop, which wastes most of that budget — a random sample of an aerial pool is mostly ground with nothing in it, drawn from whichever region happens to dominate the archive, on whichever days happened to be cloud-free.
The cold-start phase deserves its own strategy, and it is a different strategy: coverage rather than uncertainty. With no model, the only thing worth optimising is how much of the variation the first batch contains. This topic covers stratifying the pool, sampling for spread inside each stratum, using a zero-shot model to cut the tracing cost, freezing an evaluation set that active learning can never contaminate, and the measurable condition for handing over to uncertainty sampling.
Prerequisites & Toolchain Alignment
pip install geopandas==0.14.4 shapely==2.0.6 pyproj==3.6.1 \
scikit-learn==1.5.1 numpy==1.26.4 rasterio==1.3.10 torch==2.3.1
What the cold-start phase needs that later rounds do not:
- Pool metadata before pixels. Region, acquisition date, sensor and cloud cover per tile, which comes from the STAC catalog rather than from reading imagery. Stratification runs on the metadata alone.
- A cheap embedding, optionally. A small pretrained encoder over a downsampled tile gives a feature vector for diversity sampling. It is not required — geographic spread alone is a decent proxy — but it catches the case where two distant tiles look identical.
- A frozen evaluation plan. Decide the blocked evaluation set before labelling anything, following reproducible train/validation splits. Doing it afterwards means the evaluation set is drawn from ground the active loop has already picked over.
Cold-Start Workflow
Step 1 — Stratify the Pool on Metadata
Strata are the axes along which the model will be expected to generalise. For most aerial and satellite projects those are geography, time of year and sensor — and the pool is almost never balanced across them.
import geopandas as gpd
import pandas as pd
def build_strata(pool: gpd.GeoDataFrame, region_col: str = "admin_area",
date_col: str = "acquired", sensor_col: str = "platform") -> pd.Series:
"""A stratum key per tile, from metadata only — no pixels are read."""
season = pd.to_datetime(pool[date_col]).dt.quarter.map(
{1: "winter", 2: "spring", 3: "summer", 4: "autumn"})
return (pool[region_col].astype(str) + "|" + season + "|" + pool[sensor_col].astype(str))
def stratum_report(pool: gpd.GeoDataFrame, strata: pd.Series) -> pd.DataFrame:
"""What the pool actually contains — usually a surprise."""
counts = strata.value_counts().rename("tiles").to_frame()
counts["share"] = (counts["tiles"] / counts["tiles"].sum()).round(4)
return counts
Run the report before deciding anything. A pool that is 71% one city in summer is normal, and it is exactly why a random first batch teaches the model that summer in that city is what the world looks like.
Step 2 — Sample for Spread, Not Proportionally
Proportional sampling reproduces the pool’s imbalance in the labelled set. For a first batch, allocate closer to evenly across strata — capped by what each stratum can supply — and inside each stratum pick tiles that are far apart.
import numpy as np
def allocate(strata_counts: pd.Series, batch_size: int, evenness: float = 0.7) -> dict[str, int]:
"""Blend proportional and uniform allocation. evenness=1.0 is fully uniform."""
prop = strata_counts / strata_counts.sum()
unif = pd.Series(1.0 / len(strata_counts), index=strata_counts.index)
mix = (1 - evenness) * prop + evenness * unif
alloc = (mix / mix.sum() * batch_size).round().astype(int)
return {k: int(min(v, strata_counts[k])) for k, v in alloc.items()}
def spread_sample(tiles: gpd.GeoDataFrame, n: int, seed: int = 0) -> list:
"""Greedy farthest-point sampling on tile centroids: maximal geographic spread."""
if len(tiles) <= n:
return list(tiles.index)
pts = np.column_stack([tiles.geometry.centroid.x, tiles.geometry.centroid.y])
rng = np.random.default_rng(seed)
chosen = [int(rng.integers(len(pts)))]
dist = np.linalg.norm(pts - pts[chosen[0]], axis=1)
for _ in range(n - 1):
nxt = int(np.argmax(dist))
chosen.append(nxt)
dist = np.minimum(dist, np.linalg.norm(pts - pts[nxt], axis=1))
return [tiles.index[i] for i in chosen]
Farthest-point sampling is the cheap version of diversity sampling and needs no model. It is also the step that stops the first batch being forty tiles of the same industrial estate, which is what uniform random sampling within a small stratum tends to produce.
Step 3 — Pre-Label the Geometry, Not the Class
A zero-shot segmentation model cuts the expensive part of the first batch — tracing — without pretending to know your taxonomy. The annotator’s job becomes accept, adjust or delete, plus assigning the class.
from dataclasses import dataclass
@dataclass(frozen=True)
class SeedProposal:
tile_id: str
geometry_wkt: str
source: str # "sam" | "manual"
class_name: None # deliberately empty: the model does not get a vote
def to_prelabels(masks, tile_id: str, transform, crs: str) -> list[SeedProposal]:
"""Vectorise zero-shot masks into class-less proposals for human classification."""
from rasterio.features import shapes
from shapely.geometry import shape
out: list[SeedProposal] = []
for mask in masks:
for geom, value in shapes(mask.astype("uint8"), mask=mask, transform=transform):
if value != 1:
continue
poly = shape(geom).simplify(0.3).buffer(0)
if poly.is_empty or poly.area < 20.0: # square metres, projected CRS
continue
out.append(SeedProposal(tile_id, poly.wkt, "sam", None))
return out
Setting class_name to None is a design decision, not an omission. A pre-label that arrives with a confident wrong class costs more than one that arrives with none: annotators accept plausible defaults, and the resulting bias is invisible in throughput metrics. The mechanics of running the model over tiles are covered in automating pre-labeling with foundation models.
Step 4 — Freeze the Evaluation Set First
The evaluation set must be drawn before the active loop starts and never touched by it. Draw it as blocked random ground, label it, freeze it.
def carve_seed_eval(pool: gpd.GeoDataFrame, salt: str, share: float = 0.15) -> gpd.GeoDataFrame:
"""Blocked random evaluation ground, chosen before any model exists."""
from pipeline.split import block_id, assign_split # see the splits topic
blocks = [block_id(g.centroid.x, g.centroid.y, block_m=1200.0) for g in pool.geometry]
pool = pool.assign(block=blocks)
pool["split"] = [assign_split(b, salt=salt, val_share=share, test_share=0.0) for b in blocks]
return pool[pool["split"] == "val"]
Label this set with the same care as training data — better, if anything, since every future decision is measured against it. Then leave it alone. An evaluation set that grows as the project goes on cannot be compared with itself across time, which defeats the purpose of having one.
Step 5 — Hand Over on a Measurement
The switch from coverage to uncertainty is not a milestone in a schedule; it is a property of the model. Uncertainty sampling only works when the model’s confidence tracks its accuracy, and that is measurable.
import numpy as np
def expected_calibration_error(conf: np.ndarray, correct: np.ndarray, bins: int = 15) -> float:
"""Standard ECE over equal-width confidence bins."""
edges = np.linspace(0.0, 1.0, bins + 1)
ece = 0.0
for lo, hi in zip(edges[:-1], edges[1:]):
m = (conf > lo) & (conf <= hi)
if not m.any():
continue
ece += m.mean() * abs(correct[m].mean() - conf[m].mean())
return float(ece)
def ready_for_uncertainty(conf: np.ndarray, correct: np.ndarray, threshold: float = 0.05) -> bool:
"""Is the model's confidence trustworthy enough to rank an annotation queue?"""
return expected_calibration_error(conf, correct) <= threshold
Until that returns True, keep at least half the batch on coverage. The intermediate mix costs little and protects against the common failure of a project that switched to uncertainty at 500 labels, spent three rounds annotating whatever the under-trained model found confusing, and produced a training set concentrated on one kind of confusion. Fitting the temperature that makes the check meaningful is covered in calibrating confidence scores with temperature scaling.
Cold-Start Parameters & Configuration Reference
| Parameter | Typical | Notes |
|---|---|---|
| First batch size | 300 – 800 tiles | Large enough to train something, small enough to redirect after |
evenness |
0.6 – 0.8 | 1.0 over-weights strata with 30 tiles in them |
| Seed eval share | 10 – 15% of the pool’s blocks | Frozen before the first training run |
| Pre-label min area | 20 m² | Below this, correcting the proposal costs more than drawing it |
| ECE handover threshold | ≤ 0.05 | Measured on a held-out slice, after temperature fitting |
| Coverage share during transition | 50%, decaying | Reaches zero when ECE has been under threshold for two rounds |
Edge Cases & Gotchas
A stratum with almost nothing in it. Three winter tiles is not a stratum; it is a gap in the archive. Allocating it 60 tiles is impossible and allocating it three tells the model nothing. Record it as a known coverage gap and, if the model must work there, acquire imagery rather than pretending.
Pre-labels that make annotators lazy. Proposal acceptance rates above about 90% usually mean people are accepting rather than checking. Sample a slice for adjudication as described in annotation quality metrics and agreement, and watch whether the corrected fraction changes when the model improves.
Farthest-point sampling on a lopsided extent. A pool that includes one tile from a distant island will always pick it first, and possibly a few of its neighbours, because it is far from everything. Run the spread sample inside strata rather than over the whole pool.
Evaluation ground labelled by the least experienced annotator. Because the seed evaluation set is labelled first, it is often labelled by whoever is available before the team is trained. That set then defines the ceiling for every future measurement. Label it last within the cold-start phase, or re-adjudicate it once the guide has stabilised — and version that change explicitly.
Switching to uncertainty because a round number was hit. “We have 2 000 labels, turn on active learning” is the failure this phase exists to prevent. Use the calibration check.
Integration & Automation Hooks
The cold-start selection is a batch job that runs once per round and writes a task list. It fits the same Airflow DAG as the harvest, as an upstream task that reads the pool metadata and writes the next batch’s tile ids. Two details make it safe to re-run: the selection is seeded, so the same round produces the same batch; and already-labelled tiles are excluded by a join rather than by exclusion lists that drift.
Once the handover happens, the same task keeps running with a different scorer — coverage becomes one term in the score rather than the whole of it — which is why it is worth building the selection as a scoring function from the start rather than as a script that shuffles.
Validation & Testing
def test_allocation_never_exceeds_supply() -> None:
counts = pd.Series({"a": 500, "b": 40, "c": 12})
alloc = allocate(counts, batch_size=300, evenness=0.8)
for k, v in alloc.items():
assert v <= counts[k], f"asked for {v} from a stratum holding {counts[k]}"
def test_spread_sample_is_deterministic() -> None:
a = spread_sample(tiles, n=25, seed=7)
b = spread_sample(tiles, n=25, seed=7)
assert a == b
def test_uncertainty_handover_refuses_an_uncalibrated_model() -> None:
"""The check must refuse the model this phase exists to protect against."""
conf = np.full(500, 0.95) # confident everywhere
correct = np.random.default_rng(0).random(500) < 0.60 # right 60% of the time
assert not ready_for_uncertainty(conf, correct)
The last test feeds the machinery exactly the over-confident under-trained model that a young project produces, and asserts the handover check refuses it.
Frequently Asked Questions
# Can I skip the cold start by fine-tuning a published model?
Often yes, and it changes the phase rather than removing it. A model fine-tuned from an open building-footprint checkpoint starts with usable uncertainty far sooner, so the coverage phase can be one batch instead of four. Run the calibration check anyway: a transferred model is frequently confident and wrong on a new sensor, which is precisely the state that makes uncertainty ranking useless.
# How do I choose strata when the project spans several countries?
Use the coarsest axis that changes what the imagery looks like. Administrative boundaries matter less than the built form, sensor and season, so a stratification on those three is usually better than one on country. Where building style genuinely differs across a border, it will show up as a region term anyway.
# Is diversity sampling worth the embedding model?
For the first batch, geographic spread captures most of it and costs nothing. Feature-space diversity earns its keep later, when the loop is ranking by uncertainty and neighbouring tiles all score alike — the problem described in prioritizing tiles by model disagreement.
# What if the pool grows after the cold-start batches?
New tiles join the pool and the strata are recomputed, which may create strata that did not exist. That is a coverage gap and should trigger a coverage batch even if the project has already moved to uncertainty sampling — the model has no basis for being uncertain about ground it has never seen.
Related
- Uncertainty Sampling for Geospatial Active Learning — the phase this one hands over to, and the scoring it depends on
- Automating Pre-Labeling with Foundation Models — how the zero-shot proposals in Step 3 are generated over a tiled scene
- Reproducible Train/Validation Splits for Spatial Data — the blocked assignment the frozen evaluation set is drawn from
- Calibrating Confidence Scores with Temperature Scaling — the fit that makes the handover check meaningful
Cold start is the first stage of the broader Active Learning & Model Feedback Loops cycle, which takes over once the model has something to be uncertain about.