Annotation Quality Metrics & Inter-Annotator Agreement
A geospatial annotation project can report 96% agreement between its annotators and still be shipping unusable labels. The number is real; it is measured on a class distribution where nine in ten features are buildings, so two annotators who default to building on everything ambiguous agree almost perfectly while contributing nothing. Meanwhile the footprints they draw differ by three metres along every wall, which no percentage-agreement figure asks about. This topic covers the metrics that separate those two failures — delineation quality and label agreement — measures each in units that mean something on the ground, and wires the result into an adjudication queue rather than a dashboard nobody acts on.
The distinction matters because the two failures have different fixes. Sloppy delineation is a tooling and training problem: snapping is off, the imagery is being annotated at the wrong zoom, or nobody told the team that eaves are not walls. Label disagreement is a taxonomy problem: two classes are not actually separable on this imagery, or the guide never said which one wins. Reporting a single “quality score” mixes the two into a number that tells you something is wrong and nothing about what to do.
Prerequisites & Toolchain Alignment
The measurement code needs a geometry stack and nothing else — no model, no GPU, and no annotation platform API. Pin the versions, because agreement numbers computed under different shapely releases are not comparable across a project’s history.
pip install geopandas==0.14.4 shapely==2.0.6 pyproj==3.6.1 \
scikit-learn==1.5.1 numpy==1.26.4 pandas==2.2.2
You also need three things that are not packages:
- A deliberate overlap set. Agreement can only be measured where two people annotated the same ground. Five to ten percent of tiles assigned to two annotators is the standard allocation; the assignment must be made when the batch is created, because retrofitting an overlap set later biases it toward whatever was easy to re-open.
- A metric CRS for the working area. Every metric below is a distance or an area, and both are meaningless in degrees. Choose the projection once per batch as described in coordinate reference systems in annotation pipelines and record it in the batch manifest.
- A stable feature identity. Matching two annotators’ work requires knowing which of their features are meant to be the same object. Nothing in the file says so, which is why the matching step below is a real algorithm rather than a join.
Core Measurement Workflow
Step 1 — Load Both Annotators and Project to Metres
Read each annotator’s GeoJSON, confirm the declared CRS, and reproject both to the batch’s working projection before a single metric is computed. Doing it in one place means no downstream function has to ask what units it is holding.
from pathlib import Path
import geopandas as gpd
def load_pair(path_a: Path, path_b: Path, work_crs: str) -> tuple[gpd.GeoDataFrame, gpd.GeoDataFrame]:
"""Load two annotators' features for the same tile, both in a metric CRS."""
a = gpd.read_file(path_a)
b = gpd.read_file(path_b)
for name, gdf in (("a", a), ("b", b)):
if gdf.crs is None:
raise ValueError(f"annotator {name}: no CRS declared in {path_a if name == 'a' else path_b}")
a = a.to_crs(work_crs)
b = b.to_crs(work_crs)
# Repair before measuring: an invalid ring makes every area below meaningless.
a["geometry"] = a.geometry.make_valid()
b["geometry"] = b.geometry.make_valid()
return a, b
Step 2 — Match Features by Greatest Overlap
Two annotators produce two independent feature sets with no shared identifiers. The match is greedy on IoU: for each of A’s features, take B’s feature with the largest overlap, provided that overlap clears a low matching floor. The floor exists to stop a building being “matched” to the road beside it; it is deliberately much lower than the quality threshold, because a badly drawn pair is still a pair.
import geopandas as gpd
import pandas as pd
MATCH_FLOOR = 0.30 # low on purpose: this decides identity, not quality
def match_features(a: gpd.GeoDataFrame, b: gpd.GeoDataFrame) -> pd.DataFrame:
"""Greedy one-to-one matching of A's features to B's by intersection over union."""
pairs: list[dict] = []
taken: set[int] = set()
sindex = b.sindex
for ia, ga in a.geometry.items():
best_j, best_iou = None, 0.0
for ib in sindex.query(ga, predicate="intersects"):
if ib in taken:
continue
gb = b.geometry.iloc[ib]
inter = ga.intersection(gb).area
if inter == 0.0:
continue
iou = inter / (ga.area + gb.area - inter)
if iou > best_iou:
best_j, best_iou = ib, iou
if best_j is not None and best_iou >= MATCH_FLOOR:
taken.add(best_j)
pairs.append({"ia": ia, "ib": best_j, "iou": best_iou})
matched_a = {p["ia"] for p in pairs}
matched_b = {p["ib"] for p in pairs}
return pd.DataFrame(pairs), sorted(set(a.index) - matched_a), sorted(set(range(len(b))) - matched_b)
The two lists the function returns alongside the pairs are not bookkeeping. Unmatched features from A are objects B never drew, and unmatched features from B are objects A never drew — recall and precision failures respectively. A quality report that silently drops them describes only the features both annotators found, which is the easiest subset to agree on.
Step 3 — Score Delineation with Boundary IoU
Plain IoU is dominated by area. A 900 m² warehouse traced three metres wide of its true wall still scores about 0.94; a 40 m² shed traced within half a metre scores 0.88. Ranking by that number puts effort in the wrong place. Boundary IoU compares only a band around each polygon’s edge, so the score reflects the tracing rather than the size.
from shapely.geometry.base import BaseGeometry
def boundary_iou(g1: BaseGeometry, g2: BaseGeometry, band_m: float = 1.0) -> float:
"""IoU restricted to a band of `band_m` metres around each polygon's boundary.
Both geometries must already be in a projected CRS whose unit is the metre.
"""
if g1.is_empty or g2.is_empty:
return 0.0
b1 = g1.boundary.buffer(band_m)
b2 = g2.boundary.buffer(band_m)
inter = b1.intersection(b2).area
if inter == 0.0:
return 0.0
return inter / (b1.area + b2.area - inter)
Set band_m from the ground sample distance rather than by taste: roughly three pixels is the width within which two competent annotators genuinely cannot be told apart. At 30 cm imagery that is 0.9 m; at 10 cm drone imagery, 0.3 m. Using one fixed band across sensors makes the same team look better on coarse imagery, which is the opposite of what the metric is for.
Step 4 — Score Label Agreement with Cohen’s Kappa
Geometry agreement says nothing about whether the two annotators called the object the same thing. Over the matched pairs — and only the matched pairs, since an unmatched feature has no counterpart class — compute Cohen’s kappa on the class labels.
import pandas as pd
from sklearn.metrics import cohen_kappa_score, confusion_matrix
def label_agreement(a: gpd.GeoDataFrame, b: gpd.GeoDataFrame, pairs: pd.DataFrame,
class_field: str = "class_name") -> dict:
"""Cohen's kappa and a confusion matrix over the matched pairs."""
ya = [a.loc[int(r.ia), class_field] for r in pairs.itertuples()]
yb = [b.iloc[int(r.ib)][class_field] for r in pairs.itertuples()]
labels = sorted(set(ya) | set(yb))
return {
"kappa": float(cohen_kappa_score(ya, yb, labels=labels)),
"raw_agreement": float(sum(x == y for x, y in zip(ya, yb)) / max(len(ya), 1)),
"labels": labels,
"matrix": confusion_matrix(ya, yb, labels=labels).tolist(),
"n": len(ya),
}
Report kappa and raw_agreement side by side, always. The gap between them is the class imbalance in your batch, and a large gap is itself a finding: it means the headline agreement number is being carried by one dominant class.
Step 5 — Route the Disagreements
Metrics that end in a report change nothing. Every matched pair below the geometry floor for its class, and every pair whose labels disagree, becomes a queue item with both geometries attached and the tile it came from. The reviewer’s ruling does two things: it fixes the feature, and it produces a sentence for the annotation guide.
from dataclasses import dataclass, asdict
import json
@dataclass(frozen=True)
class Dispute:
tile_id: str
kind: str # "geometry" | "class"
class_a: str
class_b: str
score: float
threshold: float
def build_queue(tile_id: str, a, b, pairs, floors: dict[str, float],
class_field: str = "class_name") -> list[Dispute]:
out: list[Dispute] = []
for r in pairs.itertuples():
ga, gb = a.loc[int(r.ia)], b.iloc[int(r.ib)]
ca, cb = ga[class_field], gb[class_field]
if ca != cb:
out.append(Dispute(tile_id, "class", ca, cb, float(r.iou), 1.0))
continue # a class dispute makes the geometry moot
floor = floors.get(ca, 0.65)
biou = boundary_iou(ga.geometry, gb.geometry)
if biou < floor:
out.append(Dispute(tile_id, "geometry", ca, cb, biou, floor))
return out
def write_queue(path: str, disputes: list[Dispute]) -> None:
with open(path, "w", encoding="utf-8") as fh:
json.dump([asdict(d) for d in disputes], fh, indent=2, sort_keys=True)
fh.write("\n")
The continue after a class dispute is deliberate. When two annotators disagree about what an object is, their boundaries disagree for a reason — one is tracing a roof and the other a parcel — and scoring that disagreement as a delineation failure files the dispute under the wrong heading.
Spatial Parameters & Configuration Reference
| Parameter | Typical value | Units | What it controls |
|---|---|---|---|
| Overlap allocation | 5 – 10% of tiles | share of batch | How much duplicate labelling funds the agreement estimate |
MATCH_FLOOR |
0.30 | IoU | Whether two features are the same object at all |
| Boundary band | 3 × GSD (0.3 – 1.0 m) | metres | The width inside which tracing differences are not real |
| Boundary IoU floor, crisp classes | 0.70 – 0.80 | boundary IoU | Adjudication trigger for buildings, roads, solar arrays |
| Boundary IoU floor, fuzzy classes | 0.45 – 0.60 | boundary IoU | Adjudication trigger for wetland, burn scar, canopy edge |
| Kappa floor per class | 0.60 | kappa | Below this the class pair is a taxonomy problem, not an annotator one |
| Hausdorff alarm | 3 × boundary band | metres | A single corner far out of place, which mean-based scores hide |
Edge Cases & Spatial Gotchas
One annotator splits what the other merges. A row of terraced houses is one feature to A and six to B. Greedy one-to-one matching pairs A’s single polygon with B’s largest and reports five false positives, which reads as B inventing objects. Detect it by checking whether the unmatched features on one side fall inside a matched feature on the other, and treat the case as a taxonomy dispute — the guide has not said whether the unit is the building or the block.
Agreement measured on the easy tiles. If the overlap set is assigned by letting annotators pick a second tile, they pick tiles they found straightforward. The overlap must be sampled by the batch builder, ideally stratified by class so rare classes appear in it at all.
Kappa on two classes with one absent. When a tile’s overlap set contains only buildings, kappa is undefined or degenerate even though both annotators did perfect work. Aggregate kappa at the batch level, not per tile, and report the pair count alongside it so a number computed on eleven features is visibly weaker evidence than one computed on nine hundred.
Self-agreement drift. The same annotator re-labelling the same tile a month later will not reproduce their own work exactly. Measuring that intra-annotator agreement once, on a handful of tiles, gives you the ceiling: no pair of different people will agree better than one person agrees with themselves, and setting thresholds above that ceiling guarantees permanent adjudication traffic.
Boundary band smaller than the coordinate precision. If labels were serialised with five decimal places in EPSG:4326, coordinates are quantised to roughly a metre, and a 0.3 m boundary band measures quantisation noise. Precision has to be fixed at export time — see how to structure GeoJSON for ML training datasets — before any of these thresholds mean what they say.
Integration & Automation Hooks
The measurement runs as a batch job, not in the annotator’s loop. Two hooks matter.
A nightly job over the overlap set. It reads yesterday’s completed overlap tiles, writes a per-class agreement table and a dispute queue, and posts the classes whose kappa fell below the floor. Because it consumes only completed work, it never blocks anybody.
def summarise_batch(tiles: list[str], floors: dict[str, float]) -> pd.DataFrame:
rows = []
for tile_id in tiles:
a, b = load_pair(Path(f"a/{tile_id}.geojson"), Path(f"b/{tile_id}.geojson"), "EPSG:25832")
pairs, miss_a, miss_b = match_features(a, b)
agree = label_agreement(a, b, pairs)
rows.append({
"tile_id": tile_id,
"pairs": len(pairs),
"unmatched_a": len(miss_a),
"unmatched_b": len(miss_b),
"kappa": agree["kappa"],
"raw": agree["raw_agreement"],
"disputes": len(build_queue(tile_id, a, b, pairs, floors)),
})
return pd.DataFrame(rows)
A gate on the batch, not the feature. Individual disagreements are normal and should never block a merge. A batch whose kappa for a class has fallen below its floor is a different matter, and belongs in the same CI/CD gate that already checks geometry and schema, as a warning that names the class pair rather than an error that blocks the author.
Agreement scores also make a good confidence signal. A feature drawn on a tile whose class agreement is poor deserves a lower weight in training than one from a class two people place identically — which is what confidence scoring for geospatial labels consumes.
Validation & Testing
The metrics need their own tests, because a quality measure that silently returns optimistic numbers is worse than none.
from shapely.geometry import box
def test_identical_geometries_score_one() -> None:
g = box(0, 0, 10, 10)
assert boundary_iou(g, g) == 1.0
def test_disjoint_geometries_score_zero() -> None:
assert boundary_iou(box(0, 0, 10, 10), box(100, 100, 110, 110)) == 0.0
def test_boundary_iou_punishes_what_plain_iou_forgives() -> None:
true_ = box(0, 0, 30, 30) # 900 m²
loose = box(-3, -3, 33, 33) # 3 m out on every wall
inter = true_.intersection(loose).area
plain = inter / (true_.area + loose.area - inter)
assert plain > 0.75 # plain IoU is forgiving here
assert boundary_iou(true_, loose) < plain / 1.5
def test_kappa_collapses_on_a_dominant_class() -> None:
# Both "annotators" call everything building; raw agreement is perfect, skill is nil.
ya = ["building"] * 92 + ["road"] * 8
yb = ["building"] * 100
assert cohen_kappa_score(ya, yb, labels=["building", "road"]) <= 0.0
The last test is the one that earns its place. It feeds the machinery a case that looks perfect by raw agreement and asserts that kappa refuses it, which is the whole reason kappa is in the pipeline.
Frequently Asked Questions
# Can I measure agreement without a deliberate overlap set?
Only weakly. Proxies exist — comparing each annotator against a model’s predictions, or against the eventual adjudicated truth — but both measure agreement with something that is itself uncertain, and both are biased toward annotators whose style matches the model’s. The overlap set costs five to ten percent of the labelling budget and is the only measurement that compares two humans on identical ground.
# Which threshold should trigger re-training an annotator versus fixing the guide?
Look at whether the disagreement is concentrated. If one annotator disagrees with everyone else on many classes, that is a training conversation. If everyone disagrees with everyone on one class pair, the guide has not distinguished those classes and no amount of annotator training will fix it. The per-class confusion matrix from Step 4 separates the two cases directly.
# How does Hausdorff distance fit alongside boundary IoU?
Boundary IoU is an average-like measure over the whole outline, so a single badly placed corner can hide inside a good score. Hausdorff distance reports the single worst gap between the two boundaries, in metres. Running both means a polygon that is right everywhere except one corner is still caught — see debugging annotation drift across dataset versions for the same metric used across versions rather than across annotators.
# Should the model’s own predictions be in the agreement calculation?
Keep them separate. Model-versus-human agreement is a useful monitoring signal, but mixing it into inter-annotator agreement makes the number move when the model is retrained, which means it no longer measures the annotation team. Track them as two series over the same overlap set.
Related
- Confidence Scoring for Geospatial Labels — how the agreement scores computed here become a per-label confidence that drives QA routing and loss weighting
- Defining ROI Label Taxonomies for Aerial Imagery — the taxonomy work that a persistently low per-class kappa is telling you to do
- Human-in-the-Loop Validation Cycles — the review tiers that consume the dispute queue this topic produces
- Best Practices for Polygon vs Bounding Box Annotation — why the geometry primitive you chose sets a ceiling on the agreement you can measure
Quality measurement is one component of the broader Geospatial Annotation Fundamentals & Architecture pipeline, which covers the CRS contracts, taxonomies and formats these metrics assume.