Data Quality & Topology Validation in PyQGIS
Every analysis inherits the defects of its input. A dissolve over parcels with overlaps double-counts area; a route over a network with undershoots stops at a gap nobody can see; a join on a key with duplicates multiplies rows; a statistic over a field with impossible values reports nonsense with two decimal places. None of these errors announces itself. The map looks fine, the script runs without complaint, and the numbers are wrong.
This guide belongs to Spatial Data Processing & Automation. It is for anyone who receives data from someone else — a supplier, a field team, an open data portal, a colleague — and needs to know whether it can be used. It covers the kinds of defect that matter, how to find each one with PyQGIS, how to decide what to fix automatically, and how to turn the checks into a report that runs the same way on every delivery.
What this guide covers
- Is every shape well-formed? Build a geometry validity report with the GEOS or QGIS rule sets and decide whether automatic repair is safe.
- Does a coverage cover? Find gaps and overlaps between polygons and tell slivers from real defects.
- Is a network connected? Check line network connectivity for dangles, undershoots, unnoded crossings and isolated pieces.
- Is each object there once? Find duplicate geometries and features, exact and near, and keep the right copy.
- Are the values right? Validate attribute values against rules written as QGIS expressions.
- Can layers be made to agree? Snap geometries to a layer with a deliberate tolerance and behaviour, and audit what moved.
- Is this delivery good enough? Build a data quality report that combines all of the above into one GeoPackage, one summary and one verdict.
Geometry validity: the foundation
A geometry is valid when it follows the OGC simple-features rules: polygon rings are closed and do not cross themselves, holes lie inside their shell, multipolygon parts do not overlap, lines have at least two distinct points. Invalid geometries can usually be drawn — renderers are forgiving — but the GEOS library behind buffers, intersections, unions and spatial predicates assumes validity and gives wrong or no answers when it is missing.
Checking validity comes first because every later check depends on it: an overlap test on an invalid polygon is meaningless. The qgis:checkvalidity algorithm splits a layer into valid features, invalid features and an error point per problem; grouping those points by reason turns thousands of points into a short list of causes.
import processing
r = processing.run("qgis:checkvalidity", {
"INPUT_LAYER": "/data/delivery/parcels.gpkg", "METHOD": 2,
"VALID_OUTPUT": "memory:", "INVALID_OUTPUT": "memory:", "ERROR_OUTPUT": "memory:"})
share = r["INVALID_COUNT"] / max(r["VALID_COUNT"] + r["INVALID_COUNT"], 1)
print(f"{r['INVALID_COUNT']} invalid features ({share:.2%}), {r['ERROR_COUNT']} problems")
Breakdown: METHOD: 2 applies the GEOS rules — the ones downstream operations and PostGIS use — so a pass here means GEOS-based analysis will behave. The share of invalid features is the number to report: a fraction of a percent is a repair job, a large share means the producing process is broken. The decision between repairing and rejecting is laid out in the validity report recipe, and the repair itself in fixing invalid geometries.
Topology: how shapes relate to their neighbours
Valid shapes can still be wrong together. Topology rules describe how features in a layer — or across layers — should relate, and two kinds of data have rules that matter most.
Coverages are polygon layers where every location belongs to exactly one polygon: parcels, administrative units, land-use zones. Their rules are no overlaps and no gaps. Overlaps double-count area in every total and make point-in-polygon joins ambiguous; gaps leave locations that belong to nothing. Overlaps are found pairwise with a spatial index; enclosed gaps are the holes in the dissolved coverage; gaps reaching the edge are found by subtracting the coverage from the area it should fill.
Networks are line layers that something flows along: roads, rivers, pipes, cables. Their rule is connectivity — lines that meet in reality must share an end point or a vertex in the data. An undershoot of a few centimetres is invisible on screen and splits the network for every routing algorithm.
The same tool fixes most of these errors: snapping. Slivers between neighbouring polygons close when the layer is snapped to itself; undershoots close when line end points are snapped to the nearest line. The tolerance must come from the data's accuracy and the measured size of the defects — large enough to close the gaps found, small enough not to merge real features. Every snap should be audited by measuring how far each feature moved.
Uniqueness: one object, one feature
Duplicates arise from repeated imports, overlapping survey areas, merged extracts and copy-paste during editing. They inflate counts and distort statistics, and when they differ slightly in their attributes they create genuine ambiguity about which record is correct.
Three tests catch nearly all of them. Exact geometry duplicates hash identically. Near duplicates — the same object captured twice — sit within a small distance and are grouped with a spatial index and union-find. Key duplicates share an identifier that should be unique, regardless of location. The hard part is not finding duplicates but choosing which copy survives; a ranking rule — most complete record, most recent edit, lowest id as a tie-breaker — makes the choice consistent, and archiving the removed copies first makes it reversible.
from collections import Counter
from qgis.core import QgsFeatureRequest, QgsProject
layer = QgsProject.instance().mapLayersByName("hydrants")[0]
req = (QgsFeatureRequest().setFlags(QgsFeatureRequest.NoGeometry)
.setSubsetOfAttributes(["hydrant_no"], layer.fields()))
keys = Counter(f["hydrant_no"] for f in layer.getFeatures(req))
print(sum(n - 1 for n in keys.values() if n > 1), "extra features sharing a hydrant number")
Breakdown: Key checks are the cheapest and most reliable duplicate test because they need no geometry; reading a single column without geometry is fast even on large remote layers. The duplicate recipe combines key and distance tests, because a repeated key far apart is usually a mislabelled feature rather than a duplicate.
Attributes: present, plausible, consistent
Attribute errors outnumber geometry errors in most datasets and are invisible on a map. The rules fall into a few patterns: required fields must be present; values must come from an allowed list or a lookup table; numbers must lie in a plausible range; text must match a format; and related fields must agree with each other — an inspection after the installation, a pressure zone on every main.
Writing each rule as a QGIS expression that is true when a feature passes keeps the rules readable by the people who know the data, reuses the syntax of the field calculator and form constraints, and handles NULL in a predictable way. Validating attribute values against rules evaluates a list of such rules efficiently, checks codes against lookup tables with Python sets, and writes every violation to a table with the rule, the feature and the failing value.
A first look at an unfamiliar dataset
Before writing rules or thresholds, spend five minutes getting the measure of a new dataset. A short profiling script answers the questions that decide which checks matter: what geometry type and CRS it has, how many features, how many have no geometry, how many are invalid, and which fields are mostly empty.
from qgis.core import QgsVectorLayer, QgsAggregateCalculator, QgsWkbTypes
def profile(path):
layer = QgsVectorLayer(path, "profile", "ogr")
if not layer.isValid():
return print("cannot open", path)
n = layer.featureCount()
no_geom = sum(1 for f in layer.getFeatures() if not f.hasGeometry())
print(f"{QgsWkbTypes.displayString(layer.wkbType())} · {layer.crs().authid()} · "
f"{n} features · {no_geom} without geometry")
for field in layer.fields():
missing, _ = layer.aggregate(QgsAggregateCalculator.CountMissing, field.name())
distinct, _ = layer.aggregate(QgsAggregateCalculator.CountDistinct, field.name())
print(f" {field.name():<20} {field.typeName():<10} "
f"missing {missing / max(n, 1):>6.1%} distinct {distinct}")
profile("/data/delivery/parcels.gpkg")
Breakdown: The geometry type and CRS tell you immediately whether tolerances will be in metres or degrees and whether coverage rules even apply. The share of missing values per field shows which attributes are reliably captured and which are optional in practice — a "required" field that is 40 % empty needs a conversation before it needs a rule. The distinct count spots fields that should be unique (distinct equals feature count) and fields that should be codes (a handful of distinct values); both suggest rules to write. Run the profile before the first full report, and again whenever a supplier changes their process.
Tolerances: where the numbers come from
Almost every check in this guide takes a tolerance: the area below which an overlap is noise, the distance within which two points are the same object, the gap a snap may close. Choosing them arbitrarily produces either floods of false findings or silent acceptance of real errors.
Three sources give defensible numbers. The first is the stated accuracy of the data: parcels surveyed to 10 cm cannot meaningfully disagree by less than that, so overlaps narrower than 10 cm along a shared edge are noise. The second is the measured distribution of the defects: record the gap distance on every dangle, or the width of every sliver, and the histogram usually shows a cluster of tiny values separated from the real errors by a clear break. The third is the scale of the objects: a tolerance must stay well below the smallest real feature, or snapping starts merging neighbours and duplicate detection starts merging genuinely separate assets.
Write the chosen tolerances down next to the rules and thresholds, with the reason. When a check later produces surprising results, the first question is always whether the tolerance still fits the data, and a recorded reason makes that question quick to answer.
Preventing errors at the source
Checks downstream are a safety net; the cheapest defect is the one never captured. Several QGIS features stop errors at the moment of editing, and they can all be configured from Python as part of setting up a project for data capture:
- Snapping and topological editing in the project's snapping configuration make new vertices land exactly on existing ones and keep shared boundaries shared when one side is edited.
- Field constraints — not null, unique, expression — reject invalid values in the attribute form, as shown in setting default values and field constraints.
- Value relation and value map widgets restrict codes to a lookup table, removing the typo category of error entirely; see configuring editor widgets.
- Database constraints — primary keys, foreign keys, check constraints in PostGIS or GeoPackage triggers — enforce the same rules for every client, not just QGIS.
Preventive measures and scripted checks complement each other: prevention covers interactive editing, while the report covers imports, bulk updates and data that arrives from outside.
Fix, or send back?
Finding defects is half the job; deciding what to do with them is the other half. Three questions decide it.
Is the fix mechanical and safe? Duplicate vertices, slivers below the data's accuracy and undershoots of a few centimetres can be fixed automatically without changing what the data means — and should be, with an audit of what moved.
Does the fix require knowledge you do not have? A real overlap between two parcels needs to know which boundary is right; a code missing from the lookup needs to know what was meant. These go to whoever maintains the data, with the findings attached.
Is the defect systematic? Four hundred features failing the same format rule usually share one cause — an import that dropped a prefix — and one fix at the source is worth more than four hundred corrections downstream. A high share of invalid features or a network in dozens of pieces means the producing process needs attention, not the data.
Keeping checking and fixing separate makes this possible: the report always describes the data as received, and fixes are applied deliberately afterwards, each with its own audit.
Making it repeatable
Ad-hoc checks find problems once. A report that runs the same checks with the same thresholds on every delivery shows whether quality is improving, gives suppliers an objective acceptance test, and catches regressions the moment they appear.
The data quality report recipe gives each check the same shape — a name, a layer of findings and a count — so new checks plug in without touching the runner. Findings go into one dated GeoPackage where they can be opened on a map; counts are compared with thresholds agreed with the data producer; the verdict names the failing checks; and a script with meaningful exit codes lets the whole thing run from a scheduler or a CI pipeline. The patterns in batch processing with PyQGIS extend it to folders of deliveries.
A typical acceptance workflow
Put together, the recipes form a workflow that fits most data deliveries, whether they come monthly from a contractor or once from an open data portal:
- Profile the delivery to learn its shape, CRS, completeness and likely keys.
- Agree rules and thresholds with the producer, informed by the profile — or, for open data, decide what you can live with.
- Run the report: validity, topology for coverages or networks, duplicates, attribute rules.
- Read the verdict. A pass goes straight into use. A fail splits into mechanical defects you fix with an audited script, and knowledge-dependent or systematic defects you send back with the report GeoPackage attached.
- Re-run after fixes or a corrected delivery, and append the verdict to the log so trends become visible.
The first pass through this takes an afternoon; every later delivery takes the time the script needs to run. That asymmetry is the whole argument for automating quality control rather than eyeballing each delivery.
Key takeaways
- Check validity first; topology, overlay and join results are only meaningful on valid geometry.
- Coverages must have no overlaps and no gaps; networks must meet at shared vertices and form one component.
- Tie every tolerance to the data's accuracy and to the measured size of the defects.
- Find duplicates by key, by exact geometry and by distance, and choose survivors with an explicit ranking rule.
- Write attribute rules as QGIS expressions that are true when a feature passes, and keep them in a file.
- Fix mechanical defects automatically with an audit; send knowledge-dependent and systematic ones back.
- Make it a report: same checks, same thresholds, one GeoPackage and one verdict per delivery.
Frequently Asked Questions
Does QGIS have a topology checker I can call from Python? The Topology Checker core plugin works interactively but has no stable Python API. The Processing algorithms and scripts in this guide cover the same rules in a scriptable way, and QGIS 3.36 and later add a group of geometry checker algorithms to the toolbox.
Should I validate data in PostGIS instead?
When data already lives in PostGIS, ST_IsValid, ST_Overlaps and pgRouting's graph analysis run without moving data. PyQGIS suits files, mixed sources and deliveries that have not reached a database yet.
How strict should thresholds be? As strict as the use of the data requires. A cadastre needs zero overlaps; a volunteered points-of-interest layer can tolerate some duplicates. Agree thresholds with the data producer and write them down.
Can these checks run while people edit? Partly. Field constraints, unique constraints and digitizing snapping prevent many errors at entry. The scripted checks then act as a safety net for imports, bulk edits and external deliveries.
What about raster data quality? Rasters have their own checks — NoData coverage, value ranges, alignment and resolution — covered in the raster analysis guide through statistics and alignment recipes.
Related
- Spatial Data Processing & Automation — the section this guide belongs to
- Geometry Operations & Predicates — validity repair, spatial indexes and predicates used by these checks
- Network Analysis & Routing — what a connected network makes possible
- Attribute Tables & Field Management — fixing the values these checks flag
- Batch Processing with PyQGIS — running checks across many files
- Build a Geometry Validity Report in PyQGIS
- Find Gaps and Overlaps Between Polygons in PyQGIS
- Check Line Network Connectivity in PyQGIS
- Find Duplicate Geometries and Features in PyQGIS
- Validate Attribute Values Against Rules in PyQGIS
- Snap Geometries to a Layer in PyQGIS
- Build a Data Quality Report in PyQGIS