Find Duplicate Geometries and Features in PyQGIS
Duplicates creep into spatial data in predictable ways. An import script runs twice. Two survey teams capture the same hydrant from either side of the road. A merge of monthly extracts repeats every feature that did not change. A polygon is copied and pasted onto itself during editing. Each duplicate inflates counts, double-weights statistics and confuses joins — and the dangerous ones are not the exact copies, which are easy to spot, but the near-duplicates a few centimetres apart with slightly different attributes.
This recipe belongs to Data Quality & Topology Validation. It finds exact geometric duplicates, near-duplicate points within a distance, repeated identifiers, and features duplicated in attributes but not in location, then removes the extras with a rule that keeps the right copy.
Prerequisites
- QGIS 3.34 LTR or newer, or the QGIS 4 series.
- A layer in a projected CRS, so distance tolerances are in metres.
- A definition of "the same object" for your data — a hydrant within 50 cm, a parcel with the same cadastral number — before you start deleting anything.
Find exact geometric duplicates
Exact duplicates have byte-identical geometry. Hashing the WKB representation groups them in a single pass, without any spatial comparison.
import hashlib
from collections import defaultdict
from qgis.core import QgsProject, QgsFeatureRequest
hydrants = QgsProject.instance().mapLayersByName("hydrants")[0]
groups = defaultdict(list)
for f in hydrants.getFeatures():
if not f.hasGeometry():
continue
key = hashlib.sha1(bytes(f.geometry().asWkb())).hexdigest()
groups[key].append(f.id())
dupes = {k: ids for k, ids in groups.items() if len(ids) > 1}
extra = sum(len(ids) - 1 for ids in dupes.values())
print(f"{len(dupes)} geometries occur more than once; {extra} extra features")
Breakdown: Two geometries with the same WKB are identical in type, vertex order and every coordinate. Hashing keeps memory low on large layers — only a short digest per feature is stored. This misses polygons that are the same shape but start at a different vertex, and copies shifted by floating-point noise. For those, normalise first: g.normalize() (QGIS 3.20+) puts rings and vertices into a canonical order, after which equal shapes hash equally. Processing's native:deleteduplicategeometries does the exact-match case in one call and also reports how many it removed.
Find near-duplicate points within a tolerance
Points captured twice are rarely identical. A spatial index finds, for each point, any other point within the tolerance; grouping those pairs gives clusters of probable duplicates.
from qgis.core import QgsSpatialIndex
TOL = 0.5 # metres
index = QgsSpatialIndex(hydrants.getFeatures(),
flags=QgsSpatialIndex.FlagStoreFeatureGeometries)
points = {f.id(): f.geometry() for f in hydrants.getFeatures()}
parent = {fid: fid for fid in points}
def find(x):
while parent[x] != x:
parent[x] = parent[parent[x]]
x = parent[x]
return x
for fid, geom in points.items():
box = geom.boundingBox().buffered(TOL)
for other in index.intersects(box):
if other > fid and geom.distance(points[other]) <= TOL:
parent[find(other)] = find(fid)
clusters = defaultdict(list)
for fid in points:
clusters[find(fid)].append(fid)
near = [ids for ids in clusters.values() if len(ids) > 1]
print(len(near), "groups of near-duplicate hydrants")
Breakdown: Buffering the bounding box by the tolerance turns the index's rectangle query into a "within distance" candidate search; the exact distance test then confirms each pair. Union-find merges pairs into groups, so three points in a chain become one group even if the ends are further apart than the tolerance — usually the right behaviour for repeated captures of one object. For large layers, QgsSpatialIndex.nearestNeighbor(point, k, maxDistance) is an alternative that returns only the closest candidates.
Find repeated identifiers
A business key that should be unique — a hydrant number, a parcel ID, an address UPRN — is the most reliable duplicate test, because it does not depend on geometry at all.
from collections import Counter
counts = Counter(f["hydrant_no"] for f in hydrants.getFeatures(
QgsFeatureRequest().setFlags(QgsFeatureRequest.NoGeometry)
.setSubsetOfAttributes(["hydrant_no"], hydrants.fields())))
repeated = {k: n for k, n in counts.items() if n > 1 and k not in (None, "")}
print(len(repeated), "hydrant numbers used more than once")
for key, n in sorted(repeated.items(), key=lambda kv: -kv[1])[:10]:
print(f" {key}: {n} features")
Breakdown: Reading only the key column without geometry makes this fast even on large layers or remote databases. NULL and empty keys are excluded from the duplicate list but deserve their own count — a missing key is a separate defect. When a repeated key belongs to features in very different places, the problem is usually a data-entry error rather than a duplicate capture, and deleting one copy would lose a real feature. Combine the key check with the distance check: same key and close together is a duplicate; same key and far apart is a mislabelled feature.
Decide which copy to keep
Removing duplicates is easy; removing the right ones is the work. A keep-rule makes the choice consistent and explainable.
def null_count(f):
return sum(1 for v in f.attributes() if v is None or (hasattr(v, "isNull") and v.isNull()))
to_delete = []
for ids in near:
feats = list(hydrants.getFeatures(QgsFeatureRequest().setFilterFids(ids)))
feats.sort(key=lambda f: (null_count(f), -f["edited_on"].toMSecsSinceEpoch()
if f["edited_on"] else 0, f.id()))
keep, *rest = feats
to_delete.extend(r.id() for r in rest)
removed = hydrants.materialize(QgsFeatureRequest().setFilterFids(to_delete))
removed.setName("hydrants removed as duplicates")
QgsProject.instance().addMapLayer(removed)
print(len(to_delete), "features marked for removal")
Breakdown: Sorting by a tuple applies the rules in order: fewest NULLs first, then the most recent edit (negated so later dates sort first), then the lowest id. The id tie-breaker makes the result deterministic — run it twice and the same features survive. Materialising the losers into a separate layer before deletion means nothing is lost if the rule turns out to be wrong for some group; review that layer, then delete. edited_on here is a datetime field; adapt the rule to whatever your data records about quality and age.
Delete and verify
Deletion goes through the edit buffer so it can be undone, and a second run of the check confirms the result.
from qgis.core import edit
before = hydrants.featureCount()
with edit(hydrants):
hydrants.deleteFeatures(to_delete)
print(before, "→", hydrants.featureCount(), "features")
Breakdown: Using edit() means a failure rolls back cleanly, and an interactive user can still press undo if they ran the script from the console and immediately see a problem. Rerunning the near-duplicate search should now return no groups at the same tolerance; if it returns new ones, the union-find merged chains that span several captures and the tolerance may be too generous. For tables where duplicates differ only by attribute history — the monthly extract case — native:removeduplicatesbyattribute (3.20+) removes them by chosen fields in one Processing call.
Attribute twins at different places
The last case is features with identical attributes at different locations: two "Oak, girth 120 cm" trees in different streets. These are often legitimate, but a burst of them in one batch can reveal a copy-paste error during digitizing.
trees = QgsProject.instance().mapLayersByName("street_trees")[0]
fields_to_compare = ["species", "girth_cm", "planted"]
sig = defaultdict(list)
for f in trees.getFeatures(QgsFeatureRequest().setFlags(QgsFeatureRequest.NoGeometry)):
sig[tuple(str(f[n]) for n in fields_to_compare)].append(f.id())
twins = [ids for ids in sig.values() if len(ids) > 3]
print(len(twins), "attribute signatures shared by more than three features")
Breakdown: Grouping by an attribute signature finds repeated records regardless of where they are. A threshold above two filters out the coincidences that any large dataset contains. These groups are for review, not deletion: look at their creation dates and editors to see whether they came from one session.
QGIS version compatibility
QgsGeometry.normalize() needs QGIS 3.20 or newer; native:deleteduplicategeometries has existed since 3.0 and native:removeduplicatesbyattribute since 3.20. The rest of the code runs on any QGIS 3.x and QGIS 4 release. On QGIS 4, NULL datetime values arrive as None, which the if f["edited_on"] test already handles.
Troubleshooting
- Exact duplicates are not detected. Coordinates differ in the last decimal place; snap to a grid with
native:snappointstogridbefore hashing, or use the tolerance approach. - One huge group of near duplicates. The tolerance is larger than the spacing between real objects; reduce it.
- The wrong copy was kept. The keep rule did not reflect your data; adjust the sort key and restore from the review layer.
- Deleting is slow on PostGIS. Pass ids in one
deleteFeaturescall, not in a loop.
Conclusion
Hash WKB for exact duplicates, use a spatial index with a tolerance and union-find for near duplicates, count business keys without reading geometry, choose survivors with an explicit ranking rule, archive the losers before deleting, and treat attribute twins as a review list rather than a delete list.
Frequently Asked Questions
Is there a single Processing algorithm for all of this?
No. native:deleteduplicategeometries handles exact matches; near duplicates and keep rules need a script.
How do I find duplicate polygons that overlap almost completely? Compare intersection area to each polygon's area; a ratio above 0.95 on both sides is a near duplicate. The overlap search in finding gaps and overlaps gives you the intersections.
Should I prevent duplicates instead of cleaning them? Yes, where possible: a unique constraint on the key field rejects duplicates at entry, as covered in default values and constraints.
Does deleting duplicates break relations? It can. Check child tables for references to deleted ids and repoint them to the kept feature first.