Find Duplicate Geometries and Features in PyQGIS

Duplicates creep into spatial data in predictable ways. An import script runs twice. Two survey teams capture the same hydrant from either side of the road. A merge of monthly extracts repeats every feature that did not change. A polygon is copied and pasted onto itself during editing. Each duplicate inflates counts, double-weights statistics and confuses joins — and the dangerous ones are not the exact copies, which are easy to spot, but the near-duplicates a few centimetres apart with slightly different attributes.

This recipe belongs to Data Quality & Topology Validation. It finds exact geometric duplicates, near-duplicate points within a distance, repeated identifiers, and features duplicated in attributes but not in location, then removes the extras with a rule that keeps the right copy.

Four kinds of duplicateExact geometry duplicates share identical coordinates, found by comparing WKB. Near duplicates are the same object captured twice within a small distance, found with a spatial index and a tolerance. Key duplicates share an identifier that should be unique, found by counting values. Attribute duplicates have the same attributes but different locations, which may be legitimate or may be a copy error.Not all duplicates look the sameexact geometryidenticalcoordinatescompare WKBnear duplicatesame object,a few cm apartindex + distancekey duplicatesame ID twiceany locationcount valuesattribute twinsame values,other placemaybe legitimate

Prerequisites

  • QGIS 3.34 LTR or newer, or the QGIS 4 series.
  • A layer in a projected CRS, so distance tolerances are in metres.
  • A definition of "the same object" for your data — a hydrant within 50 cm, a parcel with the same cadastral number — before you start deleting anything.

Find exact geometric duplicates

Exact duplicates have byte-identical geometry. Hashing the WKB representation groups them in a single pass, without any spatial comparison.

import hashlib
from collections import defaultdict
from qgis.core import QgsProject, QgsFeatureRequest

hydrants = QgsProject.instance().mapLayersByName("hydrants")[0]

groups = defaultdict(list)
for f in hydrants.getFeatures():
    if not f.hasGeometry():
        continue
    key = hashlib.sha1(bytes(f.geometry().asWkb())).hexdigest()
    groups[key].append(f.id())

dupes = {k: ids for k, ids in groups.items() if len(ids) > 1}
extra = sum(len(ids) - 1 for ids in dupes.values())
print(f"{len(dupes)} geometries occur more than once; {extra} extra features")

Breakdown: Two geometries with the same WKB are identical in type, vertex order and every coordinate. Hashing keeps memory low on large layers — only a short digest per feature is stored. This misses polygons that are the same shape but start at a different vertex, and copies shifted by floating-point noise. For those, normalise first: g.normalize() (QGIS 3.20+) puts rings and vertices into a canonical order, after which equal shapes hash equally. Processing's native:deleteduplicategeometries does the exact-match case in one call and also reports how many it removed.

Find near-duplicate points within a tolerance

Points captured twice are rarely identical. A spatial index finds, for each point, any other point within the tolerance; grouping those pairs gives clusters of probable duplicates.

Grouping points within a toleranceFive points. Two of them lie 0.3 metres apart and form a group; a third lies 0.4 metres from the second and joins the same group even though it is 0.7 metres from the first. Two other points are several metres away and stay single. Groups are built with a union-find over all pairs closer than the 0.5 metre tolerance.Pairs within 0.5 m chain into groupsone group of 3singles: several metres apartunion-find over pairs closer than the tolerance

from qgis.core import QgsSpatialIndex

TOL = 0.5  # metres
index = QgsSpatialIndex(hydrants.getFeatures(),
                        flags=QgsSpatialIndex.FlagStoreFeatureGeometries)
points = {f.id(): f.geometry() for f in hydrants.getFeatures()}

parent = {fid: fid for fid in points}
def find(x):
    while parent[x] != x:
        parent[x] = parent[parent[x]]
        x = parent[x]
    return x

for fid, geom in points.items():
    box = geom.boundingBox().buffered(TOL)
    for other in index.intersects(box):
        if other > fid and geom.distance(points[other]) <= TOL:
            parent[find(other)] = find(fid)

clusters = defaultdict(list)
for fid in points:
    clusters[find(fid)].append(fid)
near = [ids for ids in clusters.values() if len(ids) > 1]
print(len(near), "groups of near-duplicate hydrants")

Breakdown: Buffering the bounding box by the tolerance turns the index's rectangle query into a "within distance" candidate search; the exact distance test then confirms each pair. Union-find merges pairs into groups, so three points in a chain become one group even if the ends are further apart than the tolerance — usually the right behaviour for repeated captures of one object. For large layers, QgsSpatialIndex.nearestNeighbor(point, k, maxDistance) is an alternative that returns only the closest candidates.

Find repeated identifiers

A business key that should be unique — a hydrant number, a parcel ID, an address UPRN — is the most reliable duplicate test, because it does not depend on geometry at all.

from collections import Counter

counts = Counter(f["hydrant_no"] for f in hydrants.getFeatures(
    QgsFeatureRequest().setFlags(QgsFeatureRequest.NoGeometry)
                       .setSubsetOfAttributes(["hydrant_no"], hydrants.fields())))
repeated = {k: n for k, n in counts.items() if n > 1 and k not in (None, "")}
print(len(repeated), "hydrant numbers used more than once")
for key, n in sorted(repeated.items(), key=lambda kv: -kv[1])[:10]:
    print(f"  {key}: {n} features")

Breakdown: Reading only the key column without geometry makes this fast even on large layers or remote databases. NULL and empty keys are excluded from the duplicate list but deserve their own count — a missing key is a separate defect. When a repeated key belongs to features in very different places, the problem is usually a data-entry error rather than a duplicate capture, and deleting one copy would lose a real feature. Combine the key check with the distance check: same key and close together is a duplicate; same key and far apart is a mislabelled feature.

Decide which copy to keep

Removing duplicates is easy; removing the right ones is the work. A keep-rule makes the choice consistent and explainable.

A keep rule for each groupWithin each group of duplicates the features are ranked by rules in order: first the most complete record with the fewest NULL attributes, then the most recent edit date, then the lowest feature id as a stable tie-breaker. The top-ranked feature is kept and the others are written to a removed layer for review before deletion.Rank, keep the first, archive the rest1. completenessfewest NULLattributes2. recencylatest editdate3. tie-breaklowest feature idstable across runskeep top · write others to a review layer

def null_count(f):
    return sum(1 for v in f.attributes() if v is None or (hasattr(v, "isNull") and v.isNull()))

to_delete = []
for ids in near:
    feats = list(hydrants.getFeatures(QgsFeatureRequest().setFilterFids(ids)))
    feats.sort(key=lambda f: (null_count(f), -f["edited_on"].toMSecsSinceEpoch()
                              if f["edited_on"] else 0, f.id()))
    keep, *rest = feats
    to_delete.extend(r.id() for r in rest)

removed = hydrants.materialize(QgsFeatureRequest().setFilterFids(to_delete))
removed.setName("hydrants removed as duplicates")
QgsProject.instance().addMapLayer(removed)
print(len(to_delete), "features marked for removal")

Breakdown: Sorting by a tuple applies the rules in order: fewest NULLs first, then the most recent edit (negated so later dates sort first), then the lowest id. The id tie-breaker makes the result deterministic — run it twice and the same features survive. Materialising the losers into a separate layer before deletion means nothing is lost if the rule turns out to be wrong for some group; review that layer, then delete. edited_on here is a datetime field; adapt the rule to whatever your data records about quality and age.

Delete and verify

Deletion goes through the edit buffer so it can be undone, and a second run of the check confirms the result.

from qgis.core import edit

before = hydrants.featureCount()
with edit(hydrants):
    hydrants.deleteFeatures(to_delete)
print(before, "→", hydrants.featureCount(), "features")

Breakdown: Using edit() means a failure rolls back cleanly, and an interactive user can still press undo if they ran the script from the console and immediately see a problem. Rerunning the near-duplicate search should now return no groups at the same tolerance; if it returns new ones, the union-find merged chains that span several captures and the tolerance may be too generous. For tables where duplicates differ only by attribute history — the monthly extract case — native:removeduplicatesbyattribute (3.20+) removes them by chosen fields in one Processing call.

Attribute twins at different places

The last case is features with identical attributes at different locations: two "Oak, girth 120 cm" trees in different streets. These are often legitimate, but a burst of them in one batch can reveal a copy-paste error during digitizing.

trees = QgsProject.instance().mapLayersByName("street_trees")[0]
fields_to_compare = ["species", "girth_cm", "planted"]
sig = defaultdict(list)
for f in trees.getFeatures(QgsFeatureRequest().setFlags(QgsFeatureRequest.NoGeometry)):
    sig[tuple(str(f[n]) for n in fields_to_compare)].append(f.id())
twins = [ids for ids in sig.values() if len(ids) > 3]
print(len(twins), "attribute signatures shared by more than three features")

Breakdown: Grouping by an attribute signature finds repeated records regardless of where they are. A threshold above two filters out the coincidences that any large dataset contains. These groups are for review, not deletion: look at their creation dates and editors to see whether they came from one session.

QGIS version compatibility

QgsGeometry.normalize() needs QGIS 3.20 or newer; native:deleteduplicategeometries has existed since 3.0 and native:removeduplicatesbyattribute since 3.20. The rest of the code runs on any QGIS 3.x and QGIS 4 release. On QGIS 4, NULL datetime values arrive as None, which the if f["edited_on"] test already handles.

Troubleshooting

  • Exact duplicates are not detected. Coordinates differ in the last decimal place; snap to a grid with native:snappointstogrid before hashing, or use the tolerance approach.
  • One huge group of near duplicates. The tolerance is larger than the spacing between real objects; reduce it.
  • The wrong copy was kept. The keep rule did not reflect your data; adjust the sort key and restore from the review layer.
  • Deleting is slow on PostGIS. Pass ids in one deleteFeatures call, not in a loop.

Conclusion

Hash WKB for exact duplicates, use a spatial index with a tolerance and union-find for near duplicates, count business keys without reading geometry, choose survivors with an explicit ranking rule, archive the losers before deleting, and treat attribute twins as a review list rather than a delete list.

Frequently Asked Questions

Is there a single Processing algorithm for all of this? No. native:deleteduplicategeometries handles exact matches; near duplicates and keep rules need a script.

How do I find duplicate polygons that overlap almost completely? Compare intersection area to each polygon's area; a ratio above 0.95 on both sides is a near duplicate. The overlap search in finding gaps and overlaps gives you the intersections.

Should I prevent duplicates instead of cleaning them? Yes, where possible: a unique constraint on the key field rejects duplicates at entry, as covered in default values and constraints.

Does deleting duplicates break relations? It can. Check child tables for references to deleted ids and repoint them to the kept feature first.