Build a Data Quality Report in PyQGIS
Individual checks find individual problems. A quality report answers the question people actually ask: is this delivery good enough to use? It runs a fixed set of checks against a dataset, stores every finding where it can be inspected on a map, summarises the result in a table anyone can read, and does all of this identically each time — so the report on March's delivery can be compared with February's, and a supplier can be shown exactly what failed.
This recipe belongs to Data Quality & Topology Validation. It assembles the checks from the other recipes in that guide into one script with a common shape, writes all findings and a summary into a single GeoPackage, renders a short HTML overview, and sets thresholds that turn the result into a pass or a fail.
Prerequisites
- QGIS 3.34 LTR or newer, or the QGIS 4 series. The report runs inside QGIS or as a standalone script; for unattended runs see handling errors and logging in unattended scripts.
- The individual checks you need. This page reuses the approach of the validity report, duplicate search and attribute rules.
Give every check the same shape
The runner is simple when every check is a function that takes the layer and returns the same thing: a name, a layer of findings (possibly empty), and a count. Each check can be as elaborate as it needs to be inside.
from dataclasses import dataclass
import processing
from qgis.core import QgsVectorLayer, QgsFeature, QgsProject
@dataclass
class CheckResult:
name: str
description: str
findings: QgsVectorLayer # memory layer, may be empty
count: int
def check_validity(layer):
r = processing.run("qgis:checkvalidity", {
"INPUT_LAYER": layer, "METHOD": 2,
"VALID_OUTPUT": "memory:", "INVALID_OUTPUT": "memory:",
"ERROR_OUTPUT": "memory:"})
return CheckResult("invalid_geometry", "GEOS-invalid geometries",
r["ERROR_OUTPUT"], r["INVALID_COUNT"])
def check_exact_duplicates(layer):
r = processing.run("native:deleteduplicategeometries", {
"INPUT": layer, "OUTPUT": "memory:", "DUPLICATES": "memory:"})
return CheckResult("duplicate_geometry", "features with identical geometry",
r["DUPLICATES"], r["DUPLICATE_COUNT"])
Breakdown: The dataclass is a contract: the runner never needs to know what a check does, only how to read its result. Processing algorithms already return findings layers and counts, so wrapping them is a few lines; native:deleteduplicategeometries has a DUPLICATES output with the removed copies and a DUPLICATE_COUNT. Custom checks — gaps, network dangles, attribute rules — build their own memory layer of findings as in the other recipes and return it in the same shape. Keep each check focused on one question, so the summary reads as a list of clear yes/no answers.
Expression-based checks — anything that can be phrased as "features matching this filter are wrong" — share one small helper:
from qgis.core import QgsFeatureRequest
def _request(expression):
return QgsFeatureRequest().setFilterExpression(expression)
def check_missing_key(layer, field="parcel_id"):
found = layer.materialize(_request(f'"{field}" IS NULL OR trim("{field}") = \'\''))
return CheckResult("missing_key", f"{field} empty", found, found.featureCount())
Breakdown: materialize returns a memory copy of the matching features, ready to be written to the report with their geometry — so a reviewer can see on the map where the keyless parcels are. Expression-based checks are the cheapest kind to add: one line of expression, one line of wrapper.
Run the checks and collect findings
The runner calls each check, writes non-empty findings to the report GeoPackage as their own table, and keeps the counts for the summary.
from datetime import datetime
from pathlib import Path
def run_report(layer, checks, out_dir="/data/qa"):
stamp = datetime.now().strftime("%Y%m%d_%H%M")
report = str(Path(out_dir) / f"{layer.name()}_quality_{stamp}.gpkg")
results, first = [], True
for check in checks:
res = check(layer)
results.append(res)
if res.count and res.findings.featureCount():
processing.run("native:savefeatures", {
"INPUT": res.findings, "OUTPUT": report, "LAYER_NAME": res.name,
"ACTION_ON_EXISTING_FILE": 0 if first else 1})
first = False
print(f"{res.name:<22} {res.count:>7}")
return report, results
parcels = QgsProject.instance().mapLayersByName("parcels")[0]
report, results = run_report(parcels, [check_validity, check_exact_duplicates,
check_missing_key])
Breakdown: The first write creates the GeoPackage (ACTION_ON_EXISTING_FILE: 0) and later ones add tables to it. Checks with no findings write nothing, keeping the file small, but still appear in the summary with a count of zero — a passed check is a result worth recording. The timestamp in the file name keeps successive reports side by side. Running each check in sequence keeps memory use bounded; for very large layers, checks that do not depend on each other can run as background tasks in parallel.
Turn counts into a verdict
A count is not a judgement until it is compared with a threshold. Thresholds encode the agreement with whoever produces the data: zero invalid geometries, fewer than ten duplicates, no missing keys.
THRESHOLDS = {"invalid_geometry": 0, "duplicate_geometry": 10, "missing_key": 0}
summary = QgsVectorLayer(
"None?field=check:string(40)&field=description:string(120)"
"&field=count:integer&field=threshold:integer&field=status:string(4)",
"summary", "memory")
rows, failed = [], []
for res in results:
limit = THRESHOLDS.get(res.name, 0)
status = "PASS" if res.count <= limit else "FAIL"
if status == "FAIL":
failed.append(res.name)
f = QgsFeature(summary.fields())
f.setAttributes([res.name, res.description, res.count, limit, status])
rows.append(f)
summary.dataProvider().addFeatures(rows)
processing.run("native:savefeatures", {"INPUT": summary, "OUTPUT": report,
"LAYER_NAME": "summary",
"ACTION_ON_EXISTING_FILE": 1})
verdict = "PASS" if not failed else "FAIL: " + ", ".join(failed)
print(verdict)
Breakdown: Keeping thresholds in a dictionary separate from the checks lets the same checks serve different agreements — a strict threshold for the cadastre, a lenient one for a volunteered dataset. A check missing from the dictionary defaults to zero, the strictest reading, so forgetting a threshold never lets problems through. The verdict names the failing checks, which is what the reader of a one-line notification needs.
Render a short HTML overview
Not everyone opens GeoPackages. A small HTML page with the verdict and the summary table can be attached to an email or published next to the data.
import html
def write_html(path, layer_name, results, thresholds, verdict, report_path):
rows = "".join(
f"<tr class='{'fail' if r.count > thresholds.get(r.name, 0) else 'pass'}'>"
f"<td>{html.escape(r.name)}</td><td>{html.escape(r.description)}</td>"
f"<td>{r.count}</td><td>{thresholds.get(r.name, 0)}</td></tr>"
for r in results)
page = f"""<!doctype html><meta charset="utf-8"><title>{html.escape(layer_name)} quality</title>
<style>body{{font-family:sans-serif;max-width:46rem;margin:2rem auto}}
.fail td{{color:#b91c1c;font-weight:bold}}td,th{{padding:.3rem .8rem;text-align:left}}</style>
<h1>{html.escape(layer_name)} — quality report</h1>
<p><strong>{html.escape(verdict)}</strong> · {datetime.now():%Y-%m-%d %H:%M}</p>
<table><tr><th>check</th><th>description</th><th>count</th><th>threshold</th></tr>{rows}</table>
<p>Detailed findings: {html.escape(report_path)}</p>"""
Path(path).write_text(page, encoding="utf-8")
write_html(report.replace(".gpkg", ".html"), parcels.name(), results,
THRESHOLDS, verdict, report)
Breakdown: Escaping every value keeps a stray < in a description from breaking the page. The HTML is deliberately plain — it needs to survive email clients and intranet pages. For a richer report with maps of the findings, a QGIS print layout can be generated from the same GeoPackage, as in generating a report with QgsReport.
Run it against every delivery
The value of a quality report grows with repetition. Wrapping the runner in a script that takes a path, exits with a non-zero status on failure and appends the verdict to a log makes it schedulable and usable in CI.
import sys
def main(path):
layer = QgsVectorLayer(path, Path(path).stem, "ogr")
if not layer.isValid():
print(f"cannot open {path}")
return 2
report, results = run_report(layer, [check_validity, check_exact_duplicates,
check_missing_key])
failed = [r.name for r in results if r.count > THRESHOLDS.get(r.name, 0)]
with open("/data/qa/quality_log.csv", "a", encoding="utf-8") as log:
log.write(f"{datetime.now():%Y-%m-%d %H:%M},{path},{'FAIL' if failed else 'PASS'},"
f"{'|'.join(failed)}\n")
return 1 if failed else 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1]))
Breakdown: Exit codes are how schedulers and CI systems understand success: 0 for pass, 1 for a failed check, 2 for a delivery that could not even be opened. The CSV log grows one line per run, which is enough to chart quality over time. Run it from cron or a task scheduler as in scheduling PyQGIS scripts, or from qgis_process if the checks are wrapped as a Processing script.
QGIS version compatibility
All algorithms used — qgis:checkvalidity, native:deleteduplicategeometries, native:savefeatures — are present on QGIS 3.34 LTR, 3.40 LTR and QGIS 4 with the same parameters. The DUPLICATES output of native:deleteduplicategeometries was added in the 3.x series; on older releases compare feature counts before and after instead. Python dataclasses need Python 3.7 or newer, which every supported QGIS ships.
Troubleshooting
- The report GeoPackage has only the last table.
ACTION_ON_EXISTING_FILEstayed at 0 for every write. - A check takes far longer than the others. It iterates without a request filter or spatial index; profile it with the techniques in profiling slow PyQGIS code.
- Counts differ between runs on the same data. A check depends on a user setting, such as
METHOD: 0in the validity check. - The HTML shows garbled characters. The file was opened without UTF-8; keep the
meta charset.
Conclusion
Give every check the same shape — name, findings layer, count — run them in sequence into one dated GeoPackage, compare counts with agreed thresholds for a named verdict, render a plain HTML summary, and wrap it all in a script with meaningful exit codes so every delivery gets the same report.
Frequently Asked Questions
Can I add a check written as a Processing model?
Yes. Run the model with processing.run("model:…") inside a check function and return its output layer and feature count.
Should the report modify the data? No. Keep checking and fixing separate, so the report always describes the delivery as received.
How do I report on several layers at once?
Call run_report per layer and combine the summaries into one table with a layer column.
Where should thresholds live? In a small configuration file versioned with the data specification, so changes to the agreement are visible.