Validate Attribute Values Against Rules in PyQGIS
Geometry gets most of the attention in spatial quality control, but attributes are where most errors live. A pipe with a diameter of 0, a tree planted in 2087, a road classified as "Primay", an inspection date before the installation date, a zoning code that exists in no lookup table — none of these show up on a map, and all of them break the analysis that relies on them. Checking attributes is straightforward once the rules are written down; the trick is to write them in a form that is easy to read, easy to extend, and produces a list of violations someone can act on.
This recipe belongs to Data Quality & Topology Validation. It expresses rules as QGIS expressions, evaluates them efficiently against every feature, checks references against lookup tables, and writes the violations to a table with the rule, the feature and the offending value.
Prerequisites
- QGIS 3.34 LTR or newer, or the QGIS 4 series.
- The rules your data should satisfy, ideally agreed with the people who maintain it. Rules nobody agreed to produce reports nobody acts on.
- Familiarity with evaluating QGIS expressions in PyQGIS.
Write rules as expressions
Expressing each rule as a QGIS expression that is true when a feature passes has three advantages: the same syntax as the field calculator and filters, so non-programmers can read and propose rules; NULL handling built in; and access to functions for dates, strings, regular expressions and geometry.
RULES = [
("R01", "diameter must be positive", "error",
'"diameter_mm" > 0'),
("R02", "material from the allowed list", "error",
'"material" IN (\'PE\', \'PVC\', \'DI\', \'CI\', \'ST\')'),
("R03", "installation year not in the future", "error",
'"install_year" <= year(now())'),
("R04", "installation year plausible", "warning",
'"install_year" >= 1850'),
("R05", "asset id format WP-000000", "error",
'regexp_match("asset_id", \'^WP-[0-9]{6}$\')'),
("R06", "inspection after installation", "error",
'"last_inspected" IS NULL OR year("last_inspected") >= "install_year"'),
("R07", "mains need a pressure zone", "warning",
'"pipe_type" <> \'main\' OR "pressure_zone" IS NOT NULL'),
]
Breakdown: Each rule carries an id for reports, a sentence a person can understand, a severity, and the expression. Allowed-value lists use IN, ranges use comparisons, formats use regexp_match, and cross-field logic like R06 and R07 reads naturally as "either the condition does not apply, or it holds". NULL needs care: "diameter_mm" > 0 is NULL — not true — when the diameter is missing, so a missing value fails R01. That is usually what you want for a required field; for optional ones, write "x" IS NULL OR … explicitly, as R06 does.
Evaluate the rules efficiently
Each expression is parsed and prepared once, then evaluated for every feature with a shared context. Preparing against the layer's fields lets QGIS resolve field references to indexes in advance.
from qgis.core import (QgsProject, QgsExpression, QgsExpressionContext,
QgsExpressionContextUtils)
pipes = QgsProject.instance().mapLayersByName("water_pipes")[0]
context = QgsExpressionContext()
context.appendScopes(QgsExpressionContextUtils.globalProjectLayerScopes(pipes))
compiled = []
for rid, text, severity, expr in RULES:
e = QgsExpression(expr)
if e.hasParserError():
raise ValueError(f"{rid}: {e.parserErrorString()}")
e.prepare(context)
compiled.append((rid, text, severity, e, e.referencedColumns()))
violations = []
for f in pipes.getFeatures():
context.setFeature(f)
for rid, text, severity, e, cols in compiled:
ok = e.evaluate(context)
if e.hasEvalError():
violations.append((rid, f["asset_id"], f"eval error: {e.evalErrorString()}", "error"))
elif not ok:
shown = ", ".join(f"{c}={f[c]}" for c in sorted(cols) if c in f.fields().names())
violations.append((rid, f["asset_id"], shown, severity))
print(len(violations), "violations")
Breakdown: Checking hasParserError() up front turns a typo in a rule into an immediate, named failure rather than thousands of evaluation errors. prepare() once per rule and setFeature() once per feature is the efficient pattern; reconstructing expressions inside the loop is the usual cause of slow validators. referencedColumns() tells you which fields each rule reads, so the violation can show exactly the values that failed rather than the whole row. An evaluation error — a function applied to the wrong type — is recorded as a violation in its own right, because it usually means the data contains something the rule's author never expected.
Check references against lookup tables
Many attributes are codes that must exist in another table: a zoning code, a species code, a street id. An expression can do this with attribute(get_feature(…)), but for large tables a Python set is far faster.
species = QgsProject.instance().mapLayersByName("species_codes")[0]
valid_codes = {str(f["code"]).strip().upper() for f in species.getFeatures()}
trees = QgsProject.instance().mapLayersByName("street_trees")[0]
for f in trees.getFeatures():
code = f["species_code"]
if code is None or (hasattr(code, "isNull") and code.isNull()) or not str(code).strip():
violations.append(("R10", f["tree_id"], "species_code missing", "warning"))
elif str(code).strip().upper() not in valid_codes:
violations.append(("R10", f["tree_id"], f"species_code={code!r} not in lookup", "error"))
Breakdown: Normalising both sides — stripping whitespace and upper-casing — prevents a trailing space from producing hundreds of false violations, which is the fastest way to make people ignore a report. Separating missing from unknown matters because they have different causes: missing usually means incomplete capture, unknown usually means a typo or an outdated code list. Where the lookup is a related table in the project, defining layer relations and a value-relation widget prevent these errors at entry.
Write violations to a table
A list in the console is fine for a first look; a table is what a team works through. An attribute-only memory layer turns the list into something that can be filtered, sorted, joined back to the source layer and saved.
from qgis.core import QgsVectorLayer, QgsFeature
table = QgsVectorLayer(
"None?field=rule_id:string(8)&field=rule:string(120)&field=asset_id:string(20)"
"&field=detail:string(250)&field=severity:string(10)", "violations", "memory")
rule_text = {rid: text for rid, text, *_ in RULES}
rule_text["R10"] = "species code exists in lookup"
rows = []
for rid, key, detail, severity in violations:
f = QgsFeature(table.fields())
f.setAttributes([rid, rule_text.get(rid, ""), str(key), detail[:250], severity])
rows.append(f)
table.dataProvider().addFeatures(rows)
QgsProject.instance().addMapLayer(table)
Breakdown: Keeping the rule text in each row makes the table self-explanatory when it is exported to a spreadsheet for someone who never sees the script. Truncating details protects against unexpectedly long values breaking a fixed-length field. To see violations on the map, join the table to the source layer on asset_id — the attribute join recipe shows how — or select the offending features by key and zoom to them.
Summarise by rule and severity
The summary is what goes into an email or a dashboard; the table is what people fix from.
from collections import Counter
by_rule = Counter((rid, sev) for rid, _, _, sev in violations)
features_failing = len({key for _, key, _, sev in violations if sev == "error"})
print(f"{features_failing} of {pipes.featureCount()} features fail at least one error rule")
for (rid, sev), n in sorted(by_rule.items(), key=lambda kv: -kv[1]):
print(f"{rid:<4} {sev:<8} {n:>6} {rule_text.get(rid, '')}")
Breakdown: The number of features failing at least one error rule is the headline that best describes overall quality; total violations overstates it when a single bad feature fails several rules. Sorting rules by count shows where effort pays off: a format rule failing on four hundred features usually has one systematic cause — an import that dropped the prefix — that can be fixed with a single update rather than four hundred edits.
Keep rules outside the script
Rules change more often than code. Storing them in a CSV or a table in the project's GeoPackage lets data stewards maintain them without touching Python.
import csv
def load_rules(path):
with open(path, newline="", encoding="utf-8") as fh:
return [(r["id"], r["description"], r["severity"], r["expression"])
for r in csv.DictReader(fh) if r.get("active", "1") == "1"]
RULES = load_rules("/data/qa/rules_water_pipes.csv")
Breakdown: A rules file with an active column lets a rule be switched off temporarily without deleting it, which is useful while a known problem is being fixed. Because the expressions are plain text, a rules file can be reviewed in a pull request, versioned with the data model, and reused across layers that share fields. The data quality report recipe reads rules this way and combines them with geometry and topology checks.
QGIS version compatibility
Expression evaluation and regexp_match behave the same on QGIS 3.x and QGIS 4. QgsExpressionContextUtils.globalProjectLayerScopes exists since 3.0. On QGIS 4, NULL attribute values arrive as None, which the checks above already handle.
Troubleshooting
- Every feature fails a rule. The field name is misspelled; an unknown field evaluates to NULL. Check
referencedColumns()againstlayer.fields().names(). - Date rules raise evaluation errors. The field is text, not a date; wrap it in
to_date()in the rule. - The validator is slow. Expressions are being created inside the loop, or a rule uses
get_featureagainst a large table; use a Python set for lookups. - Hundreds of false lookup violations. Codes differ in case or whitespace; normalise both sides.
Conclusion
Write each rule as a QGIS expression that is true when a feature passes, prepare them once and evaluate with a shared context, check codes against lookup sets, collect violations into a table with rule, key, value and severity, summarise by features failing and by rule, and keep the rules in a file data stewards can edit.
Frequently Asked Questions
Can I validate while people edit instead of afterwards? Yes. Field constraints with expressions enforce rules in the attribute form; scripted validation then becomes a safety net for imports and bulk edits.
How do I validate geometry and attributes together?
Use geometry functions inside rules, for example $area > 10 or is_valid($geometry).
Should warnings block a data release? Usually not. Errors block; warnings are reviewed and either fixed or accepted with a note.
Can rules reference other layers?
Yes, through aggregate() or get_feature() in the expression, at a performance cost on large layers.