Calculate a Field with an Expression in PyQGIS

The field calculator is one of the most used tools in QGIS, and the expression language behind it is far more capable than the dialog suggests: conditional logic, string formatting, date arithmetic, geometry measurements, aggregates over other layers. In scripts, the same expressions fill fields in bulk — population density from population and area, a category from a numeric range, a label from three attributes — consistently and reproducibly.

This recipe belongs to Attribute Tables & Field Management. It calculates fields into a new layer with Processing, updates an existing layer in place, writes conditional and NULL-safe expressions, uses geometry and aggregate functions, and creates virtual fields that never go out of date.

Three ways to calculate a fieldThree routes for the same expression. native:fieldcalculator writes a new layer with the computed field, leaving the source untouched. An in-place update evaluates the expression per feature and writes values into an existing field through the edit buffer. A virtual field stores only the expression and computes values on the fly whenever they are read.New layer, in place, or virtualfieldcalculatornew layersource untouchedscripts, modelsin-place updateexisting fieldedit bufferundoablevirtual fieldexpression storedcomputed on readnever stale

Prerequisites

  • QGIS 3.34 LTR or newer, or the QGIS 4 series.
  • A test of each expression in the QGIS expression builder before scripting it — the builder shows a preview and explains errors, which is faster than debugging in a loop. See evaluating expressions in PyQGIS for the API.

Calculate into a new layer

native:fieldcalculator evaluates an expression for every feature and writes the result into a new or existing field of an output layer.

import processing
from qgis.core import QgsProject

districts = QgsProject.instance().mapLayersByName("districts")[0]
result = processing.run("native:fieldcalculator", {
    "INPUT": districts,
    "FIELD_NAME": "density_km2",
    "FIELD_TYPE": 0,              # 0 = decimal, 1 = integer, 2 = text, 3 = date ...
    "FIELD_LENGTH": 12,
    "FIELD_PRECISION": 1,
    "FORMULA": '"population" / ($area / 1e6)',
    "OUTPUT": "memory:districts_density",
})["OUTPUT"]

for f in list(result.getFeatures())[:3]:
    print(f["name"], f["density_km2"])

Breakdown: The output is a copy of the input with the new field added; the source layer is not modified, which keeps scripts safe to re-run. $area measures each polygon using the project's ellipsoid and area unit settings; dividing by a million converts square metres to square kilometres. In a standalone script without a project, set the ellipsoid explicitly or use area($geometry) for planar area in layer units. If FIELD_NAME matches an existing field, the algorithm overwrites its values in the output.

Update an existing layer in place

When the result belongs in the source layer — filling a column that already exists, or adding one to a GeoPackage table — evaluate the expression per feature and write through the edit buffer.

from qgis.core import (QgsExpression, QgsExpressionContext, QgsExpressionContextUtils,
                       QgsField, edit)
from qgis.PyQt.QtCore import QVariant

if districts.fields().indexOf("density_km2") < 0:
    districts.dataProvider().addAttributes([QgsField("density_km2", QVariant.Double)])
    districts.updateFields()
idx = districts.fields().indexOf("density_km2")

expr = QgsExpression('"population" / ($area / 1e6)')
ctx = QgsExpressionContext(QgsExpressionContextUtils.globalProjectLayerScopes(districts))
expr.prepare(ctx)

with edit(districts):
    for f in districts.getFeatures():
        ctx.setFeature(f)
        value = expr.evaluate(ctx)
        if expr.hasEvalError():
            raise ValueError(f"{f.id()}: {expr.evalErrorString()}")
        districts.changeAttributeValue(f.id(), idx, value)

Breakdown: The expression context supplies variables, the layer's fields and — through the project scope — the ellipsoid that $area uses. Preparing once and setting the feature per iteration is the efficient pattern. Stopping on the first evaluation error with the feature id is better than writing NULLs silently. Writing inside edit() makes the whole update one transaction that rolls back on error; for a person at the desk, wrapping it in an edit command makes it one undo step, as described in undoing edits with the edit buffer.

Write conditional logic with CASE

Classification — turning a number into a category — is the most common conditional calculation. CASE reads top to bottom and returns the first matching branch.

CASE evaluates top to bottomA CASE expression classifies density. The first condition, density below 100, returns rural. The second, below 1000, returns suburban; it does not need a lower bound because rural values have already matched. The ELSE branch returns urban. A NULL density matches no WHEN branch and falls to ELSE unless a WHEN IS NULL branch comes first.First matching branch winsCASE WHEN "density" IS NULL THEN 'unknown'WHEN "density" < 100 THEN 'rural'WHEN "density" < 1000 THEN 'suburban'ELSE 'urban' END

classified = processing.run("native:fieldcalculator", {
    "INPUT": result, "FIELD_NAME": "settlement", "FIELD_TYPE": 2, "FIELD_LENGTH": 10,
    "FORMULA": """CASE
        WHEN "density_km2" IS NULL THEN 'unknown'
        WHEN "density_km2" < 100 THEN 'rural'
        WHEN "density_km2" < 1000 THEN 'suburban'
        ELSE 'urban'
    END""",
    "OUTPUT": "memory:districts_classified"})["OUTPUT"]

Breakdown: Because branches are tested in order, each condition only needs an upper bound — values below 100 never reach the second test. Testing for NULL first is important: a NULL density is neither below 100 nor anything else, so without that branch it falls through to ELSE and is silently labelled "urban". Multi-line expressions are fine; whitespace inside the formula is ignored.

Keep arithmetic NULL-safe

Any arithmetic involving NULL produces NULL. One missing value in a sum wipes out the result for that feature, which is correct in some analyses and wrong in many.

formula = 'coalesce("pop_0_14", 0) + coalesce("pop_15_64", 0) + coalesce("pop_65_plus", 0)'
total = processing.run("native:fieldcalculator", {
    "INPUT": districts, "FIELD_NAME": "pop_check", "FIELD_TYPE": 1,
    "FORMULA": formula, "OUTPUT": "memory:"})["OUTPUT"]

ratio = processing.run("native:fieldcalculator", {
    "INPUT": total, "FIELD_NAME": "old_share", "FIELD_TYPE": 0, "FIELD_PRECISION": 3,
    "FORMULA": '"pop_65_plus" / nullif("pop_check", 0)', "OUTPUT": "memory:"})["OUTPUT"]

Breakdown: coalesce substitutes zero for missing values in a sum — appropriate when a missing count genuinely means zero, misleading when it means "not collected". nullif turns a zero denominator into NULL, so the ratio is NULL instead of an error for empty districts. Choose deliberately per field; the NULL handling recipe explains the semantics.

Build text and dates

Many calculated fields are text for people: a label combining a code and a name, a formatted measurement, a date written the local way. The expression language has functions for all of these, and doing the formatting once in a field keeps labels, reports and exports consistent.

labels = processing.run("native:fieldcalculator", {
    "INPUT": classified, "FIELD_NAME": "label", "FIELD_TYPE": 2, "FIELD_LENGTH": 80,
    "FORMULA": """ "code" || ' – ' || title("name") || ' (' ||
                   format_number("density_km2", 0) || ' /km²)' """,
    "OUTPUT": "memory:"})["OUTPUT"]

ages = processing.run("native:fieldcalculator", {
    "INPUT": QgsProject.instance().mapLayersByName("inspections")[0],
    "FIELD_NAME": "days_since", "FIELD_TYPE": 1,
    "FORMULA": "day(age(now(), to_date(\"inspected_on\")))",
    "OUTPUT": "memory:"})["OUTPUT"]

Breakdown: || concatenates strings, title() capitalises names consistently, and format_number adds thousands separators and fixes decimals. If any part of a concatenation is NULL the whole result is NULL; wrap optional parts in coalesce("field", ''). For dates, age() returns an interval between two dates and day() converts it to days; to_date parses text dates in ISO format, so it also repairs dates stored as strings. Computing "days since inspection" as a stored field is a snapshot; as a virtual field, it stays current every day.

Use geometry and aggregate functions

Expressions can measure geometry and summarise other layers, which turns many multi-step workflows into a single calculation.

schools = QgsProject.instance().mapLayersByName("schools")[0]
enriched = processing.run("native:fieldcalculator", {
    "INPUT": districts, "FIELD_NAME": "schools_n", "FIELD_TYPE": 1,
    "FORMULA": f"""aggregate(layer:='{schools.id()}', aggregate:='count',
                   expression:="school_id",
                   filter:=intersects($geometry, geometry(@parent)))""",
    "OUTPUT": "memory:"})["OUTPUT"]

perimeter = processing.run("native:fieldcalculator", {
    "INPUT": enriched, "FIELD_NAME": "compactness", "FIELD_TYPE": 0, "FIELD_PRECISION": 3,
    "FORMULA": "4 * pi() * $area / ($perimeter ^ 2)", "OUTPUT": "memory:"})["OUTPUT"]

Breakdown: aggregate() evaluates an expression over another layer, with filter restricting it to features related to the current one — here schools intersecting the district. Inside the filter, $geometry refers to the school and geometry(@parent) to the district. Using the layer id rather than its name avoids ambiguity when names repeat. Aggregates are convenient but slow on large layers, because the filter runs for every feature; for big joins, a spatial join or counting points in polygons is faster. The compactness index uses $perimeter, the ellipsoidal counterpart of $area.

Create a virtual field

A virtual field stores only the expression; its values are computed whenever they are read. It never goes stale when the inputs change, which makes it ideal for derived values in layers people edit.

Stored versus virtualA stored field holds values written once; if population changes later, density is out of date until recalculated. A virtual field holds the expression; reading the attribute table, labelling or styling evaluates it, so density always reflects the current population. Virtual fields live in the project, not in the data file.Values written once, or computed every timestored fieldwritten oncegoes stale on editsin the data filevirtual fieldexpression storedalways currentin the project

from qgis.core import QgsField
from qgis.PyQt.QtCore import QVariant

field = QgsField("density_live", QVariant.Double)
districts.addExpressionField('"population" / ($area / 1e6)', field)
print(districts.fields().names()[-1], districts.fields().fieldOrigin(districts.fields().count() - 1))

Breakdown: addExpressionField adds a field whose origin is "expression"; it appears in the attribute table, can be used in labels, styles and filters, and recalculates automatically after edits. The trade-off is that it exists only in the project — exporting the layer writes its current values as ordinary fields, but the source file itself does not contain it. Use virtual fields for display and interactive work; write stored fields for data that leaves QGIS.

QGIS version compatibility

native:fieldcalculator replaced qgis:fieldcalculator in QGIS 3.18 and is available with the parameters shown on 3.34 LTR, 3.40 LTR and QGIS 4. Field type codes in the algorithm are 0 decimal, 1 integer, 2 text, 3 date, 4 time, 5 datetime, 6 boolean, 7 binary on current releases. On QGIS 4, QgsField takes QMetaType.Type values.

Troubleshooting

  • All results are NULL. A field name is misspelled or a field contains NULLs in arithmetic; check names and use coalesce.
  • Areas are wildly wrong. No ellipsoid in a standalone script; set one or use area($geometry) in a projected CRS.
  • Text values are truncated. FIELD_LENGTH is too short for the output.
  • Aggregates are very slow. Replace with a spatial join or a joined summary layer.

Conclusion

Calculate into new layers with native:fieldcalculator, update in place through the edit buffer with a prepared expression and context, write CASE with NULL checked first, use coalesce and nullif deliberately, reach for geometry and aggregate functions where they save steps, and use virtual fields for values that must track edits.

Frequently Asked Questions

Can I use Python functions in expressions? Yes — register a custom function, as in registering a custom expression function.

Is the field calculator faster than a Python loop? For simple expressions, Processing is typically faster and needs less code; Python loops win when logic is complex.

How do I calculate only selected features? Pass QgsProcessingFeatureSourceDefinition(layer.id(), selectedFeaturesOnly=True) as input, or filter in the in-place loop.

Can expressions reference the previous feature? Not directly. Order-dependent calculations such as running totals need a Python loop.