Add an Auto-Increment ID Field in PyQGIS

Every dataset that people refer to needs an identifier they can quote: "manhole MH-00412", "tree 17 in street 3", "parcel 2026-0098". QGIS's internal feature ids look like a ready-made answer, but they are the wrong one — shapefile ids change when the file is rewritten, memory layer ids restart, and none of them survive export to another format. A dedicated identifier field, numbered deliberately and protected from duplication, is what links a map feature to a work order, a photo or a row in someone else's spreadsheet.

This recipe belongs to Attribute Tables & Field Management. It numbers existing features with Processing, controls order and grouping, formats codes with prefixes and padding, sets a default so new features get the next number, uses database sequences where they exist, and protects uniqueness.

Feature ids versus business idsA feature id is assigned by the data provider: in a shapefile it is the row number and changes when the file is rewritten; in a memory layer it restarts per session; in GeoPackage it is the fid primary key, stable within that file only. A business id is a field you control, with a format people can quote, that stays with the feature through exports, edits and copies.Ids people can quote must be yoursfeature id ($id)shapefile: row numbermemory: restartsGeoPackage: this file onlychanges on exportbusiness id fieldyour format: MH-00412survives export, copyprotected by constraintquotable, stable

Prerequisites

  • QGIS 3.34 LTR or newer, or the QGIS 4 series.
  • A decision on the identifier's format and scope: one sequence for the whole layer, or numbering within groups such as streets or districts.

Number existing features with Processing

native:addautoincrementalfield adds an integer field counting up from a start value, optionally in a chosen order and restarting per group.

import processing
from qgis.core import QgsProject

manholes = QgsProject.instance().mapLayersByName("manholes")[0]
numbered = processing.run("native:addautoincrementalfield", {
    "INPUT": manholes,
    "FIELD_NAME": "mh_no",
    "START": 1,
    "MODULUS": 0,
    "GROUP_FIELDS": [],
    "SORT_EXPRESSION": "x($geometry)",
    "SORT_ASCENDING": True,
    "SORT_NULLS_FIRST": False,
    "OUTPUT": "memory:manholes_numbered",
})["OUTPUT"]

for f in list(numbered.getFeatures())[:3]:
    print(f["mh_no"])

Breakdown: Without a sort expression, numbers follow the provider's iteration order, which is arbitrary — run it twice on a rewritten file and features may get different numbers. Sorting by an expression makes numbering deterministic and meaningful: here west to east, so neighbouring numbers are neighbouring manholes. Any expression works, such as "install_date" for chronological numbering or y($geometry) * -1 for north to south. MODULUS makes numbers wrap — rarely needed for identifiers.

Number within groups

Many identifiers are scoped: trees numbered within each street, rooms within each building, plots within each block. Group fields restart the count for each distinct group.

Numbering restarts per groupTrees in two streets. With GROUP_FIELDS set to street, numbering restarts at 1 for each street and follows the sort expression within it, so each tree's id is the pair of street and number, such as Elm Street 1, 2, 3 and Oak Lane 1, 2. A composite code joins them into one quotable string.Restart the count for every streetElm StreetELM-001 · ELM-002ELM-003 · ELM-004west → eastOak LaneOAK-001 · OAK-002OAK-003west → east

trees = QgsProject.instance().mapLayersByName("street_trees")[0]
by_street = processing.run("native:addautoincrementalfield", {
    "INPUT": trees, "FIELD_NAME": "tree_no", "START": 1,
    "GROUP_FIELDS": ["street_code"],
    "SORT_EXPRESSION": "line_locate_point(geometry(get_feature('streets', 'code', \"street_code\")), $geometry)",
    "SORT_ASCENDING": True, "SORT_NULLS_FIRST": False,
    "OUTPUT": "memory:"})["OUTPUT"]

Breakdown: With GROUP_FIELDS, each street's trees are numbered from 1 independently. The sort expression orders trees by their position along their street's centreline — line_locate_point returns the distance along the line closest to the tree — so numbers run in walking order. That is the kind of numbering field crews appreciate. The expression looks up the street geometry with get_feature, which is fine for thousands of trees; for very large layers, precompute the chainage with points along lines and linear referencing instead.

Write numbers back into the source layer

native:addautoincrementalfield produces a new layer. When the numbers belong in the existing table — so its relations, styles and forms stay intact — compute them in the same order and write them through the edit buffer.

from qgis.core import QgsField, QgsFeatureRequest, edit
from qgis.PyQt.QtCore import QVariant

if manholes.fields().indexOf("mh_no") < 0:
    manholes.dataProvider().addAttributes([QgsField("mh_no", QVariant.Int)])
    manholes.updateFields()
idx = manholes.fields().indexOf("mh_no")

ordered = QgsFeatureRequest().addOrderBy("x($geometry)", ascending=True)
with edit(manholes):
    manholes.beginEditCommand("Number manholes west to east")
    for n, f in enumerate(manholes.getFeatures(ordered), start=1):
        manholes.changeAttributeValue(f.id(), idx, n)
    manholes.endEditCommand()

Breakdown: Ordering the feature request by the same expression used in the Processing call gives identical numbering, but written into the source table instead of a copy. Wrapping the loop in an edit command makes the whole numbering one undo step. Only run this once per layer — renumbering features that people already reference breaks those references; for later additions rely on the default value described below. For large PostGIS tables, a single SQL UPDATE with row_number() OVER (ORDER BY …) is much faster than a loop.

Format codes with prefixes and padding

Integers are the stored value; people quote formatted codes. Padding keeps codes the same length and sortable as text, and a prefix says what kind of object it is.

coded = processing.run("native:fieldcalculator", {
    "INPUT": by_street, "FIELD_NAME": "tree_code", "FIELD_TYPE": 2, "FIELD_LENGTH": 16,
    "FORMULA": "upper(\"street_code\") || '-' || lpad(to_string(\"tree_no\"), 3, '0')",
    "OUTPUT": "memory:trees_coded"})["OUTPUT"]
print([f["tree_code"] for f in list(coded.getFeatures())[:4]])

Breakdown: lpad pads the number to three digits with zeros, so "ELM-007" sorts before "ELM-012" as text. Choose a width with headroom — if a street may one day have a thousand trees, use four digits now, because changing code lengths later breaks every printed label and spreadsheet. Keep the integer field too: it is what the next number is computed from.

Number new features automatically

Numbering existing features is a one-off; new features digitized later need numbers too. A default value expression on the field assigns the next number whenever a feature is created.

A default value for new featuresThe field's default value expression computes the maximum existing number plus one. When a user digitizes a feature or a script creates one through createFeature, the default is evaluated and the new feature receives the next number. A unique constraint on the field rejects duplicates if two editors create features at the same time.Next number, every timenew featuredigitized or scripteddefault valuemaximum("mh_no") + 1unique constraintduplicates rejected

from qgis.core import QgsDefaultValue, QgsFieldConstraints

idx = manholes.fields().indexOf("mh_no")
manholes.setDefaultValueDefinition(idx, QgsDefaultValue('coalesce(maximum("mh_no"), 0) + 1'))
manholes.setFieldConstraint(idx, QgsFieldConstraints.ConstraintUnique,
                            QgsFieldConstraints.ConstraintStrengthHard)
manholes.setFieldConstraint(idx, QgsFieldConstraints.ConstraintNotNull,
                            QgsFieldConstraints.ConstraintStrengthHard)

Breakdown: maximum("mh_no") is an aggregate over the whole layer, so the default is always one more than the highest existing number; coalesce handles an empty layer. The unique and not-null constraints, set to hard strength, make the attribute form refuse to save a duplicate or empty id. These settings are part of the layer's style and project; save them, as described in setting default values and field constraints. Two people editing the same file simultaneously could both take the same "next" number — the constraint then rejects the second save, which is the right outcome.

Use database sequences where available

In PostgreSQL, a sequence-backed column is the robust answer for multi-user editing: the database hands out numbers atomically, so concurrent editors never collide.

from qgis.core import QgsProviderRegistry

md = QgsProviderRegistry.instance().providerMetadata("postgres")
conn = md.findConnection("assets_db")
conn.executeSql("""
    CREATE SEQUENCE IF NOT EXISTS assets.manhole_no_seq START 1000;
    ALTER TABLE assets.manholes
        ALTER COLUMN mh_no SET DEFAULT nextval('assets.manhole_no_seq');
    ALTER TABLE assets.manholes ADD CONSTRAINT manholes_mh_no_key UNIQUE (mh_no);
""")

Breakdown: A column default of nextval(...) is evaluated by the database on insert, and QGIS shows it as the provider default in forms, so new features get their number on save. Starting the sequence above the highest existing number avoids collisions with features numbered earlier. Sequences can leave gaps — a rolled-back insert consumes a number — which is harmless for identifiers and should not be "fixed". The database connections API recipe covers running SQL through stored connections.

Check uniqueness in existing data

Before relying on an identifier, verify it: no duplicates, no NULLs, no gaps you did not expect.

from collections import Counter

values = [f["mh_no"] for f in manholes.getFeatures()]
counts = Counter(values)
dupes = {v: n for v, n in counts.items() if n > 1}
nulls = sum(1 for v in values if v is None or (hasattr(v, "isNull") and v.isNull()))
print(f"{len(values)} features, {len(dupes)} duplicated ids, {nulls} without id")

Breakdown: Duplicates usually come from copy-pasting features in the editor, which copies the id along with everything else, or from merging two deliveries numbered independently. The duplicate detection recipe covers deciding which copy keeps the number. Run this check after every bulk edit or import.

QGIS version compatibility

native:addautoincrementalfield with group fields and sort expression works on QGIS 3.34 LTR, 3.40 LTR and QGIS 4. Default value definitions and field constraints are stable across these releases; on QGIS 4, constraint enums are QgsFieldConstraints.Constraint.ConstraintUnique and QgsFieldConstraints.ConstraintStrength.ConstraintStrengthHard.

Troubleshooting

  • Numbers differ between runs. No sort expression was given; numbering followed arbitrary provider order.
  • New features get no number. The default value is set on the wrong field index, or the feature was added through the provider, bypassing defaults; use QgsVectorLayerUtils.createFeature.
  • Duplicates after copy-paste. The id was copied; enforce uniqueness with a hard constraint.
  • Numbers jump after a database rollback. Normal sequence behaviour; leave the gap.

Conclusion

Never rely on feature ids for identifiers people quote; add a dedicated field with native:addautoincrementalfield, sorted deterministically and grouped where scope demands, format it into padded codes, give new features the next number through a default value or a database sequence, and protect it with unique and not-null constraints.

Frequently Asked Questions

Can I use UUIDs instead of numbers? Yes — uuid() as a default value gives globally unique ids that never collide, at the cost of being unquotable by people. Many systems store both.

Should identifiers encode meaning, like the district? Only if that meaning never changes; a tree moved to a renamed street should not need a new id.

How do I renumber after deletions? Usually you should not — gaps are harmless and renumbering breaks every external reference.

Can I number by distance along a route? Yes, with a sort expression using line_locate_point, as shown for trees.