Generate Random Points and Spatial Samples in PyQGIS
Randomness has two jobs in spatial work. The first is sampling: choosing field survey sites, plots for a vegetation inventory, households for an interview, or a subset of features for manual quality checks — in a way nobody can accuse of bias. The second is comparison: generating the random pattern that an observed pattern is tested against, so that "these points are clustered" means "more clustered than chance would produce here". QGIS has Processing algorithms for both, and a few lines of Python make them reproducible.
This recipe belongs to Spatial Statistics & Pattern Analysis. It generates random points in polygons, on lines and in an extent, samples existing features at random and by stratum, enforces minimum distances, fixes seeds for reproducibility and uses random points as a baseline for a simple Monte Carlo test.
Prerequisites
- QGIS 3.34 LTR or newer, or the QGIS 4 series.
- Layers in a projected CRS so that minimum distances and densities are in metres.
- A sampling design decided before you start: how many points, where, with what spacing, stratified by what.
Random points inside polygons
The most common request is a fixed number of points inside each polygon — survey plots per forest stand, sample sites per field. native:randompointsinpolygons does this with a count per polygon, a minimum distance and a seed.
import processing
from qgis.core import QgsProject
stands = QgsProject.instance().mapLayersByName("forest_stands")[0]
plots = processing.run("native:randompointsinpolygons", {
"INPUT": stands,
"POINTS_NUMBER": 5,
"MIN_DISTANCE": 50, # metres between plots
"MIN_DISTANCE_GLOBAL": 0,
"MAX_TRIES_PER_POINT": 50,
"SEED": 20261002,
"INCLUDE_POLYGON_ATTRIBUTES": True,
"OUTPUT": "memory:survey_plots",
})
layer = plots["OUTPUT"]
print(layer.featureCount(), "plots;", plots["POINTS_MISSED"], "could not be placed")
QgsProject.instance().addMapLayer(layer)
Breakdown: POINTS_NUMBER accepts a fixed count or, through a data-defined override, an expression — for example ceil($area / 100000) for one plot per ten hectares. MIN_DISTANCE applies within each polygon; MIN_DISTANCE_GLOBAL across all of them. The algorithm tries up to MAX_TRIES_PER_POINT random locations before giving up on a point, and POINTS_MISSED reports how many could not be placed — a non-zero value means the spacing is too tight for small polygons. The seed makes the result repeatable: the same seed, inputs and parameters give the same points.
Random points along lines and in an extent
Transect surveys sample along roads, rivers or hedgerows; pattern tests often need random points in a rectangle. Both have their own algorithms with the same seed and spacing options.
rivers = QgsProject.instance().mapLayersByName("rivers")[0]
river_sites = processing.run("native:randompointsonlines", {
"INPUT": rivers, "POINTS_NUMBER": 3, "MIN_DISTANCE": 200,
"SEED": 7, "INCLUDE_LINE_ATTRIBUTES": True,
"OUTPUT": "memory:river_sites"})["OUTPUT"]
extent_points = processing.run("native:randompointsinextent", {
"EXTENT": stands.extent(), "POINTS_NUMBER": 500, "MIN_DISTANCE": 0,
"TARGET_CRS": stands.crs(), "SEED": 7,
"OUTPUT": "memory:extent_points"})["OUTPUT"]
Breakdown: Points on lines are placed uniformly along length, so a long river gets the same number of points as a short stream when the count is per feature — use a data-defined count based on $length for proportional sampling. Random points in an extent fill a rectangle and ignore where features are; they are mostly useful as a baseline or for synthetic tests. Both accept seeds.
Enforce spacing without losing points
Minimum spacing prevents plots from clumping, which matters for field logistics and for statistical independence. Too tight a spacing in small polygons, however, means points cannot be placed.
from qgis.core import QgsProperty, QgsProcessingFeatureSourceDefinition
params = {
"INPUT": stands,
"POINTS_NUMBER": QgsProperty.fromExpression("max(1, min(5, floor($area / 20000)))"),
"MIN_DISTANCE": 50, "SEED": 20261002, "INCLUDE_POLYGON_ATTRIBUTES": True,
"OUTPUT": "memory:survey_plots_scaled"}
res = processing.run("native:randompointsinpolygons", params)
print(res["POINTS_MISSED"], "missed with area-scaled counts")
Breakdown: Passing a QgsProperty instead of a number makes the count data-defined: here one point per two hectares, between one and five per stand. Scaling counts by area keeps sampling intensity even across stands and almost always eliminates missed points. If some stands are still too small, the honest choice is to report them as unsampled or to merge them with neighbours of the same type, rather than silently dropping the spacing rule.
Sample existing features at random
When the population already exists — addresses to survey, parcels to audit, assets to inspect — the task is to pick a random subset rather than generate locations.
addresses = QgsProject.instance().mapLayersByName("addresses")[0]
sample = processing.run("native:randomextract", {
"INPUT": addresses, "METHOD": 0, # 0 = number of features, 1 = percentage
"NUMBER": 400, "OUTPUT": "memory:address_sample"})["OUTPUT"]
print(sample.featureCount(), "addresses sampled")
Breakdown: native:randomextract returns a new layer with the chosen features and all their attributes, which is what a survey team needs. METHOD: 1 takes a percentage instead. Unlike the point generators, this algorithm has no seed parameter, so a repeatable sample needs Python's own random module, shown below.
Stratified samples
A simple random sample of 400 addresses may contain three from a small district and ninety from a large one. Stratified sampling draws a fixed number — or a fixed share — from each stratum, guaranteeing coverage.
import random
from collections import defaultdict
from qgis.core import QgsFeatureRequest
rng = random.Random(20261002)
by_district = defaultdict(list)
for f in addresses.getFeatures(QgsFeatureRequest().setFlags(QgsFeatureRequest.NoGeometry)):
by_district[f["district"]].append(f.id())
PER_STRATUM = 100
chosen, weights = [], {}
for district, ids in sorted(by_district.items()):
k = min(PER_STRATUM, len(ids))
picked = rng.sample(sorted(ids), k)
chosen.extend(picked)
weights[district] = len(ids) / k # each response represents this many
stratified = addresses.materialize(QgsFeatureRequest().setFilterFids(chosen))
stratified.setName("address sample (stratified)")
QgsProject.instance().addMapLayer(stratified)
print({d: round(w, 1) for d, w in weights.items()})
Breakdown: A dedicated random.Random instance with a fixed seed makes the sample reproducible without affecting any other code that uses the random module. Sorting ids before sampling matters: without it, the result depends on iteration order, which can differ between providers or runs. The design weight per stratum — population divided by sample size — is what you later multiply responses by to estimate totals for the whole area, correcting for the deliberate over-representation of small districts.
Use random points as a baseline
Random points in the same area are the reference against which an observed pattern is judged. A simple Monte Carlo test repeats the random placement many times and asks how often chance produces a value as extreme as the observed one.
from qgis.core import QgsSpatialIndex
def mean_nn(layer):
geoms = {f.id(): f.geometry() for f in layer.getFeatures()}
index = QgsSpatialIndex(layer.getFeatures())
total = 0.0
for fid, g in geoms.items():
other = [n for n in index.nearestNeighbor(g.asPoint(), 2) if n != fid][0]
total += g.distance(geoms[other])
return total / len(geoms)
observed_layer = QgsProject.instance().mapLayersByName("cafes")[0]
city = QgsProject.instance().mapLayersByName("city_boundary")[0]
observed = mean_nn(observed_layer)
n = observed_layer.featureCount()
sims = []
for seed in range(99):
r = processing.run("native:randompointsinpolygons", {
"INPUT": city, "POINTS_NUMBER": n, "SEED": seed + 1,
"OUTPUT": "memory:"})["OUTPUT"]
sims.append(mean_nn(r))
rank = sum(1 for s in sims if s <= observed)
print(f"observed {observed:.0f} m; {rank} of 99 random runs were as clustered (p ≈ {(rank + 1) / 100:.2f})")
Breakdown: Each simulation places the same number of points at random inside the same city boundary and computes the same statistic. If the observed mean nearest distance is smaller than all 99 simulated ones, the pattern is more clustered than chance at roughly the 1 % level. This avoids the formula assumptions of the nearest neighbour index — irregular boundaries and edge effects are automatically accounted for, because the simulations experience them too. Using the seed as the loop variable makes the whole test reproducible.
QGIS version compatibility
native:randompointsinpolygons and native:randompointsonlines with seeds and POINTS_MISSED arrived in QGIS 3.22 and are present on 3.34 LTR, 3.40 LTR and QGIS 4; older releases had qgis:randompointsinsidepolygons without a seed. native:randomextract and native:randompointsinextent are available throughout. Check processing.algorithmHelp(...) for exact parameter names on your version.
Troubleshooting
- Many points missed. Minimum distance is too large for small polygons; scale counts by area.
- The sample changes every run. No seed was set, or ids were not sorted before
rng.sample. - Points appear in lakes. The polygon layer includes water; subtract it before sampling.
- Line points cluster on long features. Counts are per feature; use a length-based expression.
Conclusion
Generate random locations with the native algorithms and a seed, scale counts by area or length and keep minimum distances realistic, sample existing features with native:randomextract or a seeded Python sampler, stratify when coverage matters and record design weights, and use repeated random placements as the baseline for honest pattern tests.
Frequently Asked Questions
Is a regular grid better than random points for surveys? A grid gives even coverage and is easy to navigate; random points avoid alignment with regular features such as tree rows. Systematic sampling with a random start combines both.
Can I sample raster cells at random? Generate random points in the raster extent and sample values at them, or read pixels into NumPy and sample indices.
How many simulations does a Monte Carlo test need? 99 gives a resolution of 1 %; 999 gives 0.1 % at ten times the cost.
Does QGIS's random generator give the same points on every platform? For a given seed and version, yes in practice; across QGIS versions the implementation may change, so record the version with the seed.