Run Batch Jobs from a CSV Parameter Table in PyQGIS

Some batch jobs vary in more than their input file. Buffer each protected area by its own legally defined distance; clip each municipality's data to its boundary and write it to its own folder; run a suitability model with different weights for each scenario. The cleanest way to express such jobs is a table: one row per job, one column per parameter. Planners and analysts can maintain the table in a spreadsheet, and a short script runs whatever it describes — validating every row first, so a typo in row 47 does not fail an hour into the run.

This recipe belongs to Batch Processing with PyQGIS. It designs a parameter table, reads and validates it, converts text values to the types Processing expects, runs the jobs, writes status and outputs back into a results table, and supports dry runs.

A table of jobsA CSV file with one row per job: an id, the algorithm, the input path, parameters such as distance, and the output path. The script validates every row first — files exist, algorithm exists, values parse — and reports all problems together. Valid rows run one by one, and a results table records status, feature count, duration and any error per job.Rows in, validated, run, recordedjobs.csvid · algorithminput · distanceoutputvalidate allfiles exist?values parse?report all errorsrun + recordstatus · countduration · errorresults.csv

Prerequisites

  • QGIS 3.34 LTR or newer, or the QGIS 4 series.
  • A CSV of jobs saved as UTF-8. Spreadsheets export it with "CSV UTF-8"; check the delimiter, which is a semicolon in many European locales.

Design the parameter table

One row per job, one column per parameter, and a few columns for bookkeeping. Column names should be readable by the people maintaining the table.

job_id;enabled;algorithm;input;distance_m;dissolve;output
PA-001;1;native:buffer;/data/protected/pa_001.gpkg;250;1;/data/out/pa_001_buffer.gpkg
PA-002;1;native:buffer;/data/protected/pa_002.gpkg;500;1;/data/out/pa_002_buffer.gpkg
PA-003;0;native:buffer;/data/protected/pa_003.gpkg;250;0;/data/out/pa_003_buffer.gpkg

Breakdown: A stable job_id identifies each row in reports and outputs. An enabled column lets maintainers switch jobs off without deleting rows. Units in column names (distance_m) prevent the classic mistake of entering kilometres. Keeping the algorithm in a column makes a table that mixes several kinds of job possible, though tables are easier to validate when every row uses the same algorithm.

Read and validate every row first

Validation runs over the whole table before any processing, collecting every problem so they can be fixed in one pass.

Validate the whole table, then runValidating row by row while running stops at the first bad row after earlier rows have already run. Validating the whole table first collects every problem — a missing file in row 12, a non-numeric distance in row 30, an unknown algorithm in row 41 — reports them all, and runs nothing until the table is clean.Fail before running, not duringvalidate while runningrow 12 fails after11 jobs already ranpartial resultsvalidate firstall problems listednothing ran yetfix, then run

import csv
from pathlib import Path
from qgis.core import QgsApplication

registry = QgsApplication.processingRegistry()

def load_jobs(path, delimiter=";"):
    with open(path, newline="", encoding="utf-8-sig") as fh:
        return list(csv.DictReader(fh, delimiter=delimiter))

def validate(jobs):
    problems, seen = [], set()
    for n, job in enumerate(jobs, start=2):           # row 1 is the header
        jid = job.get("job_id", "").strip()
        if not jid or jid in seen:
            problems.append(f"row {n}: missing or duplicate job_id '{jid}'")
        seen.add(jid)
        if job.get("enabled", "1").strip() not in ("0", "1"):
            problems.append(f"row {n}: enabled must be 0 or 1")
        if registry.algorithmById(job.get("algorithm", "")) is None:
            problems.append(f"row {n}: unknown algorithm '{job.get('algorithm')}'")
        if not Path(job.get("input", "")).exists():
            problems.append(f"row {n}: input not found: {job.get('input')}")
        try:
            if float(job.get("distance_m", "").replace(",", ".")) <= 0:
                problems.append(f"row {n}: distance_m must be positive")
        except ValueError:
            problems.append(f"row {n}: distance_m is not a number: '{job.get('distance_m')}'")
    return problems

jobs = load_jobs("/data/jobs/buffers.csv")
problems = validate(jobs)
if problems:
    print("\n".join(problems))
    raise SystemExit(f"{len(problems)} problems — nothing was run")

Breakdown: utf-8-sig handles the byte-order mark some spreadsheet exports add, which otherwise corrupts the first column name. Row numbers in messages refer to spreadsheet rows, header included, so maintainers can jump straight to the problem. Every rule appends a message instead of raising, so one run of the validator reports every issue. Accepting a decimal comma in numbers spares European users a confusing error. Nothing runs until the table is clean.

Convert text to the types algorithms expect

CSV values are strings; Processing parameters need numbers, booleans, enums and paths. A small conversion layer keeps the job loop simple.

def to_bool(text):
    return text.strip().lower() in ("1", "true", "yes", "y")

def to_float(text):
    return float(text.strip().replace(",", "."))

def params_for(job):
    return {
        "INPUT": job["input"],
        "DISTANCE": to_float(job["distance_m"]),
        "SEGMENTS": 16,
        "DISSOLVE": to_bool(job.get("dissolve", "0")),
        "OUTPUT": job["output"],
    }

print(params_for(jobs[0]))

Breakdown: Explicit converters make the mapping from table to algorithm visible and testable. Values not in the table — segments here — are set once in code, so the table only contains what actually varies. For algorithms with enum parameters, map readable words in the table ("round", "flat") to the enum numbers in the converter rather than asking maintainers to type numbers. Check parameter names with processing.algorithmHelp.

Run the jobs and record results

Each enabled row runs once; status, feature count, duration and any error are written to a results table that mirrors the job table.

import time
import processing
from datetime import datetime
from qgis.core import QgsVectorLayer

results = []
for job in jobs:
    if job["enabled"].strip() != "1":
        results.append({**job, "status": "disabled"})
        continue
    Path(job["output"]).parent.mkdir(parents=True, exist_ok=True)
    start = time.perf_counter()
    try:
        out = processing.run(job["algorithm"], params_for(job))["OUTPUT"]
        count = QgsVectorLayer(out, "check", "ogr").featureCount()
        status, error = "ok", ""
    except Exception as err:
        count, status, error = 0, "failed", str(err)[:300]
    results.append({**job, "status": status, "features": count,
                    "seconds": round(time.perf_counter() - start, 1), "error": error,
                    "run_at": datetime.now().isoformat(timespec="seconds")})
    print(f"{job['job_id']:<8} {status:<8} {count:>7}")

with open("/data/jobs/buffers_results.csv", "w", newline="", encoding="utf-8") as fh:
    w = csv.DictWriter(fh, fieldnames=list(results[0].keys()) + ["features", "seconds", "error", "run_at"],
                       delimiter=";", extrasaction="ignore")
    w.writeheader()
    w.writerows(results)

Breakdown: Copying every job column into its result row means the results table can be opened next to the job table in the same spreadsheet and read row for row. Disabled jobs appear with their status, so the results account for every row. Truncating error messages keeps cells readable. Opening each output to count features confirms that something was actually written — a successful run that produced an empty layer is worth noticing. For large tables, combine with parallelising batch jobs.

Rerun only what failed

After fixing problems, rerunning everything wastes time. Reading the previous results and running only failed or new jobs makes the table a living to-do list.

Rerun only what needs itBefore running, the previous results table is read. Jobs whose last status was ok and whose output still exists are skipped. Jobs that failed, are new, or whose parameters changed since the last run are run again. The table can be edited and rerun repeatedly until every job is ok.Skip what is done and unchangedok + unchangedskipfailedrun againnew or changedrun

import hashlib

def job_hash(job):
    keys = ("algorithm", "input", "distance_m", "dissolve", "output")
    return hashlib.sha1("|".join(job.get(k, "") for k in keys).encode()).hexdigest()

previous = {}
prev_path = Path("/data/jobs/buffers_results.csv")
if prev_path.exists():
    for r in load_jobs(prev_path):
        previous[r["job_id"]] = r

todo = [j for j in jobs
        if not (previous.get(j["job_id"], {}).get("status") == "ok"
                and previous[j["job_id"]].get("hash") == job_hash(j)
                and Path(j["output"]).exists())]
print(len(todo), "of", len(jobs), "jobs need running")

Breakdown: A hash of the parameters that affect the output detects edited rows: change a distance and the hash changes, so the job reruns even though it succeeded before. Store the hash in the results table (add it to the result rows) so the next run can compare. Checking that the output file still exists catches outputs deleted since. The same idea, applied to long single runs, is covered in resuming and checkpointing a batch run.

Support a dry run

A dry run validates and prints what would happen without running anything — useful before long batches and for reviewing a colleague's table.

import sys

DRY_RUN = "--dry-run" in sys.argv
if DRY_RUN:
    for job in jobs:
        state = "would run" if job["enabled"].strip() == "1" else "disabled"
        print(f"{job['job_id']:<8} {state:<10} {job['algorithm']} → {job['output']}")
    raise SystemExit(0)

Breakdown: A command-line flag switches between dry and real runs without editing the script. Printing algorithm and output per job shows exactly what a run would touch, including outputs that would overwrite existing files. Running the dry run after every table edit is a cheap habit that catches mistakes such as two jobs writing to the same output.

QGIS version compatibility

The approach uses only processing.run and the algorithm registry, which work on QGIS 3.34 LTR, 3.40 LTR and QGIS 4. Parameter names come from each algorithm and are stable across these releases for the native algorithms shown.

Troubleshooting

  • The first column name is garbled. The file has a byte-order mark; read with utf-8-sig.
  • Every row is one big column. The delimiter is wrong; spreadsheets in many locales use semicolons.
  • Numbers fail to parse. Decimal commas or units typed into cells; normalise in the converter or tighten the table.
  • Two jobs overwrite each other. Duplicate outputs; add a uniqueness check on the output column to validation.

Conclusion

Describe jobs as rows with readable, unit-bearing column names and an enabled flag, validate the whole table before running and report every problem, convert text to typed parameters in one place, run enabled jobs while recording status, counts, timing and errors to a results table, rerun only failed, new or changed jobs, and offer a dry run.

Frequently Asked Questions

Can the table live in a spreadsheet instead of CSV? Yes — read .xlsx with pandas or openpyxl; CSV is simpler to diff and version.

Can a row reference a QGIS project layer instead of a file? Yes, by layer name; look it up with QgsProject.instance().mapLayersByName in the converter.

What about Processing's own batch interface? It loads and saves parameter tables in its own JSON format and suits interactive work; a script suits scheduled and versioned runs.

Should the job table be versioned? Yes. Keep it in the same repository as the script; a diff of the table shows exactly which jobs changed between runs, which is the audit trail most teams need.

How do I run different algorithms per row? Keep an algorithm column and one converter per algorithm, chosen by the row's value.