Convert a QGIS Layer to a GeoPandas GeoDataFrame
GeoPandas is where much of Python's spatial analysis happens: vectorised operations on whole tables, group-bys, merges, and direct access to the scientific stack. QGIS is where the data is loaded, cleaned, styled and inspected. Moving a layer from one to the other is a common step in scripts that use both — and there is more than one way to do it, with very different speed and fidelity.
This recipe belongs to PyQGIS and the Python Data Stack. It shows the fast route of reading the layer's source directly, the general route of converting features through WKB, how to keep the CRS and handle NULLs and dates, and how to choose between them.
Prerequisites
- QGIS 3.34 LTR or newer, or the QGIS 4 series, with GeoPandas and shapely 2 installed into QGIS's Python. Installing Python packages into QGIS covers the per-platform steps; on Linux and conda installations it is usually one command.
- Check the versions from the QGIS Python console before starting:
import geopandas, shapely; print(geopandas.__version__, shapely.__version__).
Route 1: read the layer's source directly
If the layer is a plain file or database table, the fastest conversion is not a conversion at all: ask the layer where its data lives and let GeoPandas read it with its own I/O, which is vectorised and much faster than iterating features.
import geopandas as gpd
from qgis.core import QgsProject, QgsProviderRegistry
layer = QgsProject.instance().mapLayersByName("buildings")[0]
parts = QgsProviderRegistry.instance().decodeUri(layer.providerType(), layer.source())
print(parts) # {'path': '/data/city.gpkg', 'layerName': 'buildings', 'subset': None, ...}
if layer.providerType() == "ogr" and not layer.isModified():
gdf = gpd.read_file(parts["path"], layer=parts.get("layerName"))
if parts.get("subset"):
print("warning: layer has a subset filter, apply it in pandas:", parts["subset"])
print(len(gdf), "rows", gdf.crs)
Breakdown: decodeUri splits the provider's source string into its parts — path, layer name, subset — without fragile string parsing, as explained in decoding and building data source URIs. GeoPandas reads GeoPackage, shapefile, GeoJSON and FlatGeobuf through pyogrio or fiona, typically ten to a hundred times faster than a Python loop over features. Two conditions must hold for the result to match what QGIS shows: the layer must have no unsaved edits (isModified()), and any subset filter or joined fields must be reproduced on the pandas side. When those conditions fail, use the second route.
Route 2: convert features through WKB
Memory layers, layers with pending edits, virtual layers and layers with joined or expression fields only exist inside QGIS. For them, iterate the features and convert geometry through WKB — the binary format both libraries understand natively.
import pandas as pd
import shapely
def layer_to_gdf(layer, request=None):
from qgis.core import QgsFeatureRequest
request = request or QgsFeatureRequest()
names = layer.fields().names()
rows, wkbs = [], []
for f in layer.getFeatures(request):
rows.append(f.attributes())
wkbs.append(bytes(f.geometry().asWkb()) if f.hasGeometry() else None)
df = pd.DataFrame(rows, columns=names)
geoms = shapely.from_wkb([w for w in wkbs], on_invalid="warn")
return gpd.GeoDataFrame(df, geometry=geoms, crs=layer.crs().toWkt())
gdf = layer_to_gdf(layer)
print(gdf.geom_type.value_counts(), gdf.crs.to_epsg())
Breakdown: Collecting attributes and WKB in two lists, then converting all geometry in one shapely.from_wkb call, is much faster than creating a shapely object per feature inside the loop. None entries become missing geometries in the frame. Passing the CRS as WKT preserves custom and compound systems that have no EPSG code; for standard systems layer.crs().authid() is shorter and equally correct. Joined and virtual fields are included automatically because they are part of layer.fields().
Read only what you need
The feature route reads every attribute and every geometry unless told otherwise. For wide tables or analyses that need a few columns, a trimmed request speeds conversion and reduces memory.
from qgis.core import QgsFeatureRequest
req = (QgsFeatureRequest()
.setSubsetOfAttributes(["building_id", "height_m", "use"], layer.fields())
.setFilterExpression('"use" IN (\'residential\', \'mixed\')'))
small = layer_to_gdf(layer, req)
small = small[["building_id", "height_m", "use", "geometry"]]
print(small.memory_usage(deep=True).sum() / 1e6, "MB")
Breakdown: With an attribute subset, the provider leaves other columns as NULL, so they still appear in the frame — selecting the wanted columns afterwards drops them. The expression filter is pushed to the provider where possible, so only matching features are read. For attribute-only analysis, add the NoGeometry flag and build a plain pandas DataFrame instead, as in analysing an attribute table with pandas.
Clean up NULLs and dates
QGIS returns empty attributes as NULL QVariants on QGIS 3 and dates as QDate or QDateTime. Pandas does not understand either, so columns can end up with object dtype holding Qt objects.
from qgis.PyQt.QtCore import QDate, QDateTime
def py(v):
if v is None or (hasattr(v, "isNull") and v.isNull()):
return None
if isinstance(v, QDateTime):
return v.toPyDateTime()
if isinstance(v, QDate):
return v.toPyDate()
return v
def layer_to_gdf_clean(layer, request=None):
gdf = layer_to_gdf(layer, request)
for col in gdf.columns.drop("geometry"):
if gdf[col].dtype == object:
gdf[col] = gdf[col].map(py)
return gdf.convert_dtypes()
gdf = layer_to_gdf_clean(layer)
print(gdf.dtypes)
Breakdown: Mapping only object-dtype columns keeps numeric columns untouched and fast. After conversion, convert_dtypes() lets pandas choose nullable types — Int64 for integer columns with missing values instead of falling back to float, string for text, proper datetime types for dates converted to Python objects. The NULL rules behind this are explained in handling NULL values and QVariant.
Keep the CRS right
A GeoDataFrame without a CRS silently breaks every later reprojection and overlay. Always carry the layer's CRS across, and check it after conversion.
crs = layer.crs()
if not crs.isValid():
raise ValueError(f"{layer.name()} has no CRS; set it before converting")
gdf = layer_to_gdf(layer)
gdf = gdf.set_crs(crs.authid() if crs.authid() else crs.toWkt(), allow_override=True)
assert gdf.crs is not None
print(gdf.crs.name, gdf.crs.axis_info[0].unit_name)
Breakdown: Checking validity first catches layers loaded from files without a .prj or with an unrecognised definition; fixing those is covered in handling missing CRS. The unit name confirms whether distances in GeoPandas will be metres or degrees. Both libraries delegate to PROJ, so transformations done in either agree, provided both use the same PROJ data files — on most installations they do.
Convert every layer in a project
Analyses that combine several layers — buildings, parcels, flood zones — often start by pulling each of them into a frame. A small loop over the project's vector layers, choosing the route per layer, does it in one go and reports what it did.
from qgis.core import QgsVectorLayer
frames = {}
for lyr in QgsProject.instance().mapLayers().values():
if not isinstance(lyr, QgsVectorLayer) or not lyr.isSpatial():
continue
src = QgsProviderRegistry.instance().decodeUri(lyr.providerType(), lyr.source())
direct = (lyr.providerType() == "ogr" and not lyr.isModified()
and not src.get("subset") and lyr.fields().count() == lyr.dataProvider().fields().count())
frames[lyr.name()] = (gpd.read_file(src["path"], layer=src.get("layerName"))
if direct else layer_to_gdf_clean(lyr))
print(f"{lyr.name():<24} {'read_file' if direct else 'features':<9} {len(frames[lyr.name()]):>8} rows")
Breakdown: The direct route is taken only when it is guaranteed to match what QGIS shows: an OGR source, no pending edits, no subset filter, and no joined or expression fields — detected by comparing the layer's field count with the provider's. Everything else goes through the feature route. Keying the dictionary by layer name keeps later code readable (frames["buildings"]), though names are not guaranteed unique in a project; key by lyr.id() if yours repeat.
Decide where the analysis should run
Converting to GeoPandas is worthwhile when the next steps benefit from it: group-bys and merges on attributes, vectorised arithmetic across columns, statistical models, or libraries that expect a frame. For standard geoprocessing — buffers, overlays, dissolves on large layers — QGIS's Processing algorithms are just as fast or faster, keep data in QGIS's own types, and appear in the Processing history. A good rule is to stay in QGIS until a step genuinely needs pandas, convert once, do the pandas work, and bring the result back — rather than shuttling data back and forth between every step.
Measure and choose
Which route is faster depends on the layer size and source. A quick timing on your own data settles it.
import time
t = time.perf_counter()
a = gpd.read_file(parts["path"], layer=parts.get("layerName"))
t1 = time.perf_counter() - t
t = time.perf_counter()
b = layer_to_gdf(layer)
t2 = time.perf_counter() - t
print(f"read_file {t1:.2f} s · feature route {t2:.2f} s · {len(a)} rows")
Breakdown: For file-backed layers of more than a few thousand features, read_file usually wins by a wide margin because the whole read happens in compiled code. The feature route's cost grows with column count and geometry complexity; it remains the only option for data that exists only in QGIS. For very large datasets where even read_file is slow, read with use_arrow=True (pyogrio) or filter with a bbox or where argument at read time.
QGIS version compatibility
The conversion code is plain PyQGIS plus GeoPandas and works on QGIS 3.34 LTR, 3.40 LTR and QGIS 4. Shapely 2's vectorised from_wkb is required; on shapely 1.8, use shapely.wkb.loads per feature. On QGIS 4, NULL values arrive as None and the cleanup function leaves them as is.
Troubleshooting
ModuleNotFoundError: geopandas. It is installed in a different Python; install into QGIS's interpreter.- Columns of
QVariantobjects. NULL values were not converted; use the cleanup function. - Row count differs from QGIS. The layer has a subset filter or unsaved edits that
read_filecannot see. - The frame has no CRS. The layer CRS was invalid; fix it in QGIS before converting.
Conclusion
Read the layer's source with geopandas.read_file when it is a plain, unedited file or table; otherwise iterate features with a trimmed request, convert geometry through WKB in one vectorised call, clean up NULLs and Qt dates, and always carry the CRS across.
Frequently Asked Questions
Does editing the GeoDataFrame change the QGIS layer? No. The frame is a copy. Write results back as described in loading a GeoDataFrame into QGIS.
Can I convert a PostGIS layer with read_file?
Use gpd.read_postgis with an SQL query and a database connection; the layer's URI gives you the connection details.
Is there a built-in QGIS function for this? Not in core. Small helper functions like the ones above are the standard approach.
What about very large layers? Filter at the source, read only needed columns, and consider Dask-GeoPandas or processing in chunks.