Blog/
GeoJSON vs shapefile for building footprint delivery
If you're specifying a delivery format for a national footprint layer, the shapefile-versus-GeoJSON question usually comes up after someone has already tried to load a 40GB .shp into QGIS and watched it choke, or tried to open a GeoJSON in a legacy desktop tool that only knows ESRI's sidecar files. Both formats carry polygon geometry fine. Where they diverge is attribute handling, file structure, and what your ingest pipeline is built to read.
Shapefile: still the default in a lot of shops
A shapefile is actually a bundle of files: .shp for geometry, .dbf for attributes, .shx for the index, and usually a .prj for the coordinate reference system. Lose one and the whole set is unreadable, which matters when a footprint dataset gets handed off between departments or vendors and someone forgets to zip the .prj along with the rest.
The .dbf format caps field names at 10 characters and has limited support for wide text fields, so attribute schemas get abbreviated in ways that aren't always obvious six months later ("BLDG_TYP_1" is not self-documenting). Shapefile also has a 2GB per-file size ceiling, which sounds generous until you're vectorizing every structure across a state or province and the file has to split into tiles anyway.
None of that makes shapefile obsolete. A lot of municipal GIS, legacy ETL scripts, and older desktop software still expect it as input, and if your downstream consumers are on that stack, converting away from shapefile just adds a step nobody asked for.
GeoJSON: easier for modern pipelines, worse at scale without tiling
GeoJSON is a single text file, geometry and attributes together, readable by anything that parses JSON. That's the appeal for teams feeding a PostGIS database, a web map, or a Python/Node ingest script: no sidecar files, no field-name truncation, and attribute values can carry full text without the schema gymnastics shapefile forces.
The tradeoff is file size and parsing speed at national scale. A plain-text format with no compression and no spatial indexing gets unwieldy fast once you're past a few hundred thousand polygons, which a building footprint layer for any sizeable region will clear quickly. Most teams handle this by tiling the output (by county, by grid cell, by administrative boundary) rather than shipping one GeoJSON for an entire country. That's a delivery and QA decision as much as a format one.
Coordinate reference system is a second gotcha. GeoJSON's spec technically wants WGS84, but a lot of production data gets written in whatever CRS the processing pipeline used, and if the receiving system assumes WGS84 without checking, footprints land in the wrong place with no error thrown. Confirm CRS on receipt every time, regardless of which format shows up.
What decides the format
For a base layer that every other property analytic sits on top of, parcel matching, flood exposure, solar potential, the format choice mostly comes down to what's already consuming the data. If your valuation models, your parcel-matching scripts, or your GIS team's existing tooling expect shapefile, take shapefile. If you're building a modern ingest pipeline into PostGIS or a cloud warehouse and want to skip the sidecar-file problem, GeoJSON is the lighter lift, as long as you tile it.
Neither format changes what's underneath it: a national-scale set of building polygons that needs annual refresh as new construction and demolitions change the map. Building Footprints delivers that layer as GeoJSON or Shapefile from VHR imagery, so the format question is a delivery setting rather than a reason to keep digitizing footprints by hand.
If you're specifying delivery for a footprint base layer, ask about format options before the imagery sourcing.