Who owns the supermarkets in every German district
The YAML file with all German districts turned out to be useful for more than a cycling score.
It has an osm_id, inhabitants and area per district, so anything counted per district can be normalized.
Counting supermarkets per brand is not very interesting, Aldi and Lidl are everywhere. I wanted to know who owns them, because a lot of the different signs in the German grocery trade belong to the same few groups.
Getting the shops
shop=supermarket in Germany is about 34,000 objects.
I fetch them from my selfhosted Overpass one federal state at a time, into one cache file per state.
The state boundary comes from ISO3166-2, so there are no ids to look up:
[out:json][timeout:900]; relation["boundary"="administrative"]["admin_level"="4"]["ISO3166-2"="DE-BW"]; map_to_area->.state; nwr["shop"="supermarket"](area.state); out tags center;
That is 19 MB of JSON for all 16 states.
Assigning each shop to a district is the same trick as in the cycling post: fetch the 400 boundary relations by osm_id, stitch the member ways with linemerge and polygonize, subtract the inner ways so an enclaved kreisfreie Stadt does not count twice, then a shapely STRtree over the polygons.
The boundaries are the biggest download at about 117 MB, cached as a 55 MB GeoJSON.
From brand to owner
My first version matched the brand tag against a list of chain names.
That list grew to 40 entries, and every one of them is a decision I had to make myself: whether E-Center counts as Edeka, or whether a bare Netto is the unrelated Danish chain.
And it says nothing about who owns what.
The better key was already in the data: 81.7% of the shops carry brand:wikidata.
Wikidata answers the ownership question with owned by (P127) and parent organization (P749), in one query for all 64 QIDs that appear.
SELECT ?brand ?brandLabel ?ownerLabel WHERE { VALUES ?brand { wd:Q701755 wd:Q879858 ... } OPTIONAL { ?brand wdt:P127|wdt:P749 ?owner. } SERVICE wikibase:label { bd:serviceParam wikibase:language "en,de". } }
Wikidata names the regional cooperative that formally owns a brand, so Edeka arrives as "Edeka Minden-Hannover" and "Edeka Südwest", and those get folded onto the group. Aldi Nord, Aldi Süd, Norma and Globus have no owner recorded at all, they are the group themselves.
For the 18.3% without a QID the chain-name matching is still useful as a fallback, and it recovers 941 shops: mostly Edeka, Norma, Penny and Lidl branches where the mapper typed the name and skipped the identifier.
The remaining 5,265 stay not identifiable, and they are the long tail: 4,775 distinct names, 4,367 of which appear exactly once, i.e. Tante-M, Dorfladen, Ihr Kaufmann or Mein Markt.
Shops per group
group shops share % ---------------- ----- ------- Edeka 9982 34.9 Rewe 6370 22.3 Schwarz-Gruppe 4042 14.1 Aldi Nord 2203 7.7 Aldi Süd 2013 7.0 Norma 1347 4.7 Dennree 380 1.3 Salling Group 343 1.2 Migros 297 1.0 not identifiable 5265 -
Shares are over the 28,580 shops that can be attributed to a group, 84.4% of the total. The five big groups hold 24,610 of those, 86.1%.
Edeka's 34.9% is 5,181 shops under its own QID -- Edeka, E-Center, nah und gut -- plus 4,319 Netto Marken-Discount. Netto Marken-Discount is Edeka's discounter, so it looks like a competitor in the shop but belongs to the same company. Without that one ownership edge Edeka and Rewe would be a lot closer.
The leading group per district
group districts led ---------------- ------------- Edeka 315 Rewe 73 Schwarz-Gruppe 4 Aldi Nord 2 feneberg 2 K+K Klaas & Kock 1 Migros 1 Norma 1 V-MARKT 1
Four of the groups at the bottom are purely regional: feneberg in Kempten and Landkreis Oberallgäu, V-Markt in Kaufbeuren, K+K in Landkreis Grafschaft Bentheim, Norma in Fürth. Migros leads one district, Landkreis Fulda, where tegut has 20 of the 95 shops. Fulda is where tegut comes from.
Leading a district says nothing about how big the lead is. In 36 districts the top group holds more than half the shops, and the extreme is Landkreis Straubing-Bogen in Bavaria: 31 of 37 identifiable shops are Edeka group, 22 with an Edeka sign and 9 Netto Marken-Discount. That is a Herfindahl index of 7,093 on the 0--10,000 scale, where a competition authority calls anything above 2,500 highly concentrated. The median district sits at 2,421.
Store counts are a rough proxy for market share -- a Kaufland hypermarket and a Penny count as one shop each.
Supermarkets per inhabitant
Germany has 40.9 supermarkets per 100,000 inhabitants. Per federal state that runs from 53.2 in Mecklenburg-Vorpommern down to 34.6 in Hamburg, and per district from 73.0 in Landkreis Landsberg am Lech to 27.0 in Bottrop.
Both ends of both lists are the wrong way round from what I assumed. The correlation between population density and shops per 100,000 inhabitants is -0.33: the denser a district, the fewer supermarkets per person. Rural districts under 150 inhabitants per km² have a median of 46.6 per 100,000, urban districts over 1,500 have 37.7.
A rural district needs a shop in a lot of small towns, each of them serving a few thousand people, while a city can put one large store where 20,000 people walk past it. The per-capita number counts shops and says nothing about their size or how far away the next one is.
How good is the data
All of the above rests on OSM being evenly mapped, and it is not.
brand:wikidata coverage runs from 52% in Delmenhorst to 98% in Landkreis Oberspreewald-Lausitz, and 58 of 400 districts are below 75%.
Per state the spread is smaller: Brandenburg 90%, Sachsen-Anhalt 89%, Mecklenburg-Vorpommern 88% at the top, Bremen 70%, Baden-Württemberg 77% and Hamburg 79% at the bottom. Coverage correlates -0.30 with population density -- the east German rural districts are the best-tagged part of the country, the western cities the worst.
Tagging quality does not explain the density result, though. Coverage against shops per 100,000 inhabitants correlates -0.04, so effectively not at all. Poorly tagged districts report the same number of shops, with less information attached.
For the concentration numbers the missing shops do matter. An unidentified shop is far more likely to be an independent than a chain, so leaving them out pushes every group's share up. The 86% for the big five is an upper bound on store count, not a measured market share.