Bài viết này chưa được dịch sang Tiếng Việt — bạn đang đọc bản gốc bằng English. Cũng có bằng:Deutsch, English, Українська
Which Countries Actually Publish Surname Registries — and Which Ones Can't
If you need to know how common a surname is, there are exactly two kinds of answers you can get. In a handful of countries, a government agency publishes a named list of surnames with a bearer count next to each one. You look up the name and you have a number. Everywhere else, you get an estimate: someone's reading of phone books, church records, sample surveys, or a scraped website — and the number you get is an opinion with a decimal point on it.
That distinction matters more than it sounds. A registry is a measurement. Without one, surname frequency is an expert estimate, and estimates inherit every bias of whatever proxy stood in for the population. The gap is not academic: it decides whether you can verify a dataset, or only argue about it.
This article goes country by country through the registries we actually pulled and used — Taiwan, Spain, France, South Korea, China, Vietnam — and then through the two cases where the honest answer is that no such data exists and cannot be made to exist: Italy and Ukraine. All figures below come from the sources listed at the end, checked on 2026-07-17.
What counts as a registry
A surname registry, for our purposes, means a published table with three properties:
- Named entries. Not "the top 10," not "surnames ending in -enko are common." Each surname appears as a row.
- Counts, not ranks. A rank tells you the order. A count tells you the distance between ranks — and the distance is what any weighting or sampling work actually needs.
- Depth past the head. Everyone knows the top ten. A registry earns its name in the tail, where estimates fall apart.
Almost every country publishes some onomastic material. Very few publish all three.
The comparison
| Country | Registry? | Depth published | Cut-off | Top-10 share | Key caveat |
|---|---|---|---|---|---|
| Taiwan | Yes, open data | 2,731 surname entries × 3 age bands | None — 1 bearer is listed | 52.79% of population | Variant graphs counted separately (温 / 溫 are different rows) |
| Spain | Yes, national stats | 86,492 surnames | ≥ 20 bearers (confidential) | 17.55% of positions / 34.36% of bearers | Two surnames per person — read the right column |
| France | Yes, national stats | 218,982 surnames × 11 decades | ≥ 30 births per decade | 1.71% | Counts births, not living bearers; births abroad excluded |
| South Korea | Yes, census | Named counts per surname | Census-level | 63.9% (hanja) / 65.8% (hangul) | Hanja and hangul give different denominators |
| China | Partial | Top 100 with shares | Nothing below rank 100 | 42.9% (top-10); 84.77% (top-100) | No published table exists for ranks 100–300 |
| Vietnam | No — large samples | Sample frequency lists | Sample-dependent | ~71% split / ~76% merged (SG01) | Not a state registry; north/south variants split or merged changes the answer |
| Italy | No | None — ISTAT does not collect surnames | — | ~2–3% (estimate) | Italian surnames are, by expert account, uncountable |
| Ukraine | No | Scholarly top-100 only | — | ~4–6% (estimate) | 707,685 unique surnames, ~64 bearers each; no list to check against |
The countries that publish
Taiwan — the cleanest registry in the world
Taiwan's Ministry of the Interior publishes surname rankings and counts as an open dataset, broken out by three age bands, as of 30 June 2023. It is the only source in this list where every weight in a dataset can come from the registry itself, with no modelled curve in between.
What makes it exceptional is not the depth but the absence of a privacy floor. Taiwan publishes surnames with a single bearer — 643 of them, plus 406 with two bearers. That has a consequence most people miss: if a surname is not in the Taiwanese registry, it genuinely has zero bearers. There is no "hidden below the threshold" excuse. That single property turns the registry into a falsification tool. We used it to delete 33 surnames that had been taken from the 百家姓 — an 11th-century classical text — and which have no bearers in Taiwan at all. And it saved 澹臺, which looks exactly like the same kind of archaic ghost but sits in the registry with two real bearers. Looks cannot tell those apart. Only a counter can.
Top-10 concentration: 52.79% of the population, with a head unlike mainland China's — 陳 11.2%, 林 8.3%, and 王 only 4.1% at rank 6.
Spain — the registry that solves the two-surname problem for you
Spain's INE publishes an Excel file of every surname with 20 or more bearers: 86,492 of them, from the census dated 1 January 2025. The ≥20 cut-off is an explicit statistical confidentiality rule, so unlike Taiwan, absence from the Spanish file proves nothing.
Spain's unique contribution is structural. Spaniards carry two surnames, which means naive frequency shares sum to roughly 200% and every downstream calculation quietly breaks. INE sidesteps this by publishing separate columns for first surname, second surname, and both — the source itself disambiguates what everyone else leaves you to guess. GARCIA: 2,915,761 across both positions.
The two denominators give very different-looking numbers for the same reality: the top 10 surnames account for 17.55% of surname positions but 34.36% of people. Neither is wrong. Quoting one while meaning the other is where the errors come from.
France — a birth registry wearing a surname registry's clothes
INSEE publishes 218,982 surnames with at least 30 recorded births, split across 11 decades from 1891 to 2000. It is by far the deepest file here, and it is the one most likely to be misread.
It counts births, not living bearers. And people born abroad are excluded — which means the entire immigrant layer of the French surname stock is systematically understated. If you need "which surnames do people in France have today," this file does not answer that question; it answers "which surnames were registered at birth in France, by decade." For historical and generational work, that decade split is a gift. For a snapshot of the living population, it is a trap.
France is also the flattest registry in this list: the top 10 cover just 1.71%.
South Korea — a census with named counters
KOSTAT's 2015 census (인구주택총조사, surname statistics) publishes per-surname counts. Korea is the most concentrated case here: 김 21.5%, 이 14.7%, 박 8.4%. Top 10 reaches 63.9% by hanja or 65.8% by hangul — the two are different objects, because distinct hanja clans collapse into the same hangul spelling.
That concentration has a practical edge. 김 at 21.5% against surnames with 7 bearers is a ratio of roughly 1,527,138:1. Any fixed-precision weighting scale simply cannot express that; the answer is to accept the shortfall rather than delete real surnames to make a target.
China — a top-100, and then a wall
The Ministry of Public Security's 2007 tabulation gives the top 100 surnames with shares: 王 7.25%, 李 7.19%, 张 6.83%. Those top 100 cover 84.77% of the population; the top 10 cover 42.9%.
Below rank 100, the data stops. No published table of shares for ranks 100–300 exists that we could find. This is worth stating plainly because it is the single most common place people invent numbers: the head is so well documented that the absence of a tail feels like an oversight rather than a boundary. It is a boundary.
Vietnam — no registry, but honest samples
Vietnam has no state surname registry. What it has is large samples published by hoten.org: VNTH01 (n = 1,682,729) and SG01 (n = 241,000, Ho Chi Minh City). Two samples with opposite regional biases converging independently is a strong argument, and they converge: Nguyễn is about 31%, not the 38% that circulates everywhere. The 38% traces back to a 1992 sample of 1,941 people — repeated endlessly not because it is good, but because it was the only one published.
A widely-cited figure with one small primary source is not a measurement. It is a rumour with a citation.
The countries that can't
Italy — "we cannot count Italian surnames," and that's the expert position
This is not our inference. Italy's leading onomastician, Enzo Caffarelli, published an article for Treccani in 2023 titled — literally — "Perché non possiamo contare i cognomi italiani" ("Why we cannot count Italian surnames"). His answer is that no amount of databases, computing, or AI fixes it.
The mechanics: ISTAT collects first names of newborns and does not collect surnames at all. ANPR, the national population register, is closed. The only national count ever made was Emidio De Felice's in the 1970s, built from telephone subscribers — SIP, later a SEAT database of roughly 24 million lines. That channel then closed permanently: mobile phones, unlisted numbers, and the split across carriers ended telephone directories as a demographic proxy. Caffarelli adds a methodological layer that survives even a perfect database: are accented and unaccented forms one surname or two (Barilla / Barillà)? Are double surnames counted apart from their components?
Commercial sites like cognomix.it still publish maps and counts. Read their basis: they count families in telephone listings, not people. It is the De Felice proxy, decades after the proxy stopped working.
Italy is the proof that "developed country + good statistics agency" does not imply "surname data exists."
Ukraine — no registry to check against
There is no Ukrainian surname registry available for verification. ridni.org holds a real snapshot with genuine counters, but from 2011–13, and it is not published as a dataset; forebears.io and stats.ridni.org likewise offer no downloadable file. Scholarly work (Yu. Pradid and others) publishes a top-100 and stops.
The consequence is concrete. Ukraine has 707,685 unique surnames, averaging 64 bearers each. In a distribution like that, "rare and strange-sounding" is the norm, not a defect — so with no registry, cleaning a long tail means deleting names by how they sound. We tried that heuristic on deliberately suspicious candidates and it produced a 21% false-positive rate: real surnames borne by real, identifiable people.
Compare directly with Taiwan. Same-looking archaic entry, opposite correct action — because Taiwan hands you a counter and Ukraine does not. The difference is never the obviousness of the defect. It is the existence of the registry.
What this means if you need surname data
- Check for a per-name counter before anything else. Rank lists and "top 10" articles are not registries.
- Read the cut-off rule. Taiwan has none (absence = zero bearers). Spain's ≥20 is a confidentiality rule (absence = unknown). Those are opposite meanings for the same empty cell.
- Read the unit. Births ≠ living bearers (France). Positions ≠ people (Spain). Families ≠ individuals (Italian commercial sites). Hanja ≠ hangul (Korea).
- Where no registry exists, prefer large independent samples that disagree in their biases and still converge (Vietnam), over one famous number from a tiny sample.
- Where nothing exists, say so. "No data" is a finding. An estimate presented as a count is not.
Methodology / data as of 2026-07-17
All figures were pulled and checked on 17 July 2026, in the course of building weighted surname corpora for 63 locales. ⚠ The corpus held 63 locales when these figures were measured; en_IE was added on 18 July 2026 and it holds 64 today. Percentages describe the share of the population unless stated otherwise. Registry files were read directly rather than via secondary summaries; Vietnamese figures come from published samples, not a state source; Italian and Ukrainian entries are negative findings — repeated attempts to locate a per-name national source found none available.
Sources:
- Taiwan, Ministry of the Interior — surname rankings and counts by age band, 30 June 2023: data.gov.tw/dataset/126774
- Spain, INE — surnames with ≥20 bearers, census of 1 January 2025: ine.es/daco/daco42/nombyapel/… (methodology: ine.es/daco/daco42/nombyapel/…)
- France, INSEE — surnames by decade of birth, 1891–2000: insee.fr/fr/statistiques/fichie…
- South Korea, KOSTAT — 2015 Population and Housing Census, surname statistics (인구주택총조사 성씨 통계)
- China, Ministry of Public Security — 2007 top-100 surname tabulation
- Vietnam — hoten.org VNTH01 (n = 1,682,729) and SG01 (n = 241,000) frequency datasets, CC BY 4.0 as declared in the page text
- Italy — Enzo Caffarelli, "Perché non possiamo contare i cognomi italiani", Treccani, 2023: treccani.it/magazine/lingua_italia…
- Ukraine — ridni.org (2011–13 snapshot, not published as a dataset); scholarly top-100 lists