Frequently asked questions

Can’t find what you need? Contact us directly — we read every message.

What exactly is a “toponym,” and what’s in each record?

A toponym is just a fancy word for a place name: the name people actually use when they talk about where they live. Each record in Toponomical holds one place name in its original language paired with its translation or transliteration in one of 22 languages. We also include the script it’s written in (Arabic, Cyrillic, Kanji, etc.), a BCP 47 locale tag so you know exactly which variant you’re getting, what type of translation it is (official government name, transliteration, something people just call it colloquially), and a confidence score from 0.0 to 1.0 so you know how much to trust it.

How is the confidence score calculated?

We cross-reference multiple independent geographic and linguistic sources rather than relying on a single authority. When sources agree on an established, official name, confidence is 1.0. As we move down through rule-based translation, common-but-variable usage, best-effort picks among multiple valid forms, and retained-Latin-script fallbacks, the score drops accordingly, down to 0.5 for uncertain or rarely used names. We report disagreement as lower confidence instead of silently picking a winner.

Surely I can just use AI to get all this information?

We tested that. A lot. Here’s what we found: AI is great at generating plausible-sounding translations, but it’s not great at being consistent or verifiable. Ask ChatGPT for Beijing in 22 languages five times and you’ll get five slightly different answers. Ask it for the same city in the same language tomorrow and you might get something different again. For production systems, that’s a problem. What Toponomical does that AI doesn’t: we cross-reference multiple independent geographic and linguistic sources for every single translation. When they disagree (and they do), we report that as a lower confidence score instead of just picking one and hoping we’re right. We’re also versioned: you know exactly what data you shipped with, so if something breaks, you can trace it. That said, customers building their own on-premise or open-source LLMs find Toponomical incredibly useful as training data. You get a clean, structured, multi-sourced dataset to fine-tune on, instead of bootstrapping from internet scrapes that might be stale or inconsistent.

Which 22 languages are covered?

We picked the top 10 languages by GDP (covers the economies doing the most cross-border commerce) and the top 12 by native speakers (reaches the most people on the planet). The top 12 alone reach 3.28 billion native speakers, about 40.55% of the world’s population. The top 10 economies represent $76.27 trillion in GDP, roughly 68.30% of global economic output. That’s our full list: Mandarin Chinese, Spanish, English, Hindi, Portuguese, Bengali, Russian, Japanese, Punjabi, Turkish, Korean, and French for speakers; plus Arabic, German, Indonesian, Italian, Dutch, Norwegian, Polish, Swedish, Thai, and Urdu for GDP coverage and business needs. If you need something outside this list, enterprise customers can request custom packs. Reach out to us.

What’s the difference between the API and the Enterprise license?

Pick API if you want to ship fast and don’t have massive scale yet: 100 free requests a month, no credit card, low friction. Pick Enterprise if you need everything local (no API calls, no rate limits, no dependencies on us being up), or if you’re running 50K+ queries a month and paying per-request gets expensive. Enterprise gets you the full SQLite database (~113MB) plus CSVs, quarterly updates, and we’ll set you up with a support contact.

How current is the data, and how often does it update?

Releases ship quarterly. Each release is versioned so you can see exactly what changed: new coverage, corrected translations, confidence-score adjustments. You’re never surprised by a silent update.

Does the dataset include administrative hierarchy (state, province, district)?

Yes. Every city record carries foreign keys into three additional levels: country, admin level 1 (e.g. state or province), and admin level 2 (e.g. county or district). Each is translated into all 22 languages. Resolve a full localized address chain in a single query.

Can I use this data to power search, not just display?

Yes. That’s one of the most common use cases. Indexing every language variant means a search for “München,” “Munich,” or “慕尼黑” all resolve to the same city. No more users searching and getting nothing because they used a different name than you indexed.

Do you offer custom language packs beyond the 22?

Enterprise customers can request custom language packs. Talk to sales about your specific language and coverage needs.

What does the free tier actually include?

100 requests per month, no credit card required, full access to all 22 languages, complete country/admin hierarchy in each response, and confidence scores on every translation, the same data quality as paid tiers, just rate-limited. Evaluate fully before you buy.