I used to work for an NLP startup, we focused on stuff you could do with Romanized names -- names that were original not written in the Latin alphabet and ended up being written in the Latin alphabet using some kind of transliteration scheme.
For example, we could take a name and generate a pretty comprehensive, and culturally aware, list of variants.
Jennifer -> Jenifer, Jen, Jenny, Jennie, etc.
Richard -> Rich, Richie, Dick, Dickie, Ritchard, etc.
Rho -> No, Lo, Loh, Noh, Roh, Ro, Nho, etc.
The intention of course was to build up lists of name variants that could be used during identification checks.
We also had some pretty significant statistical models that could guess Gender and provide a descending list with confidence levels of the most likely country of origin for a name. It was surprisingly accurate and could account for different Romanization schemes popular in different countries. It could even guess if a name was a surname or a given name.
What did we build the models on? Somehow, one of the founders was able to swing access to U.S. Border Control Data. Even though it was names and country of origin data, it's de-identified (having a list of names doesn't mean we know who the names belong to). There was something north of a billion names in the collection, and included place of birth, country of origin, gender, etc. Names were mined for digraphs so we could build CFGs that could be walked to generate variants. There was lots of manual work as well. Endless regex writing and testing, QA, that sort of thing.
For some countries, we had pretty poor data to be honest. I think we had a couple dozen North Koreans, but for most of the world, our coverage was surprisingly good. It turns out all that work boiled down into a surprisingly small library just a couple dozen megabytes in size and was pretty fast -- I don't remember how fast, but something like a few thousand names per hour. It was pretty niche, but eventually the company was acquired and I went on my way.
I always assumed that technology like that would find its way into more applications, but I'm constantly surprised it hasn't.
>'I always assumed that technology like that would find its way into more applications, but I'm constantly surprised it hasn't.'
Many years ago, I was working on a large project for an organization nothing apparently consistent between half a dozen systems with tens of thousands of users each except names. Naturally, those names were full of exactly the kind of variations you're describing.
When I went looking for a solution to do exactly what you're describing I ran into solutions that were both vague about their functionality and expensive. Like you say, pretty niche - it seemed that everyone was used to selling very specific 'solutions' not a library/API.
I ended up hacking together a very basic script to accomplish the same. It took days to run thanks to my non-existent coding skills, but the accuracy was pretty good.
What it couldn't line up was solved by later decoding and discovering correlations between the long forgotten conventions used for unique IDs in the various systems.
For example, we could take a name and generate a pretty comprehensive, and culturally aware, list of variants.
Jennifer -> Jenifer, Jen, Jenny, Jennie, etc.
Richard -> Rich, Richie, Dick, Dickie, Ritchard, etc.
Rho -> No, Lo, Loh, Noh, Roh, Ro, Nho, etc.
The intention of course was to build up lists of name variants that could be used during identification checks.
We also had some pretty significant statistical models that could guess Gender and provide a descending list with confidence levels of the most likely country of origin for a name. It was surprisingly accurate and could account for different Romanization schemes popular in different countries. It could even guess if a name was a surname or a given name.
What did we build the models on? Somehow, one of the founders was able to swing access to U.S. Border Control Data. Even though it was names and country of origin data, it's de-identified (having a list of names doesn't mean we know who the names belong to). There was something north of a billion names in the collection, and included place of birth, country of origin, gender, etc. Names were mined for digraphs so we could build CFGs that could be walked to generate variants. There was lots of manual work as well. Endless regex writing and testing, QA, that sort of thing.
For some countries, we had pretty poor data to be honest. I think we had a couple dozen North Koreans, but for most of the world, our coverage was surprisingly good. It turns out all that work boiled down into a surprisingly small library just a couple dozen megabytes in size and was pretty fast -- I don't remember how fast, but something like a few thousand names per hour. It was pretty niche, but eventually the company was acquired and I went on my way.
I always assumed that technology like that would find its way into more applications, but I'm constantly surprised it hasn't.