disarm

Confusables

Check confusable characters

Paste text to fold homoglyphs toward Latin under both of disarm's digit policies at once. They are not two spellings of one answer — where they disagree, one reading is a number and the other is a word, and which you want depends on what you are about to do with it.

The tool

Three lines, each making a different point. The first spells a brand with Cyrillic letters, which fold the same way under either policy because no digits are involved. The second hides two Devanagari zeros in another brand, where the policies disagree and one reading is the brand being imitated. The third is the trio from disarm’s own documentation — ० ೦ ١ — which TR39 folds to three different letters, o, O and l, where the numeric policy gives 0, 0 and 1.

Reads as paypal.com, and is not.

6 characters here fold onto that Latin string, so it collides with the ordinary spelling in anything that compares folded forms — a username table, a filter, a lookup. Both digit policies agree here.

Which characters are impostors

Your text again, with every confusable character called out. This is the part a folded result cannot show you: раураl and paypal are drawn identically, so an output reading “paypal” looks the same whether anything was wrong or not.

рU+0440аU+0430уU+0443рU+0440аU+0430ӏU+04CF.com

Numeric policy

A non-Latin digit becomes the ASCII digit. What prose means.

paypal.com

TR39 policy

A non-Latin digit becomes a Latin letter. What a skeleton needs.

paypal.com

Character by character

CodepointName NumericTR39
U+0440 CYRILLIC SMALL LETTER ER p p
U+0430 CYRILLIC SMALL LETTER A a a
U+0443 CYRILLIC SMALL LETTER U y y
U+0440 CYRILLIC SMALL LETTER ER p p
U+0430 CYRILLIC SMALL LETTER A a a
U+04CF CYRILLIC SMALL LETTER PALOCHKA l l

Running disarm 0.14.1, compiled to WebAssembly. Your text is never uploaded — the engine is loaded into this page and runs on your machine.

A worked example

The tool above needs JavaScript. This is the same folding written out, so it is legible without running anything.

InputNumericTR39Agree?
g००gleg00glegoogleno
раураlpaypalpaypalyes
०೦١001oOlno
paypalpaypalpaypalyes — nothing to fold

The first row is the argument for having two policies. Under the numeric policy g००gle becomes g00gle, which reads as a string with two zeros in it. Under TR39 it becomes google, which collides with the brand being imitated. If you are storing the text, the first is right. If you are asking whether someone is impersonating a domain, only the second answers the question.

The second row involves no digits, so both policies agree: five Cyrillic letters fold to their Latin lookalikes either way.

Where the table comes from

The bulk of it is generated from Unicode TR39's confusables.txt, version 17.0.0. Two smaller sets are layered on top, and the second is the interesting one.

confusables_supplement.tsv adds cross-script pairs that TR39 leaves without a shared prototype. confusables_attested.tsv adds 31 codepoints attested in real attacker text, mined from the BitCore subset of the BitAbuse corpus, which TR39 does not list as sources at all. Twenty-three of those are optical twins of a Latin letter — ɴ → n, ʍ → m, ʀ → r. Eight are not: seven are glyphs an attacker used positionally rather than because they resemble the letter (ժ → d, ᚱ → r, Ⴝ → s), and one is a reading convention (щ → w). The rule for those rows is observed attacker substitution, which is wider than visual confusability, and they are marked tier 2a and 2b.

One consequence is visible in this tool: the attested rows fold to lowercase targets. ɴ is a small capital letter, and TR39's own pairing would send it to N; disarm sends it to n, because the job is recovering the word an attacker obscured rather than preserving letter case. Paste aɴd ʙig ᴇgg above and the folded column reads and big egg. The confusables guide carries the full account.

The same thing in your own code

Each code block has been compiled and verified in CI. Provided under the MIT license to illustrate disarm. disarm on GitHub →

# Fold confusables under both digit policies. They answer different questions.
#   pip install disarm
from disarm import normalize_confusables

# A brand spelled with two Devanagari zeros standing in for the letter o.
SPOOF = "g००gle"

numeric = normalize_confusables(SPOOF)                        # the default
tr39 = normalize_confusables(SPOOF, digit_policy="tr39")

# Numeric keeps a digit a digit, which is what stored text means. TR39 folds it
# to a letter, which is what makes a spoof collide with the brand it imitates.
assert numeric == "g00gle", numeric
assert tr39 == "google", tr39
assert numeric != tr39, "the policies must disagree here"

print(f'ok: numeric gives "{numeric}", tr39 gives "{tr39}"')

Choosing a policy

You areUseBecause
Cleaning text you will store or displaynumeric A digit keeps its value. Folding ० to o would turn a quantity into a word.
Comparing usernames or identifiersTR39 A skeleton only has to collide. Whether it reads sensibly is not the question being asked.
Checking a domain against a brandTR39 The spoof must land on the same skeleton as the thing it imitates.
Building a search indexnumeric Query and document should agree, and a user typing a digit means a digit.
Comparing against a published TR39 benchmarkTR39 It is the upstream mapping; anything else will differ from the reference by design.

Folding is one control, not the whole answer. It catches cross-script substitution, where a Latin word borrows a Cyrillic letter. It does not by itself separate a label written entirely in Cyrillic that skeletons to a Latin brand from a legitimate Russian word — that needs a mixed-script or whole-script check, which disarm reports separately.

Found a string this gets wrong? The confusables table grew out of exactly that kind of report. Open an issue with it.

Related tools