disarm

Combining marks

Detect zalgo text without breaking Vietnamese

Paste text to see how deep its combining marks stack, base character by base character. The naive rule — strip the marks, or allow only one — quietly corrupts Vietnamese, which needs two on a single letter.

The tool

Four lines: Vietnamese and French, which must survive any sane rule; a zalgo with two marks, which no threshold separates from the Vietnamese; and a sprawling one, which every threshold catches.

Stacked 7 marks deep — beyond any orthography.

No writing system needs this many marks on one base. disarm reports it as zalgo at the default threshold of 3.

How deep the marks stack

Every base character with the number of combining marks riding on it. Depths of one and two are unremarkable — that is where French, Vietnamese and Hebrew live. The colour only rises above that.

Tiế2ng Việ2t — Nguyễ2n Thị1 Hư1ờ2ng
café1 naï1ve
Z͓̎2algo
Hͨ̊̽ͤ4e͓͔ͭ͑ͥͪͫ7

Your text at each threshold

is_zalgo flags a base carrying more than the threshold in marks. Depths are counted after decomposition, so a precomposed ế still counts the two marks it decomposes to. A low threshold catches more abuse and starts rejecting real languages.

ThresholdVerdictCost
1 zalgo flags Vietnamese
2 zalgo the capping default
3 zalgo the default
4 zalgo
5 zalgo

Capped at two marks

strip_zalgo with disarm's default cap, which is what Vietnamese needs.

Tiếng Việt — Nguyễn Thị Hường
café naïve
Z͓̎algo
Hͨ̊e͓͔

Running disarm 0.14.1, compiled to WebAssembly. Your text is never uploaded — the engine is loaded into this page and runs on your machine.

Where depth stops working

Mark depth is the only thing a counter can see, and it has a floor. These two are the same shape:

TextDecomposes toMarks
ế — VietnameseU+0065 U+0302 U+03012
Z͓̎ — zalgoU+005A U+030E U+03532

A base and two marks in both cases, so no threshold tells them apart. Depth catches the sprawling kind reliably and the restrained kind not at all. Separating those needs to look at which marks appear and whether they form a real orthographic unit — a different question from how many.

What depth does buy you is a safe floor. Measured against disarm, with the deepest stack each sample reaches:

SampleDeepest stackFlagged at threshold 1?At 3?
Vietnamese — Tiếng Việt2yesno
Hebrew with niqqud2yesno
French — café naïve1nono
Thai1nono
Zalgo, sprawling5+yesyes

A threshold of one rejects ordinary Vietnamese and Hebrew. Three is disarm's default for detection, and two is its default when capping, because two is what Vietnamese needs.

The same thing in your own code

Each code block has been compiled and verified in CI. Provided under the MIT license to illustrate disarm. There is no C here: the C ABI exposes no is_zalgo or strip_zalgo, so there is nothing to call. disarm on GitHub →

# Cap combining-mark stacking without corrupting Vietnamese.
#   pip install disarm
from disarm import is_zalgo, strip_zalgo

# Vietnamese puts two marks on one base: ế is U+0065 U+0302 U+0301.
VIETNAMESE = "Tiếng Việt"
# Five marks on one base, which no writing system uses.
ZALGO = "Hͤͥͦͧͨ"

# The naive rule — at most one mark — rejects an ordinary Vietnamese word.
assert is_zalgo(VIETNAMESE, threshold=1), "a threshold of 1 rejects Vietnamese"
assert not is_zalgo(VIETNAMESE, threshold=3), "the default does not"
assert is_zalgo(ZALGO, threshold=3), "and still catches sprawling zalgo"

# Capping at two is what leaves Vietnamese untouched.
assert strip_zalgo(VIETNAMESE, max_marks=2) == VIETNAMESE
assert strip_zalgo(ZALGO, max_marks=2) != ZALGO

print("ok: threshold 1 rejects Vietnamese, threshold 3 does not, zalgo caught either way")

Who this hurts

Stacked marks are usually treated as a joke, which is why the damage tends to be somewhere other than where the text is displayed.

WhereWhat happensWhy depth is the wrong lever alone
Moderation and filters A banned word carrying one mark per letter is no longer that string, and still reads as it. One mark per base is depth 1, which no threshold can reject without rejecting French.
Layout Marks render outside the line box, so text climbs into the rows above and below and covers them. This is a rendering cost, not a content one: the string can be short and still wreck a page.
Storage and indexing A visually short name can be thousands of codepoints, filling a column or a token budget. Length in characters is the defence here, not depth. Cap the marks and cap the length.
Screen readers Every mark may be announced, so a short name becomes minutes of speech. An accessibility failure that a purely visual review never sees.
Copy and paste Marks travel invisibly into commit messages, tickets and logs, and break alignment there. The destination usually has no rendering budget for them at all.

The engineering conclusion the page keeps circling: cap rather than detect. strip_zalgo with a cap of 2 keeps every orthography the tool above leaves alone and removes the excess from everything else, without needing to decide whether a two-mark string was hostile. Detection is for reporting; capping is for accepting input. Restrained zalgo stays indistinguishable from real language by depth, and that is a property of the signal rather than a gap in the implementation.

Found a string this gets wrong? The confusables table grew out of exactly that kind of report. Open an issue with it.

Related tools