One of the operation's desks publishes short video clips, and no caption leaves it until a guard in the publishing code has read it for invisible characters: the no-break space, the soft hyphen, the zero-width joiner, the byte-order mark, 54 code points in all. The guard is one statement of Python, a regular expression character class, and until 6 October that class was typed as the characters themselves. On screen most of it was blank space between square brackets, broken by three hyphens.

The operation also has a rule that everything it ships goes through a strip pass last, a script that deletes invisible and format Unicode. The rule names code on purpose, on the reasoning that a zero-width space inside a JSON file or a shell command is a bug wherever it lands. The same pass sits behind the headline veto in the post on ESPN's feed terms. During the 6 October rebuild of the desks' routines it ran over the clip engine's source and rewrote the guard. The fix, committed that evening, spelled the class out in \u escapes and kept the pattern identical.

A range is two endpoints and a hyphen

Replayed this week against the pre-fix statement from the engine's git history, the strip deletes 11 characters outright, the soft hyphen, the right-to-left mark and the byte-order mark among them. It swaps five more, the no-break space and four wide spaces, for an ordinary space. The statement drops from 79 characters to 68.

Inside a character class, a deleted character can take a range's endpoint with it. The class held three ranges, and U+2000 to U+200F was the first of them. The strip turned the U+2000 endpoint into a plain space and deleted U+200F, and the hyphen stayed where it was, now joining the space to the next character that survived, U+2028. So the class ran from the space character to U+2028, which takes in every ASCII letter, digit and punctuation mark, accented Latin, and the Greek, Cyrillic, Hebrew and Arabic alphabets. Counted across the Basic Multilingual Plane, the original guard matched 54 code points. The stripped one matched 8,203, and a caption reading “hello” would have been refused.

The guard's own test still passed

The engine has a refusal test for this guard, which hands it the word “zero”, a zero-width space, then “width”, and expects the caption to be refused. The test survived the strip because it spells the character as \u200b, an escape the strip has no reason to touch. Run against a scratch copy of the engine with the stripped statement put back, it passes. A guard that refuses everything refuses the zero-width sample along with the rest.

The other tests are what break. In the same module 39 of 141 fail on the scratch copy, against one with the current statement, and 29 of the 38 new failures carry the refusal message itself, an ordinary caption turned away for invisible characters. This version fails loudly, and the next run of the suite would have caught it. A class written as a plain list would have shrunk instead, losing members one by one, and the refusal test would go red only if its own sample happened to be among the losses. The suite had already shown how far a green run can drift from the thing it checks.

A compiler gave the same advice in 2021

The wider industry met invisible characters in source code as a security problem. Nicholas Boucher and Ross Anderson of the University of Cambridge showed that Unicode bidirectional control characters can make code display one way to a reviewer and compile another way, in a paper titled “Trojan Source: Invisible Vulnerabilities”, first posted on 30 October 2021. KrebsOnSecurity reported it on 1 November and quoted Anderson on code that “appears innocuous to a human reviewer”. The flaw is tracked as CVE-2021-42574, in the US National Vulnerability Database and in Python's PEP 672, which credits Boucher and Anderson's report as its prompt.

Rust shipped version 1.56.1 the same day with two lints that refuse to compile code carrying those code points in string literals or comments. Its security advisory told developers with a legitimate use for them to replace each one with its escape sequence, and the compiler's lint reference still shows the error suggesting \u{202e} in place of the raw character.

The desk's case ran the other way round from Trojan Source. Nobody was hiding anything. A cleaner doing its job could not tell the characters a guard was looking for from the characters it existed to remove, because in the file they were the same bytes. Written as escapes, they stop being the same bytes. A reviewer can read them in a diff, and the strip passes over them.

What to change in your own code

Spell every invisible or format character in source as an escape. That covers regular expressions, string constants, test fixtures and config files, and it holds even when no cleaner runs over your repository today, since the next editor, formatter or copy-paste may do the cleaning for you. Treat a character class as grammar, since any tool that deletes or substitutes characters can move a range's endpoints and leave its hyphen standing. Give every guard a test that clean input gets through, because a refusal test cannot tell a working guard from one that refuses everything. And when a cleaning pass runs over code, run the suite afterwards, on the cleaned bytes.

The guard is still one statement of Python. It now reads as a row of backslashes and hex digits, uglier than the blank space it replaced, and the strip passes over all 174 characters of it without changing a byte.