將易混淆字元正規化為安全文字
正規化易混淆字元,就是把每個仿冒字元替換成它所模仿的一般字元,讓看起來相同的兩個字串在比較時也相等。在比較使用者名稱、比對封鎖清單或去除重複識別碼之前,這一步特別有用。
實例解析
- 輸入
- Cоnfig file: аdmin2, Noёl, café
- 偵測到的文字系統
- 拉丁字母 (Latn), 西里爾文字 (Cyrl)
- 已標記字元
- 7
| 位置 | 字元 | 碼位 | 文字系統 | 替換為 | 規則 |
|---|---|---|---|---|---|
| 0 | C | U+FF23 | 拉丁字母 | C | NFKC 正規化 |
| 1 | о | U+043E | 西里爾文字 | o | 易混淆字元對照表 |
| 7 | fi | U+FB01 | 拉丁字母 | fi | NFKC 正規化 |
| 10 | : | U+FF1A | 一般文字 | : | 易混淆字元對照表 |
| 12 | а | U+0430 | 西里爾文字 | a | 易混淆字元對照表 |
| 17 | 2 | U+FF12 | 一般文字 | 2 | 易混淆字元對照表 |
| 22 | ё | U+0451 | 西里爾文字 | ë | 易混淆字元對照表 |
- 保留可讀的 Unicode
- Config file: admin2, Noël, café
- 嚴格的 ASCII 後備
- Config file: admin2, Noel, café
運作原理
- 首先依內建對照表替換已知的仿冒字元。對照表以外的字元則以 NFKC 正規化,把全形形式、合字與其他相容字元轉換成標準的對應字元。
- 「保留可讀的 Unicode」會在對照表定義了帶重音字母時保留它,例如西里爾字母 ё → ë。「嚴格的 ASCII 後備」則改用一般 ASCII 字母(ё → e)。
- 既不在對照表中、也不會被 NFKC 改變的字母會原樣保留,因此 café 在兩種模式下都保有重音。正規化後的文字只是比較用的鍵值,而非安全判定:請同時保存原文,並檢查被標記的字元。
轉換是盡力而為:映射的易混淆項和 NFKC 折疊是確定性的,但某些合法的 Unicode 不會被標記。
您的文字
貼上或鍵入 — 結果會在您鍵入時更新(對於長輸入會稍微去抖)。
已掃描 30 個字元
7 個可疑項目
嚴格的 ASCII 後備
原文(可疑字元已標記)
原始視圖中的可疑字元帶有下劃線並標記為“可疑”。除了突出顏色。
suspicious character Csuspicious character оnfig suspicious character filesuspicious character : suspicious character аdminsuspicious character 2, Nosuspicious character ёl, café
清理後的輸出
字元分析
| 索引(從0開始) | 原字元 | 替換為 | 代碼點 | 原因 |
|---|---|---|---|---|
| 0 | C | C | U+FF23 | NFKC normalization changed this character (compatibility or width folding). |
| 1 | о | o | U+043E | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 2 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 3 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 4 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 5 | g | g | U+0067 | Not flagged as a confusable or compatibility character. |
| 6 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 7 | fi | fi | U+FB01 | NFKC normalization changed this character (compatibility or width folding). |
| 8 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 9 | e | e | U+0065 | Not flagged as a confusable or compatibility character. |
| 10 | : | : | U+FF1A | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 11 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 12 | а | a | U+0430 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 13 | d | d | U+0064 | Not flagged as a confusable or compatibility character. |
| 14 | m | m | U+006D | Not flagged as a confusable or compatibility character. |
| 15 | i | i | U+0069 | Not flagged as a confusable or compatibility character. |
| 16 | n | n | U+006E | Not flagged as a confusable or compatibility character. |
| 17 | 2 | 2 | U+FF12 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 18 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 19 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 20 | N | N | U+004E | Not flagged as a confusable or compatibility character. |
| 21 | o | o | U+006F | Not flagged as a confusable or compatibility character. |
| 22 | ё | e | U+0451 | Non-ASCII confusable mapped to a safer Latin ASCII equivalent. |
| 23 | l | l | U+006C | Not flagged as a confusable or compatibility character. |
| 24 | , | , | U+002C | Not flagged as a confusable or compatibility character. |
| 25 | U+0020 | Not flagged as a confusable or compatibility character. | ||
| 26 | c | c | U+0063 | Not flagged as a confusable or compatibility character. |
| 27 | a | a | U+0061 | Not flagged as a confusable or compatibility character. |
| 28 | f | f | U+0066 | Not flagged as a confusable or compatibility character. |
| 29 | é | é | U+00E9 | Not flagged as a confusable or compatibility character. |