Description
follow code
var_dump(grapheme_levenshtein("神", "神")); // 神 is U+FA19, but normalize to U+795E(神)
Actual result
But expected result
Because CJK Compatibility Ideographs keeps backword compatible, However, Normalized Compatible Ideograph(Unified Ideographs) users feels differents.
So I think create catch the CJK Compatible Ideographs one option in grapheme_* functions.
This behavior is ICU depends.
CJK Compatible Ideographs
U+F900-U+FA0B Pronuncation variants from KS X 1001:1998 (Korea)
U+FA0C-U+FA0D Duplicate characters from Big 5 (Taiwan)
U+FA0E-U+FA2D The IBM 32 compatibility ideographs (Japan)
U+FA2E-U+FA2F Korean compatibility ideographs (Korea)
U+FA30-U+FA6A JIS X 0213 compatibility ideographs (Japan)
U+FA30-U+UA6D ARIB compatibility ideographs (Japan)
U+FA70-U+FAD9 DPRK compatibility ideographs (North Korea)
How to catch different
- Call
ublock_getCode
- Return value of
ublock_getCode is judge to CJK Compatible Ideograph: UBLOCK_CJK_COMPATIBILITY_IDEOGRAPHS || UBLOCK_CJK_COMPATIBILITY_IDEOGRAPHS_SUPPLEMENT.
Description
follow code
Actual result
But expected result
Because CJK Compatibility Ideographs keeps backword compatible, However, Normalized Compatible Ideograph(Unified Ideographs) users feels differents.
So I think create catch the CJK Compatible Ideographs one option in grapheme_* functions.
This behavior is ICU depends.
CJK Compatible Ideographs
U+F900-U+FA0B Pronuncation variants from KS X 1001:1998 (Korea)
U+FA0C-U+FA0D Duplicate characters from Big 5 (Taiwan)
U+FA0E-U+FA2D The IBM 32 compatibility ideographs (Japan)
U+FA2E-U+FA2F Korean compatibility ideographs (Korea)
U+FA30-U+FA6A JIS X 0213 compatibility ideographs (Japan)
U+FA30-U+UA6D ARIB compatibility ideographs (Japan)
U+FA70-U+FAD9 DPRK compatibility ideographs (North Korea)
How to catch different
ublock_getCodeublock_getCodeis judge to CJK Compatible Ideograph:UBLOCK_CJK_COMPATIBILITY_IDEOGRAPHS || UBLOCK_CJK_COMPATIBILITY_IDEOGRAPHS_SUPPLEMENT.