Rewrite Rules: Fix rewrite rules breaking for taxonomies with URL-encoded Unicode slugs - #12863
Rewrite Rules: Fix rewrite rules breaking for taxonomies with URL-encoded Unicode slugs#12863SainathPoojary wants to merge 2 commits into
Conversation
…ture group misalignment
Test using WordPress PlaygroundThe changes in this pull request can previewed and tested using a WordPress Playground instance. WordPress Playground is an experimental project that creates a full WordPress instance entirely within the browser. Some things to be aware of
For more details about these limitations and more, check out the Limitations page in the WordPress Playground documentation. |
|
The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the Core Committers: Use this line as a base for the props when committing in SVN: To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook. |
irozum
left a comment
There was a problem hiding this comment.
Solid fix for a real 9-year-old bug: taxonomies/post types registered with a percent-encoded Unicode slug (WooCommerce does this for cyrillic attribute names) end up with literal %D0%A1... bytes in the permastruct, which generate_rewrite_rules() misreads as tag delimiters, shifting every $matches[n] capture. Decoding in add_permastruct() is the right choke point since it's the single entry for taxonomies, post types, and any plugin's own call, matching the reporter's original diagnosis. It's also a nice improvement over the old 2016 patch on the ticket, which called plain urldecode() — that would convert literal + in a slug into a space; rawurldecode() avoids that, and scoping the regex to uppercase hex avoids colliding with any lowercase built-in tag like %category%.
Checked out the branch and ran the full rewrite group (1388 tests, no regressions) plus lint/PHPStan on the two changed files, all clean. The new test correctly asserts both the decoded struct and that the generated rule maps to $matches[1] rather than a shifted index — the exact symptom from the ticket's screenshots.
Non-blocking: a plugin-registered tag whose name happens to be 2+ uppercase hex letters (e.g. a hypothetical %AB%) would now get silently decoded instead of matched — no core tag is shaped that way, but requiring the decoded bytes to be valid UTF-8 before substituting would close that edge case if anyone wants extra safety.
This PR intercepts the permastruct inside
WP_Rewrite::add_permastruct()and decodes the URL-encoded slug back into standard UTF-8 text before saving it.To ensure we don't accidentally corrupt real rewrite tags (where
%cain%category%could be decoded by a naiveurldecode), we use apreg_replace_callbackthat strictly decodes only uppercase hex strings ([0-9A-F]). WordPress's URL encoding always uses uppercase, while built-in rewrite tags use lowercase, making this a safe and robust fix.Trac ticket: #41791
Use of AI Tools
AI assistance: Yes
Tool(s): GitHub Copilot
Model(s): Gemini, Claude
Used for: Checking for potential edge cases, drafting the PR description, and assisting with local code review. The final implementation and testing were written and executed manually by me.
This Pull Request is for code review only. Please keep all other discussion in the Trac ticket. Do not merge this Pull Request. See GitHub Pull Requests for Code Review in the Core Handbook for more details.