Vulnerability GHSA-97jj-33gv-5xf9
Summary
league/commonmark: DisallowedRawHtml bypassed when a disallowed tag name ends the raw-HTML literal
Details
Summary
The DisallowedRawHtml extension does not escape a disallowed tag when the tag name is the last thing in the raw HTML. A Markdown line containing just <script is emitted unchanged, and the next block can supply its attributes. With the shipped GFM defaults this allows stored XSS by anyone who can post Markdown.
Details
DisallowedRawHtmlRenderer escapes tags with this regex:
/<(\/?(?:title|textarea|style|xmp|iframe|noembed|noframes|script|plaintext)[\s\/>])/i
The trailing character class requires one character after the tag name. The block parser does not: RegexHelper::PARTIAL_HTMLBLOCKOPEN accepts end of line after a tag name, so <script alone opens an HTML block. Because a rendered HtmlBlock has no trailing newline, the regex has nothing to match and the < passes through.
In the browser the newline is still present, so the tag name terminates there and whatever follows becomes attributes.
This is the same filter that GHSA-4v6x-c7xx-hw9f fixed in 2.8.1. That fix widened the character class but still requires one character, so this case was not covered.
Reproduction
Render this with GithubFlavoredMarkdownConverter and default settings:
<div>
<script
<span src="/evil.js">
Output:
<div>
<script
<span src="/evil.js">
A browser parses that as <script src="/evil.js"> with a junk <span attribute, and the script runs. <iframe with <span onload="..."> works the same way and does not need a later </script> in the page.
Control: <script src="/evil.js"></script> is correctly escaped to <script src="/evil.js"></script>.
Affected versions
1.3.0 (when the extension was added) through the current release. The </style and mid-line forms are only affected as continuation lines inside an already-open HTML block.
Preconditions
html_inputisallow(the default)- The
DisallowedRawHtmlextension is active, which the GFM extension enables automatically - Untrusted users can post Markdown
Setting html_input to escape or strip fully mitigates this.
Suggested fix
Allow end of string after the tag name:
$regex = \sprintf('/<(\/?(?:%s))([\s\/>]|$)/i', \implode('|', \array_map('preg_quote', $tags)));
return \preg_replace($regex, '<$1$2', $rendered);
This escapes every bypass shape above and leaves <div>, <scripts> and <span class="a"> untouched. The existing unit test only covers tag names followed by another character, so a case for a bare tag name should be added.
Related Vulnerabilities
Other vulnerabilities affecting the same packages