Vulnerability GHSA-fpf4-vwcp-v4hp
Summary
Pydantic AI: Event loop blocked by quadratic title extraction in `web_fetch`
Details
Summary
The local web-fetch tool (web_fetch_tool, also used as the WebFetch capability's local fallback) processed responses with several steps whose running time grows quadratically with the size of certain server-controlled inputs, and ran them on the event loop: decoding the body with whichever charset the server declared, extracting the page title with a backtracking regular expression, and converting the HTML to markdown. An application that exposes this tool to untrusted prompts can be steered to fetch an attacker-controlled page of a megabyte or two that blocks the event loop for minutes, stalling every other coroutine in the process — other agent runs, other requests being served — for the duration.
This is an availability issue only. SSRF protections and the download size limit introduced in GHSA-v2xh-2vp8-57h8 are unaffected; that limit bounds how much is downloaded, not how long the response takes to process.
Details
Title extraction used a backtracking pattern over the raw response body, so a body made of repeated unterminated tag openings cost time proportional to the square of its size. The HTML-to-markdown conversion had the same shape in three of its steps: normalizing whitespace, stripping preformatted blocks, and numbering ordered lists all took time proportional to the square of a run of spaces or a list's length. All of it ran on the event loop, and the regex steps hold the interpreter lock even when moved off it, so the whole process paid for the size of a server-controlled response.
The response body was also decoded on the event loop with the codec named by the charset parameter of the response's Content-Type, looked up in Python's codec registry. That registry includes punycode, whose decoder takes time proportional to the square of its input: a response of about one megabyte labelled charset=punycode blocked the event loop for roughly half a minute, with no HTML required. The registry also includes codecs that aren't text encodings at all, such as rot_13 and base64_codec; a response labelled with one of those raised an unexpected exception out of the tool, aborting the agent run that fetched it.
Separately, the HTML-to-markdown conversion recursed once per nested element, so a page nested a few hundred elements deep raised a RecursionError out of the tool, aborting the agent run that fetched it. A JSON response nested deeper than the interpreter allows did the same. These only affect that one run.
Who Is Affected
You are affected if your application registers the local web-fetch tool (or relies on the WebFetch capability's local fallback) and exposes the agent to untrusted prompts. allowed_domains narrows the exposure to pages on those domains but does not remove it. Applications that only fetch developer-controlled URLs are not exposed to the model-chosen attack path.
Remediation
Upgrade to a patched version. The title is now found with a single linear scan, the conversion steps above run in linear time, and decoding, title extraction and conversion all run in a worker thread. A charset naming a codec that isn't a text encoding, and a page too deeply nested to convert, are reported back to the model as a failed fetch instead of aborting the run; a JSON body too deeply nested to parse is returned as plain text.
Credits
Reported privately by @BrianWillows, whose report covered the quadratic title extraction. The response decoding, the codecs that are not text encodings, and the quadratic steps in the HTML-to-markdown conversion were found while fixing it.
Related Vulnerabilities
Other vulnerabilities affecting the same packages