Vulnerability GHSA-6fqq-452j-qhrp
Summary
Pydantic AI: Concurrency-limited models can keep their slot when a streamed request ends early
Details
This issue was posted by Codex Desktop using gpt-6.1-sol on behalf of David.
Summary
Applications that wrap a model with ConcurrencyLimitedModel or limit_model_concurrency can permanently lose shared concurrency capacity when a streamed request releases its slot from a different task than the one that acquired it. This can happen when a stream ends early, and also when a stream is fully consumed using the default stream_text() debouncing.
In an application that exposes an affected streaming endpoint to network clients and shares a long-lived model limiter across requests, a client can repeatedly start a stream and disconnect. The completed requests retain their slots, eventually preventing subsequent requests that share the limiter from proceeding.
Agent-level max_concurrency and non-streaming model requests are not affected by this defect.
Details
The built-in limiter uses anyio.CapacityLimiter, which associates each acquired slot with its borrowing task. Pydantic AI's streaming lifecycle can acquire the slot on the task consuming the stream and run cleanup on another internal task. The limiter rejects that release, so the slot remains occupied even after the request has ended. Cleanup can raise a RuntimeError; a later request on the borrowing task can also fail because that task still holds a slot.
Early termination includes stopping iteration, a consumer exception, and cancellation. Fully consuming stream_text() with its default debounce_by=0.1 can also reach the cross-task release path. Fully consumed streams must therefore not be assumed safe.
Mitigation
Upgrade to a patched release of pydantic-ai or pydantic-ai-slim. If you cannot upgrade yet, use the agent-level max_concurrency setting instead of a concurrency-limited model, or avoid streaming runs through a concurrency-limited model.
Related Vulnerabilities
Other vulnerabilities affecting the same packages