Vulnerability GHSA-3cr3-8m4c-fpxw
Summary
Docling: METS-GBS archive member limit enforced after full member enumeration (memory exhaustion during format detection)
Details
Summary
When docling detects the input format of a gzip-compressed tar archive (METS-GBS), and later when the METS-GBS backend opens it, it calls tarfile.TarFile.getmembers(). That builds the full member list in memory before the max_member_count limit is checked. A small archive with a very large number of empty members therefore makes docling allocate memory in proportion to the member count, and the limit has no effect.
Details
docling/datamodel/document.py (format detection) and docling/backend/mets_gbs_backend.py both iterate tar.getmembers() and count members inside the loop. getmembers() reads every header up front.
Measured on Python 3.12: an archive of 1,000,000 empty members compresses to about 6.2 MB and makes getmembers() hold about 408 MB (about 66 times the input size). Format detection runs for any application/gzip input before allowed_formats is applied, so the allocation happens even when METS-GBS is not an allowed format.
This check was introduced by the fix for CVE-2026-44018 (GHSA-r3xg-rg9j-67fv) in 2.91.0. The eager enumeration it relies on has been there since METS-GBS detection was added in 2.45.0.
Impact
Memory exhaustion of the converting process, proportional to the size of the input archive. Confidentiality and integrity are not affected.
Patches
Fixed in docling 2.131.0 by #4412. Format detection and the METS-GBS backend now read archive members one at a time and stop as soon as max_member_count is exceeded, including when looking up page files.
Workarounds
Upgrade to 2.131.0. For older versions:
Reject gzip or tar inputs before they reach docling when METS-GBS support is not needed, and run conversions of untrusted input with memory limits.
Related Vulnerabilities
Other vulnerabilities affecting the same packages