Vulnerability GHSA-5gfj-9q3v-qfp3
Summary
music-metadata: EBML parser trusts element lengths, allowing memory exhaustion or process abort
Details
Summary
The Matroska/WebM EBML parser decodes attacker-controlled variable-length integer (VINT) element lengths without first validating that the declared element fits within its enclosing container or the available input. Leaf lengths are then used directly for string-token and Uint8Array allocations.
A very small crafted .webm, .mkv, or .mka file can therefore cause a disproportionately large allocation, an out-of-memory denial of service, or—in one demonstrated parseFile case—an uncatchable V8 fatal abort.
Impact
Applications that parse untrusted Matroska/WebM media can be denied service. Depending on the input, parser API, and Node.js/V8 version, the demonstrated outcomes include:
- a 34-byte input causing an approximately 128 MiB allocation before end-of-input;
- a 70-byte input causing a 500 MiB allocation through a binary EBML leaf, with concurrent parses multiplying memory consumption;
- a 48-byte file causing an uncatchable V8 fatal abort through
parseFilewhen a leaf declares a length of 32 GiB.
The fatal-abort case was reproduced on music-metadata 11.15.0 with Node.js 26.7.0. It was deterministic across three runs and exited with code 133 despite a surrounding try/catch. The same reporter did not reproduce that fatal abort through parseBuffer on Node.js 20 through 25; those configurations can still encounter allocation attempts or catchable errors. The exact failure mode is therefore runtime- and tokenizer-dependent, but the underlying validation flaw is shared.
The impact is limited to availability; no confidentiality or integrity impact has been demonstrated.
Root cause
EbmlIterator.readElement() decodes an element length from an EBML VINT and converts it to a JavaScript Number:
private async readElement(): Promise<IHeader> {
const id = await this.readVintData(this.ebmlMaxIDLength);
const lenField = await this.readVintData(this.ebmlMaxSizeLength);
lenField[0] ^= 0x80 >> (lenField.length - 1);
return {
id: readUIntBE(id, id.length),
len: readUIntBE(lenField, lenField.length)
};
}
Before the fix, that value was not checked against the parent boundary or known file size before reaching the leaf readers:
private async readString(e: IHeader): Promise<string> {
const rawString = await this.tokenizer.readToken(new StringType(e.len, 'utf-8'));
return rawString.replace(/\x00.*$/g, '');
}
private async readBuffer(e: IHeader): Promise<Uint8Array> {
const buf = new Uint8Array(e.len);
await this.tokenizer.readBuffer(buf);
return buf;
}
An EBML size VINT can encode values approaching 2^56. A crafted file can select a known string, unsigned-integer, binary, or UID element and route its declared size into an allocation before the tokenizer discovers that the input is truncated. The declared length also participates in container-boundary arithmetic.
For lengths around 32 GiB, one reporter observed V8 aborting on its external-memory accounting check:
Fatal error: Check failed: change_in_bytes < kMaxReasonableBytes
process exit code 133
This is a V8 process abort rather than a JavaScript exception, so application-level error handling cannot catch it.
Attack scenarios
The issue is reachable through ordinary metadata parsing of attacker-controlled Matroska/WebM files. Reported proofs of concept exercised string, unsigned-integer, binary, and UID leaves. Depending on the tokenizer and runtime, affected entry points include parseBuffer, parseFile, parseStream, parseBlob, and parseWebStream.
One minimal buffer-based proof of concept uses a docType string with an oversized eight-byte VINT:
import { parseBuffer } from 'music-metadata';
const payload = Uint8Array.from([
0x1a, 0x45, 0xdf, 0xa3, 0x8a,
0x42, 0x82,
0x01, 0xff, 0xff, 0xff, 0xff, 0xff, 0xff, 0xff
]);
await parseBuffer(payload, { mimeType: 'video/webm' });
Another proof of concept used a valid WebM EBML header followed by an ebmlReadVersion leaf whose VINT declared 0x800000000 bytes (32 GiB), while supplying only one payload byte. Through parseFile, this reached the fatal V8 abort described above.
Fix
PR #2735 validates EBML leaf lengths before decoding or allocating. It rejects lengths that are unsafe, exceed the enclosing container, or exceed the known remaining file data. For streams whose total size is unknown, large leaves are read incrementally so truncated input fails before one large allocation is attempted. Nested-container boundaries are also constrained by ancestor and file boundaries.
At the time this advisory was updated, PR #2735 was still open and no patched release had been published.
Credits
This advisory consolidates independent reports of the same EBML length-validation flaw:
- @bg0d-glitch reported oversized VINT lengths reaching string and binary allocation paths.
- @Zwique demonstrated approximately 128 MiB of allocation from a 34-byte EBML input as part of a broader parser review.
- @offset demonstrated a 500 MiB allocation from a 70-byte binary-leaf input and the effect of concurrent parses.
- @arpitjain099 demonstrated the 48-byte
parseFilecase that triggers an uncatchable V8 fatal abort.
The reports were consolidated because they share the same vulnerable component, missing validation, allocation sinks, and fix.
Related Vulnerabilities
Other vulnerabilities affecting the same packages