Security Deep-Dive · 2026-10-01
SVG tspan dx Drift: The Per-Word Consent Text Displacement Attack
The SVG tspan element's dx attribute is relative — it adds to the running text cursor position inherited from all preceding tspans. A chain of moderate per-word offsets can quietly push the legally significant portion of a consent sentence past the SVG viewport edge. The DOM string via textContent is complete. The visible text reads naturally. Only by accumulating the cursor position across all tspans in document order does an auditor discover that "all file system and network access" has drifted off-screen.
How SVG text positioning actually works
HTML text flows automatically within block dimensions. SVG text does not — every character's position is determined by explicit coordinate attributes. The <text> element sets an initial x and y anchor. Each nested <tspan> can override the current position in two ways:
| Attribute | Semantics | Attack surface |
|---|---|---|
x, y |
Absolute: sets the text cursor to an exact SVG coordinate, discarding accumulated position | Any value outside the SVG viewBox moves the word off-screen instantly and unconditionally |
dx, dy |
Relative: shifts the cursor from its current position by this delta. Applies after the previous tspan's own positioning and text width have advanced the cursor | Each tspan's dx adds to all previous dx values — cumulative drift. No single dx value looks suspicious |
The critical property of dx: it is additive. If tspan 1 has dx="80" and tspan 2 has dx="80", the second tspan's word starts 160px right of where tspan 1 started — not 80px right of where tspan 2's word would normally begin. The cursor position going into each tspan equals the sum of all preceding tspans' dx values plus the text widths of all preceding tspans' content.
The core asymmetry
An auditor who reads dx="80" on a single tspan might interpret this as 80px of letter-spacing — suspicious but marginal. The attack does not depend on any single tspan being detectable. It depends on the accumulation of many individually defensible values. At dx=80 per tspan across 5 words starting at x=10, the fifth word's left edge is at approximately x = 10 + (width("I") + 80) + (width("authorize") + 80) + (width("all") + 80) + (width("requested") + 80) + (width("permissions") + 80) — well past a 500px viewBox. No individual span triggered an alert. The collective result is complete clipping of the latter half of the consent sentence.
Attack 1: Calibrated per-word drift
Progressive dx offsets clip "all file system and network access" past the viewport
A consent sentence is split across five tspans: "I", "authorize", "all", "file system and network", "access." The first two tspans have no dx or small dx values — "I authorize" is visible at normal position. Each subsequent tspan has dx=90, accumulated from the previous. By the third tspan the cursor has drifted ~200px right of the natural line; by the fifth it is past the 400px viewBox edge. "all file system and network access" is clipped. The user sees only "I authorize". The full consent string is intact in the DOM.
<svg viewBox="0 0 400 60" width="400" height="60" overflow="hidden">
<text x="10" y="30" font-size="14" fill="#1a1a1a">
<tspan>I </tspan>
<tspan>authorize </tspan>
<!-- dx=90 on each remaining tspan -- cumulative drift -->
<tspan dx="90">all </tspan>
<tspan dx="90">file system and network </tspan>
<tspan dx="90">access.</tspan>
</text>
</svg>
<!-- cursor after "authorize ": ~x=80 -->
<!-- cursor after "all " tspan: ~80 + 90 + width("all ") ≈ 203 -->
<!-- cursor after "file..." tspan: ~203 + 90 + width("file system and network ") ≈ 473 -->
<!-- 473 > viewBox width 400 — "file system and network" is fully off-screen -->
<!-- textContent: "I authorize all file system and network access." -- complete -->
A per-tspan dx check that only flags values above a threshold (e.g., dx > 200) misses this attack entirely — no individual dx is above 90. Only accumulating the running cursor position reveals the overflow.
The attack is particularly effective when the beginning of the sentence — "I authorize" — is designed to sound complete in isolation. A user who glimpses "I authorize" in a consent form may not notice the missing continuation, especially if the SVG element has a fixed visual boundary that makes the clipping invisible.
Attack 2: The absolute-x reset trick
tspan x=5000 jumps a key word off-screen; subsequent tspan x=28 resets the visible line
A single critical consent word — "irrevocably" or "authorize" — is placed in a tspan with an absolute x="5000". This moves only that word off-screen. The next tspan uses absolute x="28" to reset the cursor to a normal position, so the remaining visible words continue from a natural starting point. The resulting visible text reads fluently — "By clicking Agree you all requested permissions." — with no visual gap or truncation indicating that a word was removed.
<text x="10" y="30" font-size="14" fill="#1a1a1a"> By clicking Agree you <!-- "irrevocably" sent to x=5000 -- off-screen --> <tspan x="5000"> irrevocably</tspan> <!-- x=28 resets cursor to start of line -- gap is invisible --> <tspan x="28"> authorize all requested permissions.</tspan> </text> <!-- Visible: "By clicking Agree you authorize all requested permissions." --> <!-- DOM: "By clicking Agree you irrevocably authorize all requested permissions." -->
The reset tspan at x="28" creates the illusion that the sentence is complete. The visible text is grammatically correct and appears unmodified. Only comparing the visible reconstruction against the full textContent reveals the missing adverb. This matters legally: "irrevocably" changes the recoverability of the permission grant. Its removal from what the user reads is not cosmetic.
Why the reset matters: Without the x="28" reset tspan, the remaining words would continue from the cursor position of x=5000 + width("irrevocably") — which is also off-screen. Both the attack word and all subsequent words would be invisible. The reset tspan is what makes the attack practical: the attacker removes exactly one word while preserving the appearance of an intact consent sentence. Detecting the attack requires finding the off-screen tspan, not finding a truncated visible sentence.
Attack 3: Split-word attack via character-level tspans
Word-splitting positions part of a consent word off-screen, making the partial word appear complete
A key consent word is split across two tspans. The first tspan contains the first few characters (e.g., "auth"). The second tspan containing the rest of the word ("orize all permissions") is given a large dx or x that pushes it off-screen. The visible text shows the partial fragment "auth" as the last visible token in the sentence — ambiguous enough that the user might not notice the word is incomplete. The full word is intact in the DOM string.
<text x="10" y="30" font-size="14" fill="#1a1a1a"> By clicking Agree you <!-- Split "authorize all permissions" -- first 4 chars visible --> <tspan> auth</tspan> <!-- Rest of word + key scope sent off-screen via large dx --> <tspan dx="2000">orize all permissions including file and network access.</tspan> </text> <!-- Visible: "By clicking Agree you auth" --> <!-- DOM: "By clicking Agree you authorize all permissions including file and network access." -->
This variant is harder to detect than word-level attacks because "auth" is a valid string fragment, and simple word-boundary analysis of visible text might not flag an incomplete word as a missing consent term. Detection requires reconstructing complete words from consecutive tspans and checking for word completeness at viewBox boundaries.
SVG overflow interaction
The effectiveness of tspan drift attacks depends on where the overflow boundary is enforced. SVG has three relevant overflow behaviors:
| Context | overflow value | Effect on drift attack |
|---|---|---|
SVG element itself (<svg overflow="hidden">) |
hidden (default for inline SVG) |
Content outside the viewBox is clipped to the SVG element boundary. Drifted text is invisible — attack succeeds |
SVG element (<svg overflow="visible">) |
visible |
Content outside the viewBox overflows into the surrounding HTML. The off-screen words may render in an unexpected position in the page layout — visible but misplaced. Attack partially fails but consent is still confusingly positioned |
| Ancestor HTML element | overflow: hidden on a CSS ancestor |
Even if the SVG itself has overflow=visible, a CSS clip on an ancestor element clips the overflowed SVG content. The effective boundary is the nearest CSS overflow:hidden ancestor — not the SVG viewBox |
Nested SVG (<svg> inside <svg>) |
Inner SVG defaults to overflow="hidden" |
Text that drifts past the inner SVG's viewBox is clipped by the inner SVG, even if the outer SVG has overflow=visible. The attacker places the consent text in a nested SVG to guarantee clipping |
The most reliable attack configuration is to place the consent text SVG inside a CSS-clipped ancestor with overflow: hidden. This ensures that even if the browser renders the SVG with overflow="visible", the CSS container clips the overflowed text. The attacker does not need to set overflow="hidden" on the SVG element directly — the CSS clip can be on any ancestor, including the consent form container.
An auditor checking only the SVG element's overflow attribute misses the CSS clip scenario. SkillAudit evaluates both the SVG overflow attribute and the computed CSS overflow of all ancestor elements to determine the effective clipping boundary before assessing tspan drift risk.
Why textContent is not a consent verification method
The invariant that element.textContent returns the same string regardless of tspan positioning is the structural cause of this entire attack class. SVG text elements do not have a concept of "visible text" at the DOM API level — there is only text content, which is all text nodes concatenated regardless of position. Three properties make tspan drift undetectable via textContent alone:
Position independence
textContent returns the concatenation of all text node descendants regardless of their x, y, dx, dy positions. A tspan at x=5000 contributes its text to textContent identically to a tspan at x=10.
Critical gap
Overflow independence
textContent is unaffected by SVG viewBox dimensions or CSS overflow clipping. Text that is geometrically outside the viewport contributes to textContent as if it were visible.
Critical gap
Presentation independence
textContent ignores opacity, fill-opacity, visibility, and filter state. A tspan with fill-opacity=0 contributes its text to textContent as if it were fully visible.
High gap
Any consent audit system that uses textContent or innerText as its primary or sole consent verification method is bypassed by all three of these properties. The correct comparison is between (a) the text visible to the user in the rendered SVG, computed by accumulating only in-bounds, in-opacity, non-filtered tspans, and (b) the expected consent string.
The cursor-accumulation detection algorithm
Reliable detection of tspan drift attacks requires simulating the SVG text cursor: walking all tspan descendants in document order, accumulating the cursor position at each step, and flagging any tspan whose rendered position places it outside the SVG viewBox.
function detectTspanDrift(textEl, svgEl) {
const viewBox = svgEl.viewBox.baseVal;
const vbRight = viewBox.x + viewBox.width;
const vbBottom = viewBox.y + viewBox.height;
const issues = [];
// Start cursor at the text element's x/y
let cursorX = textEl.x.baseVal.length ? textEl.x.baseVal[0].value : 0;
let cursorY = textEl.y.baseVal.length ? textEl.y.baseVal[0].value : 0;
for (const tspan of textEl.querySelectorAll('tspan')) {
// Absolute x/y reset the cursor unconditionally
if (tspan.x.baseVal.length) {
cursorX = tspan.x.baseVal[0].value;
}
if (tspan.y.baseVal.length) {
cursorY = tspan.y.baseVal[0].value;
}
// Relative dx/dy shift from current cursor position
if (tspan.dx.baseVal.length) {
cursorX += tspan.dx.baseVal[0].value;
}
if (tspan.dy.baseVal.length) {
cursorY += tspan.dy.baseVal[0].value;
}
// Is the cursor currently outside the viewBox?
const outOfBounds = cursorX > vbRight || cursorX < viewBox.x
|| cursorY > vbBottom || cursorY < viewBox.y;
if (outOfBounds) {
issues.push({
tspan,
text: tspan.textContent,
cursorX,
cursorY,
severity: 'CRITICAL',
finding: 'SA-TSPAN-DRIFT: cursor position outside SVG viewBox'
});
}
// Advance cursor by the rendered text width of this tspan
// (approximate: use getComputedTextLength() in a live DOM)
const textWidth = tspan.getComputedTextLength
? tspan.getComputedTextLength()
: tspan.textContent.length * 8; // rough fallback for static analysis
cursorX += textWidth;
}
return issues;
}
The algorithm's key insight is that it evaluates cursor position before advancing by the tspan's text width — this is when the tspan's first character is rendered, and if that position is outside the viewBox the tspan's content is clipped. After finding the position, it advances cursorX by the computed text length to set the correct starting position for the next tspan.
In a static analysis context (without a live DOM), getComputedTextLength() is unavailable. SkillAudit's static engine uses font-size × approximate character width as a conservative estimate, then flags when the estimated cursor would reach 80% of the viewBox width — a threshold that catches calibrated drift attacks while remaining below the false-positive rate of a 100% threshold.
Text reconstruction: building the visible string
The complement to cursor accumulation is visible text reconstruction: building a string containing only the words that the user can actually see, then comparing it against the expected consent string.
function reconstructVisibleText(textEl, svgEl) {
const viewBox = svgEl.viewBox.baseVal;
const vbRight = viewBox.x + viewBox.width;
let cursorX = textEl.x.baseVal.length ? textEl.x.baseVal[0].value : 0;
let visibleParts = [];
for (const tspan of textEl.querySelectorAll('tspan')) {
if (tspan.x.baseVal.length) cursorX = tspan.x.baseVal[0].value;
if (tspan.dx.baseVal.length) cursorX += tspan.dx.baseVal[0].value;
// Include only tspans whose start cursor is within the viewBox
// AND whose fill-opacity is above perceptibility threshold
const fillOpacity = parseFloat(getComputedStyle(tspan).fillOpacity || '1');
if (cursorX <= vbRight && cursorX >= viewBox.x && fillOpacity > 0.1) {
visibleParts.push(tspan.textContent);
}
const textWidth = tspan.getComputedTextLength
? tspan.getComputedTextLength()
: tspan.textContent.length * 8;
cursorX += textWidth;
}
return visibleParts.join('').replace(/\s+/g, ' ').trim();
}
function checkConsentVisibility(textEl, svgEl, expectedConsent) {
const visible = reconstructVisibleText(textEl, svgEl);
const full = textEl.textContent.replace(/\s+/g, ' ').trim();
if (visible !== full) {
return {
attack: true,
visible,
full,
missing: full.replace(visible, '[HIDDEN]')
};
}
return { attack: false };
}
The visible reconstruction function has three checks per tspan: cursor position within viewBox, fill-opacity above a perceptibility threshold, and (not shown here but included in SkillAudit's full implementation) filter absence. Only tspans that pass all three checks contribute to the visible string. The comparison between reconstructVisibleText() and textContent is the definitive consent integrity check.
Remediation and detection table
| Attack variant | SA ID | Detection method | Severity |
|---|---|---|---|
| Cumulative dx drift pushes latter-half consent words off-screen | SA-TSPAN-DX-001 | Cursor accumulation: running cursorX > viewBox.width flags overflow. No single dx is above threshold — only cumulative position detects it | Critical |
| Absolute x=5000 on key word + x=28 reset on continuation | SA-TSPAN-DX-002 | Per-tspan absolute x check: any tspan x value outside the viewBox is flagged regardless of cursor position. Compare visible reconstruction against textContent to find removed words | Critical |
| Character-level split with dx=2000 on the completion tspan | SA-TSPAN-DX-003 | Same cursor accumulation + visible reconstruction. Word-boundary analysis on the visible string catches partial words at the clip boundary | High |
| CSS ancestor clip amplifying SVG overflow=visible | SA-TSPAN-DX-004 | Check computed CSS overflow on all ancestor elements. If any ancestor has overflow:hidden or overflow:clip, use the ancestor's bounding rect as the effective clipping boundary | High |
From a remediation standpoint, the reliable fix is to avoid SVG text entirely for consent disclosure language. HTML text is not subject to coordinate-based clipping and renders predictably across all browsers and viewport sizes. When SVG text is used for branding or layout reasons, the consent sentence itself should be in a plain HTML element positioned over or adjacent to the SVG, ensuring that its visibility is governed by HTML layout rules rather than SVG coordinate geometry.
SkillAudit audits every consent text element in your MCP server's rendered output. For SVG text and tspan trees, it runs cursor accumulation, visible text reconstruction, per-tspan fill-opacity, and filter analysis. The visible reconstruction is compared against the expected consent scope to identify any words removed from what the user sees. Paste your GitHub URL for a free audit.
Related findings
- SA-TSPAN: Per-tspan fill-opacity and filter attacks on individual consent keywords
- SA-FO: SVG foreignObject zero-width viewport clipping embedded consent HTML
- SVG SMIL Animation as a Consent Timing Attack — animating tspan position at interaction time
- SA-TPATH: SVG textPath routing consent text along an off-screen path