Links, images and references
Destinations, titles, reference and collapsed forms, autolinks and where a definition may live.
Generated from resources/examples/edge-cases.md and resources/examples/core.md - edit the cases there, not here. Each case links the conformance fixture it produces.
Nested brackets in link text
1 conformance fixture
Link, image, and span text may contain balanced nested brackets; the closing ] is found by balance, not at the first inner ].
[a [b] c](/u)<p><a href="/u">a [b] c</a></p>An image's alt text closes where a link's text closes
8 conformance fixtures
316-an-image-s-alt-text-closes-where-a-link-s-text-closes316-an-image-s-alt-text-closes-where-a-link-s-text-closes-2316-an-image-s-alt-text-closes-where-a-link-s-text-closes-3316-an-image-s-alt-text-closes-where-a-link-s-text-closes-4316-an-image-s-alt-text-closes-where-a-link-s-text-closes-5316-an-image-s-alt-text-closes-where-a-link-s-text-closes-6316-an-image-s-alt-text-closes-where-a-link-s-text-closes-7316-an-image-s-alt-text-closes-where-a-link-s-text-closes-8
An image has the same three forms as a link, and PART 3 says only the leading ! and the <img src> output differ. The bracketed run is not one of the things that differ: the alt text ends at the MATCHING ], by the same balanced, escape- and literal-span-aware scan that closes link text. So an alt text may hold a bracket, at any depth, in every form and in every host that re-parses the run.
What the run does NOT share with link text is its content model. alt is an HTML attribute, so nothing inside is inline-parsed: an escape stays as authored and a backtick run stays a backtick run.
a ![t[z]][r] b
[r]: /i.png<p>a <img src="/i.png" alt="t[z]"> b</p>a ![t[z]](/i.png) b<p>a <img src="/i.png" alt="t[z]"> b</p>Nesting is unbounded, and a trailing attribute block still attaches to the resolved image.
a ![t[z[q]]][r]{.c} b
[r]: /i.png<p>a <img src="/i.png" alt="t[z[q]]" class="c"> b</p>An escape and a code span both keep their ] out of the close, and both reach the attribute as the bytes the author wrote.
a ![t\]z](/i.png) b<p>a <img src="/i.png" alt="t\]z"> b</p>a ![t`]`z](/i.png) b<p>a <img src="/i.png" alt="t`]`z"> b</p>An UNBALANCED ] still closes the run where it stands, which leaves a reference tail with nothing in front of it and the whole line literal.
a ![t]z][r] b
[r]: /i.png<p>a ![t]z][r] b</p>The hosts that re-read the run agree with the inline pass. A paragraph whose whole content is one reference image is captionable however its alt text is spelled, and an image inside link text is still an image.
![t[z]][r]
^ cap
[r]: /i.png<figure>
<img src="/i.png" alt="t[z]">
<figcaption>cap</figcaption>
</figure>a [x ![t[z]][r] y](/u) b
[r]: /i.png<p>a <a href="/u">x <img src="/i.png" alt="t[z]"> y</a> b</p>An editorial comment's bracket is content, not the close
4 conformance fixtures
The close scan skips the interior of every span whose content is LITERAL: inline code, the !-prefixed inline literal, and an editorial comment. The test is the property, not the list - a ] inside any of them cannot be escaped, so ending the run there would leave no spelling that keeps both the construct and the author's text.
A comment with no bracket in it never had a say in where the run ends.
a [t{# n #}z](/u) b<p>a <a href="/u">t<span class="critic-comment"> n </span>z</a> b</p>One with a bracket in it does not either, in a link, in a span, in an inline note, or in an alt text.
a [t{# ] #}z]{.c} b<p>a <span class="c">t<span class="critic-comment"> ] </span>z</span> b</p>a ^[t{# ] #}z] b<p>a <a id="fnref1" href="#fn1" role="doc-noteref"><sup>1</sup></a> b</p>
<section role="doc-endnotes" aria-label="Footnotes">
<hr>
<ol>
<li id="fn1">
<p>t<span class="critic-comment"> ] </span>z<a href="#fnref1" role="doc-backlink" aria-label="Back to reference">↩</a></p>
</li>
</ol>
</section>a ![t{# ] #}z](/i.png) b<p>a <img src="/i.png" alt="t{# ] #}z"> b</p>Reference labels are case-sensitive
1 conformance fixture
Reference labels are matched case-sensitively (no case normalization). A label whose case does not match its definition stays unresolved and renders literally, like any other unresolved reference.
[Text][REF]
[ref]: /u<p>[Text][REF]</p>Link destination parentheses balance
4 conformance fixtures
A ( inside a (...) destination is matched against a later ), so the destination ends at the first ) that has no opener left to pair with. URLs carrying parentheses -- Wikipedia and MDN produce them constantly -- are therefore written plainly, with no escape and no second spelling. Djot and CommonMark both balance destination parentheses the same way.
[x](http://a/b(c))<p><a href="http://a/b(c)">x</a></p>Nesting is tracked to any depth, and a ) with nothing to close ends the destination -- the rest stays literal text.
[x](a(b(c))d) and [y](e)f)<p><a href="a(b(c))d">x</a> and <a href="e">y</a>f)</p>An unbalanced parenthesis that belongs inside the URL is backslash-escaped. Only \(, \) and \\ are escapes here, so a backslash in front of anything else is an ordinary character and URLs full of backslashes are unaffected.
[x](http://a/b\)c) and [y](a\\b) and [z](a\qb)<p><a href="http://a/b)c">x</a> and <a href="a\b">y</a> and <a href="a\qb">z</a></p>A newline counts as whitespace, so it ends the destination too: an unclosed ( whose run reaches the end of the line is not a link. The ( and the following text stay literal across the line break (grammar link_destination).
[t](url
more)<p>[t](url
more)</p>Empty link and image titles are preserved
1 conformance fixture
An explicit empty title ("") is kept as title="" rather than dropped -- the grammar permits an empty link_title, and all three implementations emit it identically.
[x](u "")<p><a href="u" title="">x</a></p>A backslash in a link destination is a literal character
1 conformance fixture
A link destination has no backslash escapes: url_char includes the backslash as an ordinary URL character, kept verbatim. [t](a\b) links to a\b.
[t](a\b)<p><a href="a\b">t</a></p>An empty link destination is not a link
3 conformance fixtures
A destination has at least one character, so [x]() and ![x]() stay literal text. A title with no destination before it is literal too.
[x]()<p>[x]()</p>![x]()<p>![x]()</p>[x]( "t")<p>[x]( “t”)</p>A quote is an ordinary link destination character
5 conformance fixtures
A title opens only after the space that ends the destination, so a " inside the destination is part of the href.
[x](a"b)<p><a href="a"b">x</a></p>[x](a"b "t")<p><a href="a"b" title="t">x</a></p>[x](")<p><a href=""">x</a></p>[x]("t")<p><a href=""t"">x</a></p>[x](a`b{c}d|e^f<g>h)<p><a href="a`b{c}d|e^f<g>h">x</a></p>Autolink display keeps the raw content
1 conformance fixture
An autolink's display text is the raw content between < and >: a URI autolink keeps its scheme (<mailto:a@b> shows mailto:a@b), while an email autolink (no explicit scheme) shows the address with a mailto: href.
<mailto:a@b><p><a href="mailto:a@b">mailto:a@b</a></p>Link reference definition separator must be a space
1 conformance fixture
The same rule applies to link reference definitions: [label]: must be followed by a literal space. A tab leaves the line as a paragraph, so the later [a][] has no target to resolve.
[a]: /url
[a][]<p>[a]: /url</p>
<p>[a][]</p>A reference definition cannot take its destination from the next line
3 conformance fixtures
The destination belongs to the same physical line as the label. An empty destination makes the line prose; lazy list collection cannot supply the next line as a URL and consume both as an invisible definition.
* [d]:
:::<ul>
<li>[d]:
:::</li>
</ul>Ordinary lazy text after the same empty destination remains visible too.
* [d]:
text<ul>
<li>[d]:
text</li>
</ul>A non-empty destination on the definition line remains a definition. The following below-column fence-shaped line then lies outside the item.
* [d]: /u
:::<ul>
<li></li>
</ul>
<p>:::</p>Indented image and caption stay literal
3 conformance fixtures
A lone image on its own line at column 0 becomes a <figure> when a ^ caption line follows. Indented by a space, neither the image line nor the caption is a top-level opener, so the pair folds into a literal paragraph with the raw ^ still in the text.

^ Figure 1: moon<p><img src="a.jpg" alt="Apollo">
^ Figure 1: moon</p>An indented attribute brace above the indented image and caption is likewise literal; all three lines join as one paragraph.
{.gallery}

^ Figure 1: moon<p>{.gallery}
<img src="a.jpg" alt="Apollo">
^ Figure 1: moon</p>Control - flush left at column 0 the same image and caption form the <figure>.

^ Figure 1: moon<figure>
<img src="a.jpg" alt="Apollo">
<figcaption>Figure 1: moon</figcaption>
</figure>A lone indented image is a paragraph, and its HTML cannot say so
6 conformance fixtures
411-a-lone-indented-image-is-a-paragraph-and-its-html-cannot-say-so411-a-lone-indented-image-is-a-paragraph-and-its-html-cannot-say-so-2411-a-lone-indented-image-is-a-paragraph-and-its-html-cannot-say-so-3411-a-lone-indented-image-is-a-paragraph-and-its-html-cannot-say-so-4411-a-lone-indented-image-is-a-paragraph-and-its-html-cannot-say-so-5411-a-lone-indented-image-is-a-paragraph-and-its-html-cannot-say-so-6
PART 9 §15's strict column-0 rule says a top-level block opener must start at column 0, and docs/divergence-from-djot.md gives the worked example: # H renders <p># H</p>. A block image is a top-level block construct, so an indented one is not one - the leading space cannot be inert for an image and decisive for a heading.
158-indented-image-and-caption-stay-literal pins the indented image WITH its caption, -2 the same pair under a block-attribute line, and -3 the flush-left figure. All three carry a SECOND line, and that is the whole reason this section exists. The lone-image promotion only fires on a paragraph whose entire content is one image, so a caption line under the image keeps it from firing at all - which left the one shape it does fire on pinned nowhere. carve-rs and carve-php promoted an indented lone image to a block image, carve-js did not, and every corpus document passed on all three while they disagreed (carve#1660).
THE HTML CANNOT SEE THE DIFFERENCE, and that is stated here rather than left for a reader to rediscover. A paragraph whose whole content is one image renders as a bare <img> with no <p> wrapper - the same rule -3 leans on - so the indented reading and the promoted one emit the same bytes. What differs is the tree: paragraph > image against a top-level image. The pair below is therefore an HTML control that every engine already satisfies, and the shape comparison in npm run ast:check is the reader that fails on it.
<img src="a.jpg" alt="Apollo">The reference spelling reaches the same promotion path by a different route: it is never a syntactic block image, so it arrives as a paragraph and is promoted afterwards or not at all. Indenting it must fold the same way, for the same reason.
![Apollo][moon]
[moon]: a.jpg<img src="a.jpg" alt="Apollo">Inside a container, where the pair above cannot reach
The two examples above are both TOP-LEVEL, and that is the whole extent of what this category pinned. Measured across every .crv in the corpus: the only container-hosted image anywhere was 405-a-captioned-image-inside-a-container, which is flush AND captioned - so a caption line keeps the lone-image promotion from firing at all, and no document indented an image past a container's content column. That is why npm run ast:check passed 1393 documents unanimous while carve-js emitted <blockquote><p><img></p></blockquote>: no check failed, because no document reached the shape (markup-carve/carve-js#1440, carve#1677).
The four pairs below reach it. They also settle two places where the oracle answered §1c differently depending on which container the shape sat in - a quote framed its lone-image child on one line where a div indents it, and a list item kept the <p> this section's own prose says is not written. Both were oracle defects, ruled at carve#1677; all three engines already emitted what is recorded here.
A QUOTE HOLDING NOTHING ELSE. The child is spelled as a paragraph and is not one by the time it is serialized, so the quote takes the same expanded frame a <div> takes for the identical child. The one-line form is for a child that stays a paragraph, which > t still is.
> <blockquote>
<img src="a.jpg" alt="Apollo">
</blockquote>The flush spelling is the control, and it carries the weight here: the quote's content column is where the image is ORDINARY, so a change that framed every quoted paragraph the same way would show up on this pair rather than on the indented one. Both fold to the same bytes, for the reason the top-level pair already gives - the HTML cannot see which reading produced the image.
> <blockquote>
<img src="a.jpg" alt="Apollo">
</blockquote>A LIST ITEM, INDENTED PAST ITS CONTENT COLUMN. The image uses that authored base, but §17 still classifies the same line shape as the exact-column spelling: the blank makes the item loose while §1c renders the lone image bare.
- t
<ul>
<li><p>t</p>
<img src="a.jpg" alt="Apollo">
</li>
</ul>And the flush control for the item, at the content column. It has the same structure and loose tightness as the authored-base spelling above.
- t
<ul>
<li><p>t</p>
<img src="a.jpg" alt="Apollo">
</li>
</ul>A lone reference image at column 0, in every spelling
4 conformance fixtures
411-a-lone-indented-image-is-a-paragraph-and-its-html-cannot-say-so pins what an INDENTED lone image is (carve#1660). This section pins the SPELLING axis at column 0, where nothing was pinned at all: no document in the corpus held a lone reference image at column 0, so the question of whether the reference spellings reach block position had no answer on the record and two readings could stand (carve#1663).
All three engines AGREE here, and these pairs record that rather than argue for it. On the published tree - the one a consumer receives - the direct spelling and both reference spellings are a top-level image, and an unresolved reference is a paragraph:
| source | published shape |
|---|---|
 | image |
![Apollo][moon], with [moon]: a.jpg | image |
![Apollo][], with [Apollo]: a.jpg | image |
![Apollo][nope], undefined | paragraph holding an image |
READ THE PUBLISHED EXIT, NOT THE PARSE TREE, when checking any of this. carve-js promotes inside resolve() rather than in its syntactic block-image pass, and resolve() mutates the tree in place - so parse() reports a paragraph for the middle two rows and the tree a consumer receives holds an image. Comparing that intermediate stage against engines that resolve inside their own parse compares a stage no other engine exposes; scripts/ast-conformance.mjs names the trap in its own source, from carve#486, and takes its reference tree through toAstJson(resolve(parse(x))) for exactly this reason. carve#1663 was filed and ruled on the parse-only reading and withdrawn once the published tree was measured.
THE HTML CANNOT SEE THE FIRST THREE PAIRS. A paragraph whose whole content is one image renders as a bare <img> with no <p> wrapper, so a promoted image and a paragraph holding one emit the same bytes, and every engine passes these three whatever it does with the tree. The reader that can tell them apart is the SHAPE comparison in npm run ast:check. The fourth pair is the one the HTML does report, because a reference that resolves to nothing degrades to the literal source it was written as.
![Apollo][moon]
[moon]: a.jpg<img src="a.jpg" alt="Apollo">The collapsed spelling resolves by its own derived label and reaches the same promotion, so it lands the same way.
![Apollo][]
[Apollo]: a.jpg<img src="a.jpg" alt="Apollo">The direct spelling is the one the syntactic block-image pass matches, and it is here so the three spellings are pinned together rather than one of them being inferred from the other two.
<img src="a.jpg" alt="Apollo">An unresolved reference stays a paragraph on all three, and this is where the HTML does move: nothing defines nope, so there is no destination and the construct renders as the text the author typed.
![Apollo][nope]<p>![Apollo][nope]</p>An item's attribute block moves its content column; its checkbox does not
10 conformance fixtures
413-an-item-s-attribute-block-moves-its-content-column-its-checkbox-does-not413-an-item-s-attribute-block-moves-its-content-column-its-checkbox-does-not-2413-an-item-s-attribute-block-moves-its-content-column-its-checkbox-does-not-3413-an-item-s-attribute-block-moves-its-content-column-its-checkbox-does-not-4413-an-item-s-attribute-block-moves-its-content-column-its-checkbox-does-not-5413-an-item-s-attribute-block-moves-its-content-column-its-checkbox-does-not-6413-an-item-s-attribute-block-moves-its-content-column-its-checkbox-does-not-7413-an-item-s-attribute-block-moves-its-content-column-its-checkbox-does-not-8413-an-item-s-attribute-block-moves-its-content-column-its-checkbox-does-not-9413-an-item-s-attribute-block-moves-its-content-column-its-checkbox-does-not-10
Historical category label
The heading is retained as the append-only corpus ID. The rule it originally named was reversed before release: marker-attached attributes and task checkboxes both contribute zero to the content column. The current rule and fixtures below are authoritative.
90-list-item-attributes-4 is the only document that puts attributes on a task item, and it is a single line, so it has no continuation and cannot say where the item's content begins. 06-task-lists and 363-a-task-item-s-checkbox-is-not-decided-by-its-first-block carry continuations but no attributes, so they cannot say it either. Between them the two halves of -{#k} [x] were each pinned alone and the combination was pinned nowhere - which is how three engines came to read it three ways with every gate green (carve#1692).
The normative rule is the bare marker width plus its separator. A checkbox is CONTENT (carve#1690), so it does not move the column; the attribute block is item metadata, so it does not move the column either. -{#k} [x] a therefore has the same content column as - a: column 2. A line written there is inside the item.
-{#k} [x] bare
# inside<ul>
<li id="k"><input type="checkbox" checked disabled aria-label="bare"> bare
<h1 id="inside">inside</h1>
</li>
</ul>The old full-prefix column is pinned beside it as a migration case. It is past the canonical content column but remains structural under the authored-base rule. The metadata spelling can no longer silently select between structures.
-{#k} [x] old
# outside<ul>
<li id="k"><input type="checkbox" checked disabled aria-label="old"> old
<h1 id="outside">outside</h1>
</li>
</ul>The two neighbours a wrong fix breaks. Without the attribute block the column is the bullet's alone, so column 2 is where the content is and the heading lands inside the item - the checkbox still moves nothing.
- [x] a
# h<ul>
<li><input type="checkbox" checked disabled aria-label="a"> a
<h1 id="h">h</h1>
</li>
</ul>Without the checkbox the answer is unchanged: marker-attached attributes are metadata, so the bare bullet still puts the content at column 2.
-{#k} a
# h<ul>
<li id="k">a
<h1 id="h">h</h1>
</li>
</ul>The ordered spelling is the same rule again. The attributes contribute zero, but the ordered marker itself remains variable: 1. puts content at column 3, while 10. puts it at column 4.
1.{#k} a
# h<ol>
<li id="k">a
<h1 id="h">h</h1>
</li>
</ol>A SUB-LIST READS THE COLUMN THROUGH ANOTHER DOOR, and this pair is here because the four above cannot open it. The attributed task's bare bullet column is 2, so after reaches the inner item's content column. Metadata and checkbox text do not change the nesting or looseness calculation.
- outer
-{#k} [x] inner
after<ul>
<li><p>outer</p>
<ul>
<li id="k"><input type="checkbox" checked disabled aria-label="inner"> inner</li>
</ul>
<p>after</p>
</li>
</ul>Two sibling items with different-length metadata now share the same body column. Renaming a class cannot move either heading.
-{.x} a
# one
-{.averylongclass} b
# two<ul>
<li class="x">a
<h1 id="one">one</h1>
</li>
<li class="averylongclass">b
<h1 id="two">two</h1>
</li>
</ul>The old full-prefix column is over-indented under the bare-marker rule. It stays inside the item and establishes an authored block base, so the heading remains structural and canonical output moves it to column 2.
-{.x1} a
# h<ul>
<li class="x1">a
<h1 id="h">h</h1>
</li>
</ul>Unicode metadata is an explicit non-effect. Its UTF-8, UTF-16 and codepoint lengths never enter the column calculation, and the task checkbox remains content rather than marker width.
-{title="😀"} [x] a
# h<ul>
<li title="😀"><input type="checkbox" checked disabled aria-label="a"> a
<h1 id="h">h</h1>
</li>
</ul>An attributed outer item uses the same bare column to open a nested list. This pins the direct shape separately from an attributed task nested inside a plain outer item.
-{.outer} parent
- child<ul>
<li class="outer">parent
<ul>
<li>child</li>
</ul>
</li>
</ul>Indented reference and footnote definitions stay literal
2 conformance fixtures
A reference-link definition and a footnote definition are top-level block constructs that register at column 0. Indented by a space the definition line is an ordinary paragraph: it registers nothing, so the reference or footnote that used it never resolves and both render as literal text.
An indented reference definition does not register; the link stays unresolved.
Read [intro][x].
[x]: /intro "T"<p>Read [intro][x].</p>
<p>[x]: /intro “T”</p>An indented footnote definition does not register; the footnote reference stays literal.
Note[^fn].
[^fn]: body.<p>Note[^fn].</p>
<p>[^fn]: body.</p>A repeated definition: which one wins
3 conformance fixtures
The three definition kinds do not answer this the same way, so each is pinned separately.
A repeated link reference definition is overridden by the later one.
see [t][r].
[r]: /a
[r]: /b<p>see <a href="/b">t</a>.</p>A repeated abbreviation definition behaves the same way - the later expansion wins.
*[A]: a
*[A]: b
A here.<p><abbr title="b">A</abbr> here.</p>A repeated footnote definition does not: the FIRST one wins and the later one is dropped. carve lint reports it as duplicate-footnote-definition.
see [^f].
[^f]: one
[^f]: two<p>see <a id="fnref1" href="#fn1" role="doc-noteref"><sup>1</sup></a>.</p>
<section role="doc-endnotes" aria-label="Footnotes">
<hr>
<ol>
<li id="fn1">
<p>one<a href="#fnref1" role="doc-backlink" aria-label="Back to reference">↩</a></p>
</li>
</ol>
</section>A collapsed reference is matched by the label the author wrote
2 conformance fixtures
[label][] resolves against the definition whose label is that BRACKET TEXT, whitespace-collapsed - the same spelling the definition line registers. The rendered text is a different string as soon as the label carries markup, and keying on it inverts the rule in both directions: the definition that names the label stops resolving, and a plain definition the author never referenced starts.
This engine-visible pair is what nothing pinned. carve-php stripped _ * ~ ^ + = { } [ ] ` from every collapsed label and got both halves backwards (carve-php#768); the executable spec keyed on the rendered text and got the same two answers wrong (carve#648). Three implementations agreed all along and no case could tell.
[*bold*]: /x
see [*bold*][]<p>see <a href="/x"><strong>bold</strong></a></p>The inverse: a decorated label does not reach a plain definition, because [bold] is not the label that was written. Unresolved, the construct renders as the SOURCE the author typed - [*bold*][], markers and all - not as the bracket content re-rendered. The executable spec emitted the rendered form here ([<strong>bold</strong>][]), which drops the markers that identify the construct while keeping the brackets that make it look like one.
[bold]: /x
see [*bold*][]<p>see [*bold*][]</p>A definition inside a container is collected at that container's content column
3 conformance fixtures
A definition is invisible and active wherever its container puts it. At column 0 that is settled; one container deeper it is the same rule, measured from INSIDE the container: > - a puts the item's content column at 2 of the quoted content, so a definition written there belongs to the item.
Every engine lost this in a different way and had to be fixed for it - carve-js#646, carve-php#786, carve-rs#587 - which is what a case pins against.
> - a
> [r]: /u
see [t][r]<blockquote>
<ul>
<li>a</li>
</ul>
</blockquote>
<p>see <a href="/u">t</a></p>The mirror arrangement - a quote INSIDE an item rather than an item inside a quote - reads the same way: the quote sits at the item's content column, the definition is its content, and the quote renders empty.
- a
> [r]: /u
see [t][r]<ul>
<li>a
<blockquote>
</blockquote>
</li>
</ul>
<p>see <a href="/u">t</a></p>An indented > that reaches no content column is not a container at all: the line renders as the text it looks like and defines nothing. Without this case the two above can be satisfied by stripping whitespace indiscriminately, which is exactly what three separate fixes did before it was measured.
[x][r] here.
> [r]: /u<p>[x][r] here.</p>
<p>> [r]: /u</p>Trailing attributes on a link reference definition
3 conformance fixtures
A {...} block at the end of a definition line attaches to the DEFINITION and reaches every link that resolves the label (PART 9R R1). That is what makes a definition worth writing once: without it, attributing a destination used ten times means repeating the attribute ten times.
[Example][ex] and [again][ex]
[ex]: https://example.com {.external}<p><a href="https://example.com" class="external">Example</a> and <a href="https://example.com" class="external">again</a></p>The merge is the one stacked attribute lists already use (PART 9 §15 A3): the definition's list first, the link's second, so a repeated key takes the LAST value and classes ACCUMULATE.
[Example][ex]{.internal #b}
[ex]: /u {.external #a}<p><a href="/u" class="external internal" id="b">Example</a></p>A {...} on its own line ABOVE a definition is a different construct: it floats PAST the definition to the next visible block (§15 A2a). Both can appear at once, and they do different things.
{.a}
[ex]: /u {.b}
[E][ex] and text<p class="a"><a href="/u" class="b">E</a> and text</p>An image takes a reference the way a link does
1 conformance fixture
An image resolves against the same definition table a link does (PART 3, reference_image; carve#641). Every reference FORM was pinned for links and none for images, so the three engines agreed here with nothing holding them to it - and the AST rule that a resolved reference keeps ref and rawRef beside its destination (PART 12 §3a) had no image case either.
![moon][m]
[m]: /moon.png<img src="/moon.png" alt="moon">A collapsed image reference uses its alt text as the label
1 conformance fixture
![alt][] takes the alt text as the label, and a definition title becomes the image's title.
![moon][]
[moon]: /moon.png "Title"<img src="/moon.png" alt="moon" title="Title">One definition serves a link and an image
1 conformance fixture
The definition table is shared, so the same label resolves for both - once as a destination, once as a source.
See [text][m] and ![moon][m].
[m]: /moon.png<p>See <a href="/moon.png">text</a> and <img src="/moon.png" alt="moon">.</p>An unresolved image reference stays literal
1 conformance fixture
With no matching definition the image renders as the text the author typed, exactly as an unresolved link reference does.
![moon][gone]<p>![moon][gone]</p>A reference image takes a caption
1 conformance fixture
A caption attaches to the captionable block above it, and an image written as a reference is an image. Every captioned-image case pinned the inline form, so the oracle accepted only that spelling and left the caption as literal text under a reference image while all three engines built the figure.
![a][ok]
^ cap
[ok]: /p.png<figure>
<img src="/p.png" alt="a">
<figcaption>cap</figcaption>
</figure>An unresolved reference image takes no caption
1 conformance fixture
The caption attaches to a captionable block, and a reference image that resolves to nothing is not one: the whole thing stays the text the author typed, both lines in one paragraph. The resolved form beside it becomes a figure, and the definition may sit anywhere - which is why this cannot be decided by looking at the image line alone.
![a][nope]
^ cap<p>![a][nope]
^ cap</p>An unresolved image gives its whole caption slot back, at any depth
4 conformance fixtures
Block-image status is a property of the RESOLVED tree, so the caption binds at promotion and not on the source shape (PART 9R R7). A ^ line after an image paragraph is an unbound SLOT until the promotion phase runs, and a paragraph the phase does not promote gives the slot's source lines back as its own text.
The four one-line rows are pinned already - a resolved reference image with and without a caption in categories 207 and 412, the unresolved counterparts in 201, 209 and 412. What no CORPUS document held is the give-back on a slot more than one line wide, held only against the oracle by a unit test no engine runs, and the give-back at DEPTH, held by nothing at all. Both are paths on which a line of the document can be lost.
A caption ends the way an open paragraph does (PART 2, MULTI-LINE CAPTIONS), so this slot is two lines wide. All of it comes back:
![a][r]
^ cap one
continued<p>![a][r]
^ cap one
continued</p>The resolved control takes the same two lines as one folded caption, which is what makes the row above a give-back rather than a truncation.
![a][r]
^ cap one
continued
[r]: /u<figure>
<img src="/u" alt="a">
<figcaption>cap one
continued</figcaption>
</figure>An image paragraph can sit in any container, so the phase reaches the whole tree rather than the document's top level. Inside a list item the unpromoted paragraph keeps both lines, and the tight item renders them without a wrapper.
- ![a][r]
^ cap<ul>
<li>![a][r]
^ cap</li>
</ul>The control at the same depth: the definition sits outside the list, the image promotes, and the item holds a figure.
- ![a][r]
^ cap
[r]: /u<ul>
<li>
<figure>
<img src="/u" alt="a">
<figcaption>cap</figcaption>
</figure>
</li>
</ul>An at sign is a reference label character everywhere but the first position
1 conformance fixture
reference_label = (character - ']' - '@'), {character - ']'}. The exclusion is positional: it keeps [@key] free for citations and costs nothing after the first character.
A label that STARTS with @ is a citation definition in the extension layer, so the core corpus cannot state its rendering here - the executable spec refuses that input (carve#798). The engines agree on it: [@x]: /u defines nothing and the label renders as a mention.
[a@b]: /v
see [u][a@b].<p>see <a href="/v">u</a>.</p>A label beginning with an at sign is not a reference label
2 conformance fixtures
reference_label = (character - ']' - '@'), {character - ']'} subtracts @ from the first position, and that exclusion is what keeps [@key] free for citations. So a bracket pair whose content starts with @ was never spelling a label, and the label slot of a reference link is no exception to it: the slot does not match, the production does not match, and the construct is not a reference link. The bracket run stays literal text, exactly as any other malformed reference does (carve#1302).
[t][@a]<p>[t][@a]</p>The image spelling reads the same reference_label at the same slot, so it gets the same answer rather than a parallel rule.
![t][@a]<p>![t][@a]</p>The control is what makes this a rule rather than a breakage: @ at the first position ALREADY means citation, so it cannot also mean label, and an implementation that made the slot accept @ would have to take that spelling away from citations to do it. A citation's rendering cannot be stated here - the construct is Tier-2 and the executable spec refuses it (carve#798) - so 43-citations-at-label-in-reference-position in tests/corpus-optional carries the other half: with the Citations extension on, both spellings above are still literal text while a [@a] beside them resolves to a citation.
The run declines as ONE construct and is restored as one, so the @a inside it is text and nothing else. A rescan would read [t][@a] as a [t][ run, a mention and a ], and with the extension on it would read the now tail-less [@a] as a citation - which is why the optional control renders both spellings literally on the far side of the switch.
At the HTML layer these rows cannot separate "not a reference label" from "a label nothing defines": a @-first label can never be defined, and a declining reference renders as its verbatim source either way. PART 12 §3a is where the readings would differ, all three engines publish an unresolved-reference node there, and the grammar clause records that without settling it.
A link definition written before a footnote stays before it
2 conformance fixtures
§7 orders collected definitions by source position, and PART 11 §6 binds the writer to the order the tree holds. The corpus pinned only the shape where the footnote comes first (a definition on a footnote body's continuation line), and a writer that emits footnotes in a fixed position - first or last - is correct on exactly that shape. This is its mirror: every engine's carve output for it must put the link definition first, because the author did.
see[^a] and [t][r]
[r]: /u
[^a]: note<p>see<a id="fnref1" href="#fn1" role="doc-noteref"><sup>1</sup></a> and <a href="/u">t</a></p>
<section role="doc-endnotes" aria-label="Footnotes">
<hr>
<ol>
<li id="fn1">
<p>note<a href="#fnref1" role="doc-backlink" aria-label="Back to reference">↩</a></p>
</li>
</ol>
</section>Two definitions of the SAME kind pin the other half of the rule: a writer that sorts them by label rather than by position reverses these two.
see[^b] and[^a]
[^b]: bee
[^a]: ay<p>see<a id="fnref1" href="#fn1" role="doc-noteref"><sup>1</sup></a> and<a id="fnref2" href="#fn2" role="doc-noteref"><sup>2</sup></a></p>
<section role="doc-endnotes" aria-label="Footnotes">
<hr>
<ol>
<li id="fn1">
<p>bee<a href="#fnref1" role="doc-backlink" aria-label="Back to reference">↩</a></p>
</li>
<li id="fn2">
<p>ay<a href="#fnref2" role="doc-backlink" aria-label="Back to reference">↩</a></p>
</li>
</ol>
</section>A zero-width character in a reference definition destination
2 conformance fixtures
link_destination ends at Unicode whitespace, and a ZERO WIDTH NO-BREAK SPACE is not whitespace: it is an ordinary destination character. The same rule governs a reference definition, because the definition is built from that same production - so the character is neither skipped as the separator run nor read as the end of the destination.
The inline form was already pinned. This is the definition form, where two engines kept the character and one dropped it or truncated at it, because the host language's own whitespace class holds U+FEFF and the Unicode property does not.
[r]: https://e.com/
see [x][r]<p>see <a href="https://e.com/">x</a></p>Position does not change the answer: a definition is not truncated at a zero-width character in the middle of its destination either.
[r]: https://e.com/
see [x][r]<p>see <a href="https://e.com/">x</a></p>A block image is separated from the block after it on every target
1 conformance fixture
A lone image is a BLOCK, so whatever separates two blocks on a target separates this one from what follows. The corpus held no document where a block image is followed by another block, so no gate compared the engines on it - and one of them ran the alt text straight into the next paragraph on the plain and ANSI targets (alt textfollowing paragraph), which the two repos claiming non-HTML parity could not catch because each reads its own committed snapshot rather than the other engines (carve-rs#692, carve-js#762).

following paragraph<img src="img.png" alt="alt text">
<p>following paragraph</p>A reference definition is anchored at end of line
16 conformance fixtures
266-a-reference-definition-is-anchored-at-end-of-line266-a-reference-definition-is-anchored-at-end-of-line-2266-a-reference-definition-is-anchored-at-end-of-line-3266-a-reference-definition-is-anchored-at-end-of-line-4266-a-reference-definition-is-anchored-at-end-of-line-5266-a-reference-definition-is-anchored-at-end-of-line-6266-a-reference-definition-is-anchored-at-end-of-line-7266-a-reference-definition-is-anchored-at-end-of-line-8266-a-reference-definition-is-anchored-at-end-of-line-9266-a-reference-definition-is-anchored-at-end-of-line-10266-a-reference-definition-is-anchored-at-end-of-line-11266-a-reference-definition-is-anchored-at-end-of-line-12266-a-reference-definition-is-anchored-at-end-of-line-13266-a-reference-definition-is-anchored-at-end-of-line-14266-a-reference-definition-is-anchored-at-end-of-line-15266-a-reference-definition-is-anchored-at-end-of-line-16
reference_definition ends in newline, and always has. All three engines and the executable spec nevertheless read [a]: /u zzz as a definition with trailing junk, and nothing in the grammar authorized that reading. carve#911 settled it the way the production already said: what follows the destination and the optional title makes the production FAIL, so the line is an ordinary paragraph.
The reason it matters is not tidiness. PART 7 promises that a slot which fails to match "falls back to prose rather than silently dropping metadata", and at this line there was no prose to fall back to -- the swallowing tail took whatever the slot rejected. So the clause's promised failure mode was unreachable here, and every narrowing at this line dropped metadata instead. With the line anchored the promise holds, and the tab and cardinality rules at both slots follow from the general rule with no special case.
[a]: /u zzz
[a][]<p>[a]: /u zzz</p>
<p>[a][]</p>A quoted run after a title is junk in the same way. The title is read, and then the line fails on what is left:
[a]: /u "T" zzz
[a][]<p>[a]: /u “T” zzz</p>
<p>[a][]</p>The tab, at both slots
The title slot and the trailing-attributes slot take space under PART 7, like every other padding slot. Both sit on this line, and both were left unpinned by carve#907 for the reason above: with the line unanchored, a tab there dropped the metadata rather than producing the visible failure the clause names. Now it produces the failure, so it can be pinned.
Each slot carries the tab-first form and BOTH mixed runs. A rule about a run written as "the first character must be a space" passes the tab-first fixture and admits <SP><TAB>; written as "the last character must be a space" it admits <TAB><SP> instead. Both spellings have been written for real in this org, in three languages, on one day.
The title slot:
[a]: /u "T"
[a][]<p>[a]: /u “T”</p>
<p>[a][]</p>[a]: /u "T"
[a][]<p>[a]: /u “T”</p>
<p>[a][]</p>[a]: /u "T"
[a][]<p>[a]: /u “T”</p>
<p>[a][]</p>The trailing-attributes slot:
[a]: /u {.c}
[a][]<p>[a]: /u {.c}</p>
<p>[a][]</p>[a]: /u {.c}
[a][]<p>[a]: /u {.c}</p>
<p>[a][]</p>[a]: /u {.c}
[a][]<p>[a]: /u {.c}</p>
<p>[a][]</p>What the anchor does NOT reject
The line ending is whitespace -- a space or a tab -- which is the same terminal blank_line = {whitespace} takes (PART 1, carve#890). So trailing spaces and a trailing tab are still a line ending rather than content, and the definition stands; a trailing NO-BREAK SPACE, EN QUAD or FORM FEED is content under that same ruling, so a line ending in one is not a definition. A trailing ZERO-WIDTH character is a third answer again: U+200B and U+FEFF are not whitespace at all, so they never reach the line ending -- link_destination reads them, the definition stands, and the character is in the href. The trailing attribute block is peeled by a scan that trims the line ending first, so that scan has to trim the SAME run the anchor accepts, or the same character answers one way with a block on the line and another without one. Those shapes are pinned in tests/separator-role-split.test.mjs rather than here, because a trailing whitespace run in a reviewable Markdown source file is one editor save from vanishing -- and because what a document does with trailing whitespace is its own question (carve#926).
And the glued form is untouched, because nothing is left over: link_destination simply reads the braces.
[a]: /u{.c}
[a][]<p><a href="/u{.c}">a</a></p>A definition still INTERRUPTS, attribute block and all
The anchor changes what the pattern matches, and the pattern is read in nine places rather than one: eight of them ask "is this line a definition" to decide paragraph interruption, lazy continuation, the def-list fold, the container scan, the item fold and the marker scan. While the pattern ended in a swallow-everything tail those eight could test the RAW line and be right by accident, because [a]: /u {.c} matched it raw. Anchored, they cannot: the trailing attribute block has to be split off first, or a definition carrying one stops interrupting anything and folds into the paragraph above it.
Nothing pinned that. Reverting all eight to the raw line left the entire suite green, and a differential sweep then found 42 of 72 generated shapes moving. These three are the sweep's representatives -- top level, inside a list item, and inside a definition-list description.
text
[a]: /u {.c}
[a][]<p>text</p>
<p><a href="/u" class="c">a</a></p>- text
[a]: /u {.c}
[a][]<ul>
<li>text</li>
</ul>
<p><a href="/u" class="c">a</a></p>:: term
: def
[a]: /u {.c}
[a][]<dl>
<dt>term</dt>
<dd>def</dd>
</dl>
<p><a href="/u" class="c">a</a></p>The fourth is the one the other three cannot reach. A block quote's open paragraph asks the same question in a SECOND place -- the lazy-continuation test, not the paragraph collector -- and reverting only that one site left even the differential sweep above unmoved. Here the definition has to interrupt, or the line below it lazily continues INSIDE the quote:
> text
[a]: /u {.c}
more
[a][]<blockquote><p>text</p></blockquote>
<p>more</p>
<p><a href="/u" class="c">a</a></p>Two more, found the same way: each of the eight sites was reverted on its own, and three of them survived every shape above. A definition INSIDE a quote is the third site -- it is what leaves the quote with no open paragraph, so a following lazy line starts its own block instead of folding in:
> text
> [a]: /u {.c}
lazy<blockquote><p>text</p></blockquote>
<p>lazy</p>And a definition after a blank line inside a list item is the fourth and fifth. An invisible construct is not a second paragraph, so the list stays TIGHT (PART 9 §17 L1/L2); with the raw predicate it loosens:
- text
[a]: /u {.c}
[a][]<ul>
<li>text</li>
</ul>
<p><a href="/u" class="c">a</a></p>The CONTROLS. Every legal shape of the line still is one:
[a]: /u "T" {.c}
[a][]<p><a href="/u" title="T" class="c">a</a></p>An autolink body admits non-ASCII and excludes format characters
10 conformance fixtures
272-an-autolink-body-admits-non-ascii-and-excludes-format-characters272-an-autolink-body-admits-non-ascii-and-excludes-format-characters-2272-an-autolink-body-admits-non-ascii-and-excludes-format-characters-3272-an-autolink-body-admits-non-ascii-and-excludes-format-characters-4272-an-autolink-body-admits-non-ascii-and-excludes-format-characters-5272-an-autolink-body-admits-non-ascii-and-excludes-format-characters-6272-an-autolink-body-admits-non-ascii-and-excludes-format-characters-7272-an-autolink-body-admits-non-ascii-and-excludes-format-characters-8272-an-autolink-body-admits-non-ascii-and-excludes-format-characters-9272-an-autolink-body-admits-non-ascii-and-excludes-format-characters-10
url_char was an enumerated ASCII set, so read as written an autolink admitted no non-ASCII at all - and two engines linked internationalized domains anyway (markup-carve/carve#860). The rule now reads: outside ASCII, url_char admits any character that is not whitespace and not a FORMAT character (General_Category Cf).
The deciding asymmetry is that the same destination written as an inline link already links everywhere, because link_destination admits unicode_url_char. One destination cannot answer two ways on the character set depending on which spelling the author reached for.
<https://例.jp/><p><a href="https://例.jp/">https://例.jp/</a></p>A non-ASCII PATH links on the same terms:
<https://example.com/café><p><a href="https://example.com/café">https://example.com/café</a></p>And so does a non-ASCII character that is not a LETTER. This is the row that separates the rule from "Unicode letters are letters": the executable spec's urlChar used ohm's built-in letter, which is Unicode-aware, so it linked café for free and left a currency sign, a CJK comma or an emoji literal.
<https://example.com/€10><p><a href="https://example.com/€10">https://example.com/€10</a></p>A format character is not a URL character
The exclusion is the half that is new rather than permissive. A format character is invisible by definition, so a host carrying one renders as the host WITHOUT it and links somewhere else - a spoofing surface, not an authoring convenience. The next document has a U+FEFF BYTE ORDER MARK between the e and the .com, and it is not an autolink:
<https://e.com/><p><https://e.com/></p>A LEADING one, before the scheme, is literal too - there for a different reason, since a scheme starts with a letter and this is not one:
<https://e.com/><p><https://e.com/></p>A U+200B ZERO WIDTH SPACE is a format character as well, despite the name - Unicode moved it out of the space categories in 4.0.1. Spelling the rule as a PROPERTY is what makes these two documents answer alike: a host language whose own whitespace class happens to hold U+FEFF gets the first one right for the wrong reason and this one wrong.
<https://e.com/><p><https://e.com/></p>A U+00A0 NO-BREAK SPACE is excluded by the OTHER half of the rule - it is whitespace - and was already the answer under both readings:
<https://e .com/><p><https://e .com/></p>What did not change
CONTROL. The ASCII exclusions are untouched. ", \, a backtick, {, }, |, ^, < and > are still not url_chars, and all four artifacts already agreed on them - which is why the rule is spelled unicode_url_char - format_char rather than "any non-whitespace, non-control character". The latter would re-admit these nine and move every implementation on a question nobody asked. (The straight quotes additionally pick up smart-quote typography, which is what makes the output curly.)
<https://example.com/"q"><p><https://example.com/“q”></p>CONTROL. link_destination is a different production and is unchanged, so a format character in an INLINE destination is still an ordinary destination character. The pair below and the interior-BOM autolink above are the same character in the same position, answering differently because the two spellings are two productions:
[t](https://e.com/)<p><a href="https://e.com/">t</a></p>CONTROL. Only the BODY admits non-ASCII. A scheme is letter, {letter | digit | '+' | '-' | '.'} and letter is the enumerated ASCII alphabet, so a scheme written in another script opens no autolink. The executable spec accepted one until this landed, for the same reason it linked café: ohm's letter is Unicode-aware.
<例://example.com/><p><例://example.com/></p>Unresolved reference link
1 conformance fixture
A reference with no matching definition renders as literal text.
A [missing][nope] ref stays literal.<p>A [missing][nope] ref stays literal.</p>Bare URLs stay literal
1 conformance fixture
A bare URL is not auto-linked (matching djot); wrap it in <…> to link.
see https://example.com now<p>see https://example.com now</p>A link inside a span's label keeps its destination
9 conformance fixtures
525-a-link-inside-a-span-s-label-keeps-its-destination525-a-link-inside-a-span-s-label-keeps-its-destination-2525-a-link-inside-a-span-s-label-keeps-its-destination-3525-a-link-inside-a-span-s-label-keeps-its-destination-4525-a-link-inside-a-span-s-label-keeps-its-destination-5525-a-link-inside-a-span-s-label-keeps-its-destination-6525-a-link-inside-a-span-s-label-keeps-its-destination-7525-a-link-inside-a-span-s-label-keeps-its-destination-8525-a-link-inside-a-span-s-label-keeps-its-destination-9
The never-nest constraint binds a link's TEXT. A span's label is the same bracket run, but a span is not a link, so a link written in a label nests into nothing and its destination stands. Every link spelling is that same case: an inline link, a reference link, an autolink, and a link under emphasis.
A link inside a LINK label still unwraps, at any depth, including one wrapped in a span on the way down. An image is not a link and survives in either label.
[[t](/v)]{.c}<p><span class="c"><a href="/v">t</a></span></p>[[t][r]]{.c}
[r]: /v<p><span class="c"><a href="/v">t</a></span></p>[<https://e.com>]{.c}<p><span class="c"><a href="https://e.com">https://e.com</a></span></p>[/[t](/v)/]{.c}<p><span class="c"><em><a href="/v">t</a></em></span></p>[]{.c}<p><span class="c"><img src="/i" alt="a"></span></p>An empty attribute block is a span too, so it holds a link on the same terms.
[[t](/v)]{}<p><span><a href="/v">t</a></span></p>The link-label half, unchanged: the inner destination goes and the outer one applies.
[[t](/v)](/u)<p><a href="/u">t</a></p>[[[t](/v)]{.d}](/u)<p><a href="/u"><span class="d">t</span></a></p>[](/u)<p><a href="/u"><img src="/i" alt="a"></a></p>A tab does not open the title slot
3 conformance fixtures
link_title spells its padding space. The slot sits after the first non-whitespace character of the line, where a tab carries no syntax (PART 7), so a tab leaves the destination closed and the whole line is prose. Nothing else about the line changes: the quotes are ordinary text and smart typography reads them. image_title = link_title, so an image refuses it on the same terms.
[a](/u "t")<p>[a](/u “t”)</p><p></p>A single space is the spelling that opens it.
[a](/u "t")<p><a href="/u" title="t">a</a></p>Links
5 conformance fixtures
A backslash-escaped delimiter inside a title is a literal quote (CommonMark-style), so \" does not end a double-quoted title — it renders as a literal " (") in the title attribute (grammar link_title).
[t](/url "ti\"tle")<p><a href="/url" title="ti"tle">t</a></p>An empty link text is allowed and produces an empty anchor — useful as a target for a styled link.
[](https://example.com)<p><a href="https://example.com"></a></p>Links never nest. When a link's text contains another link, the inner link is replaced by its text and only the outer destination applies, so explicit nesting collapses to a single anchor.
[[x](y)](z)<p><a href="z">x</a></p>The same rule covers an autolink that lands inside a link's text: the autolink becomes plain text, never a nested anchor.
[pre <http://h> post](/u)<p><a href="/u">pre http://h post</a></p>It also covers a crossref, which only becomes a link once it resolves: inside a link's text it contributes its resolved text, not a second anchor.
# H
[see </#H>](/outer)<section id="H">
<h1>H</h1>
<p><a href="/outer">see H</a></p>
</section>The rule is about link content in the parsed tree. One renderer-level case is out of scope: a footnote reference inside a link label still renders its own doc-noteref anchor inside the outer anchor. Putting a footnote inside a link is unusual, and the footnote body and endnote are unaffected.
Reference link
5 conformance fixtures
"Anywhere in the document" includes inside a container: a reference definition written inside a blockquote or a list item is still an invisible block that is collected into the document-wide definition table, so a reference elsewhere resolves against it. The container keeps only its visible content (here, none).
> [ref]: /url
See [it][ref].<blockquote>
</blockquote>
<p>See <a href="/url">it</a>.</p>- [ref]: /url
See [it][ref].<ul>
<li></li>
</ul>
<p>See <a href="/url">it</a>.</p>A reference definition shares the link_destination rule, so its URL ends at the first whitespace. The definition is also ANCHORED AT END OF LINE, so what follows the destination is not ignored: it makes the production fail, and the line is an ordinary paragraph. [r]: a b c is therefore not a definition at all, and the reference below it does not resolve.
This is carve#911's ruling. The line used to be read as a definition with trailing junk, in all three engines and in the executable spec, and nothing in the grammar authorized that reading -- reference_definition has always ended in newline. It also made PART 7's promised fallback unreachable here: a title or attribute slot that failed had no prose to fall back to, so its metadata was silently dropped instead.
[r][r]
[r]: a b c<p>[r][r]</p>
<p>[r]: a b c</p>A backslash-escaped quote inside the title is a literal quote, exactly as in an inline link title (link_title). The escaped \" does not end the quoted run; it renders as a literal " (") in the title attribute.
[x][y]
[y]: /u "a\"b\"c"<p><a href="/u" title="a"b"c">x</a></p>A reference definition requires a non-empty destination (grammar reference_definition). A [r]: with nothing after the colon — or only trailing whitespace — is not a definition; the line stays literal text.
[r]:<p>[r]:</p>Autolinks
1 conformance fixture
url_char also excludes " \ ` { } | ^, so a double quote inside the brackets breaks the autolink. The whole run is literal text; the straight quotes additionally pick up smart-quote typography.
<http://a.com/"q"><p><http://a.com/“q”></p>