Skip to content

Inline spans ​

Emphasis flanking, mentions and tags, smart typography, inline literals and escaping.

Generated from resources/examples/edge-cases.md, resources/examples/core.md and resources/examples/extensions.md - edit the cases there, not here. Each case links the conformance fixture it produces.

Mention ignores email addresses ​

1 conformance fixture

@ starts a mention only at a word boundary, so an email address is left untouched.

carve
Write me@example.com or ping @markus.
html
<p>Write me@example.com or ping <span class="mention"><strong>@markus</strong></span>.</p>

Tag requires a word boundary ​

2 conformance fixtures

# starts a tag only at a word boundary; foo#bar is literal text.

carve
A #tag here, but not in foo#bar.
html
<p>A <span class="tag"><strong>#tag</strong></span> here, but not in foo#bar.</p>

A tag name may be all digits, so #123 is a tag (not literal) — Issue #123 tags the number. Only a leading word boundary is required, not a leading letter.

carve
Issue #123 and #v2 here.
html
<p>Issue <span class="tag"><strong>#123</strong></span> and <span class="tag"><strong>#v2</strong></span> here.</p>

Smart typography escapes and code ​

1 conformance fixture

A backslash keeps the literal sequence; code spans and blocks are never transformed.

carve
Escaped \-> and \... stay; code `a -- b ...` stays.
html
<p>Escaped -&gt; and ... stay; code <code>a -- b ...</code> stays.</p>

The doubled run is the canonical arrow, in both families ​

1 conformance fixture

The single-line family was complete and the double-line family was not: ->, <- and <-> all existed, => existed, and there was no <= arrow, because <= is the comparison. So <=> rendered ≤>.

<= keeps ≤, which forces the left double arrow to grow a character, and once it does the family is spelled at one width throughout. The single-hyphen forms still render, so a document written before this goes on working. => is removed rather than deprecated: key => value and Some(x) => x are ordinary prose about code, and every one of them silently became ⇒.

carve
Canonical: <-- --> <--> and <== ==> <=>

Deprecated but still rendered: <- -> <->

Not an arrow: key => value stays literal, and p <= q is a comparison.
html
<p>Canonical: ← → ↔ and ⇐ ⇒ ⇔</p>
<p>Deprecated but still rendered: ← → ↔</p>
<p>Not an arrow: key =&gt; value stays literal, and p ≤ q is a comparison.</p>

Empty delimiters ​

1 conformance fixture

A delimiter pair with no content is literal text, not emphasis.

carve
** and // and ^^
html
<p>** and // and ^^</p>

Escape coverage ​

2 conformance fixtures

A backslash escapes any ASCII punctuation character to its literal form. This pins the full ascii_punctuation matrix (&, :, ;, ? included); <, >, & are then HTML-escaped in the output.

carve
\!\"\#\$\%\&\'\(\)\*\+\,\-\.\/\:\;\<\=\>\?\@\[\\\]\^\_\`\{\|\}\~ done
html
<p>!"#$%&amp;'()*+,-./:;&lt;=&gt;?@[\]^_`{|}~ done</p>

A backslash before a non-ASCII character or a letter is literal; \\ is a single backslash.

carve
\a and \« and a\\b
html
<p>\a and \« and a\b</p>

Emphasis edge cases ​

4 conformance fixtures

Two emphasis spans of the same kind sit side by side without merging.

carve
*a* and *b*
html
<p><strong>a</strong> and <strong>b</strong></p>

A code span inside emphasis is preserved.

carve
*a `x` b*
html
<p><strong>a <code>x</code> b</strong></p>

Different-kind delimiters sit adjacent without interfering.

carve
~old~ =new=
html
<p><s>old</s> <mark>new</mark></p>

Trailing punctuation after a closer is literal.

carve
*a, b*!
html
<p><strong>a, b</strong>!</p>

Doubled emphasis delimiters ​

1 conformance fixture

A bare single-character emphasis delimiter immediately adjacent to the same delimiter does not open a span, so a doubled delimiter is literal text. This "no nesting of same type" rule is uniform across all seven single-character delimiters: **, ~~, ^^, ==, and ,, stay literal exactly like // and __.

carve
**a** ~~b~~ ^^c^^
html
<p>**a** ~~b~~ ^^c^^</p>

Two-char delimiter runs ​

1 conformance fixture

Every bare delimiter is single-char. A doubled (or longer) run of any delimiter is literal by the same-delimiter-adjacency rule, so ==x== and ~~y~~ are doubled = / ~ and render literal, while the single-char =z= and ~w~ mark.

carve
==x== ~~y~~ =z= ~w~
html
<p>==x== ~~y~~ <mark>z</mark> <s>w</s></p>

Mention and tag name boundaries ​

1 conformance fixture

A mention or tag name runs over letters, digits, _, -, and interior dots (a dot followed by another name character, as in @john.doe or #release-1.0). A dot at the end of the run is sentence punctuation, not part of the name; other punctuation ends the name and stays literal (an apostrophe becomes a typographic quote).

carve
Ping @john-doe, @john_doe and @john.doe about #release-1.0 today.

Reach @john. That is @john's idea, @john!
html
<p>Ping <span class="mention"><strong>@john-doe</strong></span>, <span class="mention"><strong>@john_doe</strong></span> and <span class="mention"><strong>@john.doe</strong></span> about <span class="tag"><strong>#release-1.0</strong></span> today.</p>
<p>Reach <span class="mention"><strong>@john</strong></span>. That is <span class="mention"><strong>@john</strong></span>’s idea, <span class="mention"><strong>@john</strong></span>!</p>
1 conformance fixture

A *[ at an emphasis-opening position is a bold span whose content begins with a link - only a line-start *[ followed by term]: is an abbreviation definition.

carve
See *[the docs](url) for more* info.
html
<p>See <strong><a href="url">the docs</a> for more</strong> info.</p>

Literal less-than in prose ​

1 conformance fixture

A < that is neither an autolink, a crossref, nor a smart-typography arrow stays literal text (HTML-escaped on output).

carve
Check if (x < 5) holds, and 3<4 too.
html
<p>Check if (x &lt; 5) holds, and 3&lt;4 too.</p>

Editorial markup takes a trailing attribute ​

1 conformance fixture

An addition {+...+} or deletion {-...-} is an ordinary inline node, so a trailing {...} attribute block attaches to its <ins> / <del>, exactly like a span, code span, link, or emphasis (PART 9 §22 / §15). The markers are single-character: the doubled form {++a++} is not special — the outer + are the delimiters and +a+ is literal content, so it yields <ins>+a+</ins> (as the example below shows).

carve
{++a++}{.a}
html
<p><ins class="a">+a+</ins></p>

Emphasis opener slash-adjacency ​

4 conformance fixtures

A / immediately before a bare delimiter suppresses an italic / or underline _ opener there: / never opens after / (same-delimiter adjacency) and _ never opens after / (the extra cross-delimiter guard, path protection). So the underscore in a_/_a_ stays literal.

carve
a_/_a_
html
<p>a_/_a_</p>

An underline opener directly after a slash is literal on its own, too.

carve
a/_y_
html
<p>a/_y_</p>

A path-like /a/ opens italic; the following _b_ does not open, because its opening _ sits immediately after the closing /.

carve
/a/_b_
html
<p><em>a</em>_b_</p>

The guard is specific to / and _: the other delimiters *, ~, = DO open after a /, so a preceding slash does not suppress them.

carve
a/~y~ a/=y=
html
<p>a/<s>y</s> a/<mark>y</mark></p>

Bold-italic delimiter needs content ​

4 conformance fixtures

A bold-italic run /*...*/ collapses to <strong><em>...</em></strong> only when it wraps content. With nothing between the delimiters, the inner ** is literal and only the outer /.../ italic applies -- so /**/ is an emphasized **, not empty bold-italic.

carve
/**/
html
<p><em>**</em></p>

A single space is content for the outer italic but not for the bold pair, so the * * stays literal inside one <em>.

carve
/* */
html
<p><em>* *</em></p>

With real content, the full bold-italic collapse still applies.

carve
/*x*/
html
<p><strong><em>x</em></strong></p>

The bold-italic pair /*...*/ has no word-boundary condition on its outer /: the combined two-character opener wins over the bare /-then-* parse even when a word character sits directly before /* or directly after the closing */, so it opens and closes intraword.

carve
a/*y*/b
html
<p>a<strong><em>y</em></strong>b</p>

A run of asterisks inside a combined token is content ​

2 conformance fixtures

/*...*/ takes the run's first * as content and the last two characters as the closer, so a three-asterisk run is a combined span over one literal *. The two-asterisk /**/ of Bold-italic delimiter needs content is the neighbour where there is nothing left for content, and it stays an emphasized ** (markup-carve/carve#2135).

carve
/***/
html
<p><strong><em>*</em></strong></p>
carve
/****/
html
<p><strong><em>**</em></strong></p>

Any character is content of the combined bold-italic token ​

7 conformance fixtures

bi_content is the content every other emphasis kind has, so a tab, a character outside the grammar's named classes, an escaped space, a hard break or a braced comment inside /*...*/ leaves the combined token whole (markup-carve/carve#2159). The first document holds a tab between the letters.

carve
/*a	b*/ y
html
<p><strong><em>a	b</em></strong> y</p>

Characters no other content rule names.

carve
/*a<b*/ y

/*a # b*/ y

/*a{b}c*/ y

/*a€b*/ y

/*a b*/ y
html
<p><strong><em>a&lt;b</em></strong> y</p>
<p><strong><em>a # b</em></strong> y</p>
<p><strong><em>a{b}c</em></strong> y</p>
<p><strong><em>a€b</em></strong> y</p>
<p><strong><em>a&nbsp;b</em></strong> y</p>

An escaped space is a no-break space.

carve
/*a\ b*/ y
html
<p><strong><em>a&nbsp;b</em></strong> y</p>

A backslash at the end of a line is a hard break.

carve
/*a\
b*/ y
html
<p><strong><em>a<br>
b</em></strong> y</p>

A braced comment renders nothing.

carve
/*a{%c%}b*/ y
html
<p><strong><em>ab</em></strong> y</p>

CONTROL. A tab before the closer is refused like a space, so the token falls through to an italic around literal asterisks. The third character is a tab.

carve
/*a	*/ y
html
<p><em>*a	*</em> y</p>

A %% comment runs to the end of its line, and the token closes on the next.

carve
/*a %%c
b*/ y
html
<p><strong><em>a
b</em></strong> y</p>

A form feed or a no-break space is content wherever whitespace is tested ​

12 conformance fixtures

Carve's whitespace is four characters: U+0020, U+0009, U+000A and U+000D (CARVE-P7-003). A form feed, a vertical tab, a no-break space and the other Unicode spaces are content in every construct, including the ones that only test for whitespace (markup-carve/carve#2157). In the documents below the invisible character is a form feed unless the text says otherwise.

A form feed after an opener or before a closer does not refuse the delimiter.

carve
a /x/ b

a /x/ b
html
<p>a <em>x</em> b</p>
<p>a <em>x</em> b</p>

A code span holding a form feed between two spaces drops the spaces.

carve
`  `
html
<p><code></code></p>

A form feed after the closing pipe is content after the row, so the line is prose.

carve
|a|b|
html
<p>|a|b|</p>

A term marker followed by a form feed has content, so it is a term.

carve
:: 
: d
html
<dl>
  <dt></dt>
  <dd>d</dd>
</dl>

A trailing form feed stays in the term.

carve
:: t
: d
html
<dl>
  <dt>t</dt>
  <dd>d</dd>
</dl>

A + followed by a form feed is not a continuation marker, in a description or an item, nested or not.

carve
:: t
: +

b
html
<dl>
  <dt>t</dt>
  <dd>+</dd>
</dl>
<p>b</p>
carve
- +

b
html
<ul>
  <li>+</li>
</ul>
<p>b</p>
carve
- a
  - +

b
html
<ul>
  <li>a
    <ul>
      <li>+</li>
    </ul>
  </li>
</ul>
<p>b</p>

A typed frontmatter opener may end in whitespace, not in a form feed.

carve
---yaml
a: 1
---

t
html
<p>---yaml
a: 1</p>
<hr>
<p>t</p>

A heading reference does not fold a no-break space into a space. The heading holds one between the letters.

carve
# a b

[a b][]
html
<section id="a b">
  <h1>a&nbsp;b</h1>
  <p>[a b][]</p>
</section>

Nor does it trim one. This heading starts with a no-break space.

carve
#  ab

[ab][]
html
<section id=" ab">
  <h1>&nbsp;ab</h1>
  <p>[ab][]</p>
</section>

A form feed after the combined /* opener does not refuse it either.

carve
/*x*/ y
html
<p><strong><em>x</em></strong> y</p>

Emphasis span closes before a following delimiter ​

1 conformance fixture

A completed emphasis span closes at its valid closer regardless of what follows, per the §9 close-first rule: a valid closer closes the nearest matching open entry. So _z_ closes into <u>z</u> even when more bare delimiters come right after it. The trailing /y/ stays literal, because a / opener is suppressed immediately after the closing _ (the slash-adjacency guard above).

carve
_z_/y/
html
<p><u>z</u>/y/</p>

Unclaimed openers stay literal ​

1 conformance fixture

Two forms that were proposed and then not adopted have no meaning in Carve, and both are pinned here so no engine can quietly start claiming them.

[>content] was the proposed sidenote form. It was dismissed, not deferred: a margin note is footnote content positioned by CSS, so it needs no syntax. The [> opener is unclaimed and the whole thing is literal text.

{:name:} was a proposed braced symbol form for intraword use. The brace is not part of any construct: it merely satisfies the symbol boundary guard, and the braces themselves stay literal. The name inside is a normal symbol, so with no symbols map configured it falls back to the literal :name: and the line renders exactly as written. See dismissed syntax.

carve
[>foo]

{:tada:}
html
<p>[&gt;foo]</p>
<p>{:tada:}</p>

Inline literal ​

3 conformance fixtures

A ! prefix on a verbatim span is an inline literal (PART 9 §27): the content is captured verbatim like a code span, but it renders as ordinary prose — no <code> wrapper — and is HTML-escaped and emitted by every renderer. It mirrors the $-math prefix, and exists so notation that collides with the bare emphasis delimiters (phonemic transcription /kaet/, glob patterns, paths) can be written without escaping each character.

carve
The word cat is !`/kaet/` in IPA.
html
<p>The word cat is /kaet/ in IPA.</p>

A trailing {…} is the ordinary inline attribute block, so an attributed literal renders a <span> carrying it. The content is HTML-escaped, and no inline construct inside it is parsed.

carve
!`/kaet/`{.ipa} and !`a<b>` and !`*not bold*`
html
<p><span class="ipa">/kaet/</span> and a&lt;b&gt; and *not bold*</p>

! still opens an image before [, and stays literal text everywhere else. A literal ! immediately before a backtick span is written \!.

carve
\!`x` is a bang before code.
html
<p>!<code>x</code> is a bang before code.</p>

All-space verbatim content ​

3 conformance fixtures

The single-space strip on a verbatim span (PART 3 code_span) drops one leading and one trailing space, but not when the content consists entirely of space characters - those spans keep every space. Without this guard a formatter round-trip loses the content, since a span stripped to empty has no writable source spelling.

carve
A single space ` ` and two spaces `  ` are preserved.
html
<p>A single space <code> </code> and two spaces <code>  </code> are preserved.</p>

Ordinary content still strips one space from each side, so the guard is narrow.

carve
But ` a ` strips one space from each side.
html
<p>But <code>a</code> strips one space from each side.</p>

The sigil-prefixed verbatim forms - the inline literal (section 27) and math (section 18) - share the same strip rule, so they keep all-space content too.

carve
Literal !`  ` and math $`  ` keep their spaces.
html
<p>Literal    and math <span class="math inline" role="math">\(  \)</span> keep their spaces.</p>

A code span closes only on a run of its own length, whatever the length ​

10 conformance fixtures

The opener is a maximal backtick run and the closer is a maximal run of the same count (PART 3 code_span), for a run of any length. A run of another length inside the span is content, and a run the span cannot close on leaves it open to the end of the block (markup-carve/carve#2144).

carve
a ````b```` c
html
<p>a <code>b</code> c</p>
carve
a ````b``` c
html
<p>a <code>b``` c</code></p>
carve
a ````b```c```` d
html
<p>a <code>b```c</code> d</p>
carve
a `b``c` d
html
<p>a <code>b``c</code> d</p>
carve
a ````````b```````` c
html
<p>a <code>b</code> c</p>

The same pairing holds wherever a code span is skipped whole: inside a forced span, in a substitution or an insertion, and in an image's alt text, where it decides whether a caption line below finds its image.

carve
{*a ````b```` c*} d
html
<p><strong>a <code>b</code> c</strong> d</p>
carve
{~a ````~>```` b~>c~} d
html
<p><del>a <code>~&gt;</code> b</del><ins>c</ins> d</p>
carve
{~a ````~}```` b~>c~} d
html
<p><del>a <code>~}</code> b</del><ins>c</ins> d</p>
carve
{+a ````+}```` b+} c
html
<p><ins>a <code>+}</code> b</ins> c</p>
carve
![a ````]```` b](i)
^ cap
html
<figure>
  <img src="i" alt="a ````]```` b">
  <figcaption>cap</figcaption>
</figure>

Adjacent slash and underscore emphasis nest ​

1 conformance fixture

/ and _ open immediately after each other when the preceding delimiter is a true opener, so adjacent pairs nest (they only stay literal as path protection when the preceding delimiter is a closer, e.g. /a/_b_).

carve
/_x_/ and _/x/_
html
<p><em><u>x</u></em> and <u><em>x</em></u></p>

A delimiter after an underscore or slash opens only when that one pairs ​

6 conformance fixtures

A delimiter never opens directly after _, nor / or _ directly after /, unless that _ or / itself opens a span that closes (CARVE-P3-013, markup-carve/carve#2156). Only pairing decides that, so an unpaired guard keeps the delimiter after it literal even at the start of a line, whatever the delimiter.

carve
_*x* q

_/x/ q

/_x_ q

_~x~ q

_=x= q
html
<p>_*x* q</p>
<p>_/x/ q</p>
<p>/_x_ q</p>
<p>_~x~ q</p>
<p>_=x= q</p>

The same after a space in mid-line.

carve
a _*x* q
html
<p>a _*x* q</p>

CONTROL. A guard that pairs lets the delimiter after it open, so the spans nest.

carve
_*x*_ q

/_x_/ q

_*x* y_ q
html
<p><u><strong>x</strong></u> q</p>
<p><em><u>x</u></em> q</p>
<p><u><strong>x</strong> y</u> q</p>

An unpaired guard leaves no open strong span to hold a forced one literal.

carve
_*{*x*} q
html
<p>_*<strong>x</strong> q</p>

A caption's # placeholder is outside a span the guard keeps from opening.

carve
![p](p.png)
^ Figure _*# x* q
html
<figure>
  <img src="p.png" alt="p">
  <figcaption>Figure _*1 x* q</figcaption>
</figure>

Settling the guard does not move the quotes after it. The apostrophe follows a straight quote, so it closes.

carve
_*x* \"'a' "
html
<p>_*x* "’a’ “</p>

Quote flanking after an escaped character ​

1 conformance fixture

A backslash-escaped character still flanks as the character it is, so a quote that follows it decides direction from the literal. \{ is an opening bracket and opens the quote; \< and \* are not, so the quote closes - the same decision the unescaped character produces. Escaping changes what the character is, not what it flanks like.

Worth pinning because the escape is consumed before the quote is resolved, so an implementation that loses the literal at that point silently flips the direction.

carve
\{"quoted"\} and \<"q"\> and \*'q'\*

A \{ before a quote still opens it: \{"open"

Unescaped for contrast: {"open"}
html
<p>{“quoted”} and &lt;”q”&gt; and *’q’*</p>
<p>A { before a quote still opens it: {“open”</p>
<p>Unescaped for contrast: {“open”}</p>

A quote after a bare delimiter follows what that delimiter does ​

7 conformance fixtures

A smart quote takes its side from the character before it (PART 3). A delimiter that opens a span puts the quote at the start of that span's content, which opens the quote; a delimiter that opens nothing is the preceding character, and none of *, _ and ~ is an opening context (markup-carve/carve#2164).

carve
*"x"

a *"x" b

a **"x" b
html
<p>*”x”</p>
<p>a *”x” b</p>
<p>a **”x” b</p>

CONTROL. The same delimiters where they do open a span.

carve
a *"x"* b

a ~*"x"* b
html
<p>a <strong>“x”</strong> b</p>
<p>a ~<strong>“x”</strong> b</p>

What stands before the opener does not reach the quote.

carve
a.*"x"* b

a,_"x"_ b
html
<p>a.<strong>“x”</strong> b</p>
<p>a,<u>“x”</u> b</p>

A forced span's own delimiter opens it, so a quote at its content start opens.

carve
x{*"y"*}
html
<p>x<strong>“y”</strong></p>

Inside it that delimiter is literal, and a quote after one closes.

carve
{*a *"x"*}
html
<p><strong>a *”x”</strong></p>

A quote after a closed one follows the glyph that one took.

carve
*"'x'
html
<p>*”’x’</p>

CONTROL. An intraword delimiter opens nothing, and / and = are opening characters in their own right.

carve
a*"x"* b

/"x"

="x"
html
<p>a*”x”* b</p>
<p>/“x”</p>
<p>=“x”</p>

A caption's placeholder is any # that does not begin a tag ​

9 conformance fixtures

CARVE-P2-022 calls the first bare # in a caption's top-level text the number placeholder, and bare is "a # that does NOT begin a tag": one followed by whitespace, :, ., the caption's end, or any character a tag name does not take. A tag name takes letters, digits, _ and -, so every other character leaves the # bare (markup-carve/carve#2165).

carve
![p](p.png)
^ Figure #* q
html
<figure>
  <img src="p.png" alt="p">
  <figcaption>Figure 1* q</figcaption>
</figure>
carve
![p](p.png)
^ Figure #, q
html
<figure>
  <img src="p.png" alt="p">
  <figcaption>Figure 1, q</figcaption>
</figure>
carve
![p](p.png)
^ Figure ## q
html
<figure>
  <img src="p.png" alt="p">
  <figcaption>Figure 1# q</figcaption>
</figure>
carve
![p](p.png)
^ Figure #é q
html
<figure>
  <img src="p.png" alt="p">
  <figcaption>Figure 1é q</figcaption>
</figure>

CONTROL. A # that does begin a tag is a tag, and the caption is unnumbered.

carve
![p](p.png)
^ Figure #1 q
html
<figure>
  <img src="p.png" alt="p">
  <figcaption>Figure <span class="tag"><strong>#1</strong></span> q</figcaption>
</figure>
carve
![p](p.png)
^ Figure #-a q
html
<figure>
  <img src="p.png" alt="p">
  <figcaption>Figure <span class="tag"><strong>#-a</strong></span> q</figcaption>
</figure>

Only the first bare # is the placeholder.

carve
![p](p.png)
^ Figure x #* y #: z
html
<figure>
  <img src="p.png" alt="p">
  <figcaption>Figure x 1* y #: z</figcaption>
</figure>

A # glued to the word before it is still bare when a tag name cannot follow.

carve
![p](p.png)
^ Figure#* q
html
<figure>
  <img src="p.png" alt="p">
  <figcaption>Figure1* q</figcaption>
</figure>

CONTROL. With a tag name after it the # is glued into the word instead.

carve
![p](p.png)
^ Figure x#tag q
html
<figure>
  <img src="p.png" alt="p">
  <figcaption>Figure x#tag q</figcaption>
</figure>

A comment inside a forced span or the combined token ends at its closer ​

4 conformance fixtures

CARVE-P9-042 names an explicit closer, not only a table cell's | or link text's ], as a boundary a %% comment does not cross: a forced span {X...X} at its X}, and the combined /*...*/ token at its */ (markup-carve/carve#2167). The comment still runs to a line break first where one comes sooner.

carve
{*a %% b*} y
html
<p><strong>a</strong> y</p>

A different forced delimiter, the same bound.

carve
{_a %% b_} y
html
<p><u>a</u> y</p>

The combined token, on one line and across a soft break.

carve
/*a %% b*/ y
html
<p><strong><em>a</em></strong> y</p>
carve
/*a %%c
b*/ y
html
<p><strong><em>a
b</em></strong> y</p>

An unresolved reference's literal source is HTML-escaped like any other text ​

3 conformance fixtures

PART 10 SS2 escapes &, < and > in text content, and a no-break space is the one exception: it serializes as &nbsp;, matching an escaped space and a line block's preserved indentation (markup-carve/carve#2168). An unresolved reference falls back to its literal source (PART 9R R1), and that source is text content like any other -- the fallback used to splice it in unescaped.

carve
[ x][]

![ x][]
html
<p>[&nbsp;x][]</p>
<p>![&nbsp;x][]</p>

The same escaping reaches the label and an attribute block on an unresolved reference, not only the bracket text.

carve
[x][a & b]

[x][]{title="a & b"}
html
<p>[x][a &amp; b]</p>
<p>[x][]{title="a &amp; b"}</p>

CONTROL. Plain source needs no escaping and renders unchanged.

carve
[x][]
html
<p>[x][]</p>

Adjacent strong spans use HTML only where their delimiters merge ​

1 conformance fixture

PART 11 §8c requires Markdown delimiters where they read back as the same construct. Two adjacent strong spans need an HTML fallback for the second span because the two delimiter runs otherwise merge into one strong span.

carve
{*a*}{*b*}
html
<p><strong>a</strong><strong>b</strong></p>

A quote after an escaped quote closes ​

5 conformance fixtures

A smart quote takes its side from the character before it (PART 3). An escaped quote renders as the straight character, which is not an opening context, so the quote after it closes, whichever quote came before the escape (markup-carve/carve#2158).

carve
"a\""

"\""
html
<p>“a"”</p>
<p>“"”</p>

The same for the single quote.

carve
'a\''
html
<p>‘a'’</p>

The escape and the quote after it need not be the same character.

carve
"a\''

'a\""
html
<p>“a'’</p>
<p>‘a"”</p>

An apostrophe after an escaped quote inside a quotation closes too.

carve
"\"'q'"
html
<p>“"’q’”</p>

CONTROL. After an unescaped opening quote a quote opens, so quotations nest.

carve
"'q'"
html
<p>“‘q’”</p>

A combined bold-italic span may cross a line ​

1 conformance fixture

/*…*/ is one construct (bold_italic), and like any inline span it folds across the lines of its paragraph. The nesting is the same either way: strong outside emphasis. Nothing pinned the multi-line form, and the oracle inverted it there - pairing the two delimiters separately gives emphasis outside strong.

carve
/*multi
line*/
html
<p><strong><em>multi
line</em></strong></p>

A bare closer does not reach inside a braced inline ​

3 conformance fixtures

A code span and a braced inline are opaque to a bare delimiter (PART 9 §9 E2a). A bare ~ inside {/y~/} cannot close the strike opened before the braces, so that ~ is literal content of the italic span, and the leading ~, left with no closer it can reach, is literal too.

carve
~{/x/}{/y~/}
html
<p>~<em>x</em><em>y~</em></p>

The outer span still closes at a bare closer outside the braces. Every member of the braced family is opaque, the editorial forms as well as the forced spans.

carve
~a{+b~+} c~
html
<p><s>a<ins>b~</ins> c</s></p>

The rule is the same for every bare delimiter.

carve
*a {_b* c_} d*
html
<p><strong>a <u>b* c</u> d</strong></p>
2 conformance fixtures

A link destination and an autolink are opaque to a bare delimiter (PART 9 §9 E2a), so the slashes of a URL cannot close emphasis opened before the link.

carve
/see [x](http://a.b/c) now/
html
<p><em>see <a href="http://a.b/c">x</a> now</em></p>
carve
/see <http://a.b/c> now/
html
<p><em>see <a href="http://a.b/c">http://a.b/c</a> now</em></p>

A tag inside a literal brace run is still a tag ​

1 conformance fixture

A trailing {…} on a heading is not an attribute list - Carve is djot-strict, so the braces are inline content (§15 A7). Their CONTENTS are inline content too, which is where #word is a tag (§19). Both rules apply at once: the braces stay as typed and the tag inside them renders.

carve
# H {#id .cls}
html
<section id="H-id-cls">
  <h1>H {<span class="tag"><strong>#id</strong></span> .cls}</h1>
</section>

A math span's base class keeps the class slot in place ​

2 conformance fixtures

math inline is a mandatory base class, so it is prepended INSIDE the class slot and the slot stays at the first-appearance position of a class in the author's order (PART 10 §1). An id written before any class is still serialized first.

markup-carve/carve#1168 pinned this for the generic ext-NAME fallback. The math span carries a base class the same way and was missed, because no case put an id before a class on it (markup-carve/carve#1164).

carve
$`E=mc^2`{#i .c k=v}
html
<p><span id="i" class="math inline c" k="v" role="math">\(E=mc^2\)</span></p>

With no authored class there is no slot to keep, so the base class leads.

carve
$`E=mc^2`{#i k=v}
html
<p><span class="math inline" id="i" k="v" role="math">\(E=mc^2\)</span></p>

A marker glued to a name opens nothing ​

3 conformance fixtures

The boundary rule has a right half. A # or @ written directly against the end of a tag or mention name is preceded by a WORD character - the last character of that name - so it cannot open a second one, exactly as a#b cannot open one (PART 9 §7).

The word production absorbs a glued marker between two word characters, which is that rule stated structurally, but a run that STARTS with a marker never enters word at all - so this shape was unreachable until markup-carve/carve-js#1029 made {#i#j} literal text rather than an attribute block, and the executable spec read it as two tags (markup-carve/carve#1156).

carve
#i#j
html
<p><span class="tag"><strong>#i</strong></span>#j</p>

A mention takes the same guard.

carve
@a@b
html
<p><span class="mention"><strong>@a</strong></span>@b</p>

Separated by a space, both open normally - that is the control the guard must not break.

carve
#a #b
html
<p><span class="tag"><strong>#a</strong></span> <span class="tag"><strong>#b</strong></span></p>

An angle bracket is escaped only where it opens markup ​

1 conformance fixture

PART 11 §8a M1e. On the Markdown target a < is escaped with a backslash when the next character is an ASCII letter, /, ! or ? - the four things that can open raw HTML - and left alone otherwise. A > takes nothing: mid-line it is inert, and at the start of a line it is a block quote marker M1 already covers.

Escaping the < alone is sufficient, because a tag that cannot open cannot be closed. Every engine used to rewrite both brackets to entities, unconditionally and with no clause behind it (markup-carve/carve#1148); an entity is not an escape, since it replaces the character rather than protecting it.

carve
a < b and a <b> c and x > y
html
<p>a &lt; b and a &lt;b&gt; c and x &gt; y</p>

An underscore pair in text is escaped where the line would pair it ​

1 conformance fixture

PART 11 §8a M1b. On the Markdown target a literal _ is escaped when another live _ on the emitted line could close the emphasis it could open. A run that flanks on both sides can neither open nor close, so company_id stays bare, and so does an underscore no second one can pair with.

Adjacency alone missed this: both underscores of _y_ are text, neither stands beside a delimiter, and every reader turned the author's text into emphasis (markup-carve/carve-js#1716).

carve
/x/_y_ in company_id and a_b_c and a _ b _ c
html
<p><em>x</em>_y_ in company_id and a_b_c and a _ b _ c</p>

An underscore pair split across a line break is escaped ​

1 conformance fixture

PART 11 §8a M1b. The underscore's pair condition reads the whole inline content of the block, because a Markdown reader pairs emphasis across a line break inside a paragraph.

carve
/x/_y
z_ w
html
<p><em>x</em>_y
z_ w</p>

The round-trip comparison normalizes a named list ​

2 conformance fixtures

PART 11 §10k. A round trip preserves meaning, not one serialization, so two renderings compare equal when they differ only by a NAMED entry: <del> and <s> are one tag, and a nested emphasis of different strengths whose child spans the whole parent may commute.

The list is closed. Equal strengths are not on it, because an emphasis inside an emphasis and a single strong are different documents.

carve
{~ ~}
html
<p><s> </s></p>
carve
/*x*/
html
<p><strong><em>x</em></strong></p>

A forced opener of an open kind is literal ​

8 conformance fixtures

PART 9 §9 E3. Forced spans push and pop on the same stack as bare ones, so a {X whose kind is already open is content, and a bare X inside a forced span of that kind is content under §22.

carve
a{*{*x*}*}b
html
<p>a<strong>{*x</strong>*}b</p>
carve
*a {*b*} c*
html
<p><strong>a {*b</strong>} c*</p>
carve
{*a *b* c*}
html
<p><strong>a *b* c</strong></p>
carve
{/a *b {/c/}*/}
html
<p><em>a *b {/c</em>*/}</p>

A braced span of another kind starts its own scope, so an opener of the outer kind nests inside it, and a closer inside it cannot reach the outer span. The bare outer span behaves the same way.

carve
{*a {/b {*c*} d/} e*}
html
<p><strong>a <em>b <strong>c</strong> d</em> e</strong></p>
carve
{*a {/b *} d/} e*}
html
<p><strong>a <em>b *} d</em> e</strong></p>
carve
*a {/b *c* d/} e*
html
<p><strong>a <em>b <strong>c</strong> d</em> e</strong></p>

A lone delimiter of the span's own kind is content, whether it touches the opener or stands apart from it, so the pair opens either way.

carve
{==h==}

{//x//}

{= =h= =}

{/ /x/ /}
html
<p><mark>=h=</mark></p>
<p><em>/x/</em></p>
<p><mark> =h= </mark></p>
<p><em> /x/ </em></p>

An escaped hash keeps its escape at a container's content position ​

2 conformance fixtures

PART 11 §8b M2b is decided on the EMITTED LINE, so a container prefix is passed over before the position is read. A hash at the start of a block quote's or a list item's content opens an ATX heading exactly as one at column 0 does, and its escape is kept there. All three engines dropped it, and a round trip through the Markdown target turned the author's text into a heading (markup-carve/carve#1330).

The narrowing itself does not move, which is the half a correction here is most likely to lose. A hash the prefix does not put at the content position is still emitted bare, and so is one that stands there but opens no heading, since M2b's reading is CommonMark's and a run closed by a letter is not a heading. This pair carries both directions.

carve
> \# heading
>
> C\# is a language

- \# heading
- \#tag rest
html
<blockquote>
  <p># heading</p>
  <p>C# is a language</p>
</blockquote>
<ul>
  <li># heading</li>
  <li>#tag rest</li>
</ul>

Nesting needs no rule of its own: the prefix is whatever the writer emitted, > > included. Neither does lazy continuation, which is a parser concept - this writer re-prefixes every line of a container, so the last line below is emitted with its > and read at the content position like any other.

carve
> > \# deep

> a
\# heading
html
<blockquote>
  <blockquote><p># deep</p></blockquote>
</blockquote>
<blockquote><p>a
# heading</p></blockquote>

An unclosed inline literal reaches the end of its block ​

3 conformance fixtures

The ! prefix changes the node kind whether or not the verbatim run closes. Like an ordinary code or math span, an unclosed literal consumes through the end of its containing block and drops trailing whitespace from its content.

carve
!`unclosed
html
<p>unclosed</p>

The same extent applies across a line-block boundary.

carve
::: |
a !`b
c d
:::
html
<div class="line-block">
  <p>a b
c d</p>
</div>

A table row is its own block boundary. Its closing pipe is not content, while an interior pipe remains part of the unclosed literal.

carve
| a !`b | c d |
html
<table>
  <tbody>
    <tr><td>a b | c d</td></tr>
  </tbody>
</table>

A hyphen run opening a word after whitespace is a flag ​

2 conformance fixtures

PART 9 §8 does not convert a hyphen run that is PRECEDED by whitespace (or the start of the content) and FOLLOWED by a non-whitespace character. That shape is a long CLI flag, and converting it mangled ordinary technical prose silently and in the rendered output only - the author saw git log --oneline in the source and the reader got a command that does not run (markup-carve/carve#1443).

The rule is deliberately narrow. Every canonical dash use is unspaced on at least the left, so requiring whitespace on both sides would have removed the feature along with the damage, and requiring the two sides to match in kind would have broken a---- b----- c------, which the corpus already pins.

carve
git log --oneline and --force-with-lease stay literal.

But pages 1--10, the Mon--Fri window, a---- b----- and a -- b all convert.
html
<p>git log --oneline and --force-with-lease stay literal.</p>
<p>But pages 1–10, the Mon–Fri window, a–– b—– and a – b all convert.</p>

An HTML comment is worth pinning rather than leaving to be discovered, because neither half survives and each is lost to a different rule. The opening <!-- is preceded by !, so the flag rule does not reach it and it converts to a dash. The closing --> was flag-shaped and literal until the arrow set took the doubled run (markup-carve/carve#1442); it is now an arrow.

Ruled deliberately rather than papered over. Guarding --> for this one context would put a context-sensitive exception into a set whose argument is that it has none, and Carve escapes raw HTML by default, so a literal HTML comment in prose is already a document about HTML rather than HTML.

carve
<!-- a comment -->
html
<p>&lt;!– a comment →</p>

A braced hyphen pair is an en dash ​

2 conformance fixtures

The flanking rule above refuses a hyphen run that has whitespace before it and a non-whitespace character after it, which is the shape every long CLI flag has. It is also a shape an author sometimes means as a dash, and there was no way to say so. {--} is that way: it renders a single en dash and the braces are consumed (markup-carve/carve#1447).

carve
The bare run is flag-shaped and stays literal: a ---(p) b.

The braced form converts wherever it stands: a {--}(p) b, x{--}y, {--}start.
html
<p>The bare run is flag-shaped and stays literal: a ---(p) b.</p>
<p>The braced form converts wherever it stands: a –(p) b, x–y, –start.</p>

The spelling was an empty deletion, and it cost nothing to take: {--} is a {- opener meeting its -} closer with nothing between them, and the production never permitted that - its content slot is a one-or-more repetition. Exactly one string moves. A deletion that holds a hyphen, or any other content, is untouched.

carve
{---} and {-x-} still delete.
html
<p><del>-</del> and <del>x</del> still delete.</p>

There is no braced em dash. {---} deletes a hyphen, which is a thing an author writes, so a tight em dash - the Spanish, French and Russian dialogue convention

  • is still written as the literal character.

An empty brace pair is not a construct ​

2 conformance fixtures

An opener that meets its own closer with nothing between them opened nothing, and its characters are text. Every forced span and every editorial form carries the same one-or-more content slot, so {//}, {**}, {__}, {~~}, {^^}, {,,}, {==}, {++} and {##} all render literally (markup-carve/carve#1447).

carve
Empty pairs are text: {//} {**} {__} {~~} {^^} {,,} {==} {++} {##}.

A pair that holds something is the construct: {/i/} {*b*} {~s~} {+ins+} {# c #}.
html
<p>Empty pairs are text: {//} {**} {__} {~~} {^^} {,,} {==} {++} {##}.</p>
<p>A pair that holds something is the construct: <em>i</em> <strong>b</strong> <s>s</s> <ins>ins</ins> <span class="critic-comment"> c </span>.</p>

The rule was already written twice - the productions spell every content slot as one-or-more, and the ambiguities guide says {^^} is not a superscript - and carve-js and carve-rs already read it that way. What the empty forms did was worse than render nothing useful: the author's braces disappeared from the output while staying in the source, which is the same silent loss the hyphen flanking rule above refuses.

A fully empty substitution is the one form left alone. Its two halves are independent, and a half-empty substitution is an ordinary edit - a deletion with no replacement, or an insertion replacing nothing - so it needs its own decision.

carve
{~a~>~} deletes, {~~>b~} inserts, and {~~>~} still does both to nothing.
html
<p><del>a</del><ins></ins> deletes, <del></del><ins>b</ins> inserts, and <del></del><ins></ins> still does both to nothing.</p>

A leading escaped caret keeps its escape ​

1 conformance fixture

PART 11 §5's unconditional caret reaches the LEADING position and stops there (§10g). Here the escape is load-bearing under §2's own test: drop it and the image is promoted to a figure with the line as its caption, so the writer emits it back unchanged. This is the shape the empty caret pair (category 388) must not be read as reaching - there the caret leads nothing, dropping the escape re-derives the same document, and the bare form is canonical.

carve
![a](b.png)
\^ not a caption
html
<p><img src="b.png" alt="a">
^ not a caption</p>

An idle escape does not spread from the block that needed one ​

1 conformance fixture

The canonical writer escapes a character IF AND ONLY IF omitting the escape would change the re-parsed document (PART 11 §2). Where it cannot decide an occurrence exactly it may fall back to the conservative form, and PART 11 §2b bounds how far that fallback reaches: the SMALLEST unit whose minimal form fails to re-parse, which is the inline run or the block holding it, never the whole document.

The first paragraph below is indented, so its text is ## H and not a heading. Written back at column zero it WOULD be a heading, so that block escalates and comes back as \#\# H. The second paragraph needs nothing: plain (b) text re-parses as itself, and its parentheses stay bare.

Both spellings render the HTML below and re-parse to the same tree, so the .fmt sidecar beside this pair - not the HTML - is what measures the scope.

carve
  ## H

plain (b) text
html
<p>## H</p>
<p>plain (b) text</p>

An idle escape does not spread from the occurrence that needed one ​

1 conformance fixture

PART 11 §2 escapes a character IF AND ONLY IF omitting the escape would change the re-parsed document, and it takes that decision per OPENER OCCURRENCE. PART 11 §2b bounds how far the conservative fallback REACHES - the smallest unit whose minimal form fails, never the whole document - and corpus 396 pins that bound. It does not make the unit itself finer, and this pair is the half it leaves: WHICH occurrences inside a unit that did escalate carry a backslash.

The paragraph below is indented, so its text is {.note} and not a block attribute. Written back at column zero the first line WOULD open one, so the { is load bearing and comes back escaped. The . and the } beside it are not: neither opens anything at that position, and \{.note} re-parses as the paragraph the author wrote. So does the . ending the second line.

The unit-scoped fallback wrote \{\.note\} here - one knob per unit, so a unit that failed was written conservatively in full and every candidate in it was escaped beside the one that needed it. That spelling renders the HTML below and re-parses to the same tree, so the .fmt sidecar beside this pair - not the HTML - is what measures the occurrence.

carve
 {.note}
 This paragraph.
html
<p>{.note}
This paragraph.</p>

A boolean attribute does not start with an underscore ​

4 conformance fixtures

An identifier may start with _, so {_x_} was two constructs at once: the boolean attribute _x_, and a forced underline. Alone on a line the attribute reading won, the underline was unreachable, and with no block beneath it to attach to the line rendered nothing at all - five characters kept in the source and gone from the output. The bare attribute form gives the collision up (markup-carve/carve#1450).

carve
{_x_}

{_x_} y
html
<p><u>x</u></p>
<p><u>x</u> y</p>

Only the bare form is narrowed. An id, a class and a key/value keep their leading underscore, and none of them can be read as an underline, because none of them ends _}. HTML has no boolean attribute starting with _, and the one underscore attribute in the wild is hyperscript's, which carries a value.

carve
{#_id ._c _k=1 _="on click"}
para
html
<p id="_id" class="_c" _k="1" _="on click">para</p>

A bare {_foo} has no underline reading either, so it is text rather than something else.

carve
{_foo}
para
html
<p>{_foo}
para</p>

The writer follows, and nothing extra pins it. PART 11 §6c shortens an empty-string attribute to its bare name, and it cannot do that here - {_u=""} written as {_u} would be text, and {_x_=""} would be an underline, either way a document that no longer says what it said. §1's parse(fmt(x)) == parse(x) is gated across all three engines over this corpus, so the case below breaks each engine's own round-trip on the day it ships the reading rule and not before.

carve
[x]{_u=""}
html
<p><span _u="">x</span></p>

Substitution content is inline, and only a top-level arrow splits it ​

2 conformance fixtures

A {~ … ~} pair is a substitution when it holds a top-level ~>. The search skips verbatim content, so an arrow inside a code span, math, an inline literal or a comment does not split the pair, and neither does an escaped one. A pair with no top-level arrow is a forced strikethrough (markup-carve/carve#2083).

carve
{~`a~>b~}

{~a `x~>y` b~}

{~/old/~>/new/~}

{~a\~>b~}
html
<p><s><code>a~&gt;b</code></s></p>
<p><s>a <code>x~&gt;y</code> b</s></p>
<p><del><em>old</em></del><ins><em>new</em></ins></p>
<p><s>a~&gt;b</s></p>

Math, an inline literal and an editorial comment are skipped the same way.

carve
{~$`a~>b`~>c~}

{~!`a~>b`~>c~}

{~a{# ~> #}b~>c~}
html
<p><del><span class="math inline" role="math">\(a~&gt;b\)</span></del><ins>c</ins></p>
<p><del>a~&gt;b</del><ins>c</ins></p>
<p><del>a<span class="critic-comment"> ~&gt; </span>b</del><ins>c</ins></p>

Emphasis ​

8 conformance fixtures

The word-boundary rule applies to every bare delimiter (/ * _ ~ = — all single-char). No bare delimiter emphasizes intraword: foo*bar*baz, foo~bar~baz, snake_case, a/b/c, x = 5, key=value all stay literal. For deliberate intraword emphasis, use the forced {X … X} family (below). Superscript and subscript have no bare delimiter at all — they exist only in the braced forms {^…^} / {,…,} (see below). For any bare delimiter:

  • an opener is recognized only if it is not followed by whitespace and is preceded by the start of the line/block, whitespace, or a punctuation character (not by an alphanumeric, _, or the same delimiter) — so a/b/c, foo_bar_baz, snake_case, and //a/ stay literal, while (/x/) and a./b/ open after punctuation;
  • a closer is recognized only if it is not preceded by whitespace and not followed by an alphanumeric character — so x /a/b y stays literal.

Highlight is the single-char = delimiter; the uniform word boundary keeps x = 5, key=value, a=b literal. Every bare delimiter is single-char, so a doubled delimiter (==x==) is literal by the same-delimiter-adjacency rule, just like **x** or //x//. This is stricter than Djot, whose _/* rule is purely whitespace-flanking. ^ and , are not bare delimiters — a comma or caret in prose is always literal text (1,2,3, a,b,c, x, y, z, 2 ^ 3), and superscript/subscript are written with the braced forms only. The boundary rule still allows /usr/local/ → <em>usr/local</em>: the opening / sits at line start and the inner same-type / characters are literal content (Carve does not nest same-type emphasis). The normative rule lives in resources/grammar.ebnf PART 9 §9 and §22.

carve
foo*bar*baz and a/b/c stay literal.
html
<p>foo*bar*baz and a/b/c stay literal.</p>

An opener without a matching closer is left as a literal character.

carve
/foo bar
html
<p>/foo bar</p>

Inner slashes inside a /…/ span are literal content — a path-like span still parses as emphasis.

carve
/usr/local/
html
<p><em>usr/local</em></p>

Whitespace immediately after an opener (or before a closer) blocks emphasis — the delimiter renders literally.

carve
/ not emphasis /
html
<p>/ not emphasis /</p>

An opener may follow punctuation, not only whitespace or the line start.

carve
(/x/) and a./b/
html
<p>(<em>x</em>) and a.<em>b</em></p>

A closer is rejected when followed by an alphanumeric character, so an interrupted path stays literal.

carve
x /a/b y
html
<p>x /a/b y</p>

No bare delimiter produces intraword emphasis — * behaves like / and _.

carve
foo_bar_baz and snake_case stay literal
html
<p>foo_bar_baz and snake_case stay literal</p>

An emphasis opener (any bare delimiter) immediately preceded by the same delimiter or by a literal _ is not valid (this does not affect different-delimiter combinations like /*bold italic*/, and a _ that itself opens an underline span does not block a following opener).

carve
//a/ and snake_/case/
html
<p>//a/ and snake_/case/</p>

Inline extensions ​

5 conformance fixtures

The extension name is an identifier, which permits a leading _, so _ is a valid extension name (grammar extension_name).

carve
:_[x]
html
<p><span class="ext-_">x</span></p>

An identifier must start with a letter or _, so a digit-first name is not a valid extension. :1[x] stays literal text; :a1[x] (a digit after the first letter) is a valid extension and renders as a generic span.

carve
:1[x]
:a1[x]
html
<p>:1[x]
<span class="ext-a1">x</span></p>

The extension_content runs up to the first ] (extension_content = {character - ']'}); a nested ] therefore closes the extension early and the remainder is literal text.

carve
:foo[a [b] c]
html
<p><span class="ext-foo">a [b</span> c]</p>

The four names core does not reserve - samp, var, cite, dfn - are the SemanticSpan extension's, so a core processor leaves them as ordinary attributes.

carve
[x]{samp} [y]{dfn="a term"}
html
<p><span samp="">x</span> <span dfn="a term">y</span></p>

An explicit abbr value takes precedence over automatic abbreviation definitions, avoiding invalid nested <abbr> markup.

carve
*[HTML]: Hyper Text Markup Language

[HTML]{abbr="Custom"}
html
<p><abbr title="Custom">HTML</abbr></p>

Smart typography dashes and quotes ​

8 conformance fixtures

A quote opens (left/opening quote) when it follows start-of-content, whitespace (incl. NBSP), or one of the opening/operator characters ( [ { = : - /; otherwise it closes. So a quote right after =, :, -, /, or an opening paren still opens the first quote (grammar smart_quote).

carve
a="b"
:"q"
-"q"
/"q"
("q")
html
<p>a=“b”
:“q”
-“q”
/“q”
(“q”)</p>

When a quote does not follow one of those opening contexts it closes instead — so a quote right after a closing bracket (} ) ]), after sentence punctuation (. ,), or mid-word always becomes a right/closing quote. An empty "" opens both marks (the second " follows a ", which is not an opening context, yet there is nothing to its right to close against, so it too renders as an opening quote).

carve
}"q"
)"q"
]"q"
."q"
,"q"
a"b
""
html
<p>}”q”
)”q”
]”q”
.”q”
,”q”
a”b
””</p>

The same opening set applies to the single quote '. After (, [, =, :, -, or / a single quote opens (‘); the matching ' then closes (’).

carve
('q')
['q']
='q'
:'q'
-'q'
/'q'
html
<p>(‘q’)
[‘q’]
=‘q’
:‘q’
-‘q’
/‘q’</p>

A single quote after { opens too — shown on its own line because a trailing {…} would otherwise be read as an attribute block.

carve
{'q'}
html
<p>{‘q’}</p>

A single quote before a digit is an apostrophe (decade elision), so a digit pair becomes apostrophes on both sides; a quote before a letter in an open context opens.

carve
the '70s and '24' and 'word'
html
<p>the ’70s and ’24’ and ‘word’</p>

A run of four or more hyphens is allocated into em/en dashes (all em if divisible by 3, all en if by 2, otherwise max em-dashes with an en remainder) — matching djot.

carve
a---- b----- c------
html
<p>a–– b—– c——</p>

Longer runs follow the same allocation with no leftover hyphen: seven is one em plus two en, eight is four en, ten is five en, eleven is three em plus one en, and thirteen is three em plus two en.

carve
a------- b-------- c---------- d----------- e-------------
html
<p>a—–– b–––– c––––– d———– e———––</p>

The open/close decision reads the character before the quote. A bare emphasis delimiter is not that character - the quote sees the start of the emphasis CONTENT - and nothing at all before a quote opens it. A quote directly after another one follows whichever half that one resolved to, so a nested pair opens while an empty pair stays closed.

carve
*'q'*

"hello"

"'nested'"

a*'q'*
html
<p><strong>‘q’</strong></p>
<p>“hello”</p>
<p>“‘nested’”</p>
<p>a*’q’*</p>

Inline code ​

7 conformance fixtures

An unclosed run is opaque: an emphasis delimiter or link tail after it is verbatim content, so the surrounding construct never closes.

carve
*a ` b*
html
<p>*a <code> b*</code></p>

A forced span's closer ends an unclosed run inside it, and the run's trailing whitespace is stripped there as at the end of a block.

carve
{~` ~}
html
<p><s><code></code></s></p>

The run is over at that closer, so a link, image or span after it opens as usual. Link text is scanned from its own [.

carve
Run {*`make*}[the docs](u).

A {_``x_}![logo](p.png) here.

A {=`y=}[note]{.k} here.

{+`added+} [n](u)
html
<p>Run <strong><code>make</code></strong><a href="u">the docs</a>.</p>
<p>A <u><code>x</code></u><img src="p.png" alt="logo"> here.</p>
<p>A <mark><code>y</code></mark><span class="k">note</span> here.</p>
<p><ins><code>added</code></ins> <a href="u">n</a></p>

A backtick that is content of an earlier construct opens no run at all, so it does not stop a link after that construct.

carve
[a]{title="`"} and [n](u)

{% a ` %}[n](u)

[m](u`) and [m](`v) and [n](w)
html
<p><span title="`">a</span> and <a href="u">n</a></p>
<p><a href="u">n</a></p>
<p><a href="u`">m</a> and <a href="`v">m</a> and <a href="w">n</a></p>

The closer bounds a run behind a math or literal prefix the same way, and a literal nests in a forced span like any other construct.

carve
x{*$`a*} here

x{~$$`b~} here

x{/!`c/} here

x{_!`d`_} here
html
<p>x<strong><span class="math inline" role="math">\(a\)</span></strong> here</p>
<p>x<s><span class="math display" role="math">\[b\]</span></s> here</p>
<p>x<em>c</em> here</p>
<p>x<u>d</u> here</p>

A run whose equal-length closer sits later in the block is a closed code span, even where a forced or editorial closer falls inside it.

carve
x{*`a*} and `b`

{*x`a*}`*}

x{+`a+} and `b`

{+x`a+}`+}
html
<p>x{*<code>a*} and </code>b<code></code></p>
<p><strong>x<code>a*}</code></strong></p>
<p>x{+<code>a+} and </code>b<code></code></p>
<p><ins>x<code>a+}</code></ins></p>

The strip at a forced closer takes a trailing line break too, as at the end of a paragraph. In a line block the break is content and stays.

carve
x{*`a
*}

::: |
x{*`a
*}
:::
html
<p>x<strong><code>a</code></strong></p>
<div class="line-block">
  <p>x<strong><code>a
</code></strong></p>
</div>

Non-breaking space ​

1 conformance fixture

The non-breaking space U+00A0 serializes as &nbsp; in text and code-span output. (In a heading id it is kept as the raw byte instead - ids are not entity-encoded; see Heading IDs.)

carve
`a b`
html
<p><code>a&nbsp;b</code></p>

Superscript and subscript ​

1 conformance fixture

A ^ or , outside the braced forms is always literal text — even where a bare delimiter's word-boundary rule would have matched:

carve
typo ,oops, happens and 10^6^ things and x ^2^ y
html
<p>typo ,oops, happens and 10^6^ things and x ^2^ y</p>

Symbols ​

2 conformance fixtures

Boundary and name-shape cases — all of these stay plain text. The colon is glued to a word character in the first two, and _ cannot open a name in the third:

carve
a:b:c and 10:30: meeting, :_x: too.
html
<p>a:b:c and 10:30: meeting, :_x: too.</p>

A symbol is recognized before smart typography, so a name made of typographic punctuation is a symbol rather than a substitution: :+-: is the symbol +-, not a ± between colons. The typographic forms still apply wherever no symbol opens — a +- b is a ± b, and word:+-: has no boundary, so its +- is substituted:

carve
Vote :+1: or :-1:. Tolerance :+-: is a symbol, but a +- b and word:+-: are not.
html
<p>Vote :+1: or :-1:. Tolerance :+-: is a symbol, but a ± b and word:±: are not.</p>

Released under the MIT License.