Differences from Python re#

REAL targets parity with Python’s re, and a randomized differential fuzzer plus a large parity corpus enforce it. A handful of behaviours diverge on purpose; each is listed here with its rationale, and each is pinned by a direct, version-independent assertion in TestIntentionalDivergences (bindings/python/tests/test_real.py) so the contract cannot drift silently.

Note

Excluded by design — a closed door, not a missing feature. Backreferences, recursion, and callouts are not on a roadmap: each makes matching super-linear and would reopen the ReDoS door this engine exists to close. Refusing them is the point of a linear-time engine, so they raise a clear real::regex_error rather than being “not yet supported”. This is the one promise REAL will not trade for feature coverage.

UTF-8 character classes (code-point mode)#

In code-point mode (the default), a character class carries specific non-ASCII code-point members and ranges: [é], [éàü], [à-ÿ], [a-zé], and their negations [^é] / [^à-ÿ] all match exactly like re (the code-point oracle). They compile to the canonical UTF-8-ranges automaton, so a class never matches an overlong or surrogate byte sequence — [é] matches only C3 A9, and [^é] matches every valid code point except é. . and an ASCII-only negated class such as [^x] match any valid non-ASCII code point, and no malformed byte sequence: neither matches a lone FF or a truncated C3, as neither does in re. A malformed UTF-8 member in the pattern is a real::regex_error. In bytes mode there is no code point there to be malformed — the unit is a byte — so a non-ASCII class member is simply the bytes it is written with: re.compile(b"[\xc3\xa9]") is a two-byte class matching either byte, and REAL compiles it to exactly that, as does std::regex<char>. Raw-byte semantics on both sides, and in every spelling of it: the bytes bare, \xHH, or a backslash before each byte.

Unicode text-mode semantics: case folding, shorthands and boundaries#

Under IGNORECASE in text (code-point) mode, REAL does full Unicode simple case folding for literals, classes and ranges — identical to re.IGNORECASE: é matches É, [é]/[à-ÿ] fold, k matches Kelvin (U+212A), ß matches ẞ (but not ss — simple, not full, folding). A cased literal is promoted to a foldable code point; a \xHH or octal byte escape keeps byte provenance and never folds (\xe9 under IGNORECASE does not match É) — the same deliberate byte-vs-code-point split described in the hex-escape section below. A shorthand inside a class is not folded, as in re: (?i)[^\W\d_] is the letters, iota included, although U+0345 (a non-word mark in \W) folds to iota. A \p{…} property does fold, as in the regex module: (?i)\p{Lu} matches a.

The \w \W \d \D \s \S shorthands are Unicode in text mode (matching re): \\w matches any Unicode word code point (é, ٣, ², astral letters…), \d any Nd digit (Arabic ٣, fullwidth 9…), \s exactly Python re’s whitespace — Unicode White_Space (NBSP, U+2028…) plus the four separators U+001C–U+001F that re (via str.isspace) adds and Unicode does not. This holds in or out of a character class ([\\w], [\\d.], and [\\W]/[\\D]/[\\S] — accepted as their Unicode complements, with [^\\W] == \\w etc). The \\b \\B word boundaries (and the \\< \\> word-start/end extensions) use the same Unicode word-ness, so \\bété\\b matches. In bytes mode, and under flags::ascii (re.A), every shorthand and boundary stays ASCII (byte-for-byte std::regex<char>), and IGNORECASE folds ASCII only, so the compat layer is unaffected.

Case-insensitive matching follows CPython’s equivalences (via str.upper/lower), so the Turkish dotless/dotted I fold with I/i — (?i)I matches ı (U+0131) and İ (U+0130), exactly as stdlib re does (re.fullmatch("(?i)I", "ı") is True) and as (?i)\p{Lu} therefore matches ı. Unicode simple CaseFolding — used by RE2 and the Rust regex crate — keeps ı apart instead. This is one code point (two, with İ) where the two conventions differ; both engines are correct for their contract (REAL for re-parity, the crate for UTS#18 folding). The Rust binding’s differential masks exactly this set (its ICASE_FOLD_DELTAS, computed by asking both engines), the twin of the \\w/\\s mask above.

A known CPython 3.14 oracle bug (not a REAL divergence)#

A scoped ascii group — (?a:...), with a among the added letters — behaves inconsistently in CPython 3.14 specifically when it is the pattern’s own first construct, with nothing preceding it at all (not even a zero-width assertion or a single literal byte): a negated shorthand \S/\D/\W inside it wrongly fails to match U+001C–U+001F, the same four separators Unicode text-mode semantics: case folding, shorthands and boundaries above documents as included in text-mode \s and excluded from ascii-mode \s. Concretely, on '\x1c': re.search(r"(?a:\s)", ...) correctly returns no match (not ascii whitespace) — but re.search(r"(?a:\S)", ...) also returns no match, an outright partition violation (\x1c matches neither \s nor its own negation \S). Prepending literally anything before the group — \B, ^, a lookahead, or a single literal byte — “fixes” it on re’s side; a runtime search(text, pos, endpos) offset does not (confirmed a compile-time artifact of re’s own opcode order, not a runtime one — pos > 0 still fails). REAL has no such inconsistency: (?a:\S) and \B(?a:\S) agree, and every one of these four codepoints partitions correctly between \s/\S regardless of what, if anything, precedes the scoped group.

This is filtered out of the differential fuzzer rather than “fixed” on REAL’s side — there is nothing on REAL’s side to fix, re is the one behaving inconsistently — and pinned directly against re’s own (buggy) behaviour in TestIntentionalDivergences. test_scoped_ascii_negated_shorthand_leading_bug (bindings/python/tests/test_real.py), so a future CPython release that fixes its own bug is caught (the pin asserts re’s buggy value too, not just REAL’s correct one) rather than silently masked forever.

\xHH (HH ≥ 0x80) is byte-level#

\xHH matches the raw byte 0xHH, not the codepoint chr(HH). On a str pattern re reads \xe9 as U+00E9 (é); REAL reads it as the single byte 0xE9 — which a well-formed str’s UTF-8 of é (C3 A9) never contains. Use é for the codepoint.

One deliberate exception makes \xHH context-dependent in a str (code-point-mode) pattern: inside a character class [\xHH], a value ≥ 0x80 is the code point U+00HH (so [\xe9] == [é] == [é]), because a class member is a code point, whereas outside a class \xe9 stays the single byte 0xE9. In bytes mode \xHH is always the raw byte, in a class or not (so a bytes class is byte-for-byte std::regex<char>).

\b / \B on an empty string or region#

By definition an empty string has no word boundary, so \B (not-a-boundary) matches at position 0 while \b does not. REAL implements this by-definition behaviour, which is also re’s since Python 3.11 (re < 3.11 had the opposite quirk for \B on the empty string). The differential fuzzer therefore cannot use re as an oracle on an empty region and skips it; the assertions pin REAL’s behaviour directly.

Capture of a nullable loop’s final iteration#

The rule. For a * or + loop whose body can match the empty string, re runs one final empty iteration after the last consuming one and records its (zero-width) capture for the group; REAL records the last consuming iteration instead. The overall match — group 0, and every match span — is identical in both; only the inner group’s captured span differs.

This remains only for a loop whose body consumes on its last productive iteration (so the inner group captured that run). A loop whose body can only ever match empty (()*, ($)*) now agrees with re — the greedy-loop empty-exit fix routes its empty iteration to the loop exit, applying that iteration’s capture, exactly as re does.

The verified forms (search; re value ↦ REAL value for group 1):

Pattern

Input

re group 1

REAL group 1

(a*)*

"a"

(1, 1)

(0, 1)

diverges (consuming body)

(a*)+

"a"

(1, 1)

(0, 1)

diverges (consuming body)

(a|)*

"a"

(1, 1)

(0, 1)

diverges (consuming body)

()*

""

(0, 0)

(0, 0)

now agrees with re

($)*

""

(0, 0)

(0, 0)

now agrees with re

Lineage. This is the linear-engine convention: RE2, the Rust regex crate, and Go’s regexp share REAL’s behaviour (they do not re-enter a loop on an empty match). re and PCRE, being backtrackers, take the extra empty step. REAL sits with the linear engines by design — the same family whose linear-time guarantee it shares.

It also touches the compat contract. The same group-capture difference shows up under real::compat against the local std::regex (ECMAScript is a backtracker too). The exhaustive compat check measures it — 4 548 cases out of 3 218 434 in the tier-1 space — and tolerates only this exact signature (whole match identical, the differing group is std’s zero-width empty-final iteration), failing on any other divergence. It is documented, not routed to std, on purpose: routing nested nullable quantifiers to a backtracker would forfeit the linearity real::compat exists to keep. See COMPATIBILITY.md.

When to revisit. Only if a future goal is exact re/std capture parity for these degenerate loops, which would mean adding a trailing empty-iteration step purely to update a capture — a cost with no matching benefit for a linear engine, and none that would justify handing catastrophic-backtracking patterns to a backtracking fallback. Pinned, both values, in TestIntentionalDivergences; the compat side is guarded by the exhaustive-compat gate’s exact-signature discriminator.

Span of a loop whose body can match the empty string#

The rule. An unbounded */+ loop whose body can match the empty string — an empty or optional alternative, (|a)*, (?:a?|b)+ — can end with a different span in re and REAL. re ends the loop on every iteration that matches nothing. REAL runs a loop as a Thompson split → body → back edge, one thread per instruction and position: a path that comes back, without consuming, to an instruction already reached at that position is dropped (that is how an empty iteration cannot loop forever), except a jump back to the loop’s head, which takes the loop’s exit. So an iteration that matches nothing ends the loop when nothing else of the body was reached at that position before it, and is dropped otherwise, leaving the body’s later alternatives free to consume:

Pattern

Subject

re

REAL

(?:a?|b)*

"b"

search (0,0)

search (0,0)

(?:a?|b)+

"ab"

search (0,1)

search (0,2)

(?:\s*|,)+

"  ,  ,x"

search (0,2)

search (0,6)

(a||b)*

"ab"

finditer (0,1)(1,1)(1,2)(2,2)

finditer (0,2)(2,2)

(|a)*, (|a)+

"aa"

finditer (0,0)(0,1)(1,1)(1,2)(2,2)

finditer (0,0)(0,2)(2,2)

search, match and fullmatch can differ, not only the forced-non-empty step of finditer / sub. An empty last branch (a\|)* agrees (its empty branch is the loop’s own exit), and so do the bounded forms (\|a){2}, (\|a){1,3} (unrolled, no back edge). RE2 and Go build loops the same way and agree with REAL on these examples, though not on every pattern of the family; the regex crate does a third thing ((\|a)* on "aa": (0,0)(1,1)(2,2)), so re is the arbiter, never the crate. This is distinct from Capture of a nullable loop’s final iteration, where the match spans are identical and only a group capture differs.

Portable form. Write the body so that no iteration can match nothing — (?:a|b)* rather than (?:a?|b)*, (?:\s+|,)* rather than (?:\s*|,)+ — and every engine gives the same answer.

When to revisit. Ending the loop on every empty iteration, as re does, needs each path of a closure to know which iterations it began at this position: one more field in every frame of the Pike closure, the engine’s hottest loop, charged to every pattern for a class this rare. Weighed on 2026-10-10 and left as is; pinned in TestIntentionalDivergences and in the C++ suite.

Rejected by design#

Each of these raises a clear real::regex_error, for a reason worth keeping:

  • Backreferences ((a)\1, (?P=name)) — a backreference makes the language non-regular; supporting it would forfeit the linear-time, ReDoS-safe guarantee that is REAL’s reason to exist.

  • Conditional groups (?(id)yes|no), pattern recursion ((?R), (?1), (?-1)), subroutine calls ((?&name), (?P>name)) and callouts ((?C), (?C1)) — non-regular control flow, each super-linear in the worst case. Excluded for the same ReDoS reason (see the note at the top of this page).

Each of these is reported as error_kind::unsupported — the C ABI’s REAL_ERR_UNSUPPORTED — and names itself: callouts are not supported, not a generic “unknown extension”, which is what a mistyped extension still gets. So a binding can tell a deliberate exclusion from a malformed pattern without reading the message, and a reader is told which wall they hit.

If you must run one of these anyway, the opt-in is only where a backtracking engine is within reach and actually implements the construct, and it forfeits the linear-time guarantee for that pattern alone. That is a narrower set than “everything above”, and it differs per surface:

construct

Python fallback=True (→ re)

C++ policy::fallback (→ std::regex)

Rust fallback (→ regex)

backreference

yes

yes

no

conditional group

yes

no

no

recursion, subroutine, callout

no

no

no

The Python binding does not guess: it asks re whether it would compile this pattern with these flags, and says fallback=True does not help here: re refuses this pattern too when it would not — so the advice is never a door that is not there. Rust’s fallback reaches none of this row — it delegates to the regex crate, which is linear too and refuses a backreference just as REAL does; there it closes the \\p{...} namespace and folding gaps instead. A native real::regex has no opt-in at all: it is the linear engine or nothing, which is what makes the chosen backend (Pattern.engine, Regex::engine(), the compat policy) worth reporting.

Unlike the above, a possessive quantifier or atomic group over a compound body ((?:ab)*+, (?>ab|a)) is rejected “not supported yet”, not “by design” — see Possessive quantifiers and atomic groups (Tier 1 — a capability beyond re? no: parity, deliberately narrower for now). The linear-time argument for the general case is expected to hold (a bounded compound body looks like it can reuse the same priority-kill sub-VM technique lookaround already ships), the gap is VM-integration work not yet done, not a ReDoS concession. Recursion/callouts/conditionals above stay permanently excluded regardless — that door really is closed.

The module surface (Python binding)#

Everything above is about the pattern language. Six differences are in the module instead, and a import real as re drop-in meets them without writing a pattern at all:

  • Flags are plain int, not a RegexFlag enum. real.I | real.M is 10, where re.I | re.M reprs as re.IGNORECASE|re.MULTILINE; there is no real.RegexFlag, so an annotation spelled flags: re.RegexFlag becomes flags: int. The values themselves are re’s, so any flag expression that works there works here.

  • Pattern.flags reports only the flags passed to compile(), not inline flags in the pattern text, and does not add re.UNICODE on str patterns. re.compile('a', re.I).flags is 34 (I|UNICODE) while REAL’s is 2; re.compile('(?i)a').flags is also 34 because re folds inline flags into the field, while REAL’s stays 0. The inline flags do apply — (?i)a matches "A" — so p.flags under-reports; do not compare it to re’s echo or branch on it after inline syntax. Bytes patterns agree on both when no inline flags are used.

  • re.L / re.LOCALE and re.DEBUG raise instead of being ignored. Locale-dependent matching is excluded by design (this engine is Unicode in text mode and raw bytes otherwise, never locale-dependent); DEBUG dumps CPython’s own compiler state, which REAL does not have. An unrecognised flag bit is refused too, rather than silently dropped – and NAMED, so the message says which bit (unknown flag 0x400 passed to real.compile). One exception, and it is the only one: bit 0x1 is ACCEPTED and means nothing. re ignores it on every interpreter this binding supports – it was re.TEMPLATE before 3.12, a no-op for compilation, and is not a flag at all since – so refusing it turned real.compile("a", True), a boolean read as “switch it on”, into an error for a pattern re compiles. Pattern.flags still echoes the bit, as re does. No real.TEMPLATE constant is exported: it exists on no supported version.

  • re.Scanner is absent. It is undocumented in CPython and has no REAL equivalent.

  • A sub callable must return a string; returning None raises instead of erasing. re.sub(r'a', lambda m: None, 'xax') returns 'xx' in CPython — the None is read as an empty replacement — while REAL raises TypeError: expected str replacement. re’s own documentation says the function must return a replacement string, so that acceptance is unspecified behaviour falling out of the implementation rather than a contract; writing it down here would promise something CPython has never promised and could withdraw. return '' is the portable erase, and it behaves identically in both. A non-string that is not None (an int, say) raises in both.

  • A replacement template’s group name follows Python 3.12’s rule on every version. A name in \g<...> must be ASCII decimal digits or an identifier, and ASCII under bytes: \g<+1>, a name of non-ASCII digits and a bytes name holding a byte over 0x7F raise real.error at the name. CPython 3.11 accepts those three with a DeprecationWarning, 3.12 made them errors, so this differs from re on 3.11 only. A well-formed name no group carries raises IndexError, as in re.

re.error and re.PatternError are both spellings of one class here, as they are in re since CPython 3.13, so except on either catches what the other raises.

Possessive quantifiers and atomic groups (Tier 1 — a capability beyond re? no: parity, deliberately narrower for now)#

re (3.11+) and PCRE2 support possessive quantifiers (a*+, a++, a?+, a{n,m}+) and atomic groups (?>...) generally, over any body. Neither RE2 nor the Rust regex crate support them at all — REAL closing part of this gap is the differentiator, in the same class as REAL’s already-shipped bounded lookarounds.

What REAL accepts (Tier 1 — the dominant real-world shape):

  • A possessive quantifier over a bare single atom — a literal byte, a character class, or . — e.g. [^"]*+, \d++, x?+, [a-z]{2,4}+.

  • The same, wrapped in exactly one capturing group — (a)*+ — the group reflects the last successful repetition’s span, matching re’s own semantics for a possessive loop.

  • An atomic group (?>X*) / (?>X+) / (?>X?) / (?>X{n,m}) wrapping a bare or singly-captured atom — desugars to exactly the possessive-quantifier case above, regardless of whether the inner repeat’s own written form used the + suffix (so (?>[^"]*) and (?>\d+), the dominant real-world atomic-group shapes, are NOT rejected as “unbounded” — a compile-time detection-order requirement, not an accident).

  • An atomic group with no repeat at all, over any deterministic (split-free) body, however compound — (?>ab), (?>) — nothing to give back regardless of body shape, since it never loops; compiled inline, at zero extra VM cost.

What REAL rejects, “not supported yet”: a possessive quantifier or atomic group over a compound, repeated body — (?:ab)*+, (?:X++)*+, (?:a?+)*+ — and an atomic group over a genuine alternation, even with no repeat — (?>ab|a), (?>a|ab)b. A design-first spike (D0) confirmed the linear-time argument for the general case is sound in principle (the same isolated, bounded sub-VM technique lookaround already ships could plausibly carry it); implementing it is separately-scoped future work (D2), not a ReDoS concession — see the note in Rejected by design above.

A correctness note on the shipped Tier 1 itself, for anyone extending it: the possessive-loop opcode’s match/no-match decision is made at epsilon-closure time (add_thread, pike.hpp), not deferred to the per-byte step. An earlier design that decided it one round later, in the byte-stepping dispatch, shipped with a real bug: inside an alternation, a lower-priority sibling that resolves in a single round could claim the shared post-alternation convergence point before the (higher-priority, but multi-round) possessive thread’s step()-time exit got a chance to compete for it — plain first-felt-this-generation dedup has no notion of true priority once insertion order like that is violated. The fix — evaluating the atom test at insertion time, in the same priority-ordered closure pass as everything else — is precedented directly above it in the same closure: assert_lookaround already runs a whole sub-VM decision at closure time, so testing one byte/class/codepoint there is a direct extension of an existing pattern, not new architecture. Regression-pinned in possessive_alternation_priority_regression (tests/frontend/test_possessive_atomic.cpp) and found originally by the differential fuzzer’s own possessive-quantifier generator (test_differential_fuzz.py) before it ever shipped.

A superset of re: Unicode property classes \p{…} are built in for General_Category, Script, Script_Extensions and the standard binary properties (see Unicode property classes \p{…} (a capability beyond re)), which standard re rejects outright. Named characters \N{NAME} are also supported — the Python binding resolves the name via unicodedata and the engine takes the resulting code point — so no table lives in C++. REAL additionally accepts the scalar form \N{U+XXXX} directly (a PCRE2-style extension; re does not), and the braced hex form \u{…} (the ECMAScript / regex-crate spelling, a synonym of \x{…}; re does not).

Patterns that compile on one side only#

Everything else on this page is about what a pattern means. Three cases are about whether it compiles at all — the kind a drop-in reader has to know first, because the failure is not a wrong answer but a refusal.

They do not run in the same direction, and the direction decides the risk. A pattern REAL accepts and re refuses is safe for a drop-in — code written for re never reaches it — but it is a portability trap the other way. A pattern re accepts and REAL refuses is the one that breaks the promise: existing code stops compiling.

{n} with nothing to repeat is literal here, an error in re (an extension). {2}a matches the four characters {2}a; re raises nothing to repeat. The same holds for every brace quantifier in that position — {,}a, {,3}a, {2,}a are all literal text here and all errors there. An ill-formed {…} — a{, a{}, a{2,3x — is literal text in both, which is ECMAScript Annex B and what re does too; the divergence is only the well-formed-but-unanchored case, where re checks for a preceding atom and REAL does not.

a{,} is not in that ill-formed list, and this page used to say it was. {,n} is Python’s shorthand for {0,n} and REAL implements it, so {,} is {0,} — unbounded — exactly as a{,3} is {0,3}. Reading it as literal text made a{,} match the four characters a{,} where re matches aaa. Without a comma, a{} stays literal in both: it is the comma that makes the bounds optional rather than absent.

(?<name>…) is accepted here, rejected by re (an extension). The .NET spelling is a synonym of (?P<name>…); re reports unknown extension ?<n. Both spellings name the same group, so a pattern using the .NET form does not port back.

A counted repeat above 1000 is refused here, accepted by re (a limitation). a{1001} raises repetition count too large; re has no ceiling. This one is a design bound, not an oversight: a bounded {n} is expanded into the program, so the cap is what keeps compile time and program size proportionate to the pattern rather than to its counts. It is real::detail::max_repeat_count in config.hpp, and raising it costs compilation, not matching.

Inline flags: a different set, and a wider grammar#

REAL’s inline flag letters are not re’s, in both directions, and the flags group accepts a few spellings re refuses.

The set. REAL adds U (ungreedy — RE2’s spelling, which inverts the default greediness of every quantifier in scope) and has no u or L. So (?U)a+ is REAL-only, and (?iu)a — valid in re, where u is a no-op — reports unknown flag here. re.LOCALE is excluded by design: this engine is Unicode in text mode and raw bytes otherwise, never locale-dependent.

The grammar. re validates flag combinations and rejects several; REAL parses them and applies a rule. (?i-i:a) turns i on and then off — re calls that “flag turned on and off”, REAL applies the removal after the addition, so the group runs case-sensitive. (?i-a) turns off a, which re refuses outright (“cannot turn off flags ‘a’, ‘L’, ‘u’”) and REAL accepts as a return to the Unicode default.

Both are supersets: a pattern written for re never reaches them, so a drop-in cannot break on this. A pattern written here and moved to re can. The rule is stated rather than hidden because it is stable — removal is applied last, always — not because it is recommended.

\N{U+XXXX} scalar escape (a capability beyond re)#

REAL accepts the scalar form \N{U+XXXX} (1–6 hex, a PCRE2-style spelling) as a code point, on both surfaces — the C++ engine and the Python binding — so the two are uniform. re knows \N{…} only as a character name and rejects the U+ form (undefined character name). The name form \N{NAME} is exact re-parity (the binding resolves it via unicodedata); only the scalar form is the extension. Pinned real-accepts / re-rejects in TestIntentionalDivergences.

\u{…} braced hex escape (a capability beyond re)#

REAL accepts the braced form \u{…} (1–6 hex, the ECMAScript / regex-crate spelling) as a code point, on both surfaces — the C++ engine and every binding — so a pattern written for the regex crate’s \u{e9} compiles here. It is the same code-point path as \x{…} and \N{U+XXXX}. re rejects it (incomplete escape \u): four hex digits (\u00e9) or eight (\U0001F600), never braces. Accepting the braces is a superset: it cannot break re-compatible code (which could never use \u{…}). \U{…} is not this form — REAL’s \U stays Python’s 8 fixed digits. Pinned real-accepts / re-rejects in TestIntentionalDivergences.

Unicode property classes \p{…} (a capability beyond re)#

REAL builds in General_Category (\p{L}, \p{Lu}, \p{Nd}, and the groups \p{L}..\p{C}), Script (\p{sc=Greek}, \p{Script=Latin}) and Script_Extensions (\p{scx=Cyrl}), with short codes, the UCD long names (\p{Letter}), gc= / sc= prefixes, loose matching (UAX44-LM3), the single-letter form \pL, and negation \P{…} — on both the C++ engine and the Python binding. Standard re rejects \p entirely (a bad escape error), so accepting it is a superset: it cannot break re-compatible code (which could never use \p). The class matches as a linear code-point predicate (the same klass_cp mechanism as \w), so it stays ReDoS-safe, and flags::ascii (re.A) does not restrict it — a Unicode property is always Unicode. The tables are pinned to a Unicode version and validated exhaustively against the UCD (the regen guards). The standard binary properties (\p{Alphabetic}, \p{White_Space}, \p{Emoji}, …) are built in the same way; other UAX44 namespaces (\p{Bidi_Class=L}, Word_Break, Age, …) raise unsupported (the Rust binding can delegate those via its fallback feature). Pinned in the property and parity suites.

Lone surrogates: refused, where re matches them#

A Python str can hold a lone surrogate ("\ud800"), and re treats it as one more character: . matches it, [^a] matches it, a template can insert it. REAL matches UTF-8, which has no encoding for U+D800..U+DFFF, so it refuses one wherever it meets it. In a pattern, written bare or as the escape \ud800, it raises real.error at the surrogate’s position, which is what lets real.compile(pattern, fallback=True) delegate the pattern to re. In a subject or a str replacement template it raises UnicodeEncodeError, a ValueError; no policy delegates a subject. A bytes pattern and subject are unaffected: there the unit is a byte.

Matching them would take a code-point mode that decodes the three-byte surrogate forms UTF-8 forbids, through the decoder every Unicode scan runs. Pinned in both directions in TestIntentionalDivergences: REAL’s refusal and re’s match.

Nesting depth: a fixed cap here, the interpreter’s stack there#

REAL caps parser nesting at 200 groups (max_nesting_depth, real/core/config.hpp), on every surface. 200 compiles; 201 raises — Python real.error with msg='pattern nesting too deep' and pos=200 — so it arrives through the same handler as every other pattern fault and carries a position you can point at. re has no cap: it recurses until the interpreter’s stack runs out and raises RecursionError, which is not a re.error and carries no position, so except re.error does not catch it.

No depth is quoted for re, because it does not have one: the boundary follows sys.getrecursionlimit(). Measured at the default 1000, re compiled 400 nested groups and failed at 500; raised to 3000 it compiled 900. REAL’s 200 did not move in either run. So the practical difference is a catchable, stack-independent failure here against an uncatchable, stack-dependent one there — and re-targeted code that wraps generated patterns may be catching nothing today.

Captures inside a lookaround stay empty (re fills them)#

A capturing group inside a lookaround is accepted and numbered, but its capture does not participate in the match: the group reads None where re fills it. Everything else agrees — the span, the group count, and every group outside the lookaround. Measured:

(?=(a))a on "a" — re gives ('a',), REAL gives (None,). (?<=(a))b on "ab" — ('a',) there, (None,) here. (?=(a))(a) on "a" — ('a', 'a') there, (None, 'a') here: a group after the lookaround keeps its number and its value.

A negative lookaround is not a divergence: when (?!...) succeeds, the groups inside it are unset in re too, so both engines read None. And only bounded lookarounds compile at all, so the divergence is confined to bounded assertions.

The cause is a deliberate implementation decision, written at the site of it (capture_free in the compiler): a lookaround runs as a sub-program whose captures do not participate in the overall match, so no capture slot crosses the assertion boundary. Wiring the slots through is a design change — the sub-match would have to export slots the outer match then owns — not a bug fix, and it is not on a roadmap. Pinned in both directions in TestIntentionalDivergences: REAL’s None and re’s fill, so if either side ever moves the pin goes red and this page moves with it.

Variable-width lookbehind (a capability beyond re)#

REAL accepts any bounded lookbehind, including variable-width alternations such as (?<=a|bb), which re rejects as non-fixed-width (PCRE2 accepts a bounded one too: 10.47, measured). The bound keeps the scan linear; see How REAL Works — a guided tour for how the lookbehind sub-VM works.

\b / \B inside lookbehind. Word-boundary assertions in a lookbehind are first-class and evaluate at the lookbehind’s match position (the same rule as outside a lookaround). REAL agrees with Python re, JavaScript (V8/node), and the usual ECMAScript reading on these cases. PCRE2 has known mis-evaluations of \b in lookbehind (a few failures per ~100k differential cases in third-party audits) — treat PCRE2 as a benchmark competitor, not an oracle for \b-in-lookbehind. There is no in-tree PCRE2 differential fuzzer (unlike fuzz_compat vs std and fuzz_re2); if one is added later, those shapes belong on an allowlist the same way std’s non-spec \b edges are, not as REAL bugs.