real::compat — std::regex compatibility#

real::compat (header <real/compat/std/regex.hpp>) is a drop-in for the <regex> surface on the char path. It runs your pattern on real — linear-time and ReDoS-safe — wherever that is provably equivalent to std::regex: the ECMAScript default, and all five POSIX grammars (basic/extended/awk/grep/egrep) when the pattern translates. It falls back to std::regex everywhere else.

No accepted pattern can make matching super-linear — across all five POSIX grammars. Under the default policy::strict, every accepted pattern executes each regex_search / regex_match in time linear in the input — REAL’s ReDoS-safety guarantee, now covering the POSIX grammars, not just the ECMAScript default. A pattern the linear engine cannot represent is rejected, never silently made non-linear; policy::fallback instead delegates it to std::regex (backtracking — the guarantee forfeited, and uses_real() reports false). So (a+)+b under an egrep grammar runs regex_search in microseconds where a std::regex drop-in blows up exponentially — pinned by the Fowler/AT&T conformance gate.

regex_replace and the iterators compose up to O(n) such operations, so their worst-case total is quadratic — inherent to repeated scanning on any linear engine (RE2 and the Rust regex crate included), not a REAL limitation — but never exponential when running on REAL. A nullable pattern’s replace/iteration runs on REAL too, advancing past an empty match as the standard requires: iterating a nullable is O(n²) in the worst case even on a linear engine, so REAL promises quadratic there, not linear, and never exponential.

New here? Start with the migration tour: Drop-in for std::regex. This page is the exhaustive per-feature reference; REAL’s own differences from Python re are in the divergences page. The RE2 drop-in (real::compat::re2) is a separate compat layer — its syntax contract is documented at the top of <real/compat/re2/re2.hpp>.

The contract: behave identically to the ECMAScript spec where real can prove it; a pattern it cannot run is rejected under policy::strict (the default) or delegated to std::regex under policy::fallback — never a silent divergence. The ECMAScript spec is the primary oracle; std::regex (libstdc++/libc++) is a secondary oracle whose known deviations from the spec are catalogued below (where real, following the spec, is the correct one).

Feature status#

Per-construct status (supported / extension / excluded by design) lives in the Features matrix — the single, CI-probed status table, with each rationale linked from its row.

How a pattern is routed#

A real::compat::regex is built with flags::bytes | flags::ecma so real’s byte-oriented, ECMAScript-$ (end-only), ECMAScript-. (excludes \n and \r) semantics line up with std::basic_regex<char>. Routing:

  1. Any single POSIX grammar — extended (ERE), basic (BRE), awk, grep, egrep — → translated to REAL and run on the linear engine with leftmost-longest bounds (group 0 — the POSIX overall-match rule), when the pattern translates; otherwise std::regex. regex.posix_longest() reports this. Captures are the winning thread’s at that longest bound, not POSIX subexpression selection (maximise group 1, then group 2, …). (x|xy)(y*) on "xy" is the one-line proof: macOS’s libc reports groups xy, xy, empty; this layer reports xy, x, y — and glibc reports the same as this layer. Go’s regexp sits with glibc (golang/go#9684); it is macOS’s libc that maximises the subexpressions. Implementing POSIX submatch on a linear engine is why Go documents the gap rather than closing it, and the same reason applies here. All operations are linear for a translated non-nullable pattern — search/match via search_longest, regex_replace/iterators via find_iter_longest; a nullable one (x*, a*) keeps search on REAL but delegates its replace/iterate to std (POSIX-correct bounds; REAL does not model POSIX’s leftmost-longest advance past an empty match). Each grammar’s shape is honoured: BRE \(/\) group and \{n\} quantify while bare ( ) { } | + ? are literals; awk adds the C-escapes (\b is backspace, plus \n\t\r\f\v\a, \/, and octal \ddd); grep (BRE) and egrep (ERE) read a newline as a top-level alternation of the lines. A construct only the backtracker runs — a BRE/grep backreference \1-\9, an ECMAScript-ism, a corner the two std libraries read differently (a medial BRE ^/$) — declines to std, never a silent wrong-match.

  2. collate or nosubs → std::regex up front.

  3. Otherwise real is tried. If it rejects the pattern (a feature it cannot represent), the layer falls back to std::regex, which may accept it. A pattern invalid for both throws real::compat::regex_error (a std::regex_error) carrying std’s exact .code().

regex.uses_real() reports which backend won; regex.nullable() reports whether the pattern can match empty (the state that routes regex_replace/iterators to std — see the traversal rows below).

What runs on real (linear, ReDoS-safe)#

Literals, concatenation, alternation, . (ECMAScript), character classes & ranges, \d \w \s (+negations), ^ $ \b \B, greedy/lazy quantifiers * + ? {m,n}, groups (capturing, non-capturing, named), lookahead and lookbehind (bounded — real’s ReDoS-safe lookaround), ASCII icase, multiline. Non-ASCII literals match byte-for-byte like std::regex<char>. The POSIX extended (ERE) grammar also runs here — translated to REAL and matched with leftmost-longest (POSIX) bounds on group 0 via search_longest, so (a+)+b and friends cannot be ReDoS’d even under an ERE grammar (std would backtrack). Captures follow the winning thread, not POSIX submatch (see routing). POSIX classes [[:alpha:]]…[[:xdigit:]] become their C-locale ASCII ranges.

Native API — trailing-LA throughput surfaces (not a correctness split)#

Bounded lookaround is always correct and linear on every public match API. A narrow class of patterns — trailing-ahead lookaround on a groupless class+ body, e.g. [a-z]+(?=[a-z]) — also has a once-per-walk monomorphic fast path (the class-loop body, LA as end-scan). That path is not taken by every API:

Surface

Trailing-LA fast path?

Why

count_matches (C++ / Rust / Python Pattern.count_matches)

yes

once-per-walk dispatch; matching-only (no Match vector)

find_all / search / match / replace (Python sub)

yes

same dispatch; find_all still pays vector cost at high match counts

find_iter (and Python finditer)

no — general VM

return type keeps the trailing-lookaround fast path off so pure [a-z]+ codegen stays pristine

Correctness is identical across the table; only throughput differs. Prefer count_matches (or search/match/replace) when the shape is eligible and raw scan speed matters. Multi-engine benches must count via count_matches (matching-only), not find_all().size() — the Match vector can dominate and is not comparable to engines that only count. Python: real.count_matches / Pattern.count_matches; do not use len(findall(...)) as a throughput proxy.

What falls back to std::regex (loses real’s ReDoS-safety)#

Construct

Why

Treatment

Backreferences \1, (?P=n)

real does not implement them

real rejects → std fallback (std supports them)

Unbounded / oversized lookaround

exceeds real’s bounded-lookaround cap

real rejects → std fallback

collate or nosubs

locale-sensitive ranges / group-hiding — outside real’s model

screened to std up front (the five POSIX grammars themselves translate — see above)

A BRE backreference \1-\9

real does not implement backreferences

translator declines → std fallback (std backtracks them)

\0 followed by a digit (\00, \012)

real reads a legacy octal escape (Annex B); libstdc++ reads \0=NUL then a literal digit — strict ECMAScript makes it a syntax error, so neither is the spec answer

screened to std up front (a both-accept divergence otherwise; the fuzzer found it)

Nullable patterns in regex_replace/iterators

the standard’s advance past an empty match ([re.regiter.incr]) is real’s own (Python’s), except the retry after an iteration’s first empty match, which reads no text before it

an ECMAScript pattern that can match empty (a*, (x)?) replaces and iterates on real, which makes that retry as the standard does; a nullable POSIX pattern routes those operations to a lazily-built std::regex — per operation, so search/match keep real’s ReDoS-safety even for nullable-ReDoS like (a*)*

These patterns run on std::regex and therefore lose the linear-time guarantee — a documented, non-silent trade. Prefer ReDoS-safe equivalents for untrusted input.

The drop-in policy: strict (default) vs fallback#

Falling back to std::regex reintroduces backtracking — the ReDoS the library exists to avoid — on exactly the patterns you can least audit. So the default is strict, not silent fallback:

real::compat::regex a(R"((\w+)\1)");                            // strict (default): THROWS
//   regex_error, code == error_complexity, "real::compat (strict policy): … backreference …"
real::compat::regex b(R"((\w+)\1)", real::compat::regex_constants::ECMAScript,
                      real::compat::policy::fallback);          // opt-in: delegates to std::regex
  • policy::strict (the default) — a pattern the linear engine cannot represent is rejected with regex_error / error_complexity and a REAL-identifiable message. Every accepted pattern is then a linear-time, ReDoS-safe guarantee. A pattern that is invalid for both engines still reports std’s own error code, so a syntax error stays a syntax error — a true std::regex drop-in there.

  • policy::fallback — restores the old behaviour: an ineligible pattern is delegated to std::regex (which may accept it, forfeiting the linear-time guarantee for that pattern). The choice is explicit and per-regex.

  • Observability. regex::uses_real() / uses_fallback() and regex::policy() always tell you which engine backs a pattern — no silent surprise either way.

The Python binding follows the same policy (real.compile(pat, fallback=True), or the module-level real.fallback), with the same default: strict.

The one tolerated divergence: nullable-loop group capture#

Against a std::regex that follows the standard, there is exactly one place a real-backed pattern’s observable differs from it, and it is documented rather than routed away: regex_search/regex_match’s per-group captures (m[N], N >= 1). For a */+ loop whose body can match empty and captures ((a*)*, (.*)*, (a|)*, (ab|)+a), real records the last consuming iteration for the group while std::regex (ECMAScript, a backtracker) records an extra empty final iteration. The whole match — and every match span — is identical; only the inner group’s captured span differs, and the std value is always a zero-width capture at the loop’s end. This is the same behaviour documented against Python re (see the divergences page); the linear engines RE2, the Rust regex crate, and Go’s regexp share it. Note the std side itself is stdlib-variant: libstdc++ and libc++ record the empty final iteration (the residue described here), while MS STL keeps the last non-empty iteration — agreeing with real’s lineage, so on MSVC this divergence does not exist at all (the test suite pins each stdlib’s edge separately).

libc++ does not follow the standard’s advance past an empty match. Apple’s system libc++ drops the text between empty matches (x* over ab replaces to --- where the standard and libstdc++ give -a-b-), and LLVM’s libc++ makes the retry after an iteration’s first empty match with the text before it (\B|^a over ba replaces to b-a- where the standard gives b--). This layer follows the standard, so on those libraries a nullable pattern’s regex_replace and iteration differ from the local std::regex by construction. The exhaustive routing check asks the local std which it is and, on such a std only, tolerates exactly that signature (a nullable pattern, every observable but the replace equal): 544 175 cases of the default tier on macOS 14’s system libc++, and none on libstdc++, where any such case fails the check.

Why search/match keep it, rather than routing to std. Routing this class to std::regex would hand exactly the textbook catastrophic-backtracking patterns — nested nullable quantifiers — to a backtracking engine, which is the one thing real::compat exists to avoid. regex_search/regex_match on real is the product; screening this signature away would forfeit the linear-time guarantee for the whole (x|)+… family of patterns, not just the divergent capture. Linearity is kept deliberately; the price is a group-capture span that matches the linear-engine family instead of the backtracker. The exhaustive compat check measures this precisely (4 548 cases out of 3 218 434 in the tier-1 space) and fails on any divergence outside this exact signature — a whole-match agreement with only an empty-final-iteration group difference — so no other silent divergence can hide behind it. This is a genuine engine-semantics divergence, not a gap in real::compat’s scope — consistent with the layer’s contract: identical to std where real can prove it, routed/documented otherwise.

regex_replace/iterators do not carry this residue. A pattern whose capturing group is nullable under a quantifier routes replace/iterate to std::regex, even when the pattern as a whole is not nullable — (ab|)+a’s trailing a forces content, so the whole-pattern nullable() gate alone misses it. A dedicated hint (nullable_captured_repeat, an AST walk at compile time — group-under-quantifier is visible in the parsed tree, not in the compiled program the usual prefilter hints derive from) is a sibling of empty_match_possible/nullable() and extends uses_real_traversal to catch this shape too. So regex_replace’s $N and sregex_token_iterator’s sub-group fields converge with std for this whole class of patterns — the residue above is confined to regex_search/regex_match, by design, not by gap.

regex_replace#

The replacement format is ECMAScript: $$ → $, $& → the whole match, $` → the text since the previous match, $' → the text to the end, $N / $NN → group N (matching std::regex_replace, which the differential harness pins). format_first_only and format_no_copy are honoured. A real-backed pattern runs the substitution on real’s traversal, nullable or not (a nullable POSIX pattern excepted, which falls back to std). Whether the swap is faster depends on the std::regex it replaces: on the one measured case it is ~4× faster than libc++’s and ~0.75× — slower — than libstdc++’s (see Performance below).

The real expanders honour format_first_only, format_no_copy, format_sed (and the match_any hint). Under format_sed the format follows sed’s rules: & is the whole match, \N group N (\0 the whole match), a backslash before any other character gives that character, and $ is an ordinary character; libstdc++, libc++ and MS STL agree on every rule but one, a final lone backslash, which the first two keep and MS STL drops: real does as the native std does. A constraining match flag (match_not_bol, match_continuous, …, which the traversal cannot apply) routes the whole substitution to std::regex_replace (so compat == std). An ECMAScript format containing $0 also routes to std: $0 is platform-variant (libstdc++ = the whole match, strict-ECMAScript/MSVC = a literal $0), so real cannot pick one without risking a silent divergence.

Errors and thread-safety#

Every compat entry point that can fail throws a real::compat::regex_error (which is a std::regex_error), never a raw std one: construction (POSIX/wide/custom-traits screens, the real→std fallback), the lazy std build, and the std engine’s own failures while it matches (libc++ gives up on a pattern such as (?:a?){1000} with error_complexity during the match, not at its build).

A real-backed pattern reaches std through four routes only: regex_search/regex_match with a match flag real does not honour (see Match flags); regex_replace with such a flag, with an ECMAScript format using $0, or on a nullable POSIX pattern; a regex_iterator (and so a regex_token_iterator) with such a flag or on a nullable POSIX pattern; and std_engine() itself. A pattern real accepts but std rejects (a real superset: a lookbehind, a named group, a{,2}; on libc++ also \A = literal A) runs search/match on real, and fails only when one of those routes first builds its std engine: a late but homogeneous compat::regex_error, never a silent wrong result. A caller who wants that error at construction calls std_engine() once after building the regex; nothing pays for it otherwise.

The std engine for a real-backed pattern is built on first use under a build mutex and published once: every later call reads it without the lock, so operations on patterns that reach std do not queue behind one another, and concurrent const operations on one shared regex object are race-free for both nullable and non-nullable patterns, as std guarantees (verified under ThreadSanitizer — make tsan, which also runs in CI — including copies made while another thread builds). A copy takes the engine once it is published, else builds its own. (std::once_flag is non-copyable, and basic_regex must stay copyable like std::regex.)

An operation that runs on std::regex inherits its limits. libstdc++’s matcher recurses per character and overflows the stack on long runs: on an 8 MiB stack, from about 50 KB of one run, measured for [^,]* iterated, (?:a|b)* replaced and (?:a|b)+c searched with match_not_bol (g++ 14, arm64, 2026-09-29). Operations that stay on real do not.

Behaviour after a failed match#

A failed regex_search / regex_match leaves the match_results ready (ready() == true, size() == 0, empty()), exactly like std::regex. operator[] / position / length / str for an out-of-range group index return an end-anchored unmatched sub_match ({end, end, false}, so position() is the full sequence length and length() is 0) — never out of bounds. A token selector like {2}/{5} or a field < -1 relies on this; a field < -1 is undefined in std, and compat is safe there, yielding an unmatched token.

Platform-variant std::regex (the three implementations disagree, and not on one axis)#

std::regex is not identical across implementations, and a few of its behaviours are platform-variant. Where real::compat wraps std (a fallback pattern, a wide CharT, a constraining flag) it is ≡ the local std by construction; where it is real-backed it chooses the spec-reasonable behaviour, which may differ from a given std on those points:

  • Out-of-range / unmatched sub_match. libstdc++/libc++ anchor it at the sequence end (position() == length); MSVC-std leaves it singular (position() == 0). real::compat is real-backed, so its match_results come from real’s offsets and it is end-anchored universally — matching libstdc++/libc++, differing from MSVC on .first/.second/.position (the participation flag and str() are empty/false everywhere). This is a deliberate, documented choice, not a bug; the tests assert the end-anchored contract directly and only differ against std on the platform-invariant fields.

  • Escape strictness (\0+digit). \0 followed by a digit (\00, \012) is screened to std (see the fallback table). std itself is platform-variant: libstdc++/libc++ accept it (Annex B legacy octal, lenient), MSVC-std rejects it (error_escape, strict). real::compat defers to the local std on both sides — it throws iff std throws, and where std accepts, it runs on std and matches it. So the construction of such a pattern succeeds on Linux and throws a real::compat::regex_error on MSVC, exactly as the platform’s std::regex does.

  • A class range whose upper endpoint is >= 0x80. [A-\x80], [A-\xff], [\x7f-\x80] and friends are refused by libstdc++ (Invalid range in bracket expression) and accepted by libc++ — the only split in this section that runs between those two rather than between MSVC and them. The cause is the signedness of char: libstdc++ compares the endpoints as signed, so 0x80 is −128 and the range reads as inverted, while libc++ and real read them as byte values. real::compat is real-backed here and stays so on both platforms (uses_real() is true under either compiler), so it accepts the pattern where the local std would have refused it. Write the implementation by name here: an unattributed sentence about this shape is false on libc++, and a catalogue claim that is wrong on one implementation of a two-implementation split is worse than no claim at all.

Iteration (regex_iterator)#

real::compat::regex_iterator (with sregex_iterator / cregex_iterator) walks the non-overlapping matches like std::regex_iterator. Same per-operation routing as regex_replace: a real-backed pattern drives real’s traversal, which advances past an empty match as [re.regiter.incr] does (no empty match again where one was, and, after the iteration’s first match came out empty, a retry that reads no text before it); the std backend, a constraining flag and a nullable POSIX pattern wrap std::regex_iterator. The default-constructed iterator is the end sentinel. Constructing from a temporary regex is =deleted (it would dangle), exactly as std::regex_iterator. The differential fuzzer compares the whole span sequence (and each match’s prefix()/suffix()), not just the first match — the empty-match traversal being the risk it pins.

regex_token_iterator (with sregex_token_iterator / cregex_token_iterator) wraps that iterator, so it inherits the same routing. For each match it yields the requested fields in order: N >= 0 is capture group N (a non-participating group is an empty matched == false token), and -1 is the text before this match since the previous one (the match’s prefix()), which makes -1 a splitter. After the last match a trailing -1 field yields the final suffix only when it is non-empty (an empty field between adjacent matches is still produced — the asymmetry std pins); with -1 and no match at all, the whole sequence is the single token. The fuzzer compares the (str, matched) token sequence for the -1 and 0 fields.

Multi-element field list, no match at all — platform-variant. The “whole sequence is the single token” fallback above is itself not universal once the field list has more than one element (e.g. {0, -1}, {3, 5, -1}): libstdc++ still yields that one whole-input token whenever -1 appears anywhere in the list; libc++ yields it only when the list is literally {-1} alone — any longer list, on a subject with zero matches, gives libc++ an empty sequence instead. real::compat follows libstdc++ (its build/verification oracle, matching this project’s CI); the differential fuzzer (fuzz_compat.cpp’s S5b check) skips comparing that one cell — multi-element list AND zero matches AND -1 present — rather than asserting a libc++ behaviour real::compat was never built to match. Every other cell (a subject that DOES match, or a single-element list) is unaffected and still compared.

Match flags (match_flag_type)#

regex_search / regex_match and both iterators take an optional match_flag_type (default match_default). The rule is honor-on-real or fall back to std, never accept-then-ignore:

  • match_default and match_any keep the real backend. match_any is a non-constraining hint (return a match) that real already satisfies by returning the leftmost match, so ignoring it is sound.

  • On one regex_search / regex_match call, match_continuous and match_prev_avail also keep real. match_continuous is a match anchored at first (real’s match); under a POSIX grammar, whose search is leftmost-longest, a search with it routes to std. match_prev_avail searches from first with the character before it as context: \b, \B, a multiline ^ and a lookbehind read it (a lookbehind sees that one character, no more), and ^ outside multiline does not hold at first. As [re.matchflag] says, match_not_bol and match_not_bow are then ignored. libc++ differs on two points, both its own defects: it ignores match_prev_avail for ^, and its \b never holds on an attempt that starts at last. real::compat follows the standard and libstdc++ there.

  • match_not_null keeps real too: the search takes the leftmost position where a non-empty match starts, and there the match the priority order prefers among the non-empty ones, as libstdc++ and libc++ both do; it runs on the VM, linear. A pattern that cannot match empty ignores the flag. A nullable capturing group under a quantifier ((|a)*) and a POSIX leftmost-longest search route the call to std.

  • match_not_eol and match_not_eow keep real too, over a rewrite of the pattern built once on first use: under match_not_eol a $ holds at no end of the sequence (outside multiline never; in multiline only before a line terminator), under match_not_eow a \b holds at no end and a \B does. A pattern with nothing those flags change keeps its own engine. A $ or \b inside a lookaround, whose rewrite would nest one lookaround in another, and a POSIX grammar route the call to std.

  • Every other constraining flag — match_not_bol and match_not_bow without match_prev_avail — and any constraining flag on an iterator or a regex_replace is not expressible through real’s API, so that single operation routes to std::regex (lazy-built if the pattern is real-backed), which honors every flag by construction. The flags are translated by an exhaustive compat→std table.

This is a per-operation decision, like the nullable routing: a pattern keeps real’s ReDoS-safety for the calls real honors and only the other calls pay the std cost. The differential fuzzer generates a random flag subset and compares compat(mf) vs std(mf) on search + match + iterate, which is what proves the partition. The flags real takes are also compared with the host std over the exhaustive space (make exhaustive-compat-flags: every pattern and input of the routing check, under each of them and their pairs; the whole space in CI on every push).

Always-std parts of the surface (wregex, POSIX, nosubs)#

real runs only the char path with default traits, ECMAScript grammar, reporting every group. Everything outside that is routed to std::regex by a compile-time gate (real_eligible<CharT, Traits>) plus the option screen — real is never even tried, so these are std by construction:

  • wregex / wchar_t (and char8/16/32_t, custom Traits): the gate is constexpr, so real’s char-only code (the byte string_view, fill_from_real, next_real) is compiled out for these instantiations — the real::regex alternative of the backend variant stays dead. wregex::uses_real() is always false. The wide typedefs are provided: wregex, wsmatch/wcmatch, wssub_match/wcsub_match, wsregex_iterator/wcregex_iterator, wsregex_token_iterator/wcregex_token_iterator; regex_search/regex_match/regex_replace are templated on CharT and dispatch the wide path to std.

  • POSIX grammars (basic/extended/awk/grep/egrep): translated to REAL and run on the linear engine with leftmost-longest bounds (group 0; captures are not POSIX submatch — see routing) when the pattern translates; an untranslatable construct (a backreference, an ECMAScript-ism, a std-library-divergent corner) declines to std. collate is screened to std up front (locale-sensitive ranges are outside real’s model).

  • nosubs: std answers it by exposing only group 0, while real always reports every group — a structural both-accept divergence — so nosubs is screened to std. (Honoring it on real by truncating match_results to size 1 is a measured optimization for later, not a correctness need.)

Intentional divergences from libstdc++ std::regex (spec-correct)#

real::compat follows the ECMAScript spec; the following are libstdc++ deviations that the differential harness allowlists (the compat behavior is the spec behavior):

  • Lookbehind (?<=…) / (?<!…): ES2018 has it and real implements it (bounded, ReDoS-safe); libstdc++’s ECMAScript engine rejects it. real::compat accepts and matches it.

Syntax notes (for migrants)#

  • POSIX bracket expressions [[:digit:]], [[.a.]], [[=a=]]: the C++ standard adds them to its ECMAScript grammar ([re.grammar]), and libstdc++ and libc++ both read them. real::compat rewrites a class as its ASCII ranges (the C locale) and stays on the linear engine. A class under icase, where std tests the folded character, a collating element and an equivalence class are only std’s to read (libstdc++ puts A in [=a=], libc++ does not): the strict policy rejects them, policy::fallback routes them to std::regex.

The compat layer builds real with flags::ecma, which makes the engine follow ECMAScript grammar rather than real’s default (Python-flavoured) one. The differences it aligns — each surfaced by the differential fuzzer (517 k iterations, zero remaining both-accept divergence):

  • $ (no multiline) matches only the very end, not before a trailing \n (Python’s re default).

  • . (no dotall) excludes \n and \r (ECMAScript line terminators), not just \n.

  • With multiline, ^ and $ also match after and before a \r (ECMAScript line terminators), as libstdc++ and libc++ do, inside a lookaround too.

  • The escapes \A \Z \< \> (REAL anchors) and \a (Python bell) become identity-escape literals (A Z < >, a) — ECMAScript has no such escapes. \n \r \t \f \v \0 \xHH are unchanged.

  • A ] in the head of a class closes it: [] is the empty class, [^] matches any character (the ECMAScript “any incl. newline” idiom). Python treats a leading ] as a literal member.

  • Inline global flags (?ims) at the start of the pattern are supported; scoped groups (?i:…) are rejected (ECMAScript has no scoped inline flags either).

Boundaries / current scope#

  • Surface: basic_regex<char>, sub_match, match_results (+ smatch/cmatch), regex_error, regex_search, regex_match, regex_replace, the two iterators, the full match_flag_type, wregex, and the POSIX grammar engines. Empty-match traversal is not a fallback trigger for single search/match – only for regex_replace/iterators, where the advance-after-empty-match rule differs from ECMAScript.

  • Members: what std names, the test suite pins at compile time against std: basic_regex’s class constants (regex::icase, regex::multiline, …), every assign and operator= (an invalid pattern leaves the regex unchanged; the policy is kept), the initializer_list constructor; sub_match’s comparisons with another sub_match, a string, a C string and a character, ordering included; match_results’ format (ECMAScript or format_sed rules; $0 reads as the native std reads it), ==, swap, max_size, get_allocator; the regex_constants::error_* codes and regex_error(code); the token iterator’s C-array field list. Not supported: basic_regex::imbue and getloc – they are not provided, and no call reaches std in their place, strict policy or not (REAL’s matching does not depend on a locale)..

  • Every algorithm and iterator accepts any bidirectional iterator. A non-contiguous range (a std::deque, a std::list, a reverse iterator) is searched on REAL over one contiguous copy, made once per search or per regex_iterator and shared by its copies: the same answers and the same linear time as over a string, for O(n) extra memory; the iterators handed back are the caller’s. sub_match::view() exists only over contiguous storage (a view into the copy would dangle); str() works everywhere. A range that takes the std route (a constraining flag, wregex) runs on std::regex over the caller’s iterators.

  • libc++’s own std::regex_iterator, after an empty multiline match, reads the byte before its range; a pattern on the std route inherits it (std parity, not worked around).

  • Matching against an rvalue std::string is deleted (the result would dangle), as in real/std.

Performance (measured, real backend vs std::regex)#

Measured 2026-09-23 with make bench-percall (benchmarks/bench_percall.cpp): both engines compile the pattern once outside the timing, each call is batched until the timed region spans 50 µs, and each column is the median of 15 draws; arm64, -O2, short subjects. Nanoseconds per call; a ratio above 1 is REAL’s favour. The std::regex side depends on the standard library, so both are given:

case

libc++ (Apple clang 16): std / REAL

ratio

libstdc++ (GCC 14.3): std / REAL

ratio

match hit (anchored)

428 / 50

8.6×

195 / 49

4.0×

match reject (anchored)

98 / 33

2.9×

107 / 32

3.4×

search in (unanchored)

994 / 84

11.8×

231 / 56

4.1×

search + captures

916 / 81

11.3×

219 / 63

3.5×

replace (trim)

3359 / 798

4.2×

703 / 938

0.75× — REAL slower

One ISA and one host, in the per-call regime a validation loop pays; throughput and the two-ISA tables are in docs/BENCHMARKS.md. ReDoS is a different axis: on (a+)+b over "a"×N with no match, libstdc++’s std::regex backtracks (4.1 s at N = 26) and libc++’s doubles its time per character, then refuses the input from N = 13 on a complexity counter (docs/BENCHMARKS.md §C, REAL 2026.7.51) — while the compat layer answers in linear time on both.