real::compat — std::regex compatibility#
real::compat (header <real/compat/std/regex.hpp>) is a drop-in for the <regex> surface on the
char path. It runs your pattern on real — linear-time and ReDoS-safe — wherever that is
provably equivalent to std::regex: the ECMAScript default, and all five POSIX grammars
(basic/extended/awk/grep/egrep) when the pattern translates. It falls back to std::regex
everywhere else.
No accepted pattern can make matching super-linear — across all five POSIX grammars. Under the default
policy::strict, every accepted pattern executes each regex_search / regex_match in time linear in
the input — REAL’s ReDoS-safety guarantee, now covering the POSIX grammars, not just the ECMAScript
default. A pattern the linear engine cannot represent is rejected, never silently made non-linear;
policy::fallback instead delegates it to std::regex (backtracking — the guarantee forfeited, and
uses_real() reports false). So (a+)+b under an egrep grammar runs regex_search in microseconds
where a std::regex drop-in blows up exponentially — pinned by the Fowler/AT&T conformance gate.
regex_replace and the iterators compose up to O(n) such operations, so their worst-case total is
quadratic — inherent to repeated scanning on any linear engine (RE2 and the Rust regex crate
included), not a REAL limitation — but never exponential when running on REAL. A nullable pattern’s
replace/iteration runs on REAL too, advancing past an empty match as the standard requires: iterating a
nullable is O(n²) in the worst case even on a linear engine, so REAL promises quadratic there, not linear,
and never exponential.
New here? Start with the migration tour: Drop-in for std::regex. This page is the exhaustive per-feature reference; REAL’s own differences from Python
reare in the divergences page. The RE2 drop-in (real::compat::re2) is a separate compat layer — its syntax contract is documented at the top of<real/compat/re2/re2.hpp>.
The contract: behave identically to the ECMAScript spec where real can prove it; a pattern it
cannot run is rejected under policy::strict (the default) or delegated to std::regex under
policy::fallback — never a silent divergence. The ECMAScript spec is the primary
oracle; std::regex (libstdc++/libc++) is a secondary oracle whose known deviations from the spec
are catalogued below (where real, following the spec, is the correct one).
Feature status#
Per-construct status (supported / extension / excluded by design) lives in the Features matrix — the single, CI-probed status table, with each rationale linked from its row.
How a pattern is routed#
A real::compat::regex is built with flags::bytes | flags::ecma so real’s byte-oriented,
ECMAScript-$ (end-only), ECMAScript-. (excludes \n and \r) semantics line up with
std::basic_regex<char>. Routing:
Any single POSIX grammar —
extended(ERE),basic(BRE),awk,grep,egrep— → translated to REAL and run on the linear engine with leftmost-longest bounds (group 0 — the POSIX overall-match rule), when the pattern translates; otherwisestd::regex.regex.posix_longest()reports this. Captures are the winning thread’s at that longest bound, not POSIX subexpression selection (maximise group 1, then group 2, …).(x|xy)(y*)on"xy"is the one-line proof: macOS’s libc reports groupsxy,xy, empty; this layer reportsxy,x,y— and glibc reports the same as this layer. Go’sregexpsits with glibc (golang/go#9684); it is macOS’s libc that maximises the subexpressions. Implementing POSIX submatch on a linear engine is why Go documents the gap rather than closing it, and the same reason applies here. All operations are linear for a translated non-nullable pattern —search/matchviasearch_longest,regex_replace/iterators viafind_iter_longest; a nullable one (x*,a*) keepssearchon REAL but delegates its replace/iterate tostd(POSIX-correct bounds; REAL does not model POSIX’s leftmost-longest advance past an empty match). Each grammar’s shape is honoured: BRE\(/\)group and\{n\}quantify while bare( ) { } | + ?are literals; awk adds the C-escapes (\bis backspace, plus\n\t\r\f\v\a,\/, and octal\ddd); grep (BRE) and egrep (ERE) read a newline as a top-level alternation of the lines. A construct only the backtracker runs — a BRE/grep backreference\1-\9, an ECMAScript-ism, a corner the two std libraries read differently (a medial BRE^/$) — declines tostd, never a silent wrong-match.collateornosubs→std::regexup front.Otherwise
realis tried. If it rejects the pattern (a feature it cannot represent), the layer falls back tostd::regex, which may accept it. A pattern invalid for both throwsreal::compat::regex_error(astd::regex_error) carrying std’s exact.code().
regex.uses_real() reports which backend won; regex.nullable() reports whether the pattern can
match empty (the state that routes regex_replace/iterators to std — see the traversal rows below).
What runs on real (linear, ReDoS-safe)#
Literals, concatenation, alternation, . (ECMAScript), character classes & ranges, \d \w \s
(+negations), ^ $ \b \B, greedy/lazy quantifiers * + ? {m,n}, groups (capturing,
non-capturing, named), lookahead and lookbehind (bounded — real’s ReDoS-safe lookaround),
ASCII icase, multiline. Non-ASCII literals match byte-for-byte like std::regex<char>.
The POSIX extended (ERE) grammar also runs here — translated to REAL and matched with leftmost-longest
(POSIX) bounds on group 0 via search_longest, so (a+)+b and friends cannot be ReDoS’d even under an ERE grammar (std
would backtrack). Captures follow the winning thread, not POSIX submatch (see routing). POSIX classes [[:alpha:]]…[[:xdigit:]] become their C-locale ASCII ranges.
Native API — trailing-LA throughput surfaces (not a correctness split)#
Bounded lookaround is always correct and linear on every public match API. A narrow class of
patterns — trailing-ahead lookaround on a groupless class+ body, e.g. [a-z]+(?=[a-z]) — also has a
once-per-walk monomorphic fast path (the class-loop body, LA as end-scan). That path is not
taken by every API:
Surface |
Trailing-LA fast path? |
Why |
|---|---|---|
|
yes |
once-per-walk dispatch; matching-only (no Match vector) |
|
yes |
same dispatch; |
|
no — general VM |
return type keeps the trailing-lookaround fast path off so pure |
Correctness is identical across the table; only throughput differs. Prefer count_matches (or
search/match/replace) when the shape is eligible and raw scan speed matters. Multi-engine benches
must count via count_matches (matching-only), not find_all().size() — the Match vector can dominate
and is not comparable to engines that only count. Python: real.count_matches / Pattern.count_matches;
do not use len(findall(...)) as a throughput proxy.
What falls back to std::regex (loses real’s ReDoS-safety)#
Construct |
Why |
Treatment |
|---|---|---|
Backreferences |
|
|
Unbounded / oversized lookaround |
exceeds |
|
|
locale-sensitive ranges / group-hiding — outside |
screened to std up front (the five POSIX grammars themselves translate — see above) |
A BRE backreference |
|
translator declines → std fallback (std backtracks them) |
|
|
screened to std up front (a both-accept divergence otherwise; the fuzzer found it) |
Nullable patterns in |
the standard’s advance past an empty match ([re.regiter.incr]) is |
an ECMAScript pattern that can match empty ( |
These patterns run on std::regex and therefore lose the linear-time guarantee — a documented,
non-silent trade. Prefer ReDoS-safe equivalents for untrusted input.
The drop-in policy: strict (default) vs fallback#
Falling back to std::regex reintroduces backtracking — the ReDoS the library exists to avoid — on exactly
the patterns you can least audit. So the default is strict, not silent fallback:
real::compat::regex a(R"((\w+)\1)"); // strict (default): THROWS
// regex_error, code == error_complexity, "real::compat (strict policy): … backreference …"
real::compat::regex b(R"((\w+)\1)", real::compat::regex_constants::ECMAScript,
real::compat::policy::fallback); // opt-in: delegates to std::regex
policy::strict(the default) — a pattern the linear engine cannot represent is rejected withregex_error/error_complexityand a REAL-identifiable message. Every accepted pattern is then a linear-time, ReDoS-safe guarantee. A pattern that is invalid for both engines still reportsstd’s own error code, so a syntax error stays a syntax error — a truestd::regexdrop-in there.policy::fallback— restores the old behaviour: an ineligible pattern is delegated tostd::regex(which may accept it, forfeiting the linear-time guarantee for that pattern). The choice is explicit and per-regex.Observability.
regex::uses_real()/uses_fallback()andregex::policy()always tell you which engine backs a pattern — no silent surprise either way.
The Python binding follows the same policy (real.compile(pat, fallback=True), or the module-level
real.fallback), with the same default: strict.
The one tolerated divergence: nullable-loop group capture#
Against a std::regex that follows the standard, there is exactly one place a real-backed
pattern’s observable differs from it, and it is documented rather than routed away: regex_search/regex_match’s per-group captures
(m[N], N >= 1). For a */+ loop whose body can match empty and captures ((a*)*, (.*)*,
(a|)*, (ab|)+a), real records the last consuming iteration for the group while std::regex
(ECMAScript, a backtracker) records an extra empty final iteration. The whole match — and every
match span — is identical; only the inner group’s captured span differs, and the std value
is always a zero-width capture at the loop’s end. This is the same behaviour documented against
Python re (see the divergences page); the linear engines RE2,
the Rust regex crate, and Go’s regexp share it. Note the std side itself is
stdlib-variant: libstdc++ and libc++ record the empty final iteration (the residue described
here), while MS STL keeps the last non-empty iteration — agreeing with real’s lineage, so on
MSVC this divergence does not exist at all (the test suite pins each stdlib’s edge separately).
libc++ does not follow the standard’s advance past an empty match. Apple’s system libc++ drops the text
between empty matches (x* over ab replaces to --- where the standard and libstdc++ give -a-b-),
and LLVM’s libc++ makes the retry after an iteration’s first empty match with the text before it
(\B|^a over ba replaces to b-a- where the standard gives b--). This layer follows the standard, so
on those libraries a nullable pattern’s regex_replace and iteration differ from the local std::regex
by construction. The exhaustive routing check asks the local std which it is and, on such a std only,
tolerates exactly that signature (a nullable pattern, every observable but the replace equal):
544 175 cases of the default tier on macOS 14’s system libc++, and none on libstdc++, where any
such case fails the check.
Why search/match keep it, rather than routing to std. Routing this class to std::regex would
hand exactly the textbook catastrophic-backtracking patterns — nested nullable quantifiers — to a
backtracking engine, which is the one thing real::compat exists to avoid. regex_search/regex_match
on real is the product; screening this signature away would forfeit the linear-time guarantee for
the whole (x|)+… family of patterns, not just the divergent capture. Linearity is kept deliberately;
the price is a group-capture span that matches the linear-engine family instead of the backtracker. The
exhaustive compat check measures this precisely (4 548 cases out of 3 218 434 in the tier-1 space) and
fails on any divergence outside this exact signature — a whole-match agreement with only an
empty-final-iteration group difference — so no other silent divergence can hide behind it. This is a
genuine engine-semantics divergence, not a gap in real::compat’s scope — consistent with the layer’s
contract: identical to std where real can prove it, routed/documented otherwise.
regex_replace/iterators do not carry this residue. A pattern whose capturing group is nullable
under a quantifier routes replace/iterate to std::regex, even when the pattern as a whole is not
nullable — (ab|)+a’s trailing a forces content, so the whole-pattern nullable() gate alone misses
it. A dedicated hint (nullable_captured_repeat, an AST walk at compile time — group-under-quantifier
is visible in the parsed tree, not in the compiled program the usual prefilter hints derive from) is a
sibling of empty_match_possible/nullable() and extends uses_real_traversal
to catch this shape too. So regex_replace’s $N and sregex_token_iterator’s sub-group fields
converge with std for this whole class of patterns — the residue above is confined to
regex_search/regex_match, by design, not by gap.
regex_replace#
The replacement format is ECMAScript: $$ → $, $& → the whole match, $` → the text since
the previous match, $' → the text to the end, $N / $NN → group N (matching
std::regex_replace, which the differential harness pins). format_first_only and format_no_copy
are honoured. A real-backed pattern runs the substitution on real’s traversal, nullable or not (a
nullable POSIX pattern excepted, which falls back to std). Whether the swap is faster depends on the std::regex it replaces:
on the one measured case it is ~4× faster than libc++’s and ~0.75× — slower — than libstdc++’s
(see Performance below).
The real expanders honour format_first_only, format_no_copy, format_sed (and the match_any
hint). Under format_sed the format follows sed’s rules: & is the whole match, \N group N (\0
the whole match), a backslash before any other character gives that character, and $ is an ordinary
character; libstdc++, libc++ and MS STL agree on every rule but one, a final lone backslash, which the
first two keep and MS STL drops: real does as the native std does. A constraining
match flag (match_not_bol, match_continuous, …, which the traversal cannot apply) routes the
whole substitution to std::regex_replace (so compat == std). An ECMAScript format containing
$0 also routes to std: $0 is platform-variant (libstdc++ = the whole match, strict-ECMAScript/MSVC
= a literal $0), so real cannot pick one without risking a silent divergence.
Errors and thread-safety#
Every compat entry point that can fail throws a real::compat::regex_error (which is a
std::regex_error), never a raw std one: construction (POSIX/wide/custom-traits screens, the
real→std fallback), the lazy std build, and the std engine’s own failures while it matches (libc++
gives up on a pattern such as (?:a?){1000} with error_complexity during the match, not at its build).
A real-backed pattern reaches std through four routes only: regex_search/regex_match with a
match flag real does not honour (see Match flags); regex_replace with such a flag, with an ECMAScript format
using $0, or on a nullable POSIX pattern; a regex_iterator (and so a regex_token_iterator) with
such a flag or on a nullable POSIX pattern; and std_engine() itself. A pattern real accepts but std
rejects (a real superset: a lookbehind, a named group, a{,2}; on libc++ also \A = literal A)
runs search/match on real, and fails only when one of those routes first builds its std engine:
a late but homogeneous compat::regex_error, never a silent wrong result. A caller who wants that
error at construction calls std_engine() once after building the regex; nothing pays for it otherwise.
The std engine for a real-backed pattern is built on first use under a build mutex and published
once: every later call reads it without the lock, so operations on patterns that reach std do not
queue behind one another, and concurrent const operations on one shared regex object are race-free
for both nullable and non-nullable patterns, as std guarantees (verified under ThreadSanitizer —
make tsan, which also runs in CI — including copies made while another thread builds). A copy takes
the engine once it is published, else builds its own. (std::once_flag is non-copyable, and
basic_regex must stay copyable like std::regex.)
An operation that runs on std::regex inherits its limits. libstdc++’s matcher recurses per character
and overflows the stack on long runs: on an 8 MiB stack, from about 50 KB of one run, measured for
[^,]* iterated, (?:a|b)* replaced and (?:a|b)+c searched with match_not_bol (g++ 14, arm64,
2026-09-29). Operations that stay on real do not.
Behaviour after a failed match#
A failed regex_search / regex_match leaves the match_results ready (ready() == true,
size() == 0, empty()), exactly like std::regex. operator[] / position / length / str
for an out-of-range group index return an end-anchored unmatched sub_match ({end, end, false},
so position() is the full sequence length and length() is 0) — never out of bounds. A token
selector like {2}/{5} or a field < -1 relies on this; a field < -1 is undefined in std, and
compat is safe there, yielding an unmatched token.
Platform-variant std::regex (the three implementations disagree, and not on one axis)#
std::regex is not identical across implementations, and a few of its behaviours are
platform-variant. Where real::compat wraps std (a fallback pattern, a wide CharT, a
constraining flag) it is ≡ the local std by construction; where it is real-backed it
chooses the spec-reasonable behaviour, which may differ from a given std on those points:
Out-of-range / unmatched
sub_match. libstdc++/libc++ anchor it at the sequence end (position() == length); MSVC-stdleaves it singular (position() == 0).real::compatis real-backed, so itsmatch_resultscome fromreal’s offsets and it is end-anchored universally — matching libstdc++/libc++, differing from MSVC on.first/.second/.position(the participation flag andstr()are empty/falseeverywhere). This is a deliberate, documented choice, not a bug; the tests assert the end-anchored contract directly and only differ againststdon the platform-invariant fields.Escape strictness (
\0+digit).\0followed by a digit (\00,\012) is screened tostd(see the fallback table).stditself is platform-variant: libstdc++/libc++ accept it (Annex B legacy octal, lenient), MSVC-stdrejects it (error_escape, strict).real::compatdefers to the localstdon both sides — it throws iffstdthrows, and wherestdaccepts, it runs onstdand matches it. So the construction of such a pattern succeeds on Linux and throws areal::compat::regex_erroron MSVC, exactly as the platform’sstd::regexdoes.A class range whose upper endpoint is
>= 0x80.[A-\x80],[A-\xff],[\x7f-\x80]and friends are refused by libstdc++ (Invalid range in bracket expression) and accepted by libc++ — the only split in this section that runs between those two rather than between MSVC and them. The cause is the signedness ofchar: libstdc++ compares the endpoints as signed, so0x80is −128 and the range reads as inverted, while libc++ andrealread them as byte values.real::compatis real-backed here and stays so on both platforms (uses_real()is true under either compiler), so it accepts the pattern where the localstdwould have refused it. Write the implementation by name here: an unattributed sentence about this shape is false on libc++, and a catalogue claim that is wrong on one implementation of a two-implementation split is worse than no claim at all.
Iteration (regex_iterator)#
real::compat::regex_iterator (with sregex_iterator / cregex_iterator) walks the non-overlapping
matches like std::regex_iterator. Same per-operation routing as regex_replace: a real-backed
pattern drives real’s traversal, which advances past an empty match as [re.regiter.incr] does (no
empty match again where one was, and, after the iteration’s first match came out empty, a retry that
reads no text before it); the std backend, a constraining flag and a nullable POSIX pattern wrap
std::regex_iterator. The default-constructed iterator is the end sentinel. Constructing from a temporary
regex is =deleted (it would dangle), exactly as std::regex_iterator. The differential fuzzer
compares the whole span sequence (and each match’s prefix()/suffix()), not just the first
match — the empty-match traversal being the risk it pins.
regex_token_iterator (with sregex_token_iterator / cregex_token_iterator) wraps that iterator,
so it inherits the same routing. For each match it yields the requested fields in
order: N >= 0 is capture group N (a non-participating group is an empty matched == false
token), and -1 is the text before this match since the previous one (the match’s prefix()),
which makes -1 a splitter. After the last match a trailing -1 field yields the final suffix
only when it is non-empty (an empty field between adjacent matches is still produced — the
asymmetry std pins); with -1 and no match at all, the whole sequence is the single token. The
fuzzer compares the (str, matched) token sequence for the -1 and 0 fields.
Multi-element field list, no match at all — platform-variant. The “whole sequence is the single
token” fallback above is itself not universal once the field list has more than one element (e.g.
{0, -1}, {3, 5, -1}): libstdc++ still yields that one whole-input token whenever -1 appears
anywhere in the list; libc++ yields it only when the list is literally {-1} alone — any longer
list, on a subject with zero matches, gives libc++ an empty sequence instead. real::compat
follows libstdc++ (its build/verification oracle, matching this project’s CI); the differential fuzzer
(fuzz_compat.cpp’s S5b check) skips comparing that one cell — multi-element list AND zero matches AND
-1 present — rather than asserting a libc++ behaviour real::compat was never built to match. Every
other cell (a subject that DOES match, or a single-element list) is unaffected and still compared.
Match flags (match_flag_type)#
regex_search / regex_match and both iterators take an optional match_flag_type (default
match_default). The rule is honor-on-real or fall back to std, never accept-then-ignore:
match_defaultandmatch_anykeep therealbackend.match_anyis a non-constraining hint (return a match) thatrealalready satisfies by returning the leftmost match, so ignoring it is sound.On one
regex_search/regex_matchcall,match_continuousandmatch_prev_availalso keepreal.match_continuousis a match anchored atfirst(real’smatch); under a POSIX grammar, whose search is leftmost-longest, a search with it routes tostd.match_prev_availsearches fromfirstwith the character before it as context:\b,\B, a multiline^and a lookbehind read it (a lookbehind sees that one character, no more), and^outside multiline does not hold atfirst. As [re.matchflag] says,match_not_bolandmatch_not_boware then ignored. libc++ differs on two points, both its own defects: it ignoresmatch_prev_availfor^, and its\bnever holds on an attempt that starts atlast.real::compatfollows the standard and libstdc++ there.match_not_nullkeepsrealtoo: the search takes the leftmost position where a non-empty match starts, and there the match the priority order prefers among the non-empty ones, as libstdc++ and libc++ both do; it runs on the VM, linear. A pattern that cannot match empty ignores the flag. A nullable capturing group under a quantifier ((|a)*) and a POSIX leftmost-longest search route the call tostd.match_not_eolandmatch_not_eowkeeprealtoo, over a rewrite of the pattern built once on first use: undermatch_not_eola$holds at no end of the sequence (outside multiline never; in multiline only before a line terminator), undermatch_not_eowa\bholds at no end and a\Bdoes. A pattern with nothing those flags change keeps its own engine. A$or\binside a lookaround, whose rewrite would nest one lookaround in another, and a POSIX grammar route the call tostd.Every other constraining flag —
match_not_bolandmatch_not_bowwithoutmatch_prev_avail— and any constraining flag on an iterator or aregex_replaceis not expressible throughreal’s API, so that single operation routes tostd::regex(lazy-built if the pattern is real-backed), which honors every flag by construction. The flags are translated by an exhaustive compat→std table.
This is a per-operation decision, like the nullable routing: a pattern keeps real’s ReDoS-safety
for the calls real honors and only the other calls pay the std cost. The differential fuzzer
generates a random flag subset and compares compat(mf) vs std(mf) on search + match + iterate,
which is what proves the partition. The flags real takes are also compared with the host std over
the exhaustive space (make exhaustive-compat-flags: every pattern and input of the routing check, under
each of them and their pairs; the whole space in CI on every push).
Always-std parts of the surface (wregex, POSIX, nosubs)#
real runs only the char path with default traits, ECMAScript grammar, reporting every group.
Everything outside that is routed to std::regex by a compile-time gate (real_eligible<CharT, Traits>)
plus the option screen — real is never even tried, so these are std by construction:
wregex/wchar_t(andchar8/16/32_t, customTraits): the gate isconstexpr, soreal’s char-only code (the bytestring_view,fill_from_real,next_real) is compiled out for these instantiations — thereal::regexalternative of the backend variant stays dead.wregex::uses_real()is alwaysfalse. The wide typedefs are provided:wregex,wsmatch/wcmatch,wssub_match/wcsub_match,wsregex_iterator/wcregex_iterator,wsregex_token_iterator/wcregex_token_iterator;regex_search/regex_match/regex_replaceare templated onCharTand dispatch the wide path tostd.POSIX grammars (
basic/extended/awk/grep/egrep): translated to REAL and run on the linear engine with leftmost-longest bounds (group 0; captures are not POSIX submatch — see routing) when the pattern translates; an untranslatable construct (a backreference, an ECMAScript-ism, a std-library-divergent corner) declines tostd.collateis screened tostdup front (locale-sensitive ranges are outsidereal’s model).nosubs:stdanswers it by exposing only group 0, whilerealalways reports every group — a structural both-accept divergence — sonosubsis screened tostd. (Honoring it onrealby truncatingmatch_resultsto size 1 is a measured optimization for later, not a correctness need.)
Intentional divergences from libstdc++ std::regex (spec-correct)#
real::compat follows the ECMAScript spec; the following are libstdc++ deviations that the
differential harness allowlists (the compat behavior is the spec behavior):
Lookbehind
(?<=…)/(?<!…): ES2018 has it andrealimplements it (bounded, ReDoS-safe); libstdc++’s ECMAScript engine rejects it.real::compataccepts and matches it.
Syntax notes (for migrants)#
POSIX bracket expressions
[[:digit:]],[[.a.]],[[=a=]]: the C++ standard adds them to its ECMAScript grammar ([re.grammar]), and libstdc++ and libc++ both read them.real::compatrewrites a class as its ASCII ranges (the C locale) and stays on the linear engine. A class undericase, where std tests the folded character, a collating element and an equivalence class are only std’s to read (libstdc++ putsAin[=a=], libc++ does not): the strict policy rejects them,policy::fallbackroutes them tostd::regex.
The compat layer builds real with flags::ecma, which makes the engine follow ECMAScript grammar
rather than real’s default (Python-flavoured) one. The differences it aligns — each surfaced by the
differential fuzzer (517 k iterations, zero remaining both-accept divergence):
$(nomultiline) matches only the very end, not before a trailing\n(Python’sredefault)..(no dotall) excludes\nand\r(ECMAScript line terminators), not just\n.With
multiline,^and$also match after and before a\r(ECMAScript line terminators), as libstdc++ and libc++ do, inside a lookaround too.The escapes
\A \Z \< \>(REAL anchors) and\a(Python bell) become identity-escape literals (A Z < >,a) — ECMAScript has no such escapes.\n \r \t \f \v \0 \xHHare unchanged.A
]in the head of a class closes it:[]is the empty class,[^]matches any character (the ECMAScript “any incl. newline” idiom). Python treats a leading]as a literal member.Inline global flags
(?ims)at the start of the pattern are supported; scoped groups(?i:…)are rejected (ECMAScript has no scoped inline flags either).
Boundaries / current scope#
Surface:
basic_regex<char>,sub_match,match_results(+smatch/cmatch),regex_error,regex_search,regex_match,regex_replace, the two iterators, the fullmatch_flag_type,wregex, and the POSIX grammar engines. Empty-match traversal is not a fallback trigger for singlesearch/match– only forregex_replace/iterators, where the advance-after-empty-match rule differs from ECMAScript.Members: what std names, the test suite pins at compile time against std:
basic_regex’s class constants (regex::icase,regex::multiline, …), everyassignandoperator=(an invalid pattern leaves the regex unchanged; the policy is kept), theinitializer_listconstructor;sub_match’s comparisons with anothersub_match, a string, a C string and a character, ordering included;match_results’format(ECMAScript orformat_sedrules;$0reads as the native std reads it),==,swap,max_size,get_allocator; theregex_constants::error_*codes andregex_error(code); the token iterator’s C-array field list. Not supported:basic_regex::imbueandgetloc– they are not provided, and no call reachesstdin their place, strict policy or not (REAL’s matching does not depend on a locale)..Every algorithm and iterator accepts any bidirectional iterator. A non-contiguous range (a
std::deque, astd::list, a reverse iterator) is searched on REAL over one contiguous copy, made once per search or perregex_iteratorand shared by its copies: the same answers and the same linear time as over a string, for O(n) extra memory; the iterators handed back are the caller’s.sub_match::view()exists only over contiguous storage (a view into the copy would dangle);str()works everywhere. A range that takes the std route (a constraining flag,wregex) runs onstd::regexover the caller’s iterators.libc++’s own
std::regex_iterator, after an empty multiline match, reads the byte before its range; a pattern on the std route inherits it (std parity, not worked around).Matching against an rvalue
std::stringis deleted (the result would dangle), as inreal/std.
Performance (measured, real backend vs std::regex)#
Measured 2026-09-23 with make bench-percall (benchmarks/bench_percall.cpp): both engines compile
the pattern once outside the timing, each call is batched until the timed region spans 50 µs, and each
column is the median of 15 draws; arm64, -O2, short subjects. Nanoseconds per call; a ratio above 1
is REAL’s favour. The std::regex side depends on the standard library, so both are given:
case |
libc++ (Apple clang 16): std / REAL |
ratio |
libstdc++ (GCC 14.3): std / REAL |
ratio |
|---|---|---|---|---|
match hit (anchored) |
428 / 50 |
8.6× |
195 / 49 |
4.0× |
match reject (anchored) |
98 / 33 |
2.9× |
107 / 32 |
3.4× |
search in (unanchored) |
994 / 84 |
11.8× |
231 / 56 |
4.1× |
search + captures |
916 / 81 |
11.3× |
219 / 63 |
3.5× |
replace (trim) |
3359 / 798 |
4.2× |
703 / 938 |
0.75× — REAL slower |
One ISA and one host, in the per-call regime a validation loop pays; throughput and the two-ISA tables
are in docs/BENCHMARKS.md. ReDoS is a different axis: on (a+)+b over "a"×N with no match,
libstdc++’s std::regex backtracks (4.1 s at N = 26) and libc++’s doubles its time per character, then
refuses the input from N = 13 on a complexity counter (docs/BENCHMARKS.md §C, REAL 2026.7.51) — while
the compat layer answers in linear time on both.