|
REAL
Regular Expression Algorithmic Library — constexpr C++20 regex
|
Recursive-descent parser: a pattern string in, an ast out. More...
#include <ast.hpp>
Classes | |
| struct | loose_buf |
A loose-match key (lowercase, no _/-/space) in a fixed buffer, so parsing needs no heap or <string>. A name longer than the buffer matches nothing. More... | |
| struct | property_table_result |
The result of parse_property_table — the ranges, and whether a leading caret was stripped (\p{^L} == \P{L}, RE2/Perl). Callers XOR caret into their own negation so it composes with \P and [^...]. More... | |
| struct | quoted_span |
| What parse_quoted_span emitted: the chained all-but-last prefix plus the final atom, the caller's quantifier target. More... | |
| struct | shorthand_spec |
The classification of a \d \D \w \W \s \S shorthand: its ASCII bitmap, its Unicode range table, and whether it is the negated (uppercase) form. More... | |
Public Member Functions | |
| constexpr | parser (std::string_view pattern, flags initial_flags=flags::none) |
| Binds the parser to a pattern and the constructor flags. | |
| constexpr ast | parse () |
| Parses the whole pattern. | |
Private Member Functions | |
| constexpr flags | current_flags () const |
| The flag set in force at the current nesting level (the scope-stack top). | |
| constexpr bool | is_verbose () const |
True when verbose mode (re.X) is in force at the current scope. | |
| constexpr bool | is_ecma () const |
True in the ECMAScript grammar. Not scopable, so every scope carries it; read it here, not from ecma_, which the flag-scope ratchet counts. | |
| constexpr bool | is_icase () const |
True when icase (re.I) is in force at the current scope. | |
| constexpr bool | is_ascii_mode () const |
True when ascii (re.A) is in force at the current scope. | |
| constexpr void | skip_insignificant () |
In verbose mode, consumes insignificant whitespace and # comments. | |
| template<typename Error = regex_error> | |
| constexpr void | fail (const char *message) const |
| Aborts the parse with a real::regex_error at the current offset. | |
| template<typename = void> | |
| constexpr void | fail_unsupported (const char *message) const |
Like fail, but tags the error unsupported (well-formed but beyond the linear engine, e.g. a backreference or a nested lookaround) so a binding classifies it without reading the message. Templated for the same reason as fail. | |
| template<typename = void> | |
| constexpr void | fail_unknown_extension (std::size_t question_pos) const |
Fails an unrecognised (?… extension the way re does: naming it. | |
| constexpr bool | eof () const |
Returns true if the read offset is at or past the end of the pattern. | |
| constexpr char | peek () const |
| Returns the current character without consuming it (undefined at eof()). | |
| constexpr bool | accept (char ch) |
Consumes the current character if it equals ch. | |
| constexpr std::int32_t | add_node (ast &out, ast_node node) |
Appends node to the pool. | |
| constexpr std::int32_t | add_class_node (ast &out, const char_class &klass, bool negated, const std::vector< code_range > &ranges={}, bool codepoint_predicate=false) |
| Interns a class bitmap and appends a node_kind::klass node. | |
| constexpr std::int32_t | add_class_node (ast &out, class_def klass, bool negated) |
Interns klass and appends a node_kind::klass node. | |
| constexpr bool | text_shorthand () const |
Whether a shorthand (\d \w \s) should be a text-mode Unicode code-point predicate: true in the default text mode, false in bytes mode or under flags::ascii (re.A). | |
| constexpr void | merge_property (char_class &klass, std::vector< code_range > &ranges, const char_class &prop_ascii, std::span< const code_range > table, bool negated, bool &property_derived) const |
Merges an in-class shorthand (\w \d \s, or negated \W \D \S) into the class being built: its ASCII bitmap plus, in text mode, its non-ASCII ranges (each complemented when negated). In text mode, sets property_derived. | |
| constexpr std::vector< code_range > | shorthand_ranges (std::span< const code_range > table) const |
| The non-ASCII part of a shorthand's range table, or nothing in bytes / ASCII mode. Wholly-ASCII ranges are dropped: the bitmap already covers them. | |
| constexpr std::vector< code_range > | resolve_property (std::string_view name) const |
Resolves a \p{...} property name to its code-point ranges, or fails. | |
| constexpr property_table_result | parse_property_table () |
Consumes p/P and {Name} (or one letter), strips a leading ^ (native dialects only) and resolves the name; rejects bytes mode. Shared by the atom and the in-class paths. Entered on the p/P; leaves pos_ just past the name. | |
| constexpr void | property_ascii_high (const std::vector< code_range > &table, char_class &ascii, std::vector< code_range > &high) const |
Splits a property's ranges into its ASCII bitmap (< 0x80) and its non-ASCII ranges. Unlike \w, a Unicode property is not restricted by flags::ascii. | |
| constexpr std::int32_t | parse_unicode_property (ast &out, bool negated) |
Parses \p{Name} / \P{Name} / \pX outside a class into a code-point class node (klass_cp, as \w). Negation is the node flag XORed with the caret (\P{^L} == \p{L}). Entered on the letter after \. | |
| constexpr void | merge_unicode_property (char_class &klass, std::vector< code_range > &ranges, const std::vector< code_range > &table, bool negated, bool &property_derived) const |
Merges \p{Name} / \P{Name} into the class being built: merge_property without the flags::ascii restriction. \P merges the complement, as \W does; an enclosing [^...] negates on top ([^\P{L}] == [\p{L}]). | |
| constexpr std::int32_t | parse_alternation (ast &out) |
| Parses ‘alternation := sequence (’|' sequence)*`. | |
| constexpr std::int32_t | parse_sequence (ast &out) |
Parses sequence := (atom quantifier?)*, stopping at | or ). | |
| constexpr quoted_span | parse_quoted_span (ast &out) |
Scans a \Q...\E span (\Q already consumed) and emits its characters as literal atoms via emit_literal_codepoint, as RE2 does. | |
| constexpr std::int32_t | parse_atom (ast &out) |
Parses one atom: a literal, ., a class, a group, an anchor or an escape. | |
| constexpr std::int32_t | parse_quantifier (ast &out, std::int32_t atom) |
Wraps atom in a repeat node if a quantifier follows. | |
| constexpr bool | try_parse_braces (std::int32_t &min, std::int32_t &max) |
Tries to parse {n} / {n,} / {,m} / {n,m} starting at {. | |
| constexpr std::int32_t | parse_repeat_count () |
| Reads an optional decimal repeat count. | |
| constexpr bool | at_letter_unit () const |
Returns true if the unit at the cursor is a LETTER — an unknown flag, not a terminator. | |
| constexpr void | fail_extension (std::size_t question_pos) const |
Fails a (?… extension, naming it when it is one REAL excludes by design. | |
| constexpr flags | consume_flag_letters () |
Consumes a run of inline flag letters (imsxaU). | |
| constexpr void | fail_if_unknown_flag () |
| Fails with "unknown flag" if the next unit is a letter that is not a flag. | |
| constexpr void | require_scoped_flags_colon (bool after_removal, std::size_t open_pos) |
After a flags run, requires : (scoped body) or fails naming the terminator fault. | |
| constexpr bool | parse_global_flags_prefix (ast &out) |
Consumes a leading global-flags group – (?imsxaU), or (?flags-flags) with a removal suffix – if present. The accepted letters are i m s x a U (is_flag_letter). | |
| constexpr std::int32_t | parse_group (ast &out) |
| Parses a group construct. | |
| constexpr std::int32_t | parse_lookaround (ast &out, look_dir direction, std::size_t open_pos) |
Parses a lookaround after (?= / (?! (ahead) or (?<= / (?<! (behind) — the =/! is not yet consumed. | |
| constexpr std::int32_t | parse_atomic_group (ast &out, std::size_t open_pos) |
Parses an atomic group after (?> (the > is not yet consumed). | |
| constexpr std::int32_t | new_group (ast &out, std::size_t open_pos) |
| Allocates the next capture group number. | |
| template<typename = void> | |
| constexpr void | fail_group_name (std::size_t begin) const |
Fails a group name the way re does: quoting the name it read, at the name's start. | |
| template<typename = void> | |
| constexpr void | fail_duplicate_group_name (std::size_t begin, std::size_t end, std::int32_t group, std::int32_t previous) const |
Fails a duplicate group name the way re does, naming it and both group numbers. | |
| constexpr void | parse_group_name (ast &out, std::int32_t group) |
Parses a group name up to > and records it. | |
| constexpr std::int32_t | parse_byte_escape (std::size_t backslash) |
| Parses a single-byte escape (valid inside and outside classes). | |
| constexpr std::int32_t | parse_digit_escape () |
Parses a \<digit> escape via the shared decode_digit_escape(). | |
| template<typename = void> | |
| constexpr void | fail_incomplete_escape (std::size_t backslash_pos) const |
Fails a truncated escape the way re does (incomplete escape \x1): quoting pattern_[backslash_pos, pos_), reported at the backslash. | |
| template<typename = void> | |
| constexpr void | fail_bad_range (std::size_t begin, std::size_t end) const |
Fails a bad character-class range the way re does, quoting it (bad character range z-a). | |
| template<typename = void> | |
| constexpr void | fail_unsupported_escape (std::size_t backslash) const |
Fails an escape REAL does not implement, naming it (unsupported escape sequence \q) and reporting at the backslash, as re does; tagged unsupported. | |
| constexpr std::int32_t | hex_digit (std::size_t backslash_pos) |
| Consumes one hexadecimal digit. | |
| constexpr std::int32_t | parse_unicode_codepoint (bool capital, std::size_t backslash) |
Decodes a \uHHHH (4 hex) or \UHHHHHHHH (8 hex) code-point escape (str only), or the braced form \u{HHHHHH} (1–6 hex) — the ECMAScript / regex-crate spelling, a synonym of \x{…} via parse_braced_hex_scalar. \U{…} is not this form (\U stays 8 fixed digits). | |
| constexpr std::int32_t | parse_braced_hex_scalar () |
Decodes a braced hex scalar HHHHHH} (1–6 hex digits; the { already consumed), shared by \N{U+…}, \x{…} and \u{…}; rejects a surrogate or a value past U+10FFFF. | |
| constexpr std::int32_t | parse_named_codepoint () |
Decodes a \N{U+XXXX} escape (1–6 hex digits). re writes \N{NAME}; the Python binding rewrites a name to this form, so the engine sees only the scalar. | |
| constexpr std::int32_t | parse_braced_hex_escape () |
Decodes a \x{XXXX} escape (RE2/Perl; callers gate on !is_ecma(), ECMAScript spelling it \u{...}). Rejected in bytes mode, read from the scope stack (see parse_property_table). The backslash and x are already consumed. | |
| constexpr std::int32_t | emit_codepoint_utf8 (ast &out, std::int32_t cp) |
| Emits a code point as its 1–4 UTF-8 bytes — the same byte-level form a literal multi-byte character produces — as a single atom (a byte node, or a concat). | |
| constexpr std::int32_t | emit_literal_codepoint (ast &out, std::int32_t cp) |
Emits a code-point literal (code-point provenance: a raw character or \\u/\\U). | |
| constexpr std::int32_t | parse_escape (ast &out) |
| Parses an escape outside a character class. | |
| constexpr std::int32_t | parse_class_item (char_class &klass, std::vector< code_range > &ranges, bool &property_derived) |
| Parses one member inside a character class. | |
| constexpr std::int32_t | parse_class (ast &out) |
Parses a bracketed character class [...] or [^...]. | |
Static Private Member Functions | |
| static constexpr bool | is_ascii_alnum (char ch) |
Returns true if ch is in [0-9A-Za-z]. | |
| static constexpr char_class | space_set_text_ascii_component () |
\s's ASCII bitmap in TEXT mode: space_set() ([ \t\n\r\f\v], the ASCII-mode set) plus U+001C-U+001F. re matches FS/GS/RS/US with \s but not with (?a)\s, and shorthand_ranges drops wholly-ASCII ranges, so the bitmap itself must differ by mode. | |
| static constexpr shorthand_spec | shorthand_class (char letter, bool text_mode) |
| Maps a shorthand letter to its shorthand_spec; the one place this fact lives, shared by the atom (parse_escape) and in-class (parse_class_item) paths. | |
| static constexpr loose_buf | loose_key (std::string_view s) |
Loose-matches a property name (UAX44-LM3): drops _, - and spaces, lowercases the rest. | |
| static constexpr flags | flag_for_letter (char letter) |
| Maps a flag letter to its flags value. | |
| static constexpr bool | is_flag_letter (char letter) |
Returns true if letter is a flag letter (imsaxU). | |
| static constexpr bool | is_ascii_digit (char ch) |
Returns true if ch is an ASCII digit (0–9). | |
| static constexpr flags | without (flags value, flags bit) |
value with bit cleared, via real::flags_without (a cast narrower than the 16-bit underlying type would drop flags::ungreedy from every scope). | |
| static constexpr bool | is_name_start (char ch) |
Returns true if ch may start a group name in bytes mode. | |
| static constexpr bool | is_name_start_cp (char32_t cp) |
Returns true if cp may start a text-mode group name (str.isidentifier()). | |
Private Attributes | |
| std::string_view | pattern_ |
| The pattern being parsed. | |
| std::size_t | pos_ {} |
| Current read offset into pattern_. | |
| std::int32_t | depth_ {} |
| Current group nesting (see max_nesting_depth). | |
| std::vector< flags > | flag_scopes_ |
| Flag set in force per nesting level; the top is current (current_flags). Read flags here, never from a member, so scoped groups are honoured. | |
| bool | in_lookaround_ {} |
| True while parsing a lookaround sub-pattern (rejects nesting). | |
| bool | bytes_ {} |
In flags::bytes mode, rejects code-point escapes (\u/\U). | |
| bool | ecma_ {} |
ECMAScript grammar: \A \Z \< \> are identity-escape literals, not anchors. | |
Recursive-descent parser: a pattern string in, an ast out.
|
inlineexplicitconstexpr |
Binds the parser to a pattern and the constructor flags.
| [in] | pattern | The pattern text (borrowed, must outlive use). |
| [in] | initial_flags | Flags from the constructor; only verbose affects parsing (a leading (?x) can add it too). |
|
inlineconstexprprivate |
Consumes the current character if it equals ch.
| [in] | ch | The character to match. |
true (and advances) on a match, else false.
|
inlineconstexprprivate |
Interns klass and appends a node_kind::klass node.
| [in,out] | out | The AST being built. |
| [in] | klass | The class as written (before negation). |
| [in] | negated | Whether the class was written negated. |
|
inlineconstexprprivate |
Interns a class bitmap and appends a node_kind::klass node.
| [in,out] | out | The AST being built. |
| [in] | klass | The class bitmap as written (before negation). |
| [in] | negated | Whether the class was written negated. |
| [in] | ranges | Non-ASCII code-point ranges (code-point mode; empty otherwise). |
| [in] | codepoint_predicate | Emit as a match-time klass_cp, not the byte-NFA. |
|
inlineconstexprprivate |
Appends node to the pool.
| [in,out] | out | The AST being built. |
| [in] | node | The node to append. |
|
inlineconstexprprivate |
Returns true if the unit at the cursor is a LETTER — an unknown flag, not a terminator.
CPython's _parse_flags splits on str.isalpha() over the mode's unit: the whole code point in text mode, the byte read as latin-1 under flags::bytes. gc_property::L agrees with str.isalpha() on every code point, so (?ié is unknown flag and (?i😀) a terminator fault.
true if the cursor is on a letter, in the mode's own unit.
|
inlineconstexprprivate |
Consumes a run of inline flag letters (imsxaU).
flags::none if the run was empty).
|
inlineconstexprprivate |
The flag set in force at the current nesting level (the scope-stack top).
|
inlineconstexprprivate |
Emits a code point as its 1–4 UTF-8 bytes — the same byte-level form a literal multi-byte character produces — as a single atom (a byte node, or a concat).
| [in,out] | out | The AST being built. |
| [in] | cp | A code point in [0, 0x10FFFF]. |
|
inlineconstexprprivate |
Emits a code-point literal (code-point provenance: a raw character or \\u/\\U).
Under icase, a CASED literal is promoted to a foldable singleton class so the compiler folds it to its whole case orbit (k↦{k, K, Kelvin}, é↦{é, É}). An ASCII letter folds in any mode; a non-ASCII code point folds only in text mode (a bytes class carries no ranges). A non-cased literal, or no icase, keeps the zero-overhead byte / UTF-8 path. \\xHH has byte provenance and never routes here, so it is never folded — the deliberate provenance split.
| [in,out] | out | The AST the node is added to. |
| [in] | cp | The literal's code point. |
icase.
|
inlineconstexprprivate |
Returns true if the read offset is at or past the end of the pattern.
|
inlineconstexprprivate |
Aborts the parse with a real::regex_error at the current offset.
A template so the always-throwing body stays a legal constexpr function (the IFNDR rule spares templates); under constant evaluation the throw fails compilation with message in the trace.
| Error | The exception type to throw (defaults to regex_error). |
| [in] | message | The cause, shown in the error and the constexpr trace. |
|
inlineconstexprprivate |
Fails a bad character-class range the way re does, quoting it (bad character range z-a).
Quotes the source text: [\x7f-\x20] names \x7f-\x20, where re prints \x-\x.
| [in] | begin | Offset of the range's first byte, which is where re reports. |
| [in] | end | One past the range's last byte, as far as the caller had read. |
| real::regex_error | always. |
|
inlineconstexprprivate |
Fails a duplicate group name the way re does, naming it and both group numbers.
| [in] | begin | Offset of the offending name's first byte. |
| [in] | end | One past its last byte. |
| [in] | group | The capture number being defined now. |
| [in] | previous | The capture number that already carries this name. |
| real::regex_error | always. |
|
inlineconstexprprivate |
Fails a (?… extension, naming it when it is one REAL excludes by design.
Callouts, recursion and subroutine calls are non-regular (super-linear), so they are named and tagged unsupported (error_kind), reported at the character after (?. Anything else is fail_unknown_extension's, at the ? as in re.
| [in] | question_pos | Offset of the ? in (?, forwarded to fail_unknown_extension. |
| real::regex_error | always. |
|
inlineconstexprprivate |
Fails a group name the way re does: quoting the name it read, at the name's start.
As re: an empty name is missing group name, an unterminated one missing >, unterminated name, and a bad character quotes the whole name it read.
| [in] | begin | Offset of the name's first byte, which is where re reports. |
| real::regex_error | always. |
|
inlineconstexprprivate |
Fails with "unknown flag" if the next unit is a letter that is not a flag.
As in CPython, this precedes the terminator check (require_scoped_flags_colon). Call only once a flags group has started (a flag letter or -): the global prefix must backtrack on (?P / (?#, not call P an unknown flag.
|
inlineconstexprprivate |
Fails a truncated escape the way re does (incomplete escape \x1): quoting pattern_[backslash_pos, pos_), reported at the backslash.
| [in] | backslash_pos | Offset of the \ that opened the escape. |
| real::regex_error | always. |
|
inlineconstexprprivate |
Fails an unrecognised (?… extension the way re does: naming it.
As re, the message quotes everything consumed since the ? plus the failing character ((?z) names ?z, (?P) names ?P)) and the offset is the ?'s. At end of pattern it is unexpected end of pattern at the read offset.
| [in] | question_pos | Offset of the ? in (? — open_pos + 1 at every call site. |
| real::regex_error | always. |
|
inlineconstexprprivate |
Like fail, but tags the error unsupported (well-formed but beyond the linear engine, e.g. a backreference or a nested lookaround) so a binding classifies it without reading the message. Templated for the same reason as fail.
| [in] | message | The diagnostic text, reported at the current read offset. |
|
inlineconstexprprivate |
Fails an escape REAL does not implement, naming it (unsupported escape sequence \q) and reporting at the backslash, as re does; tagged unsupported.
| [in] | backslash | Offset of the \ that opened the escape. |
| real::regex_error | always. |
|
inlinestaticconstexprprivate |
Maps a flag letter to its flags value.
| [in] | letter | One of 'i', 'm', 's', 'x', 'a', 'U'. |
|
inlineconstexprprivate |
Consumes one hexadecimal digit.
| [in] | backslash_pos | Offset of the \ that opened the escape, so a missing digit is reported as a truncated escape at the sequence's start. |
[0, 15]. | real::regex_error | if the next character is not a hex digit. |
|
inlinestaticconstexprprivate |
Returns true if ch is in [0-9A-Za-z].
| [in] | ch | A character. |
true if ch is in [0-9A-Za-z].
|
inlinestaticconstexprprivate |
Returns true if ch is an ASCII digit (0–9).
| [in] | ch | A character. |
true if ch is an ASCII digit.
|
inlineconstexprprivate |
True when ascii (re.A) is in force at the current scope.
|
inlineconstexprprivate |
True in the ECMAScript grammar. Not scopable, so every scope carries it; read it here, not from ecma_, which the flag-scope ratchet counts.
|
inlinestaticconstexprprivate |
Returns true if letter is a flag letter (imsaxU).
| [in] | letter | A character. |
true if letter is a flag letter (imsaxU).
|
inlineconstexprprivate |
True when icase (re.I) is in force at the current scope.
|
inlinestaticconstexprprivate |
Returns true if ch may start a group name in bytes mode.
| [in] | ch | A character. |
true if ch may start a group name.
|
inlinestaticconstexprprivate |
Returns true if cp may start a text-mode group name (str.isidentifier()).
Measured against UCD 16.0.0: isidentifier() is XID_Start ∪ {U+005F}, then XID_Continue.
| [in] | cp | A code point. |
true if cp may start a name.
|
inlineconstexprprivate |
True when verbose mode (re.X) is in force at the current scope.
|
inlinestaticconstexprprivate |
Loose-matches a property name (UAX44-LM3): drops _, - and spaces, lowercases the rest.
| [in] | s | The name as written in the pattern. |
|
inlineconstexprprivate |
Merges an in-class shorthand (\w \d \s, or negated \W \D \S) into the class being built: its ASCII bitmap plus, in text mode, its non-ASCII ranges (each complemented when negated). In text mode, sets property_derived.
| [in,out] | klass | The class being built, receiving the ASCII bitmap. |
| [in,out] | ranges | The class's non-ASCII ranges, appended to in text mode. |
| [in] | prop_ascii | The shorthand's ASCII bitmap. |
| [in] | table | The shorthand's full Unicode range table. |
| [in] | negated | True for the uppercase form (\W \D \S). |
| [out] | property_derived | Set when the class must be emitted as a klass_cp. |
|
inlineconstexprprivate |
Merges \p{Name} / \P{Name} into the class being built: merge_property without the flags::ascii restriction. \P merges the complement, as \W does; an enclosing [^...] negates on top ([^\P{L}] == [\p{L}]).
| [in,out] | klass | The class being built, receiving the ASCII bitmap. |
| [in,out] | ranges | The class's non-ASCII ranges, appended to. |
| [in] | table | The property's full range table. |
| [in] | negated | True for \P{...}, merging the complement. |
| [out] | property_derived | Set so the class is emitted as a klass_cp. |
|
inlineconstexprprivate |
Allocates the next capture group number.
| [in,out] | out | The AST being built. |
| [in] | open_pos | Offset of the group's ( (for error reporting). |
| real::regex_error | beyond max_group_count. |
|
inlineconstexpr |
Parses the whole pattern.
| real::regex_error | on any unsupported or malformed syntax. |
|
inlineconstexprprivate |
Parses ‘alternation := sequence (’|' sequence)*`.
The leftmost branch is preferred (Python / Perl semantics, not longest).
| [in,out] | out | The AST being built. |
|
inlineconstexprprivate |
Parses one atom: a literal, ., a class, a group, an anchor or an escape.
| [in,out] | out | The AST being built. |
|
inlineconstexprprivate |
Parses an atomic group after (?> (the > is not yet consumed).
A non-capturing node_kind::group with possessive = true; a capture inside keeps its number. Linearity restrictions are the compiler's.
| [in,out] | out | The AST being built. |
| [in] | open_pos | Offset of the group's ( (for error reporting). |
|
inlineconstexprprivate |
Decodes a \x{XXXX} escape (RE2/Perl; callers gate on !is_ecma(), ECMAScript spelling it \u{...}). Rejected in bytes mode, read from the scope stack (see parse_property_table). The backslash and x are already consumed.
[0, 0x10FFFF] (never a surrogate). | real::regex_error | in bytes mode, or (via parse_braced_hex_scalar) on a malformed or unterminated {...}, a surrogate, or a value beyond U+10FFFF. |
|
inlineconstexprprivate |
Decodes a braced hex scalar HHHHHH} (1–6 hex digits; the { already consumed), shared by \N{U+…}, \x{…} and \u{…}; rejects a surrogate or a value past U+10FFFF.
[0, 0x10FFFF] (never a surrogate). | real::regex_error | on a missing digit run, an unterminated brace, a surrogate, or a value beyond U+10FFFF. |
|
inlineconstexprprivate |
Parses a single-byte escape (valid inside and outside classes).
Handles \n \t \r \f \v \a \0, \xHH and escaped ASCII punctuation.
| [in] | backslash | Offset of the \ that opened the escape, forwarded so a truncated \xH reports at the sequence's start the way re does. |
\d \w \s, etc.). | real::regex_error | on a malformed \x escape. |
|
inlineconstexprprivate |
Parses a bracketed character class [...] or [^...].
Supports ranges, escapes and the embedded set escapes; a ] right after [ or [^ is a literal, and a trailing - is a literal dash.
| [in,out] | out | The AST being built. |
| real::regex_error | on an unterminated class or a bad range. |
|
inlineconstexprprivate |
Parses one member inside a character class.
| [in,out] | klass | The class being built; a set member (\d etc.) is merged directly into it. |
| [in,out] | ranges | The class's non-ASCII code-point ranges; a Unicode shorthand (\d \w \s, or a negated one) appends its ranges here in text mode. |
| [in,out] | property_derived | Set when a Unicode shorthand contributed, so the whole class is emitted as a match-time klass_cp (text mode only). |
klass. | real::regex_error | on invalid UTF-8 or an unsupported escape. |
|
inlineconstexprprivate |
Parses a \<digit> escape via the shared decode_digit_escape().
Octal escapes (\0, \012, a three-octal-digit run) become one byte (value & 0xff, as \xHH). A decimal group number is a back-reference, unsupported.
| real::regex_error | on an over-long octal escape or a back-reference. |
|
inlineconstexprprivate |
Parses an escape outside a character class.
Handles the class escapes \d \D \w \W \s \S, the anchors \A \Z \b \B, and single-byte escapes.
| [in,out] | out | The AST being built. |
| real::regex_error | on a dangling or unsupported escape. |
|
inlineconstexprprivate |
Consumes a leading global-flags group – (?imsxaU), or (?flags-flags) with a removal suffix – if present. The accepted letters are i m s x a U (is_flag_letter).
As in Python 3.11+, global flags are legal only at the very start (parse_group rejects later ones). As in RE2, a -removed suffix ((?i-s), (?-s)) clears flags from the base scope.
| [in,out] | out | Receives added letters in ast::inline_flags and removed ones in ast::inline_removed — two fields, since inline_flags is OR-ed across calls; the caller adds, then clears. |
true if a flags group was consumed (position advanced), else false (position restored, for parse_group to handle).
|
inlineconstexprprivate |
Parses a group construct.
Grammar:
Also lookarounds, atomic groups, comments and scoped flags. Unsupported extensions (backreferences, conditionals, recursion, callouts) fail naming the feature. Under flags::ecma, (?#...), (?P... and (?>...) are "unknown extension". Nesting beyond max_nesting_depth is rejected.
| [in,out] | out | The AST being built. |
| real::regex_error | on an unterminated or unsupported group. |
< A (?flags:...) group pushed a scope to pop after the body.
|
inlineconstexprprivate |
Parses a group name up to > and records it.
Text mode: a Python identifier (XID_Start ∪ {_}, then XID_Continue), so (?P<é>a) is accepted. Bytes mode: [A-Za-z_][A-Za-z0-9_]*, matching re on a bytes pattern. A malformed UTF-8 sequence is "bad character in group name" at the bad byte — not a decode error, which would be a different divergence. The name is stored as a byte span into the pattern, so group("é") / groupindex resolve by the same octets.
| [in,out] | out | The AST; the name is appended to ast::names. |
| [in] | group | The capture number this name refers to. |
| real::regex_error | on a bad character or a duplicate name. |
|
inlineconstexprprivate |
Parses a lookaround after (?= / (?! (ahead) or (?<= / (?<! (behind) — the =/! is not yet consumed.
Its capture groups take numbers (so outer numbering holds) but are compiled capture-free. A nested lookaround is rejected; boundedness is the compiler's.
| [in,out] | out | The AST being built. |
| [in] | direction | Ahead or behind. |
| [in] | open_pos | Offset of the group's ( (for error reporting). |
|
inlineconstexprprivate |
Decodes a \N{U+XXXX} escape (1–6 hex digits). re writes \N{NAME}; the Python binding rewrites a name to this form, so the engine sees only the scalar.
Rejects bytes mode (as re's bad escape \N) and a malformed {U+…}. The backslash and N are already consumed.
[0, 0x10FFFF] (never a surrogate).
|
inlineconstexprprivate |
Consumes p/P and {Name} (or one letter), strips a leading ^ (native dialects only) and resolves the name; rejects bytes mode. Shared by the atom and the in-class paths. Entered on the p/P; leaves pos_ just past the name.
|
inlineconstexprprivate |
Wraps atom in a repeat node if a quantifier follows.
Grammar: ‘quantifier := (’*' | '+' | '?' | '{n}' | '{n,}' | '{,m}' | '{,}' | '{n,m}') '?'? ({,}is{0,}). As in Python, an ill-formed{...}stays literal (a{,a{},a{2,3x). A brace quantifier with no atom before it ({2}a,{,3}a) is literal here butnothing to repeat inre`: a deliberate divergence (div_compiles). A bare anchor cannot be repeated.
| [in,out] | out | The AST being built. |
| [in] | atom | Index of the atom the quantifier would apply to. |
atom unchanged if no quantifier.
|
inlineconstexprprivate |
Scans a \Q...\E span (\Q already consumed) and emits its characters as literal atoms via emit_literal_codepoint, as RE2 does.
The span ends at \E (consumed) or at end of pattern. Everything inside is literal, including whitespace in verbose mode and a backslash not followed by E (\Qa\Qb\E is a\Qb): no escape processing, no nesting. All atoms but the last are chained; the last is left for the caller's quantifier.
| [in,out] | out | The AST being built. |
\Q\E). | real::regex_error | on an invalid UTF-8 byte inside the span (text mode). |
|
inlineconstexprprivate |
Reads an optional decimal repeat count.
| real::regex_error | if the count exceeds max_repeat_count (counts are unrolled). |
|
inlineconstexprprivate |
Parses sequence := (atom quantifier?)*, stopping at | or ).
Also handles \Q...\E (RE2/Perl, not ecma) here, since it emits a sequence, not one atom. As in RE2, a following quantifier binds to the span's last character (\Qab\E+ == ab+), and an empty \Q\E is invisible (a\Q\E+ == a+): the quantifier re-binds to the previous atom (re-chained through prev), or fails "nothing to repeat" when there is none.
| [in,out] | out | The AST being built. |
|
inlineconstexprprivate |
Decodes a \uHHHH (4 hex) or \UHHHHHHHH (8 hex) code-point escape (str only), or the braced form \u{HHHHHH} (1–6 hex) — the ECMAScript / regex-crate spelling, a synonym of \x{…} via parse_braced_hex_scalar. \U{…} is not this form (\U stays 8 fixed digits).
Rejects bytes mode, a surrogate, a value past U+10FFFF and incomplete hex. The backslash and u/U are already consumed.
| [in] | capital | True for \U (8 digits), false for \u (4 digits or \u{…}). |
| [in] | backslash | Offset of the \ that opened the escape, for a truncated-escape report. |
[0, 0x10FFFF] (never a surrogate).
|
inlineconstexprprivate |
Parses \p{Name} / \P{Name} / \pX outside a class into a code-point class node (klass_cp, as \w). Negation is the node flag XORed with the caret (\P{^L} == \p{L}). Entered on the letter after \.
| [in,out] | out | The AST the class node is added to. |
| [in] | negated | True for \P, false for \p. |
|
inlineconstexprprivate |
Returns the current character without consuming it (undefined at eof()).
|
inlineconstexprprivate |
Splits a property's ranges into its ASCII bitmap (< 0x80) and its non-ASCII ranges. Unlike \w, a Unicode property is not restricted by flags::ascii.
| [in] | table | The property's full range table. |
| [out] | ascii | Bitmap receiving its members below 0x80. |
| [out] | high | Ranges receiving its members at or above 0x80. |
|
inlineconstexprprivate |
After a flags run, requires : (scoped body) or fails naming the terminator fault.
As re's _parse_flags:
: — scoped body, consumed, return) — an unscoped group not at the start (a(?i)b; a leading one was taken by parse_global_flags_prefix): global flags not at the start of the expression-flags run, anything else (including EOF) — missing :missing -, : or ) ((?i*), (?i7), (?i)| [in] | after_removal | Whether the run just consumed a -flags suffix. |
| [in] | open_pos | Offset of the group's (, where re reports the placement fault. |
|
inlineconstexprprivate |
Resolves a \p{...} property name to its code-point ranges, or fails.
A gc= / sc= / scx= prefix (or its long form) picks the namespace. As in PCRE2, a bare name tries General_Category, then Script, then a binary property, and never means Script_Extensions. A Script's ranges are gathered from the script partition; GC, binary-property and scx ranges come from their own tables (the last two are not partitions).
| [in] | name | The property name as written, prefix and all. |
|
inlinestaticconstexprprivate |
Maps a shorthand letter to its shorthand_spec; the one place this fact lives, shared by the atom (parse_escape) and in-class (parse_class_item) paths.
| [in] | letter | The shorthand letter (d D w W s S). |
| [in] | text_mode | The caller's text_shorthand; only \s/\S read it (see space_set_text_ascii_component). |
|
inlineconstexprprivate |
The non-ASCII part of a shorthand's range table, or nothing in bytes / ASCII mode. Wholly-ASCII ranges are dropped: the bitmap already covers them.
| [in] | table | The shorthand's full Unicode range table. |
>= 0x80; empty when the shorthand stays ASCII-only.
|
inlineconstexprprivate |
In verbose mode, consumes insignificant whitespace and # comments.
Called only between tokens outside classes; an escaped space (\) is a literal read by the escape parser.
|
inlinestaticconstexprprivate |
\s's ASCII bitmap in TEXT mode: space_set() ([ \t\n\r\f\v], the ASCII-mode set) plus U+001C-U+001F. re matches FS/GS/RS/US with \s but not with (?a)\s, and shorthand_ranges drops wholly-ASCII ranges, so the bitmap itself must differ by mode.
\s.
|
inlineconstexprprivate |
Whether a shorthand (\d \w \s) should be a text-mode Unicode code-point predicate: true in the default text mode, false in bytes mode or under flags::ascii (re.A).
|
inlineconstexprprivate |
Tries to parse {n} / {n,} / {,m} / {n,m} starting at {.
| [out] | min | Lower bound on success. |
| [out] | max | Upper bound on success (-1 for unbounded). |
true on a valid quantifier (position advanced); false if the braces are not a quantifier (position restored — literal text). | real::regex_error | when the bounds are impossible (min > max). |
|
inlinestaticconstexprprivate |
value with bit cleared, via real::flags_without (a cast narrower than the 16-bit underlying type would drop flags::ungreedy from every scope).
| [in] | value | The flag set to clear from. |
| [in] | bit | The flag to clear. |
value without bit.