|
REAL
Regular Expression Algorithmic Library — constexpr C++20 regex
|
A compiled regular expression, parameterized on its storage policy. More...
#include <real.hpp>
Classes | |
| struct | replacement_piece |
| One piece of a parsed replacement template: a slice of the template, or a group to copy. More... | |
Public Types | |
| using | result_type = basic_match_result< typename Storage::slot_storage > |
| This regex's match-result type. | |
| using | owning_result_type = basic_match_result< typename Storage::slot_storage, typename Storage::name_owner > |
| What a single attempt on a temporary regex yields. | |
Public Member Functions | |
| constexpr | basic_regex (std::string_view pattern, flags compile_flags=flags::none) |
Compiles pattern at run time (the real::regex constructor). | |
| constexpr | basic_regex ()=default |
| Default constructor for the stateless compile-time storage (static_regex). | |
| constexpr result_type | match (std::string_view text) const & |
Match anchored at the start of text (Python re.match). | |
| constexpr result_type | fullmatch (std::string_view text) const & |
Match the entire text (Python re.fullmatch). | |
| constexpr result_type | search (std::string_view text) const & |
Leftmost match anywhere in text (Python re.search). | |
| constexpr result_type | match (std::string_view text, std::size_t pos, std::size_t endpos=npos) const & |
Region-aware match: anchored at pos within text[0:endpos] (Python re.match with pos / endpos). Byte offsets; pos is not a slice (see run — \A fails at pos > 0); endpos defaults to the end of text. | |
| bool | can_extend (std::string_view text, std::size_t pos=0) const |
Whether match(text, pos) could come out differently if text continued past its end. | |
| constexpr result_type | fullmatch (std::string_view text, std::size_t pos, std::size_t endpos=npos) const & |
Region-aware fullmatch: the whole region [pos, endpos) must match. | |
| constexpr result_type | search (std::string_view text, std::size_t pos, std::size_t endpos=npos) const & |
Region-aware search: leftmost match within [pos, endpos). | |
| constexpr result_type | match (const char *text) const & |
match overload for string literals. | |
| constexpr result_type | fullmatch (const char *text) const & |
fullmatch overload for string literals. | |
| constexpr result_type | search (const char *text) const & |
search overload for string literals. | |
| constexpr result_type | match (const char *text, std::size_t pos, std::size_t endpos=npos) const & |
Region-aware match overload for string literals. | |
| constexpr result_type | fullmatch (const char *text, std::size_t pos, std::size_t endpos=npos) const & |
Region-aware fullmatch overload for string literals. | |
| constexpr result_type | search (const char *text, std::size_t pos, std::size_t endpos=npos) const & |
Region-aware search overload for string literals. | |
| constexpr owning_result_type | match (std::string_view text) const && |
match on a temporary regex; the result owns its name context. | |
| constexpr owning_result_type | fullmatch (std::string_view text) const && |
fullmatch on a temporary regex; the result owns its name context. | |
| constexpr owning_result_type | search (std::string_view text) const && |
search on a temporary regex; the result owns its name context. | |
| constexpr owning_result_type | match (std::string_view text, std::size_t pos, std::size_t endpos=npos) const && |
Region-aware match on a temporary regex; the result owns its name context. | |
| constexpr owning_result_type | fullmatch (std::string_view text, std::size_t pos, std::size_t endpos=npos) const && |
Region-aware fullmatch on a temporary regex; the result owns its name context. | |
| constexpr owning_result_type | search (std::string_view text, std::size_t pos, std::size_t endpos=npos) const && |
Region-aware search on a temporary regex; the result owns its name context. | |
| constexpr owning_result_type | match (const char *text) const && |
match on a temporary regex, string-literal overload. | |
| constexpr owning_result_type | fullmatch (const char *text) const && |
fullmatch on a temporary regex, string-literal overload. | |
| constexpr owning_result_type | search (const char *text) const && |
search on a temporary regex, string-literal overload. | |
| constexpr owning_result_type | match (const char *text, std::size_t pos, std::size_t endpos=npos) const && |
Region-aware match on a temporary regex, string-literal overload. | |
| constexpr owning_result_type | fullmatch (const char *text, std::size_t pos, std::size_t endpos=npos) const && |
Region-aware fullmatch on a temporary regex, string-literal overload. | |
| constexpr owning_result_type | search (const char *text, std::size_t pos, std::size_t endpos=npos) const && |
Region-aware search on a temporary regex, string-literal overload. | |
| constexpr basic_match_range< Storage > | find_iter (std::string_view text) const & |
Lazy range over all non-overlapping matches (Python re.finditer). | |
| constexpr basic_match_range< Storage > | find_iter (const char *text) const & |
find_iter overload for string literals. | |
| constexpr basic_match_range< Storage > | find_iter (const char *text, std::size_t pos, std::size_t endpos=npos) const & |
Region-aware find_iter overload for string literals. | |
| constexpr basic_match_range< Storage > | find_iter (std::string_view text, std::size_t pos, std::size_t endpos=npos) const & |
Region-aware find_iter: iterate matches within [pos, endpos) (Python finditer with pos / endpos). endpos truncates the subject to a view so iteration stops at it; pos is the start, not a slice (see run). Byte offsets; endpos defaults to the end of text. | |
| constexpr basic_match_range< Storage > | find_iter_longest (std::string_view text, std::size_t pos=0, std::size_t endpos=npos) const & |
Experimental leftmost-**longest** find_iter (POSIX bounds), the iterator twin of search_longest. Region semantics as find_iter; captures are the winning thread's, not POSIX submatch. Every fast path is bypassed. | |
| constexpr basic_match_range< Storage > | find_iter_longest (const char *text, std::size_t pos=0, std::size_t endpos=npos) const & |
Region-aware find_iter_longest overload for string literals. | |
| basic_match_range< Storage > | find_iter_longest (const std::string &&, std::size_t=0, std::size_t=npos) const &=delete |
Deleted: the range borrows the subject, so a temporary std::string would dangle. Needs the const char* forwarder above, without which a bare literal would be ambiguous. | |
| basic_match_range< Storage > | find_iter_longest (std::string_view, std::size_t=0, std::size_t=npos) const &&=delete |
Deleted: find_iter_longest on a temporary regex would dangle, at every arity: without the same defaults, a shorter call would bind the const& overload. | |
| basic_match_range< Storage > | find_iter_longest (const char *, std::size_t=0, std::size_t=npos) const &&=delete |
Deleted: same, spelled for const char* so a literal resolves HERE rather than becoming ambiguous — the reader is told the REGEX is the temporary, which is true. | |
| basic_match_range< Storage > | find_iter (std::string_view text) const &&=delete |
Deleted: find_iter on a temporary regex would dangle. | |
| basic_match_range< Storage > | find_iter (const char *text) const &&=delete |
Deleted: find_iter on a temporary regex would dangle. | |
| basic_match_range< Storage > | find_iter (const char *text, std::size_t, std::size_t=npos) const &&=delete |
Deleted: region find_iter on a temporary regex would dangle. Spelled for const char* too, so a literal resolves here and the diagnostic names the regex as the temporary. | |
| basic_match_range< Storage > | find_iter (std::string_view text, std::size_t, std::size_t=npos) const &&=delete |
Deleted: region find_iter on a temporary regex would dangle. | |
| constexpr std::size_t | count_matches (std::string_view text, std::size_t pos=0, std::size_t endpos=npos) const |
| Count non-overlapping matches without allocating result objects. | |
| constexpr std::vector< result_type > | find_all (std::string_view text) const & |
All matches, eagerly (like Python re.findall, but full results). | |
| constexpr std::vector< result_type > | find_all (const char *text) const & |
find_all overload for string literals. | |
| std::vector< result_type > | find_all (std::string_view text) const &&=delete |
Deleted: find_all on a temporary regex would dangle. | |
| std::vector< result_type > | find_all (const char *text) const &&=delete |
Deleted: find_all on a temporary regex would dangle. | |
| constexpr std::string | replace (std::string_view text, std::string_view replacement, std::size_t max_count=0) const |
Replaces matches in text (ECMAScript / std::regex_replace $1). | |
| constexpr std::vector< std::string_view > | split (std::string_view text, std::size_t max_splits=0) const |
Splits text on matches (Python re.split). | |
| constexpr std::vector< std::string_view > | split (const char *text, std::size_t max_splits=0) const |
split overload for string literals. | |
| result_type | match (const std::string &&text) const =delete |
| Deleted: temporary text would dangle. | |
| result_type | fullmatch (const std::string &&text) const =delete |
| Deleted: temporary text would dangle. | |
| result_type | search (const std::string &&text) const =delete |
| Deleted: temporary text would dangle. | |
| result_type | match (const std::string &&text, std::size_t, std::size_t=npos) const =delete |
| Deleted: temporary text would dangle. | |
| result_type | fullmatch (const std::string &&text, std::size_t, std::size_t=npos) const =delete |
| Deleted: temporary text would dangle. | |
| result_type | search (const std::string &&text, std::size_t, std::size_t=npos) const =delete |
| Deleted: temporary text would dangle. | |
| basic_match_range< Storage > | find_iter (const std::string &&text) const &=delete |
| Deleted: temporary text would dangle. | |
| basic_match_range< Storage > | find_iter (const std::string &&text, std::size_t, std::size_t=npos) const &=delete |
| Deleted: temporary text would dangle. | |
| std::vector< result_type > | find_all (const std::string &&text) const &=delete |
| Deleted: temporary text would dangle. | |
| std::vector< std::string_view > | split (const std::string &&text, std::size_t max_splits=0) const =delete |
| Deleted: temporary text would dangle. | |
| constexpr std::string_view | pattern () const |
| Returns the pattern text this regex was compiled from. | |
| constexpr flags | compile_flags () const |
The flag set in force: constructor flags, plus a leading (?imsxa) group, minus its -removal. | |
| constexpr std::size_t | group_count () const |
| Returns the number of capturing groups (excluding group 0). | |
| constexpr detail::program_view | raw_program () const |
| The raw compiled program, for embedders (advanced). | |
| constexpr bool | has_first_byte_set () const noexcept |
| Whether first-byte filtering is useful for this pattern. | |
| constexpr std::optional< unsigned char > | unique_first_byte () const noexcept |
| The single byte every non-empty match must begin with, if unique. | |
| constexpr bool | may_start_with (unsigned char byte) const noexcept |
Whether a non-empty match can begin with byte (sound, conservative). | |
| constexpr std::size_t | left_context () const noexcept |
| How far before a position a match there may read: an upper bound, in bytes. | |
| constexpr std::size_t | group_index (std::string_view name) const |
| Resolves a group name to its number. | |
| constexpr std::vector< std::pair< std::string_view, std::size_t > > | named_groups () const |
| All named groups as (name, number) pairs, in declaration order. | |
| constexpr std::size_t | named_group_count () const |
| How many named groups the pattern declares. | |
| constexpr std::pair< std::string_view, std::size_t > | named_group_at (std::size_t index) const |
The index-th named group, in declaration order. | |
| result_type | search_longest (std::string_view text) const & |
EXPERIMENTAL, opt-in: a single leftmost-**longest** search (POSIX / RE2 set_longest_match). Among matches at the leftmost start it returns the longest, so a lazy quantifier behaves greedily; captures are the leftmost-first thread's at that bound (not POSIX submatch). Runs on the general Pike loop. Not yet a stable API; its iteration twin is find_iter_longest. | |
| owning_result_type | search_longest (std::string_view text) const && |
search_longest on a temporary regex; the result owns its name context. | |
| result_type | search_longest (std::string_view text, std::size_t pos, std::size_t endpos=npos) const & |
Region-aware form of search_longest — leftmost-longest search within [pos, endpos). pos is the start (not a slice, per run); endpos truncates the subject. Byte offsets. | |
| owning_result_type | search_longest (std::string_view text, std::size_t pos, std::size_t endpos=npos) const && |
Region-aware search_longest on a temporary regex; the result owns its name context. | |
| result_type | search_longest (const char *text) const & |
search_longest overload for string literals. | |
| owning_result_type | search_longest (const char *text) const && |
search_longest on a temporary regex, string-literal overload. | |
| result_type | search_longest (const char *text, std::size_t pos, std::size_t endpos=npos) const & |
Region-aware search_longest overload for string literals. | |
| owning_result_type | search_longest (const char *text, std::size_t pos, std::size_t endpos=npos) const && |
Region-aware search_longest on a temporary regex, string-literal overload. | |
| result_type | search_longest (const std::string &&) const =delete |
| Deleted: temporary text would dangle. | |
| result_type | search_longest (const std::string &&, std::size_t, std::size_t=npos) const =delete |
| Deleted: temporary text would dangle. | |
Private Member Functions | |
| constexpr std::size_t | count_walk (std::string_view text, std::size_t pos, std::size_t endpos) const |
| count_matches's walk: the ordinary one, run matching-only. | |
| constexpr std::size_t | count_trailing_la (std::string_view region, std::size_t pos) const |
| count_matches over the trailing-lookaround walk, outlined. | |
| constexpr std::string_view | name_of (const detail::named_group &named_group) const |
| Returns its name, sliced from the pattern text. | |
| constexpr std::vector< replacement_piece > | parse_replacement (std::string_view replacement) const |
Reads replacement once into literal slices and group references. | |
| constexpr result_type | run (std::string_view text, detail::run_mode mode) const |
| Runs a single match attempt from offset 0 (backs match/search/fullmatch). | |
| constexpr result_type | run (std::string_view text, std::size_t pos, std::size_t endpos, detail::run_mode mode, match_semantics sem=match_semantics::first) const |
Region-aware single attempt: match over text[0:endpos] starting at pos. | |
| result_type | run_non_empty (std::string_view text, std::size_t pos, detail::run_mode mode) const |
run that accepts no empty match: the leftmost position where a non-empty match starts, and there the match the leftmost-first priority prefers among the non-empty ones (match_not_null). The DFAs do not model the rule, so the VM decides the search; it stays linear. Kept apart from run, whose every caller is a hot path. | |
Static Private Member Functions | |
| static constexpr owning_result_type | detach (result_type result) |
| Hands a result the name context it will need after this regex is gone. | |
Private Attributes | |
| Storage | program_ |
| The storage policy holding the compiled program. | |
Friends | |
| struct | detail::non_empty_access |
A compiled regular expression, parameterized on its storage policy.
Storage owns the program; matching allocates only per-run scratch — and nothing at all when the storage is compile-time. Use the real::regex and real::static_regex aliases rather than this template directly.
| Storage | real::detail::dynamic_storage or real::detail::static_storage. |
| using real::basic_regex< Storage >::owning_result_type = basic_match_result<typename Storage::slot_storage, typename Storage::name_owner> |
What a single attempt on a temporary regex yields.
Same spans and groups as result_type, plus ownership of the name tables the regex would otherwise lend. Bind it with auto. Spelling result_type (or real::match_result) does not compile, which is the point: there is no conversion that could drop the ownership and leave dangling views.
On static_regex the two aliases are the same type: the tables have static storage duration, so a result from a temporary is already safe.
|
inlineexplicitconstexpr |
Compiles pattern at run time (the real::regex constructor).
| [in] | pattern | The pattern text. |
| [in] | compile_flags | Optional flags (merged with a leading global-flags group, (?imsxaU) or (?flags-flags)). |
| real::regex_error | on an invalid or over-limit pattern. |
|
inline |
Whether match(text, pos) could come out differently if text continued past its end.
For text that arrives in pieces: a lexer may commit to the match at pos only once no further text can change it. [a-z]+ on "ab" could still grow; on "ab " it cannot. The end of text is treated as a place more text may follow, not as the end of the subject, so $, \b or a lookahead that read it make the answer true. Conservative: it may say true where more text would in fact change nothing, never false where it would. Runs the general matcher, not the fast paths.
| [in] | text | The text available so far. |
| [in] | pos | Byte offset the match is anchored at. |
text could change the match at pos.
|
inlineconstexpr |
The flag set in force: constructor flags, plus a leading (?imsxa) group, minus its -removal.
regex("(?-i)a", flags::icase) reports no flags::icase and matches case-sensitively — the accessor and the engine agree.
|
inlineconstexpr |
Count non-overlapping matches without allocating result objects.
Prefer this over walking find_iter when only the count is needed. Region semantics match find_iter – pos is a start offset, not a slice (\\A / ^ still see the absolute position).
| [in] | text | The subject text. |
| [in] | pos | Byte offset to begin counting from (0 = start of text). |
| [in] | endpos | Exclusive end of the region; npos = end of text. |
|
inlineconstexprprivate |
count_matches over the trailing-lookaround walk, outlined.
Outlined and cold, like basic_match_iterator::decide_batching, since inline, a second walk inside count_matches makes one unrelated branch charge byte-identical rows.
| [in] | region | The already-clamped subject. |
| [in] | pos | Where the walk starts. |
|
inlineconstexprprivate |
count_matches's walk: the ordinary one, run matching-only.
No caller can observe a group, so the walk runs with detail::pattern_hints::capture_free_walk set: the VM steps the same positions with no capture bookkeeping. The structural half is still asked (detail::capture_free_walk_structural), since a skippable save 0 gives a wrong answer. slot_count stays as is: the batched routes arm on slot_count == 2, so lowering it would reroute. The range sets the flag, keeping one program_view copy. Outlined, so count_matches carries no branch for it: one there charges byte-identical rows. noinline, not cold: this is the ordinary path.
| [in] | text | The subject. |
| [in] | pos | Where the walk starts. |
| [in] | endpos | Region end, as find_iter takes it. |
|
inlinestaticconstexprprivate |
Hands a result the name context it will need after this regex is gone.
Taken and returned by value so the prvalue from run is constructed straight into the parameter and named-returned out: the rvalue overloads pay no move for going through here.
| [in] | result | The freshly run result. |
|
inlineconstexpr |
find_all overload for string literals.
| [in] | text | NUL-terminated text. |
|
inlineconstexpr |
All matches, eagerly (like Python re.findall, but full results).
Lvalue-only, same reason as find_iter. High match counts allocate one result per hit; prefer count_matches when only the number matters, and find_iter when you can stream.
| [in] | text | The subject text (must outlive the results). |
|
inlineconstexpr |
find_iter overload for string literals.
| [in] | text | NUL-terminated text. |
|
inlineconstexpr |
Region-aware find_iter overload for string literals.
| [in] | text | NUL-terminated text (must outlive the range). |
| [in] | pos | Byte offset iteration starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
Lazy range over all non-overlapping matches (Python re.finditer).
Only callable on an lvalue regex: a C++20 range-for would dangle if the regex were a temporary (the range initializer dies before the loop body), so the rvalue overloads are deleted.
| [in] | text | The subject text (must outlive the range). |
|
inlineconstexpr |
Region-aware find_iter: iterate matches within [pos, endpos) (Python finditer with pos / endpos). endpos truncates the subject to a view so iteration stops at it; pos is the start, not a slice (see run). Byte offsets; endpos defaults to the end of text.
| [in] | text | Subject. |
| [in] | pos | Byte offset iteration starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
Region-aware find_iter_longest overload for string literals.
| [in] | text | NUL-terminated text (must outlive the range). |
| [in] | pos | Byte offset iteration starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
Experimental leftmost-**longest** find_iter (POSIX bounds), the iterator twin of search_longest. Region semantics as find_iter; captures are the winning thread's, not POSIX submatch. Every fast path is bypassed.
| [in] | text | Subject. |
| [in] | pos | Byte offset iteration starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
fullmatch overload for string literals.
| [in] | text | NUL-terminated text. |
|
inlineconstexpr |
fullmatch on a temporary regex, string-literal overload.
| [in] | text | NUL-terminated text. |
|
inlineconstexpr |
Region-aware fullmatch overload for string literals.
| [in] | text | NUL-terminated text. |
| [in] | pos | Byte offset the region starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
Region-aware fullmatch on a temporary regex, string-literal overload.
| [in] | text | NUL-terminated text. |
| [in] | pos | Byte offset the region starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
Match the entire text (Python re.fullmatch).
| [in] | text | The subject text (must outlive the result). |
|
inlineconstexpr |
fullmatch on a temporary regex; the result owns its name context.
| [in] | text | The subject text (must outlive the result). |
|
inlineconstexpr |
Region-aware fullmatch: the whole region [pos, endpos) must match.
| [in] | text | Subject. |
| [in] | pos | Byte offset the region starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
Region-aware fullmatch on a temporary regex; the result owns its name context.
| [in] | text | Subject. |
| [in] | pos | Byte offset the region starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
Returns the number of capturing groups (excluding group 0).
|
inlineconstexpr |
Resolves a group name to its number.
| [in] | name | The group name. |
|
inlineconstexprnoexcept |
Whether first-byte filtering is useful for this pattern.
true iff every non-empty match provably begins with a byte from a known set, so may_start_with can reject positions. false when a zero-length match is possible (or the set is empty) — then may_start_with is true for every byte and the filter buys nothing. This is the same set the engine's own prefilter uses, exposed for embedders (e.g. a lexer's rule dispatch).
true if the first-byte set is usable.
|
inlineconstexprnoexcept |
How far before a position a match there may read: an upper bound, in bytes.
match(text, pos) reads text from pos - left_context() on, and nothing earlier: a lookbehind reads back as far as it can consume, and \b, \B, \<, \>, a line start and \A read what precedes the position. So a caller that keeps only part of a text – a lexer reading it in pieces – keeps this many bytes before where it matches next, and the answer is the same as on the whole text. 0 when the pattern reads nothing before where it starts (nor whether it starts the text).
|
inlineconstexpr |
match overload for string literals.
| [in] | text | NUL-terminated text. |
|
inlineconstexpr |
match on a temporary regex, string-literal overload.
| [in] | text | NUL-terminated text. |
|
inlineconstexpr |
Region-aware match overload for string literals.
| [in] | text | NUL-terminated text. |
| [in] | pos | Byte offset the match is anchored at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
Region-aware match on a temporary regex, string-literal overload.
| [in] | text | NUL-terminated text. |
| [in] | pos | Byte offset the match is anchored at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
Match anchored at the start of text (Python re.match).
| [in] | text | The subject text (must outlive the result). |
matched() / operator bool).
|
inlineconstexpr |
match on a temporary regex; the result owns its name context.
| [in] | text | The subject text (must outlive the result). |
|
inlineconstexpr |
Region-aware match: anchored at pos within text[0:endpos] (Python re.match with pos / endpos). Byte offsets; pos is not a slice (see run — \A fails at pos > 0); endpos defaults to the end of text.
| [in] | text | Subject. |
| [in] | pos | Byte offset the match must start at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
pos.
|
inlineconstexpr |
Region-aware match on a temporary regex; the result owns its name context.
| [in] | text | Subject. |
| [in] | pos | Byte offset the match must start at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexprnoexcept |
Whether a non-empty match can begin with byte (sound, conservative).
A false result is a guarantee: no non-empty match of this pattern begins with byte. A true result is a conservative superset — it does not promise a match actually starts there. When first-byte filtering is not usable (has_first_byte_set is false, i.e. an empty match is possible), this returns true for every byte, so it is safe to use on its own.
| [in] | byte | The candidate leading byte. |
false only when byte can never start a non-empty match.
|
inlineconstexprprivate |
Returns its name, sliced from the pattern text.
| [in] | named_group | A named group. |
|
inlineconstexpr |
The index-th named group, in declaration order.
| [in] | index | Position in [0, named_group_count()); out of range is undefined, as for any indexed accessor on this class. |
|
inlineconstexpr |
How many named groups the pattern declares.
With named_group_at, the allocation-free way to enumerate them (named_groups builds a vector per call). A constant factor, not a complexity fix: resolving N names through a number-keyed interface is N scans of N either way.
|
inlineconstexpr |
All named groups as (name, number) pairs, in declaration order.
|
inlineconstexprprivate |
Reads replacement once into literal slices and group references.
An invalid or out-of-range reference is an error (Python's rule), not left in place, and is reported whatever the subject. The template spelling is ECMAScript / std::regex: $1, not \1.
| [in] | replacement | The replacement template ($$, $&, $1, ${name}). |
| real::regex_error | on a malformed or out-of-range reference. |
|
inlineconstexpr |
Returns the pattern text this regex was compiled from.
|
inlineconstexpr |
The raw compiled program, for embedders (advanced).
Lets an embedder (e.g. the Python binding) drive detail::pike_vm with caller-owned reusable scratch. Valid as long as this regex is alive.
|
inlineconstexpr |
Replaces matches in text (ECMAScript / std::regex_replace $1).
The replacement may reference groups: $$ → '$', $& or $0 → whole match, $1 …, and ${name}. This is not Python re.sub (\1 / \g<name>) — that spelling is the Python and Go bindings. Returns an owning string, so a temporary text is fine here.
| [in] | text | The subject text. |
| [in] | replacement | The replacement template. |
| [in] | max_count | Maximum replacements (0 = all). |
| real::regex_error | on a malformed or out-of-range group reference in replacement, whether or not anything matches (as Python's re.sub): the template is read once, before the walk. |
|
inlineconstexprprivate |
Runs a single match attempt from offset 0 (backs match/search/fullmatch).
| [in] | text | The subject text. |
| [in] | mode | The anchoring mode. |
|
inlineconstexprprivate |
Region-aware single attempt: match over text[0:endpos] starting at pos.
pos is the VM start offset, not a slice — zero-width assertions still see the absolute position, so \A and ^ (non-multiline) fail at pos > 0, matching Python re. endpos truncates the subject to a view (no copy), so $ / \Z treat it as the end. endpos is clamped to the text length; pos > endpos yields no match. Capture offsets are absolute byte offsets in text.
| [in] | text | The full subject (offsets are relative to it; must outlive the result). |
| [in] | pos | Byte offset to start matching at. |
| [in] | endpos | Byte offset of the exclusive region end; npos = end of text. |
| [in] | mode | The anchoring mode. |
| [in] | sem | Match semantics: leftmost-first (default) or the experimental leftmost-longest. |
text.
|
inlineprivate |
run that accepts no empty match: the leftmost position where a non-empty match starts, and there the match the leftmost-first priority prefers among the non-empty ones (match_not_null). The DFAs do not model the rule, so the VM decides the search; it stays linear. Kept apart from run, whose every caller is a hot path.
| [in] | text | The subject (must outlive the result). |
| [in] | pos | Byte offset to start at. |
| [in] | mode | Search, or anchored at pos (prefix). |
text.
|
inlineconstexpr |
search overload for string literals.
| [in] | text | NUL-terminated text. |
|
inlineconstexpr |
search on a temporary regex, string-literal overload.
| [in] | text | NUL-terminated text. |
|
inlineconstexpr |
Region-aware search overload for string literals.
| [in] | text | NUL-terminated text. |
| [in] | pos | Byte offset the search starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
Region-aware search on a temporary regex, string-literal overload.
| [in] | text | NUL-terminated text. |
| [in] | pos | Byte offset the search starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
Leftmost match anywhere in text (Python re.search).
| [in] | text | The subject text (must outlive the result). |
|
inlineconstexpr |
search on a temporary regex; the result owns its name context.
| [in] | text | The subject text (must outlive the result). |
|
inlineconstexpr |
Region-aware search: leftmost match within [pos, endpos).
| [in] | text | Subject. |
| [in] | pos | Byte offset the search starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
Region-aware search on a temporary regex; the result owns its name context.
| [in] | text | Subject. |
| [in] | pos | Byte offset the search starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inline |
search_longest overload for string literals.
| [in] | text | NUL-terminated text. |
|
inline |
search_longest on a temporary regex, string-literal overload.
Needs its own const&&: otherwise a literal on a temporary regex binds the borrowing const& overload and detaches nothing.
| [in] | text | NUL-terminated text. |
|
inline |
Region-aware search_longest overload for string literals.
| [in] | text | NUL-terminated text. |
| [in] | pos | Byte offset the search starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inline |
Region-aware search_longest on a temporary regex, string-literal overload.
| [in] | text | NUL-terminated text. |
| [in] | pos | Byte offset the search starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inline |
EXPERIMENTAL, opt-in: a single leftmost-**longest** search (POSIX / RE2 set_longest_match). Among matches at the leftmost start it returns the longest, so a lazy quantifier behaves greedily; captures are the leftmost-first thread's at that bound (not POSIX submatch). Runs on the general Pike loop. Not yet a stable API; its iteration twin is find_iter_longest.
| [in] | text | Subject. |
|
inline |
search_longest on a temporary regex; the result owns its name context.
Without it, a lookup by name would read the temporary's freed pattern and name table.
| [in] | text | The subject text (must outlive the result). |
|
inline |
Region-aware form of search_longest — leftmost-longest search within [pos, endpos). pos is the start (not a slice, per run); endpos truncates the subject. Byte offsets.
| [in] | text | Subject. |
| [in] | pos | Byte offset the search starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inline |
Region-aware search_longest on a temporary regex; the result owns its name context.
| [in] | text | Subject. |
| [in] | pos | Byte offset the search starts at. |
| [in] | endpos | Byte offset the region ends at; defaults to the end of text. |
|
inlineconstexpr |
split overload for string literals.
| [in] | text | NUL-terminated text. |
| [in] | max_splits | Max splits. |
|
inlineconstexpr |
Splits text on matches (Python re.split).
Each capturing group's text is inserted after its split (an unset group yields an empty view, where Python would use None).
| [in] | text | The subject text (must outlive the returned views). |
| [in] | max_splits | Maximum splits (0 = split everywhere). |
|
inlineconstexprnoexcept |
The single byte every non-empty match must begin with, if unique.
if / def); std::nullopt for zero or several.