REAL
Regular Expression Algorithmic Library — constexpr C++20 regex
Loading...
Searching...
No Matches
real::detail::pike_vm< State, StateBoundToProgram > Class Template Reference

The Pike VM, generic over the scratch-state container policy. More...

#include <pike.hpp>

Collaboration diagram for real::detail::pike_vm< State, StateBoundToProgram >:
[legend]

Classes

struct  alternation_hit
 A match the pair scan found (start == npos: none), and where the block scans stopped. More...
 
struct  cp_hi_cache_entry
 Cache entry for cp_hi_cached (thread-local, not on basic_pike_state). Keyed by a content fingerprint, never a pointer into a program: programs die while the cache lives, and a recycled cp_ranges address would return another class's table (false membership, e.g. emoji matching [\w€]). More...
 
struct  cp_span
 One buffered class-loop match: its whole-match span (fill_span_slots mirrors a wrap). More...
 
struct  scan_set_reset
 Clears scan_set_ when an inner-literal scan returns, before its lease ends. More...
 
struct  slot_pair
 A two-slot sink, for a filler that must call a route function expecting a slot container. More...
 

Public Member Functions

constexpr pike_vm (const program_view &prog, State &state)
 Binds the VM to a program and caller-owned scratch state.
 
template<bool Cascade = false, typename OutSlots >
constexpr bool run (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots, std::size_t forbid_empty_until=0, match_semantics sem=match_semantics::first)
 Runs the VM over text starting at start.
 
template<bool Cascade = false, bool Probe = false, typename OutSlots >
constexpr bool run_general (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots, std::size_t *forward_stop=nullptr)
 The general Pike VM loop, also run by the lazy-DFA route on the [s, e] window its two passes located.
 
template<typename OutSlots >
bool confirm_at (std::string_view text, std::size_t s, OutSlots &out_slots, std::size_t &stop)
 Confirm a match anchored at s: the forward DFA finds its end and the one-pass table fills the captures, as on the lazy-DFA route. Falls back to the anchored Pike when the pattern is not DFA/one-pass eligible, or when the DFA's leftmost match does not begin at s.
 
constexpr std::size_t find_on_subject (std::string_view text, std::size_t pos, std::string_view lit, std::size_t rare, bool inner) const
 The next occurrence of a literal of the pattern's hints at or after pos, by find_literal_adaptive with a density kept for the whole subject.
 
constexpr void il_reset_on_new_haystack (std::string_view text)
 Re-enables the inner-literal route and clears its density counters on a new haystack.
 
template<typename OutSlots >
bool run_inner_literal (std::string_view text, std::size_t start, OutSlots &out_slots, bool &abandon, bool density_gate=true)
 The inner-literal search: memmem a required literal, reverse-match the prefix to the match start, forward-confirm — the reverse-inner protocol (regex-automata's ReverseInner).
 
lookaround_scratch & lookaround_state ()
 The lookaround sub-scratch, built on first use.
 
template<bool Cascade, typename OutSlots >
bool run_class_loop_trailing_la (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 Trailing-lookaround class+: body scan + longest end where lookaround holds.
 
template<bool Cascade, bool Cp, typename OutSlots >
bool trailing_la_walk (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 The body of run_class_loop_trailing_la for one body kind.
 
template<bool Cascade, bool WbEdge, bool WbKept>
constexpr std::size_t fill_class_spans (std::string_view text, std::size_t start, cp_span *out, std::size_t cap)
 Fills up to cap class_loop matches from start without leaving the route.
 
constexpr std::size_t fill_single_class_spans (std::string_view text, std::size_t start, cp_span *out, std::size_t cap)
 Fills up to cap bare single byte-class matches from start without leaving the route.
 
template<bool WbEdge>
constexpr std::size_t fill_cp_class_spans (std::string_view text, std::size_t start, cp_span *out, std::size_t cap)
 Fills up to cap cp_class_loop matches from start without leaving the route.
 
template<bool WbEdge>
constexpr std::size_t fill_cp_class_spans_wrapped (std::string_view text, std::size_t start, cp_span *out, std::size_t cap)
 fill_cp_class_spans for a pattern with a kept \b/\B wrap: its spans, less those whose wrap does not hold.
 
template<typename OutSlots >
constexpr void write_cp_span_slots (OutSlots &out_slots, std::size_t s, std::size_t e)
 Writes a buffered span into a caller's slots exactly as the per-match path would.
 
template<typename OutSlots >
constexpr bool run_cp_class_loop (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 Fast path for a whole-pattern code-point class klass_cp, optionally a greedy +.
 
template<typename OutSlots , typename InClass , typename ScanEnd , typename LastWidth >
constexpr bool run_possessive_loop_generic (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots, const InClass &in_class, const ScanEnd &scan_end, const LastWidth &last_width)
 Shared driver: a possessive class+/++ loop, bare/suffixed (pattern_hints::possessive_prefix_size == 0) or delimited/"quoted" (non-zero); the body's membership comes from in_class / scan_end, so it serves byte- and code-point classes.
 
template<typename OutSlots >
constexpr bool run_possessive_byte_loop (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 Possessive literal-byte +/++ loop (byte_loop_possessive, e.g. a++), on the shared algorithm of run_possessive_loop_generic.
 
template<typename OutSlots >
constexpr bool run_possessive_class_loop (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 Possessive class+/++ loop over a BYTE class (klass_loop_possessive). See run_possessive_loop_generic for the shared algorithm.
 
template<typename OutSlots >
constexpr bool run_possessive_cp_class_loop (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 Possessive class+/++ loop over a CODE-POINT class (klass_cp_loop_possessive), on the decode/membership primitives of run_cp_class_loop (the scan predicate differs per compiler, see in_class). See run_possessive_loop_generic for the shared algorithm.
 
template<bool SkipSaves = false>
constexpr std::size_t match_byte_klass_run (std::string_view text, std::size_t pc, std::size_t s) const
 Matches the run of byte/klass instructions starting at pc, one text byte each, up to the first non-consuming op. Shared by the fixed-shape and alternation fast paths.
 
constexpr bool wb_boundaries_ok (std::size_t s, std::size_t e) const
 O(1) lead/trail \b/\B check at match bounds [s, e), for every wb-wrapping fast path (hints 0/1/2 from pattern_hints::wb_lead / pattern_hints::wb_trail).
 
template<bool SkipSaves>
constexpr std::size_t match_fixed_body_wb (std::string_view text, std::size_t s) const
 Fixed-shape body match from pattern_hints::body_pc, then B1 \b/\B wrap.
 
template<typename MatchAt , typename OutSlots >
constexpr bool fast_search (std::string_view text, std::size_t start, MatchAt match_at, OutSlots &out_slots)
 Leftmost search over the candidate starts of next_candidate, for the first one match_at accepts. Shared by the fast paths that verify a fixed shape at a position.
 
template<typename OutSlots >
bool run_pair_filtered_shape (std::string_view text, std::size_t start, OutSlots &out_slots)
 Search route for a HETEROGENEOUS fixed shape: vector-prefilter two positions, verify each survivor with the ordinary fixed-body walk, hand the sub-block tail to fast_search.
 
template<typename OutSlots >
constexpr bool run_fixed_shape (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 Fast path for a whole-pattern fixed-width byte/klass sequence.
 
template<typename OutSlots >
constexpr void fill_fixed_saves (std::size_t match_start, OutSlots &out_slots) const
 Fills the capturing-group slots of a fixed-shape match. Every consuming op is one byte wide, so each save sits at a constant offset from the match start: one linear pass, no re-match.
 
constexpr bool run_shape_atom (std::string_view text, std::size_t pc, std::size_t &at, std::size_t e) const
 Consumes the atom at pc at at, within e.
 
template<typename OutSlots >
constexpr bool match_run_shape (std::string_view text, std::size_t s, std::size_t e, OutSlots &out_slots) const
 Fills the groups of a match the DFAs found at [s, e) for a program of run shape, by one walk that takes every loop as far as it goes.
 
template<typename OutSlots >
constexpr std::size_t match_cp_shape (std::string_view text, std::size_t s, OutSlots &out_slots) const
 Verifies a fixed code-point shape forward from s, filling capture slots as it goes.
 
template<typename OutSlots >
constexpr bool fail_slots (OutSlots &out_slots, std::size_t count) const
 Clears the capture slots for a search that found nothing, and says so.
 
template<typename OutSlots >
constexpr bool fail_slots (OutSlots &out_slots) const
 fail_slots for every slot of the program.
 
constexpr std::size_t prefix_run_end (std::string_view text, std::size_t from, std::size_t limit) const
 Where the two-run shape's prefix class run, continued forward from from, stops.
 
template<typename OutSlots >
constexpr void fill_two_run_saves (std::string_view text, std::size_t s, std::size_t h, std::size_t lit_end, std::size_t e, OutSlots &out_slots) const
 Fills capture slots for a class+ <literal> class+ match, by anchor rather than by offset.
 
template<bool Cascade>
constexpr std::size_t fill_codepoint_class_spans (std::string_view text, std::size_t start, cp_span *out, std::size_t cap)
 Batched twin of run_codepoint_class, filling up to cap maximal spans in ONE call.
 
template<bool Cascade, typename OutSlots >
constexpr bool run_codepoint_class (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 Fast path for . / a negated class, optionally a greedy +.
 
template<typename OutSlots >
bool run_aho_corasick (std::string_view text, std::size_t start, OutSlots &out_slots)
 Multi-literal search via the automaton ac_ready hands back, cached per regex in detail::regex_immutables.
 
alternation_density & alternation_density_for (std::string_view text) const
 This subject's alternation density, judged anew when the state last sampled another subject.
 
const alternation_density * alternation_density_seen (std::string_view text) const
 The alternation density the state holds for this subject, without sampling it.
 
const alternation_pairs * alternation_plan (std::string_view text, std::size_t pos, std::array< std::uint8_t, 8 > mem, std::size_t cnt) const
 The alternation's probe pairs when this subject's first bytes are dense, for the block scans of run_alternation and fill_alternation_spans, else null (a compile-time storage keeps no state for them, or no plan fits).
 
bool alternation_filter_takes (std::string_view text, std::size_t start) const
 Whether run_alternation will mask this subject's blocks by its pairs or fingerprint: a dense subject, where the Aho-Corasick gate (calibrated against the first-byte scan) would pick the automaton. The filtered scan beats the automaton there.
 
const alternation_pairs * alternation_plan_decide (std::string_view text, std::size_t pos, std::array< std::uint8_t, 8 > mem, std::size_t cnt) const
 Out of line, the half of alternation_plan that resets the density on a new subject, samples it, and builds the plan.
 
bool alternation_wide_may_take (std::string_view text) const
 Whether run_alternation_wide may take this search: an alternation with more first bytes than the small set holds and no single first byte, within the fingerprint plan's branches, on a subject its sample has not refused. The cheap half, inline, so that a refused subject costs no call.
 
template<typename OutSlots >
std::optional< bool > run_alternation_wide (std::string_view text, std::size_t start, OutSlots &out_slots, cp_span *spans=nullptr, std::size_t cap=0, std::size_t *filled=nullptr)
 Search route for an alternation of literals with more first bytes than the small set holds: the blocks the nibble fingerprint marks, verified in branch order (priority unchanged), then the last bytes by the first-byte table. Taken per subject on a sample: where false candidates are dense enough that verifying them costs more than the automaton's walk, it declines to the automaton's gate. Out of line: run is shared by every route.
 
const alternation_pairs * variant_plan (std::string_view text, std::size_t start)
 The variants' fingerprint for a program that is not a fixed alternation (branches holding a case-folded i, s or k, folding to non-ASCII), when this subject's sample finds its first bytes dense and the fingerprint's candidates among them sparse. Taken only where next_candidate would scan by the first bytes; decided once per subject. The fingerprint admits every start a match can have and a few more, which the confirming anchored walk rejects. No candidate lands inside a code point: a continuation byte's high nibble is no first byte's.
 
std::size_t variant_candidate (std::string_view text, std::size_t pos, const alternation_pairs &plan) const
 The next start at or after pos the variants' fingerprint admits (the first bytes, in the last blocks it cannot read past).
 
constexpr void add_branch_nibbles (alternation_pairs &plan, std::size_t branch) const
 Adds the branch at branch to plan's nibble fingerprint, in bucket plan.count % 8: each of its first three positions admits its byte, or every member of its class; a position past the branch's end admits anything (only that bucket loses selectivity). A branch one byte wide leaves the plan without a fingerprint: its bucket would mark every start.
 
bool add_variant_nibbles (alternation_pairs &plan, std::size_t branch) const
 Adds the branch at branch to plan's fingerprint, following each byte offset 0 to 2 a path through it can reach: a byte or class admits its members at every offset reached and moves one on; a code-point class admits its ASCII members, and the UTF-8 encoding of each other member laid from each offset reached, and moves on by every width it has; where the straight line ends, every byte is admitted from each offset reached on. A superset of the bytes a match can start with.
 
alternation_pairs build_cp_alternation_plan () const
 The fingerprint of a program laid out as an alternation of straight-line branches (after leading position assertions, which only narrow where a match starts) that is not a fixed alternation: one of its branches holds a code-point class. No pairs: those need a byte at every head.
 
alternation_pairs build_alternation_pairs () const
 Each branch's first byte and its farthest byte within 15 of it (the pairs, when every branch opens on a byte), and the branches' nibble fingerprint, read from the split chain in source order as the scans' match_at reads it. A branch that opens on a class leaves the plan without pairs; the fingerprint then carries it alone, whatever its minimum of branches against the pairs.
 
template<typename OutSlots >
constexpr bool run_alternation (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 Fast path for an alternation of straight-line branches.
 
std::size_t fill_alternation_wide_spans (std::string_view text, std::size_t start, cp_span *out, std::size_t cap, bool &partial, bool &disarm)
 Fills up to cap matches of an alternation with more first bytes than the small set holds, from start, by the scan of run_alternation_wide run from each match's end, without re-entering run().
 
bool alternation_automaton_claims (std::string_view text, std::size_t start)
 Whether run()'s cascade would hand this subject's search at start to the Aho-Corasick automaton rather than to run_alternation: the gate's conditions in the same order, on the same per-subject state, whose verdicts are sticky, so that a batched walk (fill_alternation_spans) asks once and never overrules a routing decision that was measured. The fingerprint run() tries first needs more first bytes than the small set holds, which that walk's eligibility excludes.
 
constexpr std::size_t fill_alternation_spans (std::string_view text, std::size_t start, cp_span *out, std::size_t cap)
 Fills up to cap fixed_alternation matches from start without leaving the route.
 
std::size_t fill_fixed_shape_spans (std::string_view text, std::size_t start, cp_span *out, std::size_t cap)
 Fills up to cap fixed-shape matches from start without re-entering the route gate.
 
std::size_t fill_exact_literal_spans (std::string_view text, std::size_t start, cp_span *out, std::size_t cap)
 Fills up to cap exact-literal matches from start without re-entering the route gate.
 
std::size_t fill_inner_literal_spans (std::string_view text, std::size_t start, cp_span *out, std::size_t cap, bool &partial, bool &disarm)
 Fills up to cap inner-literal matches from start without re-entering the route gate.
 
std::size_t fill_lazy_dfa_spans (std::string_view text, std::size_t start, cp_span *out, std::size_t cap, bool &partial)
 Fills up to cap lazy-DFA matches from start without re-entering the route gate.
 
constexpr bool literal_at (std::string_view text, std::size_t cand, std::size_t len) const
 Tests whether the fixed literal prefix occurs at cand.
 
template<typename OutSlots >
constexpr bool replay_literal (std::size_t cand, std::size_t len, OutSlots &out_slots) const
 Fills capture slots for a literal match at cand: replays save instructions at their consumed offsets and checks the chain's zero-width assertions there.
 
template<typename OutSlots >
bool run_literal_one_search (std::string_view text, std::size_t start, std::size_t len, OutSlots &out_slots)
 The whole exact-literal search in one find_prefix, for a pattern_hints::literal_one_search program (called from run_exact_literal).
 
template<typename OutSlots >
constexpr bool run_exact_literal (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 Fast path for a pure-literal pattern.
 
constexpr std::size_t next_candidate (std::string_view text, std::size_t pos, std::size_t start) const
 First position >= pos that could start a match, per the hints: the prefilter step (literal prefix, rare or unique byte, line start, first-byte set); pos itself when nothing skips.
 
constexpr bool seed_viable (std::string_view text, std::size_t pos, std::size_t start) const
 Cheap pre-check before seeding a new thread at pos: live threads may drag the loop through positions the prefilter would skip. Also enforces code-point alignment in text mode.
 
constexpr bool assertion_holds (assert_kind kind, std::size_t pos, bool word_ness_flipped) const
 Evaluates a zero-width assertion at pos in the current text.
 
template<typename OutSlots >
bool run_bounded_backtrack (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 The general loop's answer, by backtracking under a bit per (instruction, position).
 
template<typename OutSlots >
bool backtrack_from (backtrack_frame &frame, std::size_t seed, std::size_t start, run_mode mode, bool cf, OutSlots &out_slots)
 Every branch from pc 0 at seed, in priority order – one start of run_bounded_backtrack.
 
std::int32_t backtrack_possessive (backtrack_frame &frame, const instr &instruction, std::int32_t pc, std::size_t &pos, bool cf)
 A possessive loop's step in backtrack_from. The VM decides it when the thread arrives, so a match consumes (writing the loop's capture, if it has one) and a miss leaves by the exit at the same position.
 
template<typename OutSlots >
bool extends_past_end (std::string_view text, std::size_t start, OutSlots &out_slots)
 Whether a match anchored at start could come out differently if text continued past its end: what a caller lexing text that arrives in pieces must know before committing a token.
 
constexpr bool cut_short (std::size_t pos) const
 Whether the code point at pos is not all there: past the end of the text, or a sequence the end cuts short – what a class test or a word boundary at pos would read more text to decide.
 
constexpr void probe_step (const instr &instruction, std::size_t pos)
 Probe of extends_past_end on a thread about to consume at pos: one alive at the end of the text, or at a code point the end cuts short, would read what comes next.
 
constexpr bool reads_right (const lookaround_sub &sub) const
 Whether sub holds an assertion that reads what follows where it stands: $, \Z, \z, \b, \B, \<, \>.
 
constexpr void probe_closure (const instr &instruction, std::size_t pos)
 Probe of extends_past_end on an epsilon step at pos: an assertion that looks right, a lookahead, or a possessive test whose answer the end of the text decides.
 
template<bool Probe = false, typename OutSlots >
constexpr void step (list_type &clist, list_type &nlist, std::size_t pos, run_mode mode, bool &matched, OutSlots &out_slots)
 Advances every thread of clist by the byte at pos; survivors land in nlist. A thread reaching match records its slots and cuts all lower-priority threads (leftmost-greedy order).
 
constexpr void tier1_capture_on_match (list_type &clist, std::size_t i, std::int32_t capture_start_slot, std::size_t start, std::size_t end)
 Tier 1's on-match capture write: if capture_start_slot is not -1, records [start, end) into thread i's capture block, in place.
 
template<bool Probe = false>
constexpr void advance_thread (list_type &clist, list_type &nlist, std::size_t i, std::int32_t next_pc, std::size_t next_pos)
 Advances thread i of clist by one consumed byte, seeding its continuation's closure into nlist with its own reference on the thread's capture block (shared until a save copies it).
 
constexpr const std::size_t * thread_slots (list_type &clist, std::size_t i)
 Pointer to thread i's slot_count capture values (its COW block), read by match.
 
constexpr bool cp_class_matches (const detail::cp_class &cc, char32_t cp) const
 Tests a decoded code point against a klass_cp class: ASCII bitmap below 0x80, binary search of the class's range slice above (constexpr / const paths; cp_class_matches_idx uses the cached tables). The class is already the effective set: a plain positive membership test.
 
constexpr bool cp_class_matches_idx (std::size_t cp_index, char32_t cp)
 Membership by class index (ASCII + European page + sparse hi / bsearch).
 
template<bool Probe = false>
constexpr void add_thread (list_type &list, std::int32_t pc0, std::size_t pos, std::size_t initial)
 Adds pc0 and its whole epsilon closure to list — the one closure walk (COW). Each DFS frame carries a capture-block index (in eps_entry::block) rather than mutating a shared working array, so capture state is copy-on-write and there are no slot-restore entries:
 
constexpr void cow_release_blocks (list_type &list)
 Releases the block references a list's threads hold, before the list is reset or the run returns: the one decref site paired with each step→closure incref (keep it single).
 
constexpr bool lookaround_holds (std::uint16_t sub_id, std::size_t pos)
 Evaluates a bounded lookaround at pos (true if the thread should proceed).
 
constexpr bool single_class_ahead (const instr &body, std::size_t pos)
 L1 peephole — does the single consuming op body match the code point / byte AT pos (ahead)? Mirrors the per-op logic of lookahead_matches for a one-instruction sub-program.
 
constexpr const instr * single_atom_body (const lookaround_sub &sub) const
 The one consuming instruction of a lookaround body that is a single atom, or null.
 
constexpr bool single_class_behind (const instr &body, std::size_t pos)
 L1 peephole — does body match the code point / byte ending EXACTLY at pos (behind)? The defining lookbehind trap: the match must END at pos, so the code point is the one whose aligned start s gives s + length == pos (byte mode: pos - 1).
 
constexpr bool lookahead_matches (const lookaround_sub &sub, std::size_t pos)
 Lookahead: does the sub-pattern match a prefix starting at pos? A forward Pike simulation bounded to l_max bytes, stopping at the first match (capture-free: any match is a witness).
 
constexpr bool unbounded_lookahead_matches (std::uint16_t sub_id, const lookaround_sub &sub, std::size_t pos)
 Unbounded lookahead: does the sub-pattern match a prefix of the text from pos?
 
constexpr void fill_ahead_table (const lookaround_sub &sub, lookaround_scratch::ahead_table &table)
 Fills table with every position's answer, for unbounded_lookahead_matches to read.
 
constexpr bool lookbehind_matches (std::uint16_t sub_id, const lookaround_sub &sub, std::size_t pos)
 Lookbehind: does the sub-pattern match a window ENDING EXACTLY at pos?
 
constexpr void sub_add_thread (thread_list &list, std::int32_t pc0, std::size_t pos, bool &matched)
 Epsilon-closure for the lookaround sub-VM, on the isolated sub-scratch.
 

Static Public Member Functions

static constexpr std::size_t run_shape_loop_atom (const program_view &prog, std::size_t begin, std::size_t end)
 The one atom of a loop body [begin, end) that holds nothing else but saves – a group around one atom, repeated: ([aeiou])+.
 
static constexpr bool is_run_shape (const program_view &prog)
 Whether prog is saves, atoms (a byte, a byte class, a code-point class) and greedy atom+ and atom* loops, then match: nothing else, no alternation, no lazy loop, no assertion.
 
static std::uint8_t nibble3_buckets (const alternation_pairs &plan, std::string_view text, std::size_t at)
 The fingerprint buckets the three bytes at at admit, as the vector scans compute them for a block: a bucket's bit survives where both nibbles of each byte carry it. Every bucket where fewer than three bytes remain, as the scans leave those starts to the first-byte table.
 
static constexpr bool no_class_loop_above (const pattern_hints &hints) noexcept
 Whether no class loop takes the pattern first: the byte-class loop, the code-point one and the three possessive loops sit above every literal, shape and alternation route in the cascade.
 
static constexpr bool exact_literal_is_the_route (const pattern_hints &hints) noexcept
 Is the exact-literal route the one run() would take, in its one-search subset?
 
static constexpr bool inner_literal_is_the_route (const program_view &prog) noexcept
 Is the inner-literal route the one run() would take for this program?
 
static constexpr bool fixed_shape_is_the_route (const program_view &prog) noexcept
 Is the fixed-shape route the one run() would take for this program, in a groupless search?
 
static constexpr bool lazy_dfa_is_the_route (const pattern_hints &hints) noexcept
 Is the lazy-DFA route the one run() would actually take for this program?
 
static constexpr bool window_cut_before (std::string_view text, std::size_t start, std::size_t candidate)
 Whether no whole code point lies between start and candidate: the candidate IS start, or only continuation bytes separate them.
 

Static Public Attributes

static constexpr std::size_t lazy_dfa_min_input {512}
 Below this input length the lazy-DFA route is skipped (the two-pass setup does not amortise). Public: real::basic_match_iterator reads it to decide whether to batch this route.
 

Private Types

using list_type = std::remove_reference_t< decltype(std::declval< State & >().list_a)>
 The concrete thread-list type taken from the bound State.
 
using pool_type = std::remove_reference_t< decltype(std::declval< State & >().pool)>
 The capture-block pool type of the bound State (COW): heap-backed or static_vec.
 

Private Member Functions

bool ac_candidate_completes (std::string_view text, std::size_t at)
 Does a branch of the alternation COMPLETE at at?
 
template<typename Dummy = void>
bool ac_density_favours_automaton (std::string_view text, std::size_t start)
 Decides ONCE PER HAYSTACK whether the Aho-Corasick automaton should take this alternation's searches, by sampling candidate density at the search start.
 
void ensure_immutables ()
 Build (or rebuild) the per-regex immutables, race-free: the Tier-A byte program the DFAs run over and its alphabet. Invalidated by program identity (regex_immutables::built_for == prog_.code.data()); the hot path is one atomic load. Not a once_flag: a spent one never rebuilds after assign-onto-warmed (silent 0 matches). Needs no DFA, so the anchored path can call it without the DFA build.
 
void ensure_op_table ()
 Build (or rebuild) the one-pass capture extractor, on top of ensure_immutables.
 
void ensure_set_search_dfas (detail::regex_immutables &immut, shared_dfa_set &set)
 Builds the search DFAs for immut into this thread's leased set, once.
 
void ensure_set_il_prefix_rev (detail::regex_immutables &immut, shared_dfa_set &set)
 Builds the IL-prefix reverse DFA for immut into this thread's leased set, once.
 
template<typename Fn >
bool with_search_dfas (Fn &&fn)
 Run fn with this thread's search DFAs for the regex (see dfa_lease), taking no lock.
 
template<typename Walk >
bool walk_on_scan_set (const Walk &walk)
 The forward walk of with_search_dfas, on the set an inner-literal scan already leased.
 
template<bool Cascade, typename OutSlots >
std::optional< bool > try_shared_lazy_dfa_search (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 Lazy-DFA search route on the shared confirm DFAs. noinline: inlined, its body inflates the x86 class-loop codegen of run (as ac_ready).
 
template<typename Dummy = void>
const ac_automaton * ac_ready ()
 Build (or rebuild, on a program change) the Aho-Corasick automaton for a fixed_alternation program whose branch count has reached ac_branch_threshold.
 
const alternation_pairs * alternation_pairs_ready () const
 The regex's alternation probe pairs, built once per regex in its immutables, or null when there is no per-regex cache: the caller then scans by the first bytes. Same identity discipline as ac_ready.
 
constexpr bool row_key_stale (std::int32_t have, std::int32_t want) const
 Is the state's cached row key stale for want?
 
void verify_class_row (detail::regex_immutables &cache, std::size_t class_index)
 Verifies (and if needed fills) the byte row for class_index, then caches it in the state.
 
constexpr const std::uint8_t * derive_class_table (std::size_t class_index)
 Derives the byte row into the VM state: the constant-evaluation path, where no per-regex cache exists.
 
void ensure_membership_rows (detail::regex_immutables &cache) const
 Sizes the per-regex membership rows for this program, if not already. Cold: once per regex, behind an acquire load on the hot path.
 
void fill_class_row (detail::regex_immutables &cache, std::size_t class_index) const
 Fills one byte-class row of the per-regex cache, once.
 
void fill_cp_ascii_row (detail::regex_immutables &cache, std::size_t cp_index) const
 Fills one code-point-class ASCII row of the per-regex cache, once.
 
void fill_cp_page_row (detail::regex_immutables &cache, std::size_t cp_index) const
 Fills one U+0080..U+07FF membership bitmap of the per-regex cache, once.
 
constexpr const std::uint8_t * class_table (std::size_t class_index)
 Returns a flat 256-byte membership table for class class_index (one load per byte).
 
template<bool Cascade, typename OutSlots >
constexpr bool run_class_loop_anchored (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 Cold half of the class-loop route: everything a \A/^ or \Z/$ implies.
 
template<typename OutSlots >
constexpr bool run_class_loop_end_anchored (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 X+$ / ^X+$ in search mode: the run that ENDS at the anchor, found by walking back.
 
constexpr const std::uint8_t * resolve_class_table (std::size_t class_index)
 Cold half of class_table (the storage-mode resolution), and the only path that writes the state's row cache for a byte class.
 
constexpr const std::uint8_t * cp_ascii_table (std::size_t cp_index)
 Byte-indexed membership table for the ASCII bitmap of a cp_class, as class_table for the klass_cp scan loop, keyed negatively so it never collides with a byte class.
 
constexpr bool cp_class_holds (const cp_class &cc, char32_t cp) const
 Stateless membership of cp in cc: no VM-state cache touched.
 
constexpr const std::uint64_t * cp_page_table (std::size_t cp_index)
 Builds (once, cached) the cp_class's membership bitmap over [U+0080, U+07FF]: one load instead of a range search on two-byte code points (see basic_pike_state::cp_page).
 
constexpr bool cp_member_page (std::size_t cp_index, char32_t cp)
 Page-bitmap membership for U+0080..U+07FF. Kept separate so class-loop lambdas can inline it without pulling the sparse-hi path into the European hot stream (\p{N}, accented).
 
constexpr bool cp_member_high (std::size_t cp_index, char32_t cp)
 Membership for cp > U+07FF: sparse 2-stage hi table, else bsearch (small classes / constexpr). The table build is cold-outlined; the per-cp probe is last-hit + bit test.
 
bool cp_member_high_unshared (std::size_t cp_index, char32_t cp)
 cp_member_high for the trailing-lookaround walk, written out rather than called.
 
void resolve_hi (std::size_t cp_index)
 Fills the state's sparse-hi memo for cp_index, the cold half of cp_member_high, outlined so the per-code-point path stays a class-key compare and a bit test.
 
constexpr bool cp_member_hi (std::size_t cp_index, char32_t cp)
 Non-ASCII membership: European page bitmap, then sparse 2-stage hi / bsearch.
 
template<typename OutSlots >
constexpr void fill_span_slots (OutSlots &out_slots, std::size_t match_start, std::size_t match_end) const
 Writes a class-loop fast-path result into out_slots: the whole-match span in slots 0/1, mirrored into the group's slots for a pattern wrapped in one capturing group ((\w+), ([a-z]+)): the group's span equals the whole match by construction.
 
constexpr std::size_t run_cascade_stop (std::string_view text, std::size_t from) const
 The memchr-cascade run tail: the next stop byte at or after from, or the text end. Its own function so the cascade never inlines into the per-byte loop of run_class_loop (that bloat slowed stop-dense short runs); reached only past cascade_run_threshold bytes.
 
template<bool Cascade>
constexpr std::size_t class_run_end (std::string_view text, const std::uint8_t *tbl, std::size_t match_start) const
 The end of the maximal run of tbl's members that begins with the member at match_start.
 
template<bool Cascade, typename OutSlots >
constexpr bool run_class_loop (std::string_view text, std::size_t start, run_mode mode, OutSlots &out_slots)
 Fast path for a whole-pattern "class+": a maximal run of class bytes in one scan loop, exactly the VM's greedy result, with no thread lists.
 

Static Private Member Functions

static constexpr run_mode window_mode (run_mode mode) noexcept
 The mode the VM fills a window's groups in once the DFAs proved where the match starts.
 
static void fill_byte_row (detail::regex_immutables &cache, std::size_t ready_index, const char_class &klass, std::uint8_t *row)
 Expands klass into one flat 256-byte membership row of the per-regex cache, once.
 
static const cp_hi_table * cp_hi_build (const program_view &prog, std::size_t cp_index, std::uint64_t key_fp, std::array< cp_hi_cache_entry, 8 > &cache, const cp_hi_table *&last_tab, std::uint64_t &last_fp)
 Cold path: build a sparse hi table and install it in the thread-local cache. Outlined so the hot membership check never inlines the range-walk builder.
 
static const cp_hi_table * cp_hi_cached (const program_view &prog, std::size_t cp_index)
 Thread-local sparse hi tables, keyed by cp_class::fingerprint (set once at intern), so basic_pike_state keeps its size. Hot path: a uint64 load and a sticky compare.
 
template<typename OutSlots >
static constexpr void ensure_slot_size (OutSlots &out, std::size_t n)
 Size out without a full npos fill when already sized: ensure_size, or a grow-only resize for the seam tests' std::vector.
 

Private Attributes

const program_view & prog_
 The program being executed (borrowed; a stable lvalue that outlives the VM).
 
State & state_
 Borrowed reusable scratch state.
 
shared_dfa_set * scan_set_ {nullptr}
 The set an inner-literal scan leased for its prefix reverse, or null.
 
std::string_view text_
 The subject text for the current run.
 
std::size_t forbid_empty_until_ {}
 Reject an empty match starting below this offset (CPython 3.7+ rule); 0 = none.
 
match_semantics sem_ {match_semantics::first}
 Match semantics for the current run; match_semantics::longest forces the general loop.
 
bool extends_ {false}
 Set by a probing run (extends_past_end) when more text could change the answer.
 

Static Private Attributes

static constexpr std::uint32_t il_density_probe_candidates {8}
 Inner-literal density gate: candidates sampled across the haystack before the verdict.
 
static constexpr std::size_t il_density_milli_threshold {60}
 Candidate density, in candidates per 1000 bytes, at or above which the IL route yields to the DFA.
 
static constexpr std::size_t ac_density_sample_bytes {256}
 AC routing: sample window, and the candidate-work product at or above which the automaton beats the memchr cascade.
 
static constexpr std::size_t ac_density_min_span {64}
 Shortest span an early verdict may rest on.
 
static constexpr std::size_t ac_density_work_threshold {550}
 Product (candidates per 1000 bytes) * branch_count at or above which the automaton wins, at or above ac_branch_threshold branches.
 
static constexpr std::uint16_t ac_branch_floor {4}
 Fewest branches the automaton is ever considered for; below this nothing is measured.
 
static constexpr std::size_t ac_completion_pct {15}
 Percentage of sampled candidates that may COMPLETE a branch and still leave the automaton ahead. Above it the cascade wins whatever the candidate density says.
 
static constexpr std::size_t ac_density_work_threshold_low {1400}
 The same product for ac_branch_floor .. ac_branch_threshold branches, where the safe direction is reversed.
 
static constexpr std::size_t ac_completion_walk_budget {1024}
 Branch walks the completion half of the sample may spend per decision: each verified candidate walks every branch (a thousand-word alternation spent ~2 M instructions deciding a ~4 k search). Past it, candidates count for density only; twelve branches still verify 85.
 
static constexpr std::uint16_t ac_branch_threshold {12}
 Branch count of a pattern_hints::fixed_alternation at or above which one Aho-Corasick walk beats the memchr cascade of pattern_hints::small_set (none past 8 first bytes). Just below, the automaton loses on prose; 12, not 11: AC must beat the VM-branch path here.
 
static constexpr std::uint32_t cp_page_max {0x7FFU}
 Highest code point covered by the cp_page bitmap (the 2-byte UTF-8 range).
 
static constexpr std::size_t cascade_run_threshold {32}
 Accepted-byte count after which a class+ run switches from the per-byte advance to a memchr-cascade to the next stop byte, so stop-dense short runs stay at baseline.
 
static constexpr std::uint32_t cp_hi_range_threshold {20U}
 Below this many total ranges, high-cp membership stays on bsearch (small scripts). The classes straddling the crossover pull opposite ways on both ISAs; this value sits in that gap, and moving it past either costs that class.
 

Detailed Description

template<typename State, bool StateBoundToProgram = false>
class real::detail::pike_vm< State, StateBoundToProgram >

The Pike VM, generic over the scratch-state container policy.

Template Parameters
StateA basic_pike_state instantiation (vector- or static-backed).
StateBoundToProgramThe caller guarantees this state never serves a second program (fresh per search, or owned by a walk over one regex), so the membership-row accessors skip the per-run() program-identity compare (2.9 points of a [a-z]+ walk). Defaults to false: an embedder holding a state across regexes (Python binding, meta-seam harness) needs it.

Constructor & Destructor Documentation

◆ pike_vm()

template<typename State , bool StateBoundToProgram = false>
constexpr real::detail::pike_vm< State, StateBoundToProgram >::pike_vm ( const program_view &  prog,
State &  state 
)
inlineconstexpr

Binds the VM to a program and caller-owned scratch state.

Parameters
[in]progThe compiled program to execute.
[in,out]stateReusable scratch (borrowed; must outlive the VM).

Member Function Documentation

◆ ac_candidate_completes()

template<typename State , bool StateBoundToProgram = false>
bool real::detail::pike_vm< State, StateBoundToProgram >::ac_candidate_completes ( std::string_view  text,
std::size_t  at 
)
inlineprivate

Does a branch of the alternation COMPLETE at at?

Walks the split chain in source order, asking match_byte_klass_run per branch, as run_alternation and fill_alternation_spans do: no thread lists, no capture work. A copy, not a shared call: relocating those hot bodies risks a regression costing more than this gate wins. The copies agree because they ask the same primitive in the same order.

Parameters
[in]textThe subject.
[in]atA candidate position (a branch head byte occurs there).
Returns
True when some branch matches at at.

◆ ac_density_favours_automaton()

template<typename State , bool StateBoundToProgram = false>
template<typename Dummy = void>
bool real::detail::pike_vm< State, StateBoundToProgram >::ac_density_favours_automaton ( std::string_view  text,
std::size_t  start 
)
inlineprivate

Decides ONCE PER HAYSTACK whether the Aho-Corasick automaton should take this alternation's searches, by sampling candidate density at the search start.

Sticky per subject data pointer (as pike_state::il_abandoned), since find_iter re-enters search() per match, and short per-match searches would never see the density (and would pay the sample each time). Reuses next_candidate, so "candidate" means what it means to the cascade.

Parameters
[in]textSubject.
[in]startWhere this search begins; the window is measured from here.
Returns
true when the automaton should take over. Storages without the guard fields (the compile-time scratch) answer true unconditionally, preserving their behaviour.

◆ ac_ready()

template<typename State , bool StateBoundToProgram = false>
template<typename Dummy = void>
const ac_automaton * real::detail::pike_vm< State, StateBoundToProgram >::ac_ready ( )
inlineprivate

Build (or rebuild, on a program change) the Aho-Corasick automaton for a fixed_alternation program whose branch count has reached ac_branch_threshold.

noinline, NOT cold: inlined into the dispatch chain of run() it regresses run_class_loop on x86 by presence alone, and cold would deoptimize a function called on every AC-eligible search.

Note
Cached per REGEX in detail::regex_immutables, not per state: a state is fresh per search() (a per-state automaton would be rebuilt per search). Its identity atomic is its own, never folded into built_for: only this route consults it.
Warning
The automaton scans at a flat rate: worst-case insurance, not a fast path (far slower with no match, far faster on a subject full of false starts). The subject, not the branch count, predicts the regime, hence ac_density_favours_automaton.
Template Parameters
DummyNever named by a caller: a member template with one if constexpr-gated call site is not emitted (nor counted uncovered) for instantiations that never take the route.
Returns
The automaton, or nullptr when this program has none (not a fixed alternation past the threshold, or a branch's icase-fold expansion would exceed ac_max_branch_expansion); the caller then falls back to run_alternation.

◆ add_branch_nibbles()

template<typename State , bool StateBoundToProgram = false>
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::add_branch_nibbles ( alternation_pairs &  plan,
std::size_t  branch 
) const
inlineconstexpr

Adds the branch at branch to plan's nibble fingerprint, in bucket plan.count % 8: each of its first three positions admits its byte, or every member of its class; a position past the branch's end admits anything (only that bucket loses selectivity). A branch one byte wide leaves the plan without a fingerprint: its bucket would mark every start.

Parameters
[in,out]planThe plan being built; plan.count is this branch's index.
[in]branchThe branch's first instruction.

◆ add_thread()

template<typename State , bool StateBoundToProgram = false>
template<bool Probe = false>
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::add_thread ( list_type &  list,
std::int32_t  pc0,
std::size_t  pos,
std::size_t  initial 
)
inlineconstexpr

Adds pc0 and its whole epsilon closure to list — the one closure walk (COW). Each DFS frame carries a capture-block index (in eps_entry::block) rather than mutating a shared working array, so capture state is copy-on-write and there are no slot-restore entries:

  • a frame popped from the stack owns one reference to its block;
  • split shares (incref: one ref → the two pushed frames), jump transfers it;
  • save — the ONLY write — copies-on-write first if the block is shared (capture_pool::cow_write);
  • a failed assertion or an already-seen pc releases the ref (decref);
  • a consuming/accept leaf transfers it into the thread list (one block handle per thread).
Parameters
[in,out]listThe thread list to populate (its slots hold one block index per pc).
[in]pc0The program counter to seed from.
[in]posThe current input position.
[in]initialCapture-free: group 0's START (full width, hence std::size_t rather than an eps_entry field). Otherwise the starting block, whose ref the caller already owns.

◆ add_variant_nibbles()

template<typename State , bool StateBoundToProgram = false>
bool real::detail::pike_vm< State, StateBoundToProgram >::add_variant_nibbles ( alternation_pairs &  plan,
std::size_t  branch 
) const
inline

Adds the branch at branch to plan's fingerprint, following each byte offset 0 to 2 a path through it can reach: a byte or class admits its members at every offset reached and moves one on; a code-point class admits its ASCII members, and the UTF-8 encoding of each other member laid from each offset reached, and moves on by every width it has; where the straight line ends, every byte is admitted from each offset reached on. A superset of the bytes a match can start with.

Parameters
[in,out]planThe plan being built; plan.count is this branch's bucket.
[in]branchThe branch's first instruction.
Returns
False when the branch is too short for a fingerprint (a path under two bytes) or opens on a code-point class too large to lay out (more than 64 non-ASCII members).

◆ advance_thread()

template<typename State , bool StateBoundToProgram = false>
template<bool Probe = false>
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::advance_thread ( list_type &  clist,
list_type &  nlist,
std::size_t  i,
std::int32_t  next_pc,
std::size_t  next_pos 
)
inlineconstexpr

Advances thread i of clist by one consumed byte, seeding its continuation's closure into nlist with its own reference on the thread's capture block (shared until a save copies it).

Parameters
[in]clistCurrent list, holding the thread to advance.
[in,out]nlistNext list, receiving the continuation's closure.
[in]iThread index within clist.
[in]next_pcProgram counter the thread continues at.
[in]next_posText position the thread continues at.

◆ alternation_automaton_claims()

template<typename State , bool StateBoundToProgram = false>
bool real::detail::pike_vm< State, StateBoundToProgram >::alternation_automaton_claims ( std::string_view  text,
std::size_t  start 
)
inline

Whether run()'s cascade would hand this subject's search at start to the Aho-Corasick automaton rather than to run_alternation: the gate's conditions in the same order, on the same per-subject state, whose verdicts are sticky, so that a batched walk (fill_alternation_spans) asks once and never overrules a routing decision that was measured. The fingerprint run() tries first needs more first bytes than the small set holds, which that walk's eligibility excludes.

Parameters
[in]textThe subject.
[in]startWhere the walk begins.
Returns
True when the automaton would take the subject.

◆ alternation_density_for()

template<typename State , bool StateBoundToProgram = false>
alternation_density & real::detail::pike_vm< State, StateBoundToProgram >::alternation_density_for ( std::string_view  text) const
inline

This subject's alternation density, judged anew when the state last sampled another subject.

Parameters
[in]textThe subject.
Returns
The density, kept in the state.

◆ alternation_density_seen()

template<typename State , bool StateBoundToProgram = false>
const alternation_density * real::detail::pike_vm< State, StateBoundToProgram >::alternation_density_seen ( std::string_view  text) const
inline

The alternation density the state holds for this subject, without sampling it.

Parameters
[in]textThe subject.
Returns
The density, or null when the state holds none for this subject.

◆ alternation_filter_takes()

template<typename State , bool StateBoundToProgram = false>
bool real::detail::pike_vm< State, StateBoundToProgram >::alternation_filter_takes ( std::string_view  text,
std::size_t  start 
) const
inline

Whether run_alternation will mask this subject's blocks by its pairs or fingerprint: a dense subject, where the Aho-Corasick gate (calibrated against the first-byte scan) would pick the automaton. The filtered scan beats the automaton there.

Parameters
[in]textThe subject.
[in]startWhere the search starts.
Returns
True when the alternation's block filter takes the subject.

◆ alternation_pairs_ready()

template<typename State , bool StateBoundToProgram = false>
const alternation_pairs * real::detail::pike_vm< State, StateBoundToProgram >::alternation_pairs_ready ( ) const
inlineprivate

The regex's alternation probe pairs, built once per regex in its immutables, or null when there is no per-regex cache: the caller then scans by the first bytes. Same identity discipline as ac_ready.

Returns
The plan, or null.

◆ alternation_plan()

template<typename State , bool StateBoundToProgram = false>
const alternation_pairs * real::detail::pike_vm< State, StateBoundToProgram >::alternation_plan ( std::string_view  text,
std::size_t  pos,
std::array< std::uint8_t, 8 >  mem,
std::size_t  cnt 
) const
inline

The alternation's probe pairs when this subject's first bytes are dense, for the block scans of run_alternation and fill_alternation_spans, else null (a compile-time storage keeps no state for them, or no plan fits).

Decided once per subject, from alternation_sample_bytes bytes at pos, and only when at least alternation_sample_min remain: a short subject is scanned by the first bytes, unsampled. The plan is built once per program from its branches, read as the scans' match_at reads them. Deciding is out of line (alternation_plan_decide), so the callers' first-byte loops keep their code.

Parameters
[in]textThe subject.
[in]posWhere the scan starts.
[in]memThe branches' first bytes.
[in]cntHow many of mem are valid.
Returns
The plan, or null.

◆ alternation_plan_decide()

template<typename State , bool StateBoundToProgram = false>
const alternation_pairs * real::detail::pike_vm< State, StateBoundToProgram >::alternation_plan_decide ( std::string_view  text,
std::size_t  pos,
std::array< std::uint8_t, 8 >  mem,
std::size_t  cnt 
) const
inline

Out of line, the half of alternation_plan that resets the density on a new subject, samples it, and builds the plan.

Parameters
[in]textThe subject.
[in]posWhere the scan starts.
[in]memThe branches' first bytes, padded.
[in]cntHow many of mem are valid.
Returns
The plan, or null.

◆ alternation_wide_may_take()

template<typename State , bool StateBoundToProgram = false>
bool real::detail::pike_vm< State, StateBoundToProgram >::alternation_wide_may_take ( std::string_view  text) const
inline

Whether run_alternation_wide may take this search: an alternation with more first bytes than the small set holds and no single first byte, within the fingerprint plan's branches, on a subject its sample has not refused. The cheap half, inline, so that a refused subject costs no call.

Parameters
[in]textThe subject.
Returns
False when the route certainly declines.

◆ assertion_holds()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::assertion_holds ( assert_kind  kind,
std::size_t  pos,
bool  word_ness_flipped 
) const
inlineconstexpr

Evaluates a zero-width assertion at pos in the current text.

Parameters
[in]kindThe assertion to evaluate.
[in]posThe position at which to evaluate it.
[in]word_ness_flippedFor a word assert (\b \B \< \>), whether this instruction flips the program's default word-ness — set for a scoped (?a:...) / (?-a:...) island.
Returns
true if the assertion holds there.

◆ backtrack_from()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
bool real::detail::pike_vm< State, StateBoundToProgram >::backtrack_from ( backtrack_frame &  frame,
std::size_t  seed,
std::size_t  start,
run_mode  mode,
bool  cf,
OutSlots &  out_slots 
)
inline

Every branch from pc 0 at seed, in priority order – one start of run_bounded_backtrack.

Parameters
[in,out]frameThe search's marks, leaf flags, slots and pending branches.
[in]seedThe start.
[in]startThe search's start (row 0).
[in]modeAnchoring (a full match must end at the text's end).
[in]cfThe program's capture-free walk: only group 0's start is carried.
[out]out_slotsCapture slots, filled on a match.
Returns
True when a branch matched.

◆ backtrack_possessive()

template<typename State , bool StateBoundToProgram = false>
std::int32_t real::detail::pike_vm< State, StateBoundToProgram >::backtrack_possessive ( backtrack_frame &  frame,
const instr &  instruction,
std::int32_t  pc,
std::size_t &  pos,
bool  cf 
)
inline

A possessive loop's step in backtrack_from. The VM decides it when the thread arrives, so a match consumes (writing the loop's capture, if it has one) and a miss leaves by the exit at the same position.

Parameters
[in,out]frameThe search's frame (slots, pending restores).
[in]instructionThe possessive instruction.
[in]pcIts program counter.
[in,out]posThe position; advanced by one on a match.
[in]cfThe program's capture-free walk: the loop's capture is not recorded.
Returns
The thread's next instruction.

◆ build_alternation_pairs()

template<typename State , bool StateBoundToProgram = false>
alternation_pairs real::detail::pike_vm< State, StateBoundToProgram >::build_alternation_pairs ( ) const
inline

Each branch's first byte and its farthest byte within 15 of it (the pairs, when every branch opens on a byte), and the branches' nibble fingerprint, read from the split chain in source order as the scans' match_at reads it. A branch that opens on a class leaves the plan without pairs; the fingerprint then carries it alone, whatever its minimum of branches against the pairs.

Returns
The plan; count == 0 when a branch opens on neither a byte nor a class, the branches outnumber it, or neither filter holds.

◆ build_cp_alternation_plan()

template<typename State , bool StateBoundToProgram = false>
alternation_pairs real::detail::pike_vm< State, StateBoundToProgram >::build_cp_alternation_plan ( ) const
inline

The fingerprint of a program laid out as an alternation of straight-line branches (after leading position assertions, which only narrow where a match starts) that is not a fixed alternation: one of its branches holds a code-point class. No pairs: those need a byte at every head.

Returns
The plan; count == 0 when the layout is not that, a branch declines (add_variant_nibbles), the branches outnumber the buckets' capacity, or there is no fingerprint on this target.

◆ class_run_end()

template<typename State , bool StateBoundToProgram = false>
template<bool Cascade>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::class_run_end ( std::string_view  text,
const std::uint8_t *  tbl,
std::size_t  match_start 
) const
inlineconstexprprivate

The end of the maximal run of tbl's members that begins with the member at match_start.

Template Parameters
CascadePast cascade_run_threshold bytes, hand the rest to run_cascade_stop, which is sound because a byte-class run never validates UTF-8.
Parameters
[in]textSubject.
[in]tblThe class's byte membership table.
[in]match_startA member's offset.
Returns
The first offset past the run.

◆ class_table()

template<typename State , bool StateBoundToProgram = false>
constexpr const std::uint8_t * real::detail::pike_vm< State, StateBoundToProgram >::class_table ( std::size_t  class_index)
inlineconstexprprivate

Returns a flat 256-byte membership table for class class_index (one load per byte).

always_inline is load-bearing: out of line, the call frame costs more than the inlined accessor (6.2 M instructions against 0.85 M on a 64 KiB [a-z]+ walk). It fits only with derive_class_table kept out of it.

Parameters
[in]class_indexIndex into the program's interned classes.
Returns
Pointer to a 256-entry table: 1 where the byte is in the class.

◆ confirm_at()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
bool real::detail::pike_vm< State, StateBoundToProgram >::confirm_at ( std::string_view  text,
std::size_t  s,
OutSlots &  out_slots,
std::size_t &  stop 
)
inline

Confirm a match anchored at s: the forward DFA finds its end and the one-pass table fills the captures, as on the lazy-DFA route. Falls back to the anchored Pike when the pattern is not DFA/one-pass eligible, or when the DFA's leftmost match does not begin at s.

Parameters
[in]textSubject.
[in]sCandidate match start to confirm.
[out]out_slotsCapture slots, filled on a match.
[out]stopHow far the confirm reached, for the linearity backstop.
Returns
True when a match begins exactly at s.

◆ cow_release_blocks()

template<typename State , bool StateBoundToProgram = false>
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::cow_release_blocks ( list_type &  list)
inlineconstexpr

Releases the block references a list's threads hold, before the list is reset or the run returns: the one decref site paired with each step→closure incref (keep it single).

Parameters
[in]listList whose threads' block references are dropped.

◆ cp_ascii_table()

template<typename State , bool StateBoundToProgram = false>
constexpr const std::uint8_t * real::detail::pike_vm< State, StateBoundToProgram >::cp_ascii_table ( std::size_t  cp_index)
inlineconstexprprivate

Byte-indexed membership table for the ASCII bitmap of a cp_class, as class_table for the klass_cp scan loop, keyed negatively so it never collides with a byte class.

Parameters
[in]cp_indexIndex into the program's cp_classes.
Returns
Pointer to a 256-entry table: 1 where the byte (< 0x80) is a member.

◆ cp_class_holds()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::cp_class_holds ( const cp_class &  cc,
char32_t  cp 
) const
inlineconstexprprivate

Stateless membership of cp in cc: no VM-state cache touched.

The cached paths (cp_member_page, cp_member_high) hold ONE class each; the inner-literal reverse alternates with the confirm's classes per candidate, so it reads the class directly (ASCII bitmap, else a binary search of its ranges).

Parameters
[in]ccThe code-point class.
[in]cpThe code point.
Returns
true if cp is a member.

◆ cp_class_matches()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::cp_class_matches ( const detail::cp_class &  cc,
char32_t  cp 
) const
inlineconstexpr

Tests a decoded code point against a klass_cp class: ASCII bitmap below 0x80, binary search of the class's range slice above (constexpr / const paths; cp_class_matches_idx uses the cached tables). The class is already the effective set: a plain positive membership test.

Parameters
[in]ccThe code-point class (from prog_.cp_classes).
[in]cpThe decoded code point.
Returns
Whether cp is a member.

◆ cp_class_matches_idx()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::cp_class_matches_idx ( std::size_t  cp_index,
char32_t  cp 
)
inlineconstexpr

Membership by class index (ASCII + European page + sparse hi / bsearch).

Parameters
[in]cp_indexIndex of the code-point class.
[in]cpThe decoded code point.
Returns
Whether cp is a member.

◆ cp_hi_build()

template<typename State , bool StateBoundToProgram = false>
static const cp_hi_table * real::detail::pike_vm< State, StateBoundToProgram >::cp_hi_build ( const program_view &  prog,
std::size_t  cp_index,
std::uint64_t  key_fp,
std::array< cp_hi_cache_entry, 8 > &  cache,
const cp_hi_table *&  last_tab,
std::uint64_t &  last_fp 
)
inlinestaticprivate

Cold path: build a sparse hi table and install it in the thread-local cache. Outlined so the hot membership check never inlines the range-walk builder.

Parameters
[in]progProgram owning the class.
[in]cp_indexIndex of the code-point class in prog.cp_classes.
[in]key_fpThe class's content fingerprint, the cache key.
[in,out]cacheThread-local entries, one of which is overwritten.
[out]last_tabSticky last-hit table pointer, set to the built table.
[out]last_fpSticky last-hit fingerprint, set to key_fp.
Returns
The installed table, owned by cache.

◆ cp_hi_cached()

template<typename State , bool StateBoundToProgram = false>
static const cp_hi_table * real::detail::pike_vm< State, StateBoundToProgram >::cp_hi_cached ( const program_view &  prog,
std::size_t  cp_index 
)
inlinestaticprivate

Thread-local sparse hi tables, keyed by cp_class::fingerprint (set once at intern), so basic_pike_state keeps its size. Hot path: a uint64 load and a sticky compare.

Parameters
[in]progProgram owning the class.
[in]cp_indexIndex of the code-point class in prog.cp_classes.
Returns
The class's sparse table, built on first use for this thread.

◆ cp_member_hi()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::cp_member_hi ( std::size_t  cp_index,
char32_t  cp 
)
inlineconstexprprivate

Non-ASCII membership: European page bitmap, then sparse 2-stage hi / bsearch.

Parameters
[in]cp_indexIndex of the code-point class.
[in]cpCode point at or above U+0080.
Returns
True when cp is a member.

◆ cp_member_high()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::cp_member_high ( std::size_t  cp_index,
char32_t  cp 
)
inlineconstexprprivate

Membership for cp > U+07FF: sparse 2-stage hi table, else bsearch (small classes / constexpr). The table build is cold-outlined; the per-cp probe is last-hit + bit test.

Parameters
[in]cp_indexIndex of the code-point class.
[in]cpCode point above cp_page_max.
Returns
True when cp is a member.

◆ cp_member_high_unshared()

template<typename State , bool StateBoundToProgram = false>
bool real::detail::pike_vm< State, StateBoundToProgram >::cp_member_high_unshared ( std::size_t  cp_index,
char32_t  cp 
)
inlineprivate

cp_member_high for the trailing-lookaround walk, written out rather than called.

A second call site of cp_member_high, even from this cold walk, makes GCC stop inlining it into fill_cp_class_spans, which charges the code-point class scans. Runtime only (the walk is dynamic-only).

Parameters
[in]cp_indexIndex of the code-point class.
[in]cpCode point above cp_page_max.
Returns
True when cp is a member.

◆ cp_member_page()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::cp_member_page ( std::size_t  cp_index,
char32_t  cp 
)
inlineconstexprprivate

Page-bitmap membership for U+0080..U+07FF. Kept separate so class-loop lambdas can inline it without pulling the sparse-hi path into the European hot stream (\p{N}, accented).

Parameters
[in]cp_indexIndex of the code-point class.
[in]cpCode point in U+0080..U+07FF; outside that range the bit index is meaningless.
Returns
True when cp is a member.

◆ cp_page_table()

template<typename State , bool StateBoundToProgram = false>
constexpr const std::uint64_t * real::detail::pike_vm< State, StateBoundToProgram >::cp_page_table ( std::size_t  cp_index)
inlineconstexprprivate

Builds (once, cached) the cp_class's membership bitmap over [U+0080, U+07FF]: one load instead of a range search on two-byte code points (see basic_pike_state::cp_page).

Parameters
[in]cp_indexIndex into the program's cp_classes.
Returns
Pointer to the 30-word bitmap (bit cp - 0x80).

◆ cut_short()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::cut_short ( std::size_t  pos) const
inlineconstexpr

Whether the code point at pos is not all there: past the end of the text, or a sequence the end cuts short – what a class test or a word boundary at pos would read more text to decide.

Parameters
[in]posThe position.
Returns
True when more text could change what is read at pos.

◆ derive_class_table()

template<typename State , bool StateBoundToProgram = false>
constexpr const std::uint8_t * real::detail::pike_vm< State, StateBoundToProgram >::derive_class_table ( std::size_t  class_index)
inlineconstexprprivate

Derives the byte row into the VM state: the constant-evaluation path, where no per-regex cache exists.

noinline is load-bearing: inlined, this 256-iteration loop keeps class_table out of basic_match_iterator::advance (figures at class_table).

Parameters
[in]class_indexIndex into the program's interned byte classes.
Returns
The state's table.

◆ ensure_membership_rows()

template<typename State , bool StateBoundToProgram = false>
void real::detail::pike_vm< State, StateBoundToProgram >::ensure_membership_rows ( detail::regex_immutables &  cache) const
inlineprivate

Sizes the per-regex membership rows for this program, if not already. Cold: once per regex, behind an acquire load on the hot path.

Parameters
[in,out]cacheThe per-regex immutables.

◆ ensure_op_table()

template<typename State , bool StateBoundToProgram = false>
void real::detail::pike_vm< State, StateBoundToProgram >::ensure_op_table ( )
inlineprivate

Build (or rebuild) the one-pass capture extractor, on top of ensure_immutables.

Split out: it is the expensive half (on a first capture search, more than the byte program and the lazy DFA together) and only some routes consult it. Guarded by its own regex_immutables::op_table_for, which ensure_immutables clears on rebuild, so a reassigned regex never reads a stale extractor.

◆ ensure_set_il_prefix_rev()

template<typename State , bool StateBoundToProgram = false>
void real::detail::pike_vm< State, StateBoundToProgram >::ensure_set_il_prefix_rev ( detail::regex_immutables &  immut,
shared_dfa_set &  set 
)
inlineprivate

Builds the IL-prefix reverse DFA for immut into this thread's leased set, once.

Parameters
[in]immutPer-regex immutables naming the program to build for.
[in,out]setThe leased DFA set to populate.

◆ ensure_set_search_dfas()

template<typename State , bool StateBoundToProgram = false>
void real::detail::pike_vm< State, StateBoundToProgram >::ensure_set_search_dfas ( detail::regex_immutables &  immut,
shared_dfa_set &  set 
)
inlineprivate

Builds the search DFAs for immut into this thread's leased set, once.

Parameters
[in]immutPer-regex immutables naming the program to build for.
[in,out]setThe leased DFA set to populate.

◆ ensure_slot_size()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
static constexpr void real::detail::pike_vm< State, StateBoundToProgram >::ensure_slot_size ( OutSlots &  out,
std::size_t  n 
)
inlinestaticconstexprprivate

Size out without a full npos fill when already sized: ensure_size, or a grow-only resize for the seam tests' std::vector.

Parameters
[in,out]outSlot storage to grow.
[in]nMinimum size required; out is never shrunk.

◆ exact_literal_is_the_route()

template<typename State , bool StateBoundToProgram = false>
static constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::exact_literal_is_the_route ( const pattern_hints &  hints)
inlinestaticconstexprnoexcept

Is the exact-literal route the one run() would take, in its one-search subset?

Mirrors run()'s cascade, like lazy_dfa_is_the_route: only the class loops sit above this route. literal_one_search carries the rest (no capture, assertion or anchor, a literal of >= 2 bytes, prefix_size == exact_literal_len), so the answer is find_prefix plus two stores.

Parameters
[in]hintsThe program's shape hints.
Returns
True when no earlier route in the cascade claims this shape.

◆ extends_past_end()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
bool real::detail::pike_vm< State, StateBoundToProgram >::extends_past_end ( std::string_view  text,
std::size_t  start,
OutSlots &  out_slots 
)
inline

Whether a match anchored at start could come out differently if text continued past its end: what a caller lexing text that arrives in pieces must know before committing a token.

Runs the general loop in prefix mode with probes (probe_step, probe_closure). A thread still alive when the text runs out outranks the match found (prefix mode cuts lower priorities), and so can anything that read the end as an end: an assertion looking right ($, \Z, \b, ...), a lookahead window reaching it, a code point it cuts short. Conservative: may say true where a closer look says false, never the reverse (a waiting caller loses time, not tokens).

Parameters
[in]textThe text available so far.
[in]startWhere the match is anchored.
[out]out_slotsThe match on text as it stands (prefix mode).
Returns
True when text past the end could change the match.

◆ fail_slots() [1/2]

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::fail_slots ( OutSlots &  out_slots) const
inlineconstexpr

fail_slots for every slot of the program.

Parameters
[out]out_slotsThe slots.
Returns
False, for the caller to return.

◆ fail_slots() [2/2]

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::fail_slots ( OutSlots &  out_slots,
std::size_t  count 
) const
inlineconstexpr

Clears the capture slots for a search that found nothing, and says so.

Parameters
[out]out_slotsThe slots, count of them set to real::npos.
[in]countHow many: the program's slots, or the two a groupless route writes.
Returns
False, for the caller to return.

◆ fast_search()

template<typename State , bool StateBoundToProgram = false>
template<typename MatchAt , typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::fast_search ( std::string_view  text,
std::size_t  start,
MatchAt  match_at,
OutSlots &  out_slots 
)
inlineconstexpr

Leftmost search over the candidate starts of next_candidate, for the first one match_at accepts. Shared by the fast paths that verify a fixed shape at a position.

Template Parameters
MatchAtCallable std::size_t(std::size_t pos): the match end at pos, or npos.
OutSlotsOutput slot container (already sized to two).
Parameters
[in]textThe subject text.
[in]startIndex to begin at.
[in]match_atThe per-position matcher.
[out]out_slotsReceives the (start, end) span on success.
Returns
true if a match was found.

◆ fill_ahead_table()

template<typename State , bool StateBoundToProgram = false>
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::fill_ahead_table ( const lookaround_sub &  sub,
lookaround_scratch::ahead_table &  table 
)
inlineconstexpr

Fills table with every position's answer, for unbounded_lookahead_matches to read.

Parameters
[in]subThe lookaround sub-program.
[in,out]tableThe table to fill for the current subject.

◆ fill_alternation_spans()

template<typename State , bool StateBoundToProgram = false>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::fill_alternation_spans ( std::string_view  text,
std::size_t  start,
cp_span *  out,
std::size_t  cap 
)
inlineconstexpr

Fills up to cap fixed_alternation matches from start without leaving the route.

The route is return-dominated (see run_alternation). Small-set shape only (2..8 distinct first bytes, pattern_hints::small_set_size), which the mask scan needs; other alternations are not batched, rather than growing a second scan body here (docs/MEASUREMENT.md §5.4).

The scan is a COPY of run_alternation's, not a shared call: relocating that measured hot body risks a regression worse than this filler's gain.

Parameters
[in]textThe subject.
[in]startWhere to begin.
[out]outBuffer for the spans found.
[in]capCapacity of out; the walk stops there and resumes from the last end.
Returns
How many spans were written.

◆ fill_alternation_wide_spans()

template<typename State , bool StateBoundToProgram = false>
std::size_t real::detail::pike_vm< State, StateBoundToProgram >::fill_alternation_wide_spans ( std::string_view  text,
std::size_t  start,
cp_span *  out,
std::size_t  cap,
bool &  partial,
bool &  disarm 
)
inline

Fills up to cap matches of an alternation with more first bytes than the small set holds, from start, by the scan of run_alternation_wide run from each match's end, without re-entering run().

run() hands such an alternation to that route where alternation_wide_may_take says so and the route takes the subject on its sample (a verdict sticky per subject); otherwise the automaton's gate decides. The filler asks the same questions, and where either declines it writes nothing and says so through disarm: the walk leaves every later search to run().

Parameters
[in]textThe subject.
[in]startWhere to begin.
[out]outBuffer for the spans found.
[in]capCapacity of out.
[out]partialThe fill stopped without proving the subject spent.
[out]disarmThe route declined this subject: batch no more of it.
Returns
How many spans were written.

◆ fill_byte_row()

template<typename State , bool StateBoundToProgram = false>
static void real::detail::pike_vm< State, StateBoundToProgram >::fill_byte_row ( detail::regex_immutables &  cache,
std::size_t  ready_index,
const char_class &  klass,
std::uint8_t *  row 
)
inlinestaticprivate

Expands klass into one flat 256-byte membership row of the per-regex cache, once.

Shared body of fill_class_row and fill_cp_ascii_row (cold, once per class). fill_cp_page_row is not folded in: it builds a 30-word bitmap from range pairs.

Parameters
[in,out]cacheThe per-regex immutables.
[in]ready_indexIndex of this row's ready bit.
[in]klassThe membership set to expand.
[out]rowDestination, 256 bytes.

◆ fill_class_row()

template<typename State , bool StateBoundToProgram = false>
void real::detail::pike_vm< State, StateBoundToProgram >::fill_class_row ( detail::regex_immutables &  cache,
std::size_t  class_index 
) const
inlineprivate

Fills one byte-class row of the per-regex cache, once.

Parameters
[in,out]cacheThe per-regex immutables.
[in]class_indexIndex into the program's interned byte classes.

◆ fill_class_spans()

template<typename State , bool StateBoundToProgram = false>
template<bool Cascade, bool WbEdge, bool WbKept>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::fill_class_spans ( std::string_view  text,
std::size_t  start,
cp_span *  out,
std::size_t  cap 
)
inlineconstexpr

Fills up to cap class_loop matches from start without leaving the route.

The byte-class twin of fill_cp_class_spans. On word text this route emits a match every few bytes and the scan is a table lookup per byte, so the per-match return dominates.

Template Parameters
CascadeWhether the memchr stop-tail applies, chosen once per walk by the caller.
Parameters
[in]textThe subject.
[in]startWhere to begin.
[out]outBuffer for the spans found.
[in]capCapacity of out.
Returns
How many spans were written.
Note
Each branch added to refill_batch must be re-measured on BOTH ISAs against rows that never touch it: fillers move the translation unit's inline budget (a ./negated-class filler took back most of this one's gain; docs/design.dox §10.1).

◆ fill_codepoint_class_spans()

template<typename State , bool StateBoundToProgram = false>
template<bool Cascade>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::fill_codepoint_class_spans ( std::string_view  text,
std::size_t  start,
cp_span *  out,
std::size_t  cap 
)
inlineconstexpr

Batched twin of run_codepoint_class, filling up to cap maximal spans in ONE call.

Otherwise the ./negated-class shape pays a full route entry per match, the other class routes one per sixteen. Not a flag on the existing function: widening a shared scan lambda by one branch charges the property-class rows that never use it.

Search semantics only: basic_match_iterator excludes anchored shapes from batching, and run_mode::full keeps run_codepoint_class.

Template Parameters
CascadeSelect the memchr-cascade run scan, chosen once per walk.
Parameters
[in]textThe subject.
[in]startByte offset to begin at.
[out]outReceives the spans.
[in]capCapacity of out.
Returns
How many spans were written (0 = no further match).

◆ fill_cp_ascii_row()

template<typename State , bool StateBoundToProgram = false>
void real::detail::pike_vm< State, StateBoundToProgram >::fill_cp_ascii_row ( detail::regex_immutables &  cache,
std::size_t  cp_index 
) const
inlineprivate

Fills one code-point-class ASCII row of the per-regex cache, once.

Parameters
[in,out]cacheThe per-regex immutables.
[in]cp_indexIndex into the program's code-point classes.

◆ fill_cp_class_spans()

template<typename State , bool StateBoundToProgram = false>
template<bool WbEdge>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::fill_cp_class_spans ( std::string_view  text,
std::size_t  start,
cp_span *  out,
std::size_t  cap 
)
inlineconstexpr

Fills up to cap cp_class_loop matches from start without leaving the route.

Most of a single-code-point row is the per-match return (run()'s dispatch, fill_span_slots, the iterator's re-entry), not the scan: a buffer amortises it over cap matches and hoists asc.

The caller (basic_match_iterator) guards search semantics and no \b/\B wrap (a kept one goes through fill_cp_class_spans_wrapped).

Parameters
[in]textThe subject.
[in]startWhere to begin.
[out]outBuffer for the spans found.
[in]capCapacity of out; the walk stops there and resumes from the last end.
Returns
How many spans were written.

◆ fill_cp_class_spans_wrapped()

template<typename State , bool StateBoundToProgram = false>
template<bool WbEdge>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::fill_cp_class_spans_wrapped ( std::string_view  text,
std::size_t  start,
cp_span *  out,
std::size_t  cap 
)
inlineconstexpr

fill_cp_class_spans for a pattern with a kept \b/\B wrap: its spans, less those whose wrap does not hold.

Filters the plain filler's batches (the per-match route skips a failing run whole) and refills until a span survives or the runs are spent: never an empty batch while runs remain (the iterator reads one as the end). Kept apart and cold: a template parameter on the plain filler changes GCC's inlining of its code-point lookup.

Parameters
[in]textThe subject.
[in]startWhere to begin.
[out]outBuffer for the spans found.
[in]capCapacity of out.
Returns
How many spans were written.

◆ fill_cp_page_row()

template<typename State , bool StateBoundToProgram = false>
void real::detail::pike_vm< State, StateBoundToProgram >::fill_cp_page_row ( detail::regex_immutables &  cache,
std::size_t  cp_index 
) const
inlineprivate

Fills one U+0080..U+07FF membership bitmap of the per-regex cache, once.

Parameters
[in,out]cacheThe per-regex immutables.
[in]cp_indexIndex into the program's code-point classes.

◆ fill_exact_literal_spans()

template<typename State , bool StateBoundToProgram = false>
std::size_t real::detail::pike_vm< State, StateBoundToProgram >::fill_exact_literal_spans ( std::string_view  text,
std::size_t  start,
cp_span *  out,
std::size_t  cap 
)
inline

Fills up to cap exact-literal matches from start without re-entering the route gate.

Enlarging refill_batch charges no other row; one extra comparison on the hot path of advance does (see run_literal_one_search).

No partial state: find_prefix scans to the subject's end, so an empty return proves exhaustion.

Parameters
[in]textThe subject.
[in]startWhere to begin.
[out]outBuffer for the spans found.
[in]capCapacity of out; the walk stops there and resumes from the last end.
Returns
How many spans were written.

◆ fill_fixed_saves()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::fill_fixed_saves ( std::size_t  match_start,
OutSlots &  out_slots 
) const
inlineconstexpr

Fills the capturing-group slots of a fixed-shape match. Every consuming op is one byte wide, so each save sits at a constant offset from the match start: one linear pass, no re-match.

Starts at pattern_hints::body_pc, not 1: a leading \b/\B sits at pc 1 and the walk would break on it at once, filling no group (\B(\w){2}).

Parameters
[in]match_startByte offset where the match begins.
[out]out_slotsReceives the group slots.

◆ fill_fixed_shape_spans()

template<typename State , bool StateBoundToProgram = false>
std::size_t real::detail::pike_vm< State, StateBoundToProgram >::fill_fixed_shape_spans ( std::string_view  text,
std::size_t  start,
cp_span *  out,
std::size_t  cap 
)
inline

Fills up to cap fixed-shape matches from start without re-entering the route gate.

Calls run_fixed_shape rather than copying it, as fill_inner_literal_spans calls its route: the scan and verify are the per-match walk's own, so the two cannot disagree, and what goes is the walk's return through the iterator and run()'s cascade per match – about a quarter of a dense date row.

Parameters
[in]textThe subject.
[in]startWhere the walk resumes.
[out]outThe spans found.
[in]capCapacity of out.
Returns
How many spans were written.

◆ fill_inner_literal_spans()

template<typename State , bool StateBoundToProgram = false>
std::size_t real::detail::pike_vm< State, StateBoundToProgram >::fill_inner_literal_spans ( std::string_view  text,
std::size_t  start,
cp_span *  out,
std::size_t  cap,
bool &  partial,
bool &  disarm 
)
inline

Fills up to cap inner-literal matches from start without re-entering the route gate.

The per-match route bills one engine entry per match, flat across densities: the return is the cost. Calls run_inner_literal rather than copying it (unlike fill_alternation_spans): its linearity backstop, density guard, size floor and reverse confirm are state whose duplication would make the batched and per-match walks disagree. The per-haystack reset is shared through il_reset_on_new_haystack.

partial follows the lazy-DFA filler's contract: every abandonment (density, linearity, size floor, unplaceable start) leaves matches for another route; only memmem running out proves exhaustion.

Parameters
[in]textThe subject.
[in]startWhere to begin.
[out]outBuffer for the spans found.
[in]capCapacity of out; the walk stops there and resumes from the last end.
[out]partialTrue unless the subject was proven spent; see above.
[out]disarmSet when the route has ABANDONED this haystack: the caller must stop calling this filler for the rest of the walk, or it pays a failed refill per match (3634 attempts against 7 on a dense date cell, slower than the core).
Returns
How many spans were written.

◆ fill_lazy_dfa_spans()

template<typename State , bool StateBoundToProgram = false>
std::size_t real::detail::pike_vm< State, StateBoundToProgram >::fill_lazy_dfa_spans ( std::string_view  text,
std::size_t  start,
cp_span *  out,
std::size_t  cap,
bool &  partial 
)
inline

Fills up to cap lazy-DFA matches from start without re-entering the route gate.

Patterns no shape recognizer claims ([a-z]+|[0-9]+) land here; per match, the return, not the DFA scan, is the cost. The anchored walks from candidates give way to one forward pass and reverse, as in try_shared_lazy_dfa_search.

partial: an empty return ends the walk in basic_match_iterator::advance, sound only where the scan covers the whole subject. This route can stop with matches still ahead (a DFA quit, under lazy_dfa_min_input bytes left, no shared DFAs yet), so partial stays set unless exhaustion is PROVEN, and the caller resumes on the per-match path.

Parameters
[in]textThe subject.
[in]startWhere to begin.
[out]outBuffer for the spans found.
[in]capCapacity of out; the walk stops there and resumes from the last end.
[out]partialTrue unless the subject was proven spent; see above.
Returns
How many spans were written.

◆ fill_single_class_spans()

template<typename State , bool StateBoundToProgram = false>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::fill_single_class_spans ( std::string_view  text,
std::size_t  start,
cp_span *  out,
std::size_t  cap 
)
inlineconstexpr

Fills up to cap bare single byte-class matches from start without leaving the route.

Each accepted byte is one match, otherwise a full route entry (several times slower per byte than [a-z]+); no Cascade variant. The caller (real::basic_match_iterator) guards search semantics, no anchor and no \b/\B wrap; the 4-opcode shape rules out groups and a {k,} minimum.

Parameters
[in]textThe subject.
[in]startWhere to begin.
[out]outBuffer for the spans found.
[in]capCapacity of out; the walk stops there and resumes from the last end.
Returns
How many spans were written.

◆ fill_span_slots()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::fill_span_slots ( OutSlots &  out_slots,
std::size_t  match_start,
std::size_t  match_end 
) const
inlineconstexprprivate

Writes a class-loop fast-path result into out_slots: the whole-match span in slots 0/1, mirrored into the group's slots for a pattern wrapped in one capturing group ((\w+), ([a-z]+)): the group's span equals the whole match by construction.

No npos fill: this writer covers every slot such shapes have.

Parameters
[out]out_slotsCapture slots to write.
[in]match_startWhole-match start offset.
[in]match_endWhole-match end offset.

◆ fill_two_run_saves()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::fill_two_run_saves ( std::string_view  text,
std::size_t  s,
std::size_t  h,
std::size_t  lit_end,
std::size_t  e,
OutSlots &  out_slots 
) const
inlineconstexpr

Fills capture slots for a class+ <literal> class+ match, by anchor rather than by offset.

No fixed widths, but every save lands on one of four positions, decided by where it sits relative to the two loops and the literal (one walk of a dozen instructions per match). When the literal can occur inside the prefix run (pattern_hints::il_fwd_last), the greedy prefix gives back only to the LAST occurrence that leaves the suffix a member, so the literal is moved there first.

Parameters
[in]textThe subject.
[in]sMatch start (the prefix run's beginning).
[in]hThe candidate literal's start.
[in]lit_endOne past the candidate literal.
[in]eMatch end (the suffix run's end).
[out]out_slotsSlots to fill.

◆ find_on_subject()

template<typename State , bool StateBoundToProgram = false>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::find_on_subject ( std::string_view  text,
std::size_t  pos,
std::string_view  lit,
std::size_t  rare,
bool  inner 
) const
inlineconstexpr

The next occurrence of a literal of the pattern's hints at or after pos, by find_literal_adaptive with a density kept for the whole subject.

A subject whose rarest literal byte proved common stays on the two-byte filter for later searches. A storage without the fields keeps the density per call (correct, re-learns); constant evaluation takes the plain searches.

Parameters
[in]textThe subject.
[in]posIndex to start from.
[in]litThe literal (the prefix or the inner literal).
[in]rareOffset of its rarest byte, from the hints.
[in]innerWhether lit is the inner literal (each literal keeps its own density).
Returns
The index of the occurrence, else real::npos.

◆ fixed_shape_is_the_route()

template<typename State , bool StateBoundToProgram = false>
static constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::fixed_shape_is_the_route ( const program_view &  prog)
inlinestaticconstexprnoexcept

Is the fixed-shape route the one run() would take for this program, in a groupless search?

Mirrors the cascade above run_fixed_shape, one clause per earlier route, for the same reason lazy_dfa_is_the_route states: a batched walk bypasses run(), so batching a shape an earlier route claims takes it off that route. The inner-literal route declines fixed shapes itself.

Parameters
[in]progThe program.
Returns
True when fill_fixed_shape_spans answers as the per-match walk does.

◆ il_reset_on_new_haystack()

template<typename State , bool StateBoundToProgram = false>
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::il_reset_on_new_haystack ( std::string_view  text)
inlineconstexpr

Re-enables the inner-literal route and clears its density counters on a new haystack.

One mechanism: run()'s gate and fill_inner_literal_spans must observe the same reset of the sticky per-haystack guards, or a batched walk and a per-match walk silently diverge.

Parameters
[in]textThe subject being scanned.

◆ inner_literal_is_the_route()

template<typename State , bool StateBoundToProgram = false>
static constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::inner_literal_is_the_route ( const program_view &  prog)
inlinestaticconstexprnoexcept

Is the inner-literal route the one run() would take for this program?

Mirrors that route's gate and only the routes with their own run_* body above it (class loops, possessive loops, exact literal): a scan strategy is not a route. prefix_code is required unconditionally, conservatively (the gate needs it only for reverse-confirming storages), so a static regex without a prefix program keeps the per-match walk.

Parameters
[in]progThe compiled program view.
Returns
True when no earlier route in the cascade claims this shape.

◆ is_run_shape()

template<typename State , bool StateBoundToProgram = false>
static constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::is_run_shape ( const program_view &  prog)
inlinestaticconstexpr

Whether prog is saves, atoms (a byte, a byte class, a code-point class) and greedy atom+ and atom* loops, then match: nothing else, no alternation, no lazy loop, no assertion.

For such a program the walk that takes every loop as far as its atom matches, never backing up, is the highest-priority path: at each loop it chose the preferred branch whenever that branch could be taken. So when that walk reaches match its groups are the VM's (match_run_shape), and when an atom fails it, the VM decides.

Parameters
[in]progThe program.
Returns
True for that shape.

◆ lazy_dfa_is_the_route()

template<typename State , bool StateBoundToProgram = false>
static constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::lazy_dfa_is_the_route ( const pattern_hints &  hints)
inlinestaticconstexprnoexcept

Is the lazy-DFA route the one run() would actually take for this program?

Mirrors run()'s cascade: a batched walk bypasses run(), so batching a shape an EARLIER route claims takes it off that faster route (a plain literal would lose its memmem). The conditions are stated positively, one per route above the lazy DFA, never as a residue of the shape recognizers.

Warning
Adding a route to run() above the lazy DFA means adding its hint here. Nothing enforces it; the failure is a silent slowdown on exactly the new route's shape.
Parameters
[in]hintsThe program's shape hints.
Returns
True when no earlier route in the cascade claims this shape.

◆ literal_at()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::literal_at ( std::string_view  text,
std::size_t  cand,
std::size_t  len 
) const
inlineconstexpr

Tests whether the fixed literal prefix occurs at cand.

Parameters
[in]textThe subject text.
[in]candCandidate start offset.
[in]lenLength of the literal (hints.exact_literal_len).
Returns
true if text[cand : cand+len] equals the literal.

◆ lookahead_matches()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::lookahead_matches ( const lookaround_sub &  sub,
std::size_t  pos 
)
inlineconstexpr

Lookahead: does the sub-pattern match a prefix starting at pos? A forward Pike simulation bounded to l_max bytes, stopping at the first match (capture-free: any match is a witness).

Parameters
[in]subThe lookaround sub-program.
[in]posPosition the lookaround is evaluated at.
Returns
True when the sub matches somewhere in the forward window.

◆ lookaround_holds()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::lookaround_holds ( std::uint16_t  sub_id,
std::size_t  pos 
)
inlineconstexpr

Evaluates a bounded lookaround at pos (true if the thread should proceed).

Runs a capture-free Pike simulation of the sub-program on the isolated sub-scratch (state_.lookaround): the main state_ is never touched, so an in-flight match is unaffected. Bounded to l_max bytes (linear per position); (?! / (?<! negate the result.

Parameters
[in]sub_idIndex into prog_.lookarounds.
[in]posThe text position the assertion is evaluated at.
Returns
true if the (possibly negated) assertion holds, so the thread proceeds.

◆ lookaround_state()

template<typename State , bool StateBoundToProgram = false>
lookaround_scratch & real::detail::pike_vm< State, StateBoundToProgram >::lookaround_state ( )
inline

The lookaround sub-scratch, built on first use.

Lazy: search() builds a fresh state, and an eager scratch (two thread lists and a stack) would be built and destroyed on every search of every pattern.

Returns
The scratch, engaged.

◆ lookbehind_matches()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::lookbehind_matches ( std::uint16_t  sub_id,
const lookaround_sub &  sub,
std::size_t  pos 
)
inlineconstexpr

Lookbehind: does the sub-pattern match a window ENDING EXACTLY at pos?

The match must finish precisely at pos (the defining lookbehind trap); a start may lie anywhere in [pos - l_max, pos], pos itself being the empty window. Outside byte mode a start inside a code point can only match the empty window (no sub-program consumes from a continuation byte: code-point ops decode strictly, literals begin with a lead byte, \C forces byte mode), so a thread started at every position is equivalent.

One forward walk per lookbehind (lookaround_scratch::behind_walk) starts a thread at each position and steps them together, so each byte is stepped once per search (per-start windows cost O(l_max^2) per position). A query that moves backward or leaps more than l_max ahead restarts it at pos - l_max.

Parameters
[in]sub_idIndex of the lookaround in prog_.lookarounds.
[in]subThe lookaround sub-program.
[in]posPosition the sub must end exactly at.
Returns
True when some start in the window fullmatches up to pos.

◆ match_byte_klass_run()

template<typename State , bool StateBoundToProgram = false>
template<bool SkipSaves = false>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::match_byte_klass_run ( std::string_view  text,
std::size_t  pc,
std::size_t  s 
) const
inlineconstexpr

Matches the run of byte/klass instructions starting at pc, one text byte each, up to the first non-consuming op. Shared by the fixed-shape and alternation fast paths.

Parameters
[in]textThe subject text.
[in]pcIndex of the first instruction of the run.
[in]sText offset to match from.
Returns
The end offset on a full match, or npos on a mismatch.

◆ match_cp_shape()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::match_cp_shape ( std::string_view  text,
std::size_t  s,
OutSlots &  out_slots 
) const
inlineconstexpr

Verifies a fixed code-point shape forward from s, filling capture slots as it goes.

The shape is a sequence of code-point atoms and literal bytes with no loop (pattern_hints::il_cp_shape_eligible), so one linear walk decides the whole match and every save lands on the position the walk has reached — no engine, and no separate capture pass.

Parameters
[in]textSubject.
[in]sCandidate match start.
[out]out_slotsReceives the slots (untouched unless the walk succeeds).
Returns
The match end, or real::npos if the shape does not hold at s.

◆ match_fixed_body_wb()

template<typename State , bool StateBoundToProgram = false>
template<bool SkipSaves>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::match_fixed_body_wb ( std::string_view  text,
std::size_t  s 
) const
inlineconstexpr

Fixed-shape body match from pattern_hints::body_pc, then B1 \b/\B wrap.

Template Parameters
SkipSavesStep over interleaved save instructions (a grouped shape, whose slots the caller fills from constant offsets).
Parameters
[in]textSubject.
[in]sCandidate match start.
Returns
Match end offset, or npos when the body or the boundary wrap fails.

◆ match_run_shape()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::match_run_shape ( std::string_view  text,
std::size_t  s,
std::size_t  e,
OutSlots &  out_slots 
) const
inlineconstexpr

Fills the groups of a match the DFAs found at [s, e) for a program of run shape, by one walk that takes every loop as far as it goes.

Parameters
[in]textThe subject.
[in]sMatch start.
[in]eMatch end.
[out]out_slotsSlots to fill.
Returns
True when the walk reaches match (its groups are the VM's, and it ends at e, where the VM's path ends); false leaves the answer to the VM.

◆ next_candidate()

template<typename State , bool StateBoundToProgram = false>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::next_candidate ( std::string_view  text,
std::size_t  pos,
std::size_t  start 
) const
inlineconstexpr

First position >= pos that could start a match, per the hints: the prefilter step (literal prefix, rare or unique byte, line start, first-byte set); pos itself when nothing skips.

Parameters
[in]textThe subject text.
[in]posCurrent position.
[in]startThe run's start offset (for one-shot anchored patterns).
Returns
The next candidate offset, or real::npos if none exists.

◆ nibble3_buckets()

template<typename State , bool StateBoundToProgram = false>
static std::uint8_t real::detail::pike_vm< State, StateBoundToProgram >::nibble3_buckets ( const alternation_pairs &  plan,
std::string_view  text,
std::size_t  at 
)
inlinestatic

The fingerprint buckets the three bytes at at admit, as the vector scans compute them for a block: a bucket's bit survives where both nibbles of each byte carry it. Every bucket where fewer than three bytes remain, as the scans leave those starts to the first-byte table.

Parameters
[in]planThe fingerprint (its nibbles valid).
[in]textThe subject.
[in]atThe start.
Returns
The bucket bits.

◆ no_class_loop_above()

template<typename State , bool StateBoundToProgram = false>
static constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::no_class_loop_above ( const pattern_hints &  hints)
inlinestaticconstexprnoexcept

Whether no class loop takes the pattern first: the byte-class loop, the code-point one and the three possessive loops sit above every literal, shape and alternation route in the cascade.

Parameters
[in]hintsThe program's hints.
Returns
True when none of those routes claims it.

◆ prefix_run_end()

template<typename State , bool StateBoundToProgram = false>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::prefix_run_end ( std::string_view  text,
std::size_t  from,
std::size_t  limit 
) const
inlineconstexpr

Where the two-run shape's prefix class run, continued forward from from, stops.

Parameters
[in]textThe subject.
[in]fromA position inside or at the end of the run.
[in]limitWhere to stop at the latest.
Returns
The first position at or after from that is not a member, or limit.

◆ probe_closure()

template<typename State , bool StateBoundToProgram = false>
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::probe_closure ( const instr &  instruction,
std::size_t  pos 
)
inlineconstexpr

Probe of extends_past_end on an epsilon step at pos: an assertion that looks right, a lookahead, or a possessive test whose answer the end of the text decides.

Parameters
[in]instructionThe instruction the closure walk is at.
[in]posThe position.

◆ probe_step()

template<typename State , bool StateBoundToProgram = false>
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::probe_step ( const instr &  instruction,
std::size_t  pos 
)
inlineconstexpr

Probe of extends_past_end on a thread about to consume at pos: one alive at the end of the text, or at a code point the end cuts short, would read what comes next.

Parameters
[in]instructionThe thread's instruction.
[in]posThe position it consumes at.

◆ reads_right()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::reads_right ( const lookaround_sub &  sub) const
inlineconstexpr

Whether sub holds an assertion that reads what follows where it stands: $, \Z, \z, \b, \B, \<, \>.

Parameters
[in]subThe lookaround (they do not nest, so its code is all its own).
Returns
True when the lookaround may read past the bytes it consumes.

◆ replay_literal()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::replay_literal ( std::size_t  cand,
std::size_t  len,
OutSlots &  out_slots 
) const
inlineconstexpr

Fills capture slots for a literal match at cand: replays save instructions at their consumed offsets and checks the chain's zero-width assertions there.

Template Parameters
OutSlotsOutput slot container.
Parameters
[in]candStart offset of the literal match.
[in]lenLength of the literal.
[out]out_slotsReceives the capture slots.
Returns
false (and clears out_slots) if an assertion fails here, so the caller tries the next occurrence; true otherwise.

◆ resolve_class_table()

template<typename State , bool StateBoundToProgram = false>
constexpr const std::uint8_t * real::detail::pike_vm< State, StateBoundToProgram >::resolve_class_table ( std::size_t  class_index)
inlineconstexprprivate

Cold half of class_table (the storage-mode resolution), and the only path that writes the state's row cache for a byte class.

Outlined: every byte beside the row-key compare competes for the budget that lets the accessor inline into basic_match_iterator::advance (inlined, it charged the class-scan rows).

Parameters
[in]class_indexIndex into the program's interned byte classes.
Returns
Pointer to the 256-entry membership row, also cached in the state.

◆ resolve_hi()

template<typename State , bool StateBoundToProgram = false>
void real::detail::pike_vm< State, StateBoundToProgram >::resolve_hi ( std::size_t  cp_index)
inlineprivate

Fills the state's sparse-hi memo for cp_index, the cold half of cp_member_high, outlined so the per-code-point path stays a class-key compare and a bit test.

Parameters
[in]cp_indexIndex of the code-point class to resolve.

◆ row_key_stale()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::row_key_stale ( std::int32_t  have,
std::int32_t  want 
) const
inlineconstexprprivate

Is the state's cached row key stale for want?

With StateBoundToProgram a matching key is proof on its own, and the program-identity compare (a pointer chase per run(), so per match on a walk) compiles away. Without it the compare is required: a state carried across regexes would answer from the previous program's rows.

Parameters
[in]haveThe key the state last verified (table_class or cp_page_class).
[in]wantThe key wanted now.
Returns
true if the row must be re-verified.

◆ run()

template<typename State , bool StateBoundToProgram = false>
template<bool Cascade = false, typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots,
std::size_t  forbid_empty_until = 0,
match_semantics  sem = match_semantics::first 
)
inlineconstexpr

Runs the VM over text starting at start.

On success fills out_slots with byte offsets (npos for unset capture slots; slots 0/1 are the whole match).

Template Parameters
CascadeSelect the memchr-cascade class-run variant (chosen once by the caller from stop_set_size, never per match). Off = the plain hot path, byte for byte.
OutSlotsOutput slot container (resized to the program's slot count).
Parameters
[in]textThe subject text.
[in]startIndex to begin matching/searching from.
[in]modeAnchoring mode (run_mode).
[out]out_slotsReceives the capture slots on success.
[in]forbid_empty_untilReject an empty match starting below this offset (the iterator sets it to the next codepoint boundary, CPython 3.7+ rule). 0 means no restriction.
[in]semMatch semantics: match_semantics::first (default, leftmost-first) or the experimental match_semantics::longest (which forces the general loop, off every fast path).
Returns
true if a match was found.
Note
On gcc/x86 the mere presence of the Aho-Corasick code in this TU slows a class scan that never uses it (loop alignment; instructions identical). Forcing align-loops only moves the regression between class-scan routes. Accepted.

◆ run_aho_corasick()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
bool real::detail::pike_vm< State, StateBoundToProgram >::run_aho_corasick ( std::string_view  text,
std::size_t  start,
OutSlots &  out_slots 
)
inline

Multi-literal search via the automaton ac_ready hands back, cached per regex in detail::regex_immutables.

Search mode only: the automaton's leftmost-first scan IS the candidate search. Runtime only, never instantiated for the static storage's State. noinline, NOT cold (as ac_ready), since inside the body of run() it charges the negated-class rows on x86.

Template Parameters
OutSlotsOutput slot container.
Parameters
[in]textThe subject text.
[in]startIndex to begin at.
[out]out_slotsReceives the matched span on success.
Returns
true if some branch matched.

◆ run_alternation()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_alternation ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inlineconstexpr

Fast path for an alternation of straight-line branches.

Each branch is a fixed-width byte/klass sequence: at a candidate the branches are tried in source order (read from the split chain) and the first that matches wins, as the VM's thread priority.

Template Parameters
OutSlotsOutput slot container.
Parameters
[in]textThe subject text.
[in]startIndex to begin at.
[in]modeAnchoring mode.
[out]out_slotsReceives the matched span on success.
Returns
true if some branch matched.
Note
At density this route is almost entirely per-match return (a straight line across five densities), which fill_alternation_spans batches. New filler bodies have charged unrelated rows (docs/MEASUREMENT.md §5.4, §5.5): judge changes on BOTH instruments.

◆ run_alternation_wide()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
std::optional< bool > real::detail::pike_vm< State, StateBoundToProgram >::run_alternation_wide ( std::string_view  text,
std::size_t  start,
OutSlots &  out_slots,
cp_span *  spans = nullptr,
std::size_t  cap = 0,
std::size_t *  filled = nullptr 
)
inline

Search route for an alternation of literals with more first bytes than the small set holds: the blocks the nibble fingerprint marks, verified in branch order (priority unchanged), then the last bytes by the first-byte table. Taken per subject on a sample: where false candidates are dense enough that verifying them costs more than the automaton's walk, it declines to the automaton's gate. Out of line: run is shared by every route.

Template Parameters
OutSlotsOutput slot container.
Parameters
[in]textThe subject.
[in]startWhere the search starts.
[out]out_slotsThe span, on a match (a single search).
[out]spansNon-null for a batched walk: the matches from start, each from the previous one's end, instead of one search's span.
[in]capCapacity of spans.
[out]filledWith spans: how many were written, fewer than cap only at the end of the subject.
Returns
Matched or not when the route took the search (with spans: whether any was written); empty when it declined.

◆ run_bounded_backtrack()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
bool real::detail::pike_vm< State, StateBoundToProgram >::run_bounded_backtrack ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inline

The general loop's answer, by backtracking under a bit per (instruction, position).

Walks the program depth first in the VM's priority order (a split's preferred branch first, each start in turn), marking every (instruction, position) entered; a marked pair prunes the branch, as the VM's list drops a present thread. The first match reached is the VM's answer. A jump into a loop head already entered at this position takes the loop's exit, as in the VM: this position's marks are the VM's seen set there, since only the walk at a position marks it.

Marks persist across starts, and starts the prefilter rules out are skipped. Neither changes the answer: pairs an exploration marked without matching are closed under every transition (a split holds both branches; a jump exits a loop only when its head and body were entered), so none reaches a match. Each pair is entered at most once: O(n x m), the caller holding n x m under bounded_backtrack_bits.

Parameters
[in]textSubject (already in text_).
[in]startByte offset to begin at.
[in]modeAnchoring: full, prefix or search.
[out]out_slotsCapture slots, filled on a match.
Returns
True on a match.

◆ run_cascade_stop()

template<typename State , bool StateBoundToProgram = false>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::run_cascade_stop ( std::string_view  text,
std::size_t  from 
) const
inlineconstexprprivate

The memchr-cascade run tail: the next stop byte at or after from, or the text end. Its own function so the cascade never inlines into the per-byte loop of run_class_loop (that bloat slowed stop-dense short runs); reached only past cascade_run_threshold bytes.

Parameters
[in]textSubject.
[in]fromOffset to search from.
Returns
Offset of the next stop byte, or text.size() when none remains.

◆ run_class_loop()

template<typename State , bool StateBoundToProgram = false>
template<bool Cascade, typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_class_loop ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inlineconstexprprivate

Fast path for a whole-pattern "class+": a maximal run of class bytes in one scan loop, exactly the VM's greedy result, with no thread lists.

No-lookaround path only: trailing-lookaround class+ goes to run_class_loop_trailing_la from outside run. always_inline: on x86 an out-of-line call costs a double-digit share of a match-dense find_iter walk.

Template Parameters
CascadeTake the memchr-cascade run tail (chosen once per walk from stop_set_size).
OutSlotsOutput slot container.
Parameters
[in]textThe subject text.
[in]startIndex to begin at.
[in]modeAnchoring mode.
[out]out_slotsReceives the (start, end) span on success.
Returns
true if a non-empty run was found.

◆ run_class_loop_anchored()

template<typename State , bool StateBoundToProgram = false>
template<bool Cascade, typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_class_loop_anchored ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inlineconstexprprivate

Cold half of the class-loop route: everything a \A/^ or \Z/$ implies.

Outlined so the unanchored path pays exactly one branch. \A/^ is a MODE (search becomes prefix; a region past 0 cannot match); \Z/$ is a LIMIT (run_class_loop_end_anchored). Both were peeled out of the program, so this is all that enforces them.

Template Parameters
CascadeWhether the memchr stop-tail applies.
OutSlotsOutput slot container.
Parameters
[in]textThe subject.
[in]startRegion start.
[in]modeAnchoring mode as the caller asked for it.
[out]out_slotsReceives the span on success.
Returns
true on a match.

◆ run_class_loop_end_anchored()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_class_loop_end_anchored ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inlineconstexprprivate

X+$ / ^X+$ in search mode: the run that ENDS at the anchor, found by walking back.

A trailing \Z/$ pins the end, so the leftmost match is the maximal class run finishing there: one backward walk from the limit.

\Z (kind 1) is the strict end; $ (kind 2) also matches before ONE final newline. A class holding \n (\s+$) consumes it and ends at the true end (\s$ over "ab\n" is (2, 3)), so the newline is stripped only when the class cannot hold it.

Soundness of picking one limit: this route arms only an UNBOUNDED greedy run (+, {k,}), which over a class holding \n always reaches the true end; a bounded one would not ([ \t\n]{1,2}$ over " \n\n" is (0, 2)). The route-vs-general product test catches a wrong choice.

Template Parameters
OutSlotsOutput slot container.
Parameters
[in]textThe subject.
[in]startRegion start; the match may not begin before it.
[in]modeAnchoring mode: search, prefix or full.
[out]out_slotsReceives the span on success.
Returns
true when a run ends at the anchor.

◆ run_class_loop_trailing_la()

template<typename State , bool StateBoundToProgram = false>
template<bool Cascade, typename OutSlots >
bool real::detail::pike_vm< State, StateBoundToProgram >::run_class_loop_trailing_la ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inline

Trailing-lookaround class+: body scan + longest end where lookaround holds.

Called from real.hpp / find_iter outside run, once per match. Only this selector inlines into the caller; each walk is out of line, since it must not share a body or inlining unit with run_class_loop (the hot [a-z]+ path). A cold, out-of-line selector makes every match two calls, since clang keeps the walk out of it. Dynamic-only. A code-point body (pattern_hints::trailing_la_cp) walks whole code points: a run holds only valid ones, so its candidate ends are the bytes that are not UTF-8 continuations.

Parameters
[in]textSubject.
[in]startByte offset to begin at.
[in]modeAnchoring: full, prefix or search.
[out]out_slotsCapture slots, filled on a match.
Returns
True on a match.

◆ run_codepoint_class()

template<typename State , bool StateBoundToProgram = false>
template<bool Cascade, typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_codepoint_class ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inlineconstexpr

Fast path for . / a negated class, optionally a greedy +.

Scans code points as the VM's byte-level expansion would: an ASCII byte matches the ASCII set; a valid 2–4 byte UTF-8 sequence always matches (a negated ASCII class excludes only ASCII); anything else stops, as the VM's lead/continuation branches fail. Covers .+, [^,]+, ., [^,].

Template Parameters
OutSlotsOutput slot container.
Parameters
[in]textThe subject text.
[in]startIndex to begin at.
[in]modeAnchoring mode.
[out]out_slotsReceives the matched span on success.
Returns
true if at least one codepoint matched.

◆ run_cp_class_loop()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_cp_class_loop ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inlineconstexpr

Fast path for a whole-pattern code-point class klass_cp, optionally a greedy +.

Scans code points directly against the class predicate (ASCII bitmap below 0x80, range binary search above), advancing by the code point's byte width, with no thread lists — the analog of run_class_loop for a Unicode shorthand (\w+, \d+, \s+). A malformed sequence stops the run, exactly as the VM's klass_cp fails on it.

Template Parameters
OutSlotsOutput slot container.
Parameters
[in]textThe subject text.
[in]startIndex to begin at.
[in]modeAnchoring mode.
[out]out_slotsReceives the matched span on success.
Returns
true if a non-empty run was found.

Whether a class member starts at i: membership only, no width.

A strict decode of a byte below 0x80 is {lead, 1, valid}, so asc[lead] is the answer, like in_class in run_class_loop. Kept apart from width, which extend_run needs for the length (asking width for the bit made this scan cost several times the byte-class route's).

◆ run_exact_literal()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_exact_literal ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inlineconstexpr

Fast path for a pure-literal pattern.

The prefilter locates the bytes; this replays saves directly, with no thread lists. A leading or trailing zero-width assertion (\b, ^, $ …) may fail an occurrence, so search mode scans successive occurrences until they hold (\B2 on "220").

Template Parameters
OutSlotsOutput slot container.
Parameters
[in]textThe subject text.
[in]startIndex to begin at.
[in]modeAnchoring mode.
[out]out_slotsReceives the capture slots on success.
Returns
true if a match was found.
Note
A one-byte literal is NOT redirected to the batched single-class route: dense bytes favour the batched class, sparse ones memchr. That needs a density gate (like ac_density_favours_automaton), not a recognition-time redirect.

◆ run_fixed_shape()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_fixed_shape ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inlineconstexpr

Fast path for a whole-pattern fixed-width byte/klass sequence.

A straight-line program (no branches/assertions) has one thread: each byte/klass instruction consumes one byte, verified by a single walk. Covers class{n} and \d{4}-\d{2}-\d{2}.

Template Parameters
OutSlotsOutput slot container.
Parameters
[in]textThe subject text.
[in]startIndex to begin at.
[in]modeAnchoring mode.
[out]out_slotsReceives the matched span on success.
Returns
true if the sequence matched.

◆ run_general()

template<typename State , bool StateBoundToProgram = false>
template<bool Cascade = false, bool Probe = false, typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_general ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots,
std::size_t *  forward_stop = nullptr 
)
inlineconstexpr

The general Pike VM loop, also run by the lazy-DFA route on the [s, e] window its two passes located.

Parameters
[in]textSubject.
[in]startByte offset to begin at.
[in]modeAnchoring: full, prefix or search.
[out]out_slotsCapture slots, filled on a match.
[out]forward_stopWhen non-null, receives how far the forward scan reached — the inner-literal route's linearity backstop.
Returns
True on a match.

◆ run_inner_literal()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
bool real::detail::pike_vm< State, StateBoundToProgram >::run_inner_literal ( std::string_view  text,
std::size_t  start,
OutSlots &  out_slots,
bool &  abandon,
bool  density_gate = true 
)
inline

The inner-literal search: memmem a required literal, reverse-match the prefix to the match start, forward-confirm — the reverse-inner protocol (regex-automata's ReverseInner).

Search mode only, runtime only (the reverse DFA is not constexpr). Two guards keep it linear: the reverse is bounded below by min_match_start (the previous literal's end), and a literal starting before min_pre_start (the last confirm's forward reach) abandons the scan.

Parameters
[in]textSubject.
[in]startByte offset to begin the scan at.
[out]out_slotsCapture slots, filled on a match.
[out]abandonSet when a linearity guard trips, so the caller retries the whole search on the core VM.
[in]density_gateWhether to consult the candidate-density gate. False only for the batched filler, whose candidates mostly complete: the gate counts before confirming.
Returns
True on a match; false on none, and false with abandon set when the route gave up.

◆ run_literal_one_search()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
bool real::detail::pike_vm< State, StateBoundToProgram >::run_literal_one_search ( std::string_view  text,
std::size_t  start,
std::size_t  len,
OutSlots &  out_slots 
)
inline

The whole exact-literal search in one find_prefix, for a pattern_hints::literal_one_search program (called from run_exact_literal).

noinline on a HOT path: inside run_exact_literal it grows a function sharing an inlining unit with run and the class loops, and the growth alone charges run_codepoint_class (the hazard documented on run). Out of line it costs a short literal one call.

Template Parameters
OutSlotsOutput slot container.
Parameters
[in]textThe subject text.
[in]startIndex to begin searching at.
[in]lenThe literal's length (hints.exact_literal_len, >= 2 by the hint).
[out]out_slotsReceives [cand, cand + len] on success.
Returns
true if the literal occurs at or after start.
Note
The span filler (fill_exact_literal_spans) must leave count_matches unchanged: recompiling that shared entry point (every row measures through it) charges other rows. Judge a filler change on machine code first (function sizes in the consumer unit), then on layout.

◆ run_pair_filtered_shape()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
bool real::detail::pike_vm< State, StateBoundToProgram >::run_pair_filtered_shape ( std::string_view  text,
std::size_t  start,
OutSlots &  out_slots 
)
inline

Search route for a HETEROGENEOUS fixed shape: vector-prefilter two positions, verify each survivor with the ordinary fixed-body walk, hand the sub-block tail to fast_search.

Its own route: hosted inside run_fixed_shape (inline or noinline), it cost instructions on shapes that never enter it ([0-9]{2}:[0-9]{2}). noinline so the body of run does not grow. Search mode only.

Template Parameters
OutSlotsOutput slot container.
Parameters
[in]textThe subject text.
[in]startIndex to begin searching at.
[out]out_slotsReceives the matched span on success, npos on failure (seam parity with run_fixed_shape, through fail_slots).
Returns
true if the sequence matched.

◆ run_possessive_byte_loop()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_possessive_byte_loop ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inlineconstexpr

Possessive literal-byte +/++ loop (byte_loop_possessive, e.g. a++), on the shared algorithm of run_possessive_loop_generic.

Parameters
[in]textSubject.
[in]startByte offset to begin at.
[in]modeAnchoring: full, prefix or search.
[out]out_slotsCapture slots, filled on a match.
Returns
True on a match.

◆ run_possessive_class_loop()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_possessive_class_loop ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inlineconstexpr

Possessive class+/++ loop over a BYTE class (klass_loop_possessive). See run_possessive_loop_generic for the shared algorithm.

Parameters
[in]textSubject.
[in]startByte offset to begin at.
[in]modeAnchoring: full, prefix or search.
[out]out_slotsCapture slots, filled on a match.
Returns
True on a match.

◆ run_possessive_cp_class_loop()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_possessive_cp_class_loop ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inlineconstexpr

Possessive class+/++ loop over a CODE-POINT class (klass_cp_loop_possessive), on the decode/membership primitives of run_cp_class_loop (the scan predicate differs per compiler, see in_class). See run_possessive_loop_generic for the shared algorithm.

Parameters
[in]textSubject.
[in]startByte offset to begin at.
[in]modeAnchoring: full, prefix or search.
[out]out_slotsCapture slots, filled on a match.
Returns
True on a match.

◆ run_possessive_loop_generic()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots , typename InClass , typename ScanEnd , typename LastWidth >
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_possessive_loop_generic ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots,
const InClass &  in_class,
const ScanEnd &  scan_end,
const LastWidth &  last_width 
)
inlineconstexpr

Shared driver: a possessive class+/++ loop, bare/suffixed (pattern_hints::possessive_prefix_size == 0) or delimited/"quoted" (non-zero); the body's membership comes from in_class / scan_end, so it serves byte- and code-point classes.

A possessive run never gives back, so a required literal SUFFIX (or, delimited, the closing SUFFIX) may fail to follow, with nothing to retry within the attempt. In search mode the retry skips to the failed attempt's body end: sound and linear PROVIDED the eligibility pattern_hints documents held at recognition (prefilter.hpp), since every candidate strictly inside the run fails identically.

Template Parameters
InClassbool(std::size_t) -> true if the body's class/cp-class accepts the byte/code point starting at that offset.
ScanEndstd::size_t(std::size_t from) -> end of the maximal body run starting at from (from itself when no run starts there).
LastWidthstd::size_t(std::size_t end) -> width (in bytes) of the LAST atom of a non-empty run ending at end: 1 for a byte or byte-class body, a backward UTF-8 decode (codepoint_retreat) for a code-point class. Called only with end past the run's start.
Parameters
[in]textSubject.
[in]startByte offset to begin at.
[in]modeAnchoring: full, prefix or search.
[out]out_slotsCapture slots, filled on a match.
[in]in_classMembership test, per InClass.
[in]scan_endMaximal-run scanner, per ScanEnd.
[in]last_widthLast-atom width, per LastWidth.
Returns
True on a match.

◆ run_shape_atom()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::run_shape_atom ( std::string_view  text,
std::size_t  pc,
std::size_t &  at,
std::size_t  e 
) const
inlineconstexpr

Consumes the atom at pc at at, within e.

Parameters
[in]textThe subject.
[in]pcThe atom's instruction.
[in,out]atThe position; advanced past the atom when it matches.
[in]eThe window's end.
Returns
True when the atom matched.

◆ run_shape_loop_atom()

template<typename State , bool StateBoundToProgram = false>
static constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::run_shape_loop_atom ( const program_view &  prog,
std::size_t  begin,
std::size_t  end 
)
inlinestaticconstexpr

The one atom of a loop body [begin, end) that holds nothing else but saves – a group around one atom, repeated: ([aeiou])+.

Parameters
[in]progThe program.
[in]beginThe body's first instruction (the loop split's preferred target).
[in]endThe loop split.
Returns
The atom's instruction, or real::npos when the body is anything else.

◆ seed_viable()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::seed_viable ( std::string_view  text,
std::size_t  pos,
std::size_t  start 
) const
inlineconstexpr

Cheap pre-check before seeding a new thread at pos: live threads may drag the loop through positions the prefilter would skip. Also enforces code-point alignment in text mode.

Parameters
[in]textThe subject text.
[in]posThe candidate seed position.
[in]startThe run's start offset.
Returns
true if a fresh thread should be seeded at pos.

◆ single_atom_body()

template<typename State , bool StateBoundToProgram = false>
constexpr const instr * real::detail::pike_vm< State, StateBoundToProgram >::single_atom_body ( const lookaround_sub &  sub) const
inlineconstexpr

The one consuming instruction of a lookaround body that is a single atom, or null.

Such a body is [byte | klass; match] (two slots) or, in text mode, [klass_cp; three continuation slots; match] (five: the compiler always emits the chain), so a code-point class is matched by klass_cp alone.

Parameters
[in]subThe lookaround sub-program.
Returns
The instruction to test directly, or null when the body needs the sub-simulation.

◆ single_class_ahead()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::single_class_ahead ( const instr &  body,
std::size_t  pos 
)
inlineconstexpr

L1 peephole — does the single consuming op body match the code point / byte AT pos (ahead)? Mirrors the per-op logic of lookahead_matches for a one-instruction sub-program.

Parameters
[in]bodyThe sub-program's single consuming instruction.
[in]posPosition the lookaround is evaluated at.
Returns
True when body accepts what starts at pos; false at the text end.

◆ single_class_behind()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::single_class_behind ( const instr &  body,
std::size_t  pos 
)
inlineconstexpr

L1 peephole — does body match the code point / byte ending EXACTLY at pos (behind)? The defining lookbehind trap: the match must END at pos, so the code point is the one whose aligned start s gives s + length == pos (byte mode: pos - 1).

Parameters
[in]bodyThe sub-program's single consuming instruction.
[in]posPosition the lookaround is evaluated at.
Returns
True when body accepts the atom ending at pos; false at the text start.

◆ step()

template<typename State , bool StateBoundToProgram = false>
template<bool Probe = false, typename OutSlots >
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::step ( list_type &  clist,
list_type &  nlist,
std::size_t  pos,
run_mode  mode,
bool &  matched,
OutSlots &  out_slots 
)
inlineconstexpr

Advances every thread of clist by the byte at pos; survivors land in nlist. A thread reaching match records its slots and cuts all lower-priority threads (leftmost-greedy order).

Template Parameters
OutSlotsOutput slot container.
Parameters
[in,out]clistThe current thread list (consumed).
[in,out]nlistThe next thread list (receives survivors).
[in]posThe current input position.
[in]modeAnchoring mode (affects match acceptance).
[in,out]matchedSet to true when a match is recorded.
[out]out_slotsReceives the slots of an accepted match.

◆ sub_add_thread()

template<typename State , bool StateBoundToProgram = false>
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::sub_add_thread ( thread_list &  list,
std::int32_t  pc0,
std::size_t  pos,
bool &  matched 
)
inlineconstexpr

Epsilon-closure for the lookaround sub-VM, on the isolated sub-scratch.

Parks consuming pcs in list and sets matched on reaching the sub's match. A capture-free sub emits no save and no assert_lookaround (nesting is rejected at compile time). Touches only state_.lookaround->stack. Linear: mark_seen dedups within a generation, so each assert_lookaround is evaluated at most once per position (a lookahead costs O(L) there; a lookbehind advances its walk by the bytes since its last query).

Parameters
[in,out]listThe sub thread list to populate.
[in]pc0The sub-program counter to seed from.
[in]posThe current input position (for assertions).
[in,out]matchedSet to true if the sub's match is reachable here.

◆ thread_slots()

template<typename State , bool StateBoundToProgram = false>
constexpr const std::size_t * real::detail::pike_vm< State, StateBoundToProgram >::thread_slots ( list_type &  clist,
std::size_t  i 
)
inlineconstexpr

Pointer to thread i's slot_count capture values (its COW block), read by match.

Parameters
[in]clistList holding the thread.
[in]iThread index within clist.
Returns
Pointer to the thread's first capture slot.

◆ tier1_capture_on_match()

template<typename State , bool StateBoundToProgram = false>
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::tier1_capture_on_match ( list_type &  clist,
std::size_t  i,
std::int32_t  capture_start_slot,
std::size_t  start,
std::size_t  end 
)
inlineconstexpr

Tier 1's on-match capture write: if capture_start_slot is not -1, records [start, end) into thread i's capture block, in place.

Called ONLY on a confirmed atom match, never before the test: a possessive loop always attempts one more repetition, so an early save would tear the capture into [next attempt's start, this iteration's end). See program.hpp's opcode-family note.

Parameters
[in,out]clistThe current thread list (whose slot this thread owns is updated).
[in]iIndex of the thread in clist.
[in]capture_start_slotThe start slot, or -1 for an uncaptured Tier 1 loop (a no-op).
[in]startPosition before the atom was consumed.
[in]endPosition after the atom was consumed.

◆ trailing_la_walk()

template<typename State , bool StateBoundToProgram = false>
template<bool Cascade, bool Cp, typename OutSlots >
bool real::detail::pike_vm< State, StateBoundToProgram >::trailing_la_walk ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inline

The body of run_class_loop_trailing_la for one body kind.

Out of line but not cold: optimized for size, the walk runs more instructions on GCC.

Template Parameters
CascadeWhether the byte walk may take its memchr-cascade tail.
CpThe body is a klass_cp (whole code points) rather than a klass.
Parameters
[in]textSubject.
[in]startByte offset to begin at.
[in]modeAnchoring: full, prefix or search.
[out]out_slotsCapture slots, filled on a match.
Returns
True on a match.

◆ try_shared_lazy_dfa_search()

template<typename State , bool StateBoundToProgram = false>
template<bool Cascade, typename OutSlots >
std::optional< bool > real::detail::pike_vm< State, StateBoundToProgram >::try_shared_lazy_dfa_search ( std::string_view  text,
std::size_t  start,
run_mode  mode,
OutSlots &  out_slots 
)
inlineprivate

Lazy-DFA search route on the shared confirm DFAs. noinline: inlined, its body inflates the x86 class-loop codegen of run (as ac_ready).

Parameters
[in]textSubject.
[in]startByte offset to begin at.
[in]modeAnchoring: full, prefix or search.
[out]out_slotsCapture slots, filled on a match.
Returns
matched / no-match when the route handled the search; empty when the caller must fall to Pike.

◆ unbounded_lookahead_matches()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::unbounded_lookahead_matches ( std::uint16_t  sub_id,
const lookaround_sub &  sub,
std::size_t  pos 
)
inlineconstexpr

Unbounded lookahead: does the sub-pattern match a prefix of the text from pos?

A sub-pattern with no bound (.*, +, {n,}) run forward from every position is quadratic. Whether it matches from a position depends only on the text after it, so one pass from the end answers every position: row pos says, per sub-program instruction, whether match is reachable from it at pos. A consuming instruction (a code-point test steps its continuation chain byte by byte) depends only on row pos + 1; match holds; epsilon instructions (jump, split, a position assertion at pos) propagate within the row. Once per subject, O(n x m), into lookaround_scratch::ahead_table.

Parameters
[in]sub_idIndex of the lookaround in prog_.lookarounds.
[in]subThe lookaround sub-program (l_max < 0).
[in]posPosition the lookahead is evaluated at.
Returns
True when the sub-pattern matches from pos.

◆ variant_candidate()

template<typename State , bool StateBoundToProgram = false>
std::size_t real::detail::pike_vm< State, StateBoundToProgram >::variant_candidate ( std::string_view  text,
std::size_t  pos,
const alternation_pairs &  plan 
) const
inline

The next start at or after pos the variants' fingerprint admits (the first bytes, in the last blocks it cannot read past).

Parameters
[in]textThe subject.
[in]posWhere to look from.
[in]planThe fingerprint (variant_plan).
Returns
The start, or real::npos when none is left.

◆ variant_plan()

template<typename State , bool StateBoundToProgram = false>
const alternation_pairs * real::detail::pike_vm< State, StateBoundToProgram >::variant_plan ( std::string_view  text,
std::size_t  start 
)
inline

The variants' fingerprint for a program that is not a fixed alternation (branches holding a case-folded i, s or k, folding to non-ASCII), when this subject's sample finds its first bytes dense and the fingerprint's candidates among them sparse. Taken only where next_candidate would scan by the first bytes; decided once per subject. The fingerprint admits every start a match can have and a few more, which the confirming anchored walk rejects. No candidate lands inside a code point: a continuation byte's high nibble is no first byte's.

Parameters
[in]textThe subject.
[in]startWhere the search starts.
Returns
The plan, or null to scan by the first bytes.

◆ verify_class_row()

template<typename State , bool StateBoundToProgram = false>
void real::detail::pike_vm< State, StateBoundToProgram >::verify_class_row ( detail::regex_immutables &  cache,
std::size_t  class_index 
)
inlineprivate

Verifies (and if needed fills) the byte row for class_index, then caches it in the state.

Must stay outlined: inlined, it pushes class_table past what inlines into basic_match_iterator::advance (out of line, class_table costs a tenth of a class-loop walk).

Parameters
[in,out]cacheThe per-regex immutables.
[in]class_indexIndex into the program's interned byte classes.

◆ walk_on_scan_set()

template<typename State , bool StateBoundToProgram = false>
template<typename Walk >
bool real::detail::pike_vm< State, StateBoundToProgram >::walk_on_scan_set ( const Walk &  walk)
inlineprivate

The forward walk of with_search_dfas, on the set an inner-literal scan already leased.

Parameters
[in]walkCallable taking (lazy_dfa& fwd).
Returns
True when walk ran; false when the route must stay on the Pike VM.

◆ wb_boundaries_ok()

template<typename State , bool StateBoundToProgram = false>
constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::wb_boundaries_ok ( std::size_t  s,
std::size_t  e 
) const
inlineconstexpr

O(1) lead/trail \b/\B check at match bounds [s, e), for every wb-wrapping fast path (hints 0/1/2 from pattern_hints::wb_lead / pattern_hints::wb_trail).

Parameters
[in]sMatch start (lead assert position).
[in]eMatch end (trail assert position).
Returns
true if both configured boundaries hold (or are unset).

◆ window_cut_before()

template<typename State , bool StateBoundToProgram = false>
static constexpr bool real::detail::pike_vm< State, StateBoundToProgram >::window_cut_before ( std::string_view  text,
std::size_t  start,
std::size_t  candidate 
)
inlinestaticconstexpr

Whether no whole code point lies between start and candidate: the candidate IS start, or only continuation bytes separate them.

The DROP rule's window-edge guard asks it of a search's first candidate: one reached past a whole non-class character satisfies a dropped leading \b, but when the window begins inside a code point the scan crosses only its tail, and the character before the candidate (unseen, it started before the window) may be a word character.

Parameters
[in]textThe subject.
[in]startThe window's start.
[in]candidateThe first candidate, at or after start.
Returns
True when the character before candidate is not one the scan crossed.

◆ window_mode()

template<typename State , bool StateBoundToProgram = false>
static constexpr run_mode real::detail::pike_vm< State, StateBoundToProgram >::window_mode ( run_mode  mode)
inlinestaticconstexprprivatenoexcept

The mode the VM fills a window's groups in once the DFAs proved where the match starts.

Parameters
[in]modeThe search's own mode.
Returns
run_mode::prefix for a search, anchored at the proved start; any other mode unchanged.

◆ with_search_dfas()

template<typename State , bool StateBoundToProgram = false>
template<typename Fn >
bool real::detail::pike_vm< State, StateBoundToProgram >::with_search_dfas ( Fn &&  fn)
inlineprivate

Run fn with this thread's search DFAs for the regex (see dfa_lease), taking no lock.

Parameters
[in]fnCallable taking (lazy_dfa& fwd, reverse_dfa& rev).
Returns
True when fn ran; false when the route must stay on the Pike VM (no immut / ineligible).

◆ write_cp_span_slots()

template<typename State , bool StateBoundToProgram = false>
template<typename OutSlots >
constexpr void real::detail::pike_vm< State, StateBoundToProgram >::write_cp_span_slots ( OutSlots &  out_slots,
std::size_t  s,
std::size_t  e 
)
inlineconstexpr

Writes a buffered span into a caller's slots exactly as the per-match path would.

Not fill_span_slots called directly: that writer is always_inline, and expanded into the batched walk's emission it charges Python's sub under Apple clang, which instruction counts do not show.

Parameters
[out]out_slotsSlots to fill.
[in]sMatch start.
[in]eMatch end.

Member Data Documentation

◆ ac_completion_pct

template<typename State , bool StateBoundToProgram = false>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::ac_completion_pct {15}
staticconstexprprivate

Percentage of sampled candidates that may COMPLETE a branch and still leave the automaton ahead. Above it the cascade wins whatever the candidate density says.

A false start costs the cascade (verify, reject, resume), not the automaton; a match favours the cascade (it stops there) and charges the automaton a per-match return: at fixed density the completed fraction alone flips the verdict (benchmarks/ac_regime.cpp). The more conservative ISA's balance point: it may decline a win, never take a loss.

◆ ac_density_sample_bytes

template<typename State , bool StateBoundToProgram = false>
constexpr std::size_t real::detail::pike_vm< State, StateBoundToProgram >::ac_density_sample_bytes {256}
staticconstexprprivate

AC routing: sample window, and the candidate-work product at or above which the automaton beats the memchr cascade.

The automaton scans at a flat rate while the cascade spans two orders of magnitude on the same pattern and length, and the crossover moves with branch count (the cascade tries branches in order): the rule is the product (candidates per 1000 bytes) * branch_count (benchmarks/ac_regime.cpp). Density cannot tell a false start from a completed match, which pull in opposite directions, hence ac_completion_pct too. Do not retune these constants against that sweep (it moves the error).

The constant is under the lowest crossover measured over both ISAs: at or above ac_branch_threshold branches, switching early is the safe error, switching late forfeits a win.

◆ il_density_probe_candidates

template<typename State , bool StateBoundToProgram = false>
constexpr std::uint32_t real::detail::pike_vm< State, StateBoundToProgram >::il_density_probe_candidates {8}
staticconstexprprivate

Inner-literal density gate: candidates sampled across the haystack before the verdict.

Above the crossover every candidate is a failed confirm and the core scan wins by a growing margin; the threshold sits conservatively above it, so sparse IL wins (≪ 10/1000) stay. Capture-free only (slot_count ≤ 2): with groups, IL still beat the forced DFA on dense input.

Note
Calibrated against the DFA fallback ((?:\w+)_(?:\w+)): the crossover moves with the fallback's cost, so one threshold cannot fit every fallback (the fixed-shape one is excluded by route condition). A fix needs that cost as a second variable.

The documentation for this class was generated from the following file: