The per-regex immutable cache the router shares across every find_iter on a regex: the byte program (klass_cp expanded to the deterministic trie) and, when the pattern is one-pass, the extractor table.
More...
|
| bool | row_ready (std::size_t i) const noexcept |
| | Reads the "filled" flag for flag index i, acquiring what the filling thread released.
|
| |
| void | set_row_ready (std::size_t i) noexcept |
| | Publishes the "filled" flag for flag index i.
|
| |
|
| regex_immutables ()=default |
| | An empty cache: every identity key null, nothing built.
|
| |
| | regex_immutables (const regex_immutables &) noexcept |
| | Copies as an EMPTY cache: a copied regex is an independent regex.
|
| |
|
| regex_immutables (regex_immutables &&) noexcept |
| | Moves as an empty cache, for the same reason as the copy constructor.
|
| |
| void | invalidate_all () noexcept |
| | Clears EVERY identity key, so nothing built for the old program survives an assignment.
|
| |
| regex_immutables & | operator= (const regex_immutables &) noexcept |
| | Keeps this object's cache STORAGE but marks it invalid.
|
| |
| regex_immutables & | operator= (regex_immutables &&) noexcept |
| | Invalidates as the copy assignment does, and for the same reason.
|
| |
|
constexpr | ~regex_immutables () |
| | Erases this regex's shared DFA slot at run time; constant evaluation skips the map, so dynamic_storage::compile and static_assert stay valid.
|
| |
|
|
byte_program | byte_prog |
| | klass_cp-expanded byte program (empty until built).
|
| |
|
lazy_byte_alphabet | alphabet |
| | byte-class alphabet of byte_prog, shared by both DFAs.
|
| |
|
byte_program | look_prog |
| | byte program keeping its position assertions, built only when byte_prog declined: the search DFAs run it.
|
| |
|
lazy_byte_alphabet | look_alphabet |
| | byte-class alphabet of look_prog.
|
| |
| bool | run_shape {false} |
| |
|
std::atomic< std::uint32_t > | prefix_calls {0} |
| | Anchored matches run before op_table was built for them, counted to onepass_prefix_warm_calls. Relaxed: a lost increment only delays the build.
|
| |
|
std::optional< onepass > | op_table |
| | one-pass extractor, present iff the pattern is one-pass.
|
| |
|
byte_program | il_prefix_prog |
| | IL: the inner-literal prefix's byte program (ineligible until built); per regex, so the reverse DFA over it is shared.
|
| |
| std::size_t | il_min_haystack {} |
| |
|
std::atomic< std::size_t > | il_short_bytes {0} |
| | Bytes of subjects the inner-literal route declined below its floor before the immutables were built, counted to il_short_scan_budget. Relaxed: a lost addition only delays the build.
|
| |
| std::vector< std::uint8_t > | class_rows |
| | Byte-indexed membership rows, filled on first use of each class and kept for the regex's life.
|
| |
|
std::vector< std::uint8_t > | cp_ascii_rows |
| | One 256-byte row per cp_class: its ASCII half.
|
| |
|
std::vector< std::uint64_t > | cp_page_rows |
| | One 30-word bitmap per cp_class: [U+0080, U+07FF].
|
| |
| std::atomic< std::uint64_t > | row_ready_bits {0} |
| | "Row filled" flags: one bit per row, the three runs packed into one word, plus an overflow vector for indices that do not fit.
|
| |
|
std::vector< std::atomic< char > > | row_ready_overflow |
| | Flags for row indices at or past row_ready_bit_capacity.
|
| |
|
std::size_t | cp_ascii_ready_at {0} |
| | Where the cp_ascii run starts in the flag index space.
|
| |
| std::size_t | cp_page_ready_at {0} |
| |
|
std::atomic< const void * > | rows_for {nullptr} |
| | prog.code.data() the rows above were sized for, or null. Independent of built_for, as scan routes that never build the DFA caches need the rows.
|
| |
|
std::vector< ac_automaton > | ac |
| | The multi-literal automaton for a fixed_alternation past the branch threshold, or empty when never built or declined (a pathological icase-fold expansion). Per regex, not per state: a state is fresh per search() and would rebuild it every call. A vector of at most one element, not a unique_ptr, keeps this type literal and the 432-byte header off regexes that never build one.
|
| |
|
std::atomic< const void * > | ac_for {nullptr} |
| | prog.code.data() ac was built for, or null. Not folded into built_for, since only the alternation route consults the automaton (see op_table_for).
|
| |
|
std::vector< alternation_pairs > | alt_pairs |
| | The alternation's probe pairs with their splats, or empty when never built. Per regex, not per state: in the per-search() state gcc would zero the 512 bytes of splats every call (check-state-zeroing). At most one element, on the heap, as ac.
|
| |
|
std::atomic< const void * > | alt_pairs_for {nullptr} |
| | prog.code.data() alt_pairs was built for, or null (own identity atomic, as ac_for).
|
| |
|
std::atomic< const void * > | op_table_for {nullptr} |
| | prog.code.data() op_table was built for, or null. Kept out of built_for because the extractor costs more than the byte program and lazy DFA together, and only routes that fill captures through it need it.
|
| |
|
std::atomic< const void * > | built_for {nullptr} |
| | prog.code.data() this cache was built for, or null if never built / invalidated. Hot path: one atomic load. Not once_flag — assignment reuses this object under a new program; a spent once_flag would never rebuild (silent wrong matches).
|
| |
The per-regex immutable cache the router shares across every find_iter on a regex: the byte program (klass_cp expanded to the deterministic trie) and, when the pattern is one-pass, the extractor table.
Each product is keyed by program identity (built_for and its siblings), so a const regex shared by threads builds race-free and an assignment onto a warmed regex rebuilds. The mutable lazy-DFA caches live in a process-wide side table keyed by this object's address (shared_dfa_slot), one set per scanning thread (dfa_lease), which keeps std::mutex out of this struct. Each rebuild in pike_vm::ensure_immutables drops the slot's DFAs (reset_shared_dfas); the destructor erases the map entry, so match-time caches never outlive the regex.
| std::vector<std::uint8_t> real::detail::regex_immutables::class_rows |
Byte-indexed membership rows, filled on first use of each class and kept for the regex's life.
Deriving a row into the VM state charges every short search (search() builds a fresh state); filling every row at compile charges patterns that read few of their classes (a negated class interns a dozen). Per class, on demand, per regex pays for neither.
Thread safety: rows_for keys the program the rows were sized for, as built_for does. A row's flag is release-stored after the row is filled under immut_build_mu and acquire-loaded before it is read; only the lock holder that saw the flag clear writes a row, so a published row is immutable. One 256-byte row per interned BYTE class.
| std::atomic<std::uint64_t> real::detail::regex_immutables::row_ready_bits {0} |
"Row filled" flags: one bit per row, the three runs packed into one word, plus an overflow vector for indices that do not fit.
Every real::regex construction pays for this block: separate std::vector<std::atomic<char>> runs allocate even for a pattern with no cp_class, where bits allocate nothing. Read on a VM-state miss, not per call. The overflow is a vector of atomics, not a unique_ptr array (non-constexpr destructor; this struct must stay literal), and not std::atomic_ref, which a supported libc++ lacks.