DFA construction internals: subset construction over a flattened NFA. Not a stable API.
More...
|
| class | ac_automaton |
| | Aho-Corasick automaton for a fixed_alternation program's branch set, built once per compiled program and reused across every match on it. More...
|
| |
| struct | alternation_density |
| | Whether an alternation's first bytes are dense in one subject, decided once from a sample of it. More...
|
| |
| struct | alternation_pairs |
| | Two probe bytes per branch of a literal alternation: the branch's first byte and one byte further in it, for the pair filter an alternation's block scan turns to once the first bytes prove common. More...
|
| |
| struct | anchored_walk_bill |
| | Tells when the anchored walks from candidates should give way to one forward pass and one reverse, from what the walks that found no match cost against the distance crossed. More...
|
| |
| struct | ast |
| | A parsed pattern: the node pool plus side tables. More...
|
| |
| struct | ast_node |
| | One AST node. Active fields depend on kind (noted per field). More...
|
| |
| struct | backtrack_frame |
| | The bounded backtracker's state for one search, on the caller's stack (see pike_vm::run_bounded_backtrack). More...
|
| |
| struct | basic_capture_pool |
| | Copy-on-write pool of capture blocks, the one capture-slot mechanism for both storages. More...
|
| |
| struct | basic_pike_state |
| | Reusable VM scratch state. More...
|
| |
| struct | basic_thread_list |
| | One priority-ordered list of NFA threads (leftmost-greedy semantics). More...
|
| |
| struct | binprop_alias_entry |
| | A loose-normalized (lowercase, no _/-/space) binary-property name and its value. More...
|
| |
| struct | borrowed_names |
| | The compile-time policy's name owner: there is nothing to own. More...
|
| |
| struct | byte_program |
| | A byte-level view of a Pike program for the DFA passes: every klass_cp is expanded into UTF-8 byte-range split/klass chains, so a forward DFA can represent it; the Pike program is untouched. eligible is false when an op no DFA can represent is present — the caller keeps the Pike VM. More...
|
| |
| struct | char_class |
| | A set of byte values (0–255) as a 256-bit bitmap. More...
|
| |
| struct | class_def |
| | A parsed character class: its ASCII bitmap plus its non-ASCII code-point ranges, bundled so the two cannot desynchronize. More...
|
| |
| struct | class_ref |
| | A typed reference into a possessive-loop body's operand space. More...
|
| |
| struct | code_range |
| | An inclusive code-point range [lo, hi], shared by ast.hpp's classes and the generated Unicode tables; here so those headers need not include the parser. More...
|
| |
| class | compiler |
| | Compiles an ast into a dynamic_program (NFA bytecode). More...
|
| |
| struct | cp_class |
| | A match-time code-point class for klass_cp: an ASCII bitmap below 0x80 plus a slice of sorted non-ASCII ranges in the program's flat cp_ranges, negation already applied. Unlike the byte-NFA klass, the ranges are kept and searched at match time (O(log ranges)). More...
|
| |
| struct | cp_hi_table |
| | Unicode-property sparse 2-stage membership for code points > U+07FF (page = cp>>8 → 256-bit block). A thread-local heap cache, so basic_pike_state (and the ASCII class loop) keeps its size. More...
|
| |
| struct | decoded_codepoint |
| | The result of a strict UTF-8 decode: the code point, its byte length, and validity. More...
|
| |
| struct | dfa_byte_classes |
| | Computes byte-equivalence classes: two bytes are equivalent iff they satisfy the same consuming predicates (every klass test and every byte literal). Reduces the alphabet so the DFA is built over classes, not over 256 bytes. More...
|
| |
| struct | dfa_fidelity_raw |
| | The per-pattern answer of real::dfa_faithful, before the public wrapping. More...
|
| |
| struct | dfa_instr |
| | A flattened NFA instruction (global PCs, global class index). More...
|
| |
| class | dfa_lease |
| | This thread's DFA set for one regex, for the lifetime of the lease: a scan through it takes no lock, so threads sharing a regex do not queue on its DFAs. More...
|
| |
| struct | dfa_nfa |
| | The union NFA over all the patterns, flattened into one address space. More...
|
| |
| struct | dfa_tables |
| | The baked DFA tables produced by dfa_build. More...
|
| |
| struct | digit_escape_result |
| | Result of decode_digit_escape. More...
|
| |
| struct | dynamic_program |
| | The view is copied on every find_iter and count_matches call: a fixed per-call cost under any throughput row's noise floor, so its size is guarded by counting bytes, not by timing. More...
|
| |
| struct | dynamic_storage |
| | Storage policy backing real::regex: heap, sized once at run time. More...
|
| |
| struct | eps_entry |
| | One frame on the epsilon-closure DFS stack: a pc plus the capture block its branch carries (a split shares it, a save copies it on write), so no slot-restore entry is needed. More...
|
| |
| struct | fold_entry |
| | A code point and the other members of its case-fold orbit (up to 3; orbits <= 4). More...
|
| |
| struct | gc_alias_entry |
| | A loose-normalized (lowercase, no _/-/space) General_Category name and its property. More...
|
| |
| struct | inner_literal |
| | The best required inner literal of a pattern (the memmem candidate). More...
|
| |
| struct | inner_literal_bill |
| | Tells when the inner-literal route's candidates should give way to the core search, from what reaching their starts and confirming them read against the distance crossed. More...
|
| |
| struct | instr |
| | One NFA instruction. Field meaning depends on op. More...
|
| |
| struct | lazy_byte_alphabet |
| | Byte-class alphabet over a Pike program: bytes satisfying exactly the same byte/klass predicates share a class, so the DFA transitions over classes instead of 256 raw bytes. More...
|
| |
| class | lazy_dfa |
| | A lazy priority-preserving forward DFA over a Pike program (the kFirstMatch forward pass). More...
|
| |
| struct | literal_alt_chain |
| | The literal alternation starting at a split: its exit and each branch's bytes. More...
|
| |
| struct | literal_alt_trie |
| | A literal alternation (every branch a run of byte ops converging on one exit) factored into a trie that keeps leftmost-first priority. More...
|
| |
| struct | literal_density |
| | What the adaptive literal search learned about one subject: whether the needle's rarest byte is common there. More...
|
| |
| struct | literal_memo |
| | The two literal densities of one subject, held in a std::optional built at the first literal search, so constructing a search state (every search does) writes no densities. More...
|
| |
| struct | lookaround_scratch |
| | Reusable, isolated scratch for one level of lookaround evaluation (dynamic only). More...
|
| |
| struct | lookaround_sub |
| | A bounded lookaround sub-program, referenced by assert_lookaround's arg16. More...
|
| |
| class | name_context_box |
| | A uniquely-owning, deep-copying box for owned_name_context that survives constant evaluation. More...
|
| |
| struct | named_group |
| | A named capture group. More...
|
| |
| struct | non_empty_access |
| | The searches that accept no empty match (std::regex_constants::match_not_null), for the std drop-in; not part of REAL's own interface. More...
|
| |
| class | onepass |
| | Builds and holds the one-pass classification (and table, when eligible) of a byte-program. More...
|
| |
| struct | onepass_edge |
| | One outgoing edge of a one-pass node, for a byte-class: the next node and the capture slots that take the current position as the byte is consumed. Two epsilon paths reaching the same class with a different edge is the one-pass conflict — the pattern is then rejected. More...
|
| |
| struct | onepass_node |
| | A one-pass node: one edge per byte-class, plus whether the run may end here and with what captures. Nodes are the points the automaton can be in between byte reads. More...
|
| |
| struct | onepass_step |
| | One edge of the flattened table onepass::extract walks: onepass_edge with the target given as the offset of its row, so a step is one load from one array. More...
|
| |
| struct | owned_name_context |
| | The name-resolution context a result owns when it must outlive the regex it came from. More...
|
| |
| class | parser |
| | Recursive-descent parser: a pattern string in, an ast out. More...
|
| |
| struct | pattern_hints |
| | Search-acceleration hints extracted from a compiled program by analyze_program (prefilter.hpp). They change how fast, never what matches. More...
|
| |
| struct | pc_set_cache |
| | A chained hash set of interned state ids keyed by their pc-set: maps a candidate pc-set to its state id, or not_found. All-std::vector, not std::unordered_map, so the DFAs stay literal types (a constexpr real::regex embeds one in its scratch state). More...
|
| |
| struct | pike_state |
| | VM scratch state for the dynamic storage mode, plus the lookaround sub-scratch. More...
|
| |
| class | pike_vm |
| | The Pike VM, generic over the scratch-state container policy. More...
|
| |
| struct | program_view |
| | A non-owning view of everything the engine needs to run one compiled pattern: the borrowed spans, the slot count, the mode flags, the search hints, and the per-regex cache. Both storages hand one of these to the VM, which is why the engine is storage-agnostic. Valid only as long as the program it views is alive. More...
|
| |
| struct | range_intern_table |
| | Intern table for UTF-8 edge byte ranges, keyed by the exact 16-bit (lo << 8) | hi. More...
|
| |
| struct | regex_immutables |
| | The per-regex immutable cache the router shares across every find_iter on a regex: the byte program (klass_cp expanded to the deterministic trie) and, when the pattern is one-pass, the extractor table. More...
|
| |
| class | reverse_dfa |
| | The start-finder companion to lazy_dfa. Given a match end, it finds the leftmost start (the design guide §7.6 contract). It runs the inverted program — the forward program's edges transposed, its consuming bytes kept — as a cached DFA over the text scanned right-to-left from the end, recording an accept each time it reaches the original start (reverse-kLongest: the furthest-back accept is the start). It needs no priority ordering — its states are plain unordered (sorted) PC sets and its rule is longest — so it is simpler than the forward pass. Dynamic only. More...
|
| |
| struct | script_alias_entry |
| | A loose-normalized (lowercase, no _/-/space) Script name and its value. More...
|
| |
| struct | script_range |
| | One code-point range and the Script it belongs to (the table partitions the code space). More...
|
| |
| struct | shape_close |
| | The shape_lead counterpart: optional trail \b/\B, optional \Z/$, then exactly save 1, match ending the program. More...
|
| |
| struct | shape_lead |
| | A fixed shape's lead: save 0, an optional \A/^, an optional \b/\B. More...
|
| |
| struct | shared_dfa_set |
| | One thread's lazy DFAs for one regex: the transition caches a scan fills as it walks. More...
|
| |
| struct | shared_dfa_slot |
| | Process-wide per-regex DFA state keyed by regex_immutables*: a pool of shared_dfa_set, one per thread using the regex, and the flags every thread shares. More...
|
| |
| class | small_vec |
| | Small-buffer-optimized vector for the dynamic hot paths: up to InlineCapacity elements inline, spilling to the heap beyond that. More...
|
| |
| struct | static_il_guard_fields |
| | IL: the per-haystack guard fields the inner-literal route needs, for a compile-time storage. More...
|
| |
| struct | static_no_il_guard_fields |
| | No IL fields: the route is not compiled for this pattern. More...
|
| |
| struct | static_pike_scratch |
| | Compile-time-storage VM scratch, all fixed-capacity (zero heap), keyed on DIMENSIONS ONLY. More...
|
| |
| struct | static_storage |
| | Storage policy backing real::static_regex: compile-time, stateless. More...
|
| |
| class | static_vec |
| | Fixed-capacity vector backed by an inline array (no heap): the subset of std::vector the Pike VM uses, for the static storage mode. More...
|
| |
| struct | utf8_byte_range |
| | One byte-range step [lo, hi] of a UTF-8 sequence produced by the code-point-range algorithm. More...
|
| |
| struct | utf8_byte_seq |
| | A canonical UTF-8 byte-range sequence (1–4 steps) covering part of a code-point range. More...
|
| |
| struct | utf8_second_byte_bounds |
| | [lo, hi] bounds for the FIRST continuation byte of a multi-byte UTF-8 sequence, given its lead byte — one entry of utf8_second_byte_bounds_table. More...
|
| |
| struct | utf8_trie |
| | A minimal deterministic UTF-8 trie for a code-point class. root == -1 means the class is empty. More...
|
| |
| struct | utf8_trie_node |
| | One node of a minimal deterministic UTF-8 trie for a code-point class. Its byte-range transitions are pairwise disjoint (at most one edge matches a byte), which makes the byte-program one-pass-friendly. A target >= 0 is a node id; -1 is accept (the run continues at the construct's successor). More...
|
| |
| struct | visit_marks |
| | The pcs one closure computation has entered, by generation: starting one bumps the generation instead of clearing per pc, so a cache miss costs its closure, not the program's size (a large alternation's byte program runs to hundreds of thousands of instructions). More...
|
| |
|
| enum class | ac_verdict : std::uint8_t { not_consulted = 0
, cascade
, automaton
} |
| | What the AC density gate last decided; ac_density_last_verdict() below reports it. More...
|
| |
| enum class | opcode : std::uint8_t {
byte
, klass
, klass_cp
, split
,
jump
, save
, assert_position
, match
,
assert_lookaround
, byte_loop_possessive
, klass_loop_possessive
, klass_cp_loop_possessive
} |
| | NFA instruction opcodes executed by the Pike VM. More...
|
| |
| enum class | assert_kind : std::uint8_t {
text_start
, text_end
, text_end_or_final_newline
, line_start
,
line_end
, word_boundary
, not_word_boundary
, word_start
,
word_end
, line_start_cr
, line_end_cr
} |
| | Kind of zero-width assertion carried in assert_position's arg8. Multiline and trailing-newline variants are resolved at compile time. More...
|
| |
| enum class | look_dir : std::uint8_t { ahead
, behind
} |
| | Direction of a lookaround sub-pattern. More...
|
| |
| enum class | class_kind : std::uint8_t { none
, byte
, klass
, klass_cp
} |
| | Which operand space a class_ref indexes: the literal byte (byte_loop_possessive), classes[] (klass_loop_possessive) or cp_classes[] (klass_cp_loop_possessive). none = unarmed.
|
| |
| enum class | run_mode : std::uint8_t { prefix
, full
, search
} |
| | How a VM run is anchored. More...
|
| |
| enum class | counter : std::uint8_t {
prefilter_work_units
, vm_window_runs
, batch_fills
, inner_literal_bill_trips
,
inner_literal_reverse_bytes
, inner_literal_confirm_bytes
, byte_program_builds
, il_prefix_run_walks
,
batch_handouts
, vm_reseeds
, onepass_anchored_walks
, bounded_backtrack_runs
,
dfa_quits
, literal_pair_scans
, alternation_avx2_blocks
, literal_avx2_scans
,
fixed_shape_batches
, literal_rest_scans
, alternation_pair_blocks
, alternation_nibble_blocks
,
ac_completion_walks
, alternation_variant_scans
, alternation_wide_scans
, alternation_pair_candidates
,
ahead_table_rows
, behind_walk_steps
, behind_atom_steps
, dfa_span_batches
,
dfa_leases_taken
, class_folds
, count_
} |
| | Test counters, one per mechanism, for the tests that pin when it runs. Billed through note() only under REAL_TEST_INSTRUMENT (free in production); process-wide relaxed atomics. More...
|
| |
| enum class | node_kind : std::uint8_t {
empty
, byte
, klass
, any
,
concat
, repeat
, alternation
, group
,
anchor
, lookaround
} |
| | Kind of an AST node; selects which fields of real::detail::ast_node are meaningful. More...
|
| |
| enum class | anchor_kind : std::uint8_t {
caret
, dollar
, text_start
, text_end
,
word_boundary
, not_word_boundary
, word_start
, word_end
} |
| | The specific zero-width assertion of an anchor node (see node_kind::anchor). More...
|
| |
| enum class | digit_escape_kind : std::uint8_t { octal
, group_ref
, octal_overflow
} |
| | What a \<digit> escape decoded to (see decode_digit_escape()). More...
|
| |
| enum class | binprop : std::uint8_t {
ASCII_Hex_Digit
, Alphabetic
, Bidi_Control
, Case_Ignorable
,
Cased
, Changes_When_Casefolded
, Changes_When_Casemapped
, Changes_When_Lowercased
,
Changes_When_Titlecased
, Changes_When_Uppercased
, Dash
, Default_Ignorable_Code_Point
,
Deprecated
, Diacritic
, Emoji
, Emoji_Component
,
Emoji_Modifier
, Emoji_Modifier_Base
, Emoji_Presentation
, Extended_Pictographic
,
Extender
, Grapheme_Base
, Grapheme_Extend
, Grapheme_Link
,
Hex_Digit
, Hyphen
, IDS_Binary_Operator
, IDS_Trinary_Operator
,
IDS_Unary_Operator
, ID_Compat_Math_Continue
, ID_Compat_Math_Start
, ID_Continue
,
ID_Start
, Ideographic
, Join_Control
, Logical_Order_Exception
,
Lowercase
, Math
, Modifier_Combining_Mark
, Noncharacter_Code_Point
,
Other_Alphabetic
, Other_Default_Ignorable_Code_Point
, Other_Grapheme_Extend
, Other_ID_Continue
,
Other_ID_Start
, Other_Lowercase
, Other_Math
, Other_Uppercase
,
Pattern_Syntax
, Pattern_White_Space
, Prepended_Concatenation_Mark
, Quotation_Mark
,
Radical
, Regional_Indicator
, Sentence_Terminal
, Soft_Dotted
,
Terminal_Punctuation
, Unified_Ideograph
, Uppercase
, Variation_Selector
,
White_Space
, XID_Continue
, XID_Start
, count
} |
| | A Unicode binary property: the 63 standard yes/no properties this build knows.
|
| |
| enum class | gc_property : std::uint8_t {
Lu
, Ll
, Lt
, Lm
,
Lo
, Mn
, Mc
, Me
,
Nd
, Nl
, No
, Pc
,
Pd
, Ps
, Pe
, Pi
,
Pf
, Po
, Sm
, Sc
,
Sk
, So
, Zs
, Zl
,
Zp
, Cc
, Cf
, Co
,
Cn
, L
, M
, N
,
P
, S
, Z
, C
,
count
} |
| | A Unicode General_Category property: the 29 assignable categories then the 7 groups.
|
| |
| enum class | script : std::uint8_t {
Unknown
, Adlam
, Ahom
, Anatolian_Hieroglyphs
,
Arabic
, Armenian
, Avestan
, Balinese
,
Bamum
, Bassa_Vah
, Batak
, Bengali
,
Bhaiksuki
, Bopomofo
, Brahmi
, Braille
,
Buginese
, Buhid
, Canadian_Aboriginal
, Carian
,
Caucasian_Albanian
, Chakma
, Cham
, Cherokee
,
Chorasmian
, Common
, Coptic
, Cuneiform
,
Cypriot
, Cypro_Minoan
, Cyrillic
, Deseret
,
Devanagari
, Dives_Akuru
, Dogra
, Duployan
,
Egyptian_Hieroglyphs
, Elbasan
, Elymaic
, Ethiopic
,
Garay
, Georgian
, Glagolitic
, Gothic
,
Grantha
, Greek
, Gujarati
, Gunjala_Gondi
,
Gurmukhi
, Gurung_Khema
, Han
, Hangul
,
Hanifi_Rohingya
, Hanunoo
, Hatran
, Hebrew
,
Hiragana
, Imperial_Aramaic
, Inherited
, Inscriptional_Pahlavi
,
Inscriptional_Parthian
, Javanese
, Kaithi
, Kannada
,
Katakana
, Kawi
, Kayah_Li
, Kharoshthi
,
Khitan_Small_Script
, Khmer
, Khojki
, Khudawadi
,
Kirat_Rai
, Lao
, Latin
, Lepcha
,
Limbu
, Linear_A
, Linear_B
, Lisu
,
Lycian
, Lydian
, Mahajani
, Makasar
,
Malayalam
, Mandaic
, Manichaean
, Marchen
,
Masaram_Gondi
, Medefaidrin
, Meetei_Mayek
, Mende_Kikakui
,
Meroitic_Cursive
, Meroitic_Hieroglyphs
, Miao
, Modi
,
Mongolian
, Mro
, Multani
, Myanmar
,
Nabataean
, Nag_Mundari
, Nandinagari
, New_Tai_Lue
,
Newa
, Nko
, Nushu
, Nyiakeng_Puachue_Hmong
,
Ogham
, Ol_Chiki
, Ol_Onal
, Old_Hungarian
,
Old_Italic
, Old_North_Arabian
, Old_Permic
, Old_Persian
,
Old_Sogdian
, Old_South_Arabian
, Old_Turkic
, Old_Uyghur
,
Oriya
, Osage
, Osmanya
, Pahawh_Hmong
,
Palmyrene
, Pau_Cin_Hau
, Phags_Pa
, Phoenician
,
Psalter_Pahlavi
, Rejang
, Runic
, Samaritan
,
Saurashtra
, Sharada
, Shavian
, Siddham
,
SignWriting
, Sinhala
, Sogdian
, Sora_Sompeng
,
Soyombo
, Sundanese
, Sunuwar
, Syloti_Nagri
,
Syriac
, Tagalog
, Tagbanwa
, Tai_Le
,
Tai_Tham
, Tai_Viet
, Takri
, Tamil
,
Tangsa
, Tangut
, Telugu
, Thaana
,
Thai
, Tibetan
, Tifinagh
, Tirhuta
,
Todhri
, Toto
, Tulu_Tigalari
, Ugaritic
,
Vai
, Vithkuqi
, Wancho
, Warang_Citi
,
Yezidi
, Yi
, Zanabazar_Square
, count
} |
| | A Unicode Script value; Unknown (0) is every code point no script assigns.
|
| |
|
| bool & | lazy_dfa_route_disabled () |
| | Test seam: force the matcher off the lazy-DFA route onto the pure Pike VM, so a differential can assert routed and unrouted searches agree in one binary. Not for production use (applies to every *_disabled seam below).
|
| |
| std::size_t & | lazy_dfa_byte_budget () |
| | Test seam: the byte budget the search DFAs are built with (read at each DFA's construction).
|
| |
| bool & | bounded_backtrack_route_disabled () |
| | Test seam: force the general loop off the bounded backtracker onto the Pike VM, so a differential can assert both agree on every small subject.
|
| |
| bool & | inner_literal_route_disabled () |
| | Test seam: force the matcher off the inner-literal search route onto the core search. The route cannot miss a leftmost match because its reverse bound never advances mid-search.
|
| |
| bool & | rare_disc_route_disabled () |
| | Test seam: force off the rare-discriminant prefilter (https?:// memchr-: route) onto prefix/first-byte search.
|
| |
| bool & | inner_literal_guard_disabled () |
| | Test seam: force the inner-literal small-haystack guard off, so the route fires on any size and tiny correctness inputs exercise it. The guard uses regex_immutables::il_min_haystack on the first candidate scan and il_warm_floor thereafter.
|
| |
| bool & | trailing_la_route_disabled () |
| | Test seam: force the matcher off the trailing-lookaround class+ route onto the pure Pike VM.
|
| |
| bool & | fixed_shape_pair_route_disabled () |
| | Test seam: force the matcher off the heterogeneous fixed-shape pair-filter route onto the ordinary run_fixed_shape walk. The route only filters; match_fixed_body_wb decides each candidate.
|
| |
| bool & | fixed_shape_route_disabled () |
| | Test seam: force the matcher off the fixed-shape walk (run_fixed_shape) onto the general Pike loop. inner_literal_route_disabled does not reach a fixed_shape pattern (the inner-literal gate excludes it); a differential on one needs this seam.
|
| |
| bool & | class_fastpath_disabled () |
| | Test/profile seam: skip the dedicated class-scan fast paths (byte class-loop, cp-class-loop, codepoint_class, negated-class ./[^,]+), so such a pattern falls through to lazy-DFA / general.
|
| |
| bool & | possessive_fastpath_disabled () |
| | Test/profile seam: force the matcher off the possessive-loop fast paths (bare/suffixed/delimited X*+/X++) onto the general VM.
|
| |
| bool & | aho_corasick_route_disabled () |
| | Test seam: force the matcher off the Aho-Corasick multi-literal route onto the pattern_hints::fixed_alternation run_alternation path.
|
| |
| bool & | alternation_pairs_disabled () |
| | Test seam: keep an alternation's block scans on its first bytes whatever the subject's density, so a differential can compare the pair filter with the first-byte scan.
|
| |
| bool & | alternation_nibbles_disabled () |
| | Test seam: mask a dense alternation's blocks by its byte pairs rather than by the nibble fingerprint, so a differential can compare both filters.
|
| |
| bool & | ac_density_gate_disabled () |
| | Test seam: take the Aho-Corasick density gate out, so the route is chosen on branch count alone.
|
| |
| std::atomic< bool > & | il_density_last_abandoned () |
| | Test observability: whether the inner-literal density gate last abandoned the route.
|
| |
| std::atomic< ac_verdict > & | ac_density_last_verdict () |
| | Test observability: the AC density gate's most recent verdict.
|
| |
| constexpr utf8_trie | build_utf8_trie (const cp_class &cc, std::span< const code_range > cp_ranges) |
| | Builds the minimal deterministic trie recognising a code-point class's UTF-8 byte sequences.
|
| |
| constexpr std::size_t | utf8_trie_emit_size (const utf8_trie &trie) |
| | The instruction count emit_utf8_trie writes: an empty class is one dead klass; otherwise each node is a split-guarded chain of k byte ranges (3k - 1 instructions).
|
| |
| REAL_BUILD_COLD constexpr void | emit_utf8_trie (byte_program &bp, const utf8_trie &trie, std::int32_t after, range_intern_table &seen) |
| | Emits trie into bp as a deterministic split/klass/jump fragment, interning each edge's byte range through seen.
|
| |
| bool & | alternation_trie_disabled () |
| | Test seam: build the byte program's literal alternations flat, so a differential can compare the trie against them.
|
| |
| constexpr bool | detect_literal_alt (std::span< const instr > code, std::size_t pc, literal_alt_chain &out) |
| | Whether [pc, exit) is a chain of splits whose branches are runs of byte ops jumping forward to one exit, the last branch falling through to it.
|
| |
| REAL_BUILD_COLD constexpr byte_program | build_byte_program (const program_view &prog, bool keep_assertions=false, std::size_t max_size=max_byte_program_size) |
| | Builds the byte-level DFA program for prog (see byte_program). A klass_cp at P (the op plus three utf8_cont slots) is replaced by its class's deterministic UTF-8 trie (build_utf8_trie), converging on the mapped P+4; every other op is copied with remapped targets. The first pass sizes each construct into the old→new pc map and enforces max_size; the second emits.
|
| |
| constexpr bool | is_cr_line_assert (const instr &in) |
| | Whether in is an ECMAScript line assertion, whose line ends at \r as well as \n.
|
| |
| REAL_BUILD_COLD constexpr lazy_byte_alphabet | compute_lazy_alphabet (std::span< const instr > code, std::span< const char_class > classes) |
| | Partition 0..255 by the program's consuming predicates (every klass test, every byte literal). Bytes with an identical signature collapse to one class.
|
| |
| constexpr bool | undecidable_word (assert_kind kind, bool prev_word, bool prev_nonascii, bool next_word, bool next_nonascii) |
| | Whether a word assertion needs a code point's word-ness that one byte does not give: a side it reads is a non-ASCII byte, and the ASCII side does not settle it alone.
|
| |
| constexpr bool | dfa_representable (std::span< const instr > code, bool ascii_word) |
| | Whether a lazy DFA, forward or reversed, can represent every op of code.
|
| |
| void | erase_shared_dfas (const regex_immutables *immut) |
| | Retire this regex's slot (called from ~regex_immutables). Scans still holding the slot's shared_ptr keep it alive; clearing shared_dfa_slot::owner stops their cached copy matching, so a new regex at this address is never served the retired slot.
|
| |
| std::mutex & | immut_build_mu (const regex_immutables *immut) |
| | Striped rebuild lock for pike_vm::ensure_immutables (not on regex_immutables — layout isolation). Distinct from shared_dfa_map_mu / shared_dfa_slot::pool_mu so reset_shared_dfas cannot self-deadlock. Different immutables rarely share a stripe.
|
| |
| std::mutex & | shared_dfa_map_mu () |
| | The mutex guarding insert/erase on the process-wide shared_dfa_slot map.
|
| |
| std::unordered_map< const regex_immutables *, std::shared_ptr< shared_dfa_slot > > & | shared_dfa_map () |
| | Process-wide map, deliberately never destroyed: other statics' ~regex_immutables still call erase_shared_dfas at exit. Entries are erased per destructor, so nothing accumulates.
|
| |
| shared_dfa_slot & | shared_dfa_for (regex_immutables *immut) |
| | Resolve the process-wide DFA slot for this regex (map insert under shared_dfa_map_mu).
|
| |
| void | reset_shared_dfas (regex_immutables *immut, bool keep_warm=false) |
| | Drop any DFAs cached for immut (caller holds nothing; takes map + slot locks). Invoked from pike_vm's ensure_immutables rebuild so a reused immutables address — or the same address under a new program — cannot keep a previous pattern's DFAs.
|
| |
| std::size_t | shared_dfa_map_size_for_test () |
| | Test/audit: number of live shared-DFA map entries (process-wide). Not for production.
|
| |
| constexpr std::size_t | encode_utf8_bytes (std::uint32_t cp, std::uint8_t(&out)[4]) |
| | Encodes cp to its UTF-8 bytes in out, returning the length (1–4).
|
| |
| constexpr void | utf8_push_range (std::uint32_t start, std::uint32_t end, std::vector< utf8_byte_seq > &out) |
| | Appends to out the byte-range sequences recognising exactly the UTF-8 encodings of [start, end] (RE2 / rust regex-syntax Utf8Sequences).
|
| |
| constexpr std::vector< utf8_byte_seq > | utf8_range_sequences (std::uint32_t lo, std::uint32_t hi) |
| | Canonical UTF-8 byte-range sequences for the code-point range [lo, hi], excluding the surrogate block [U+D800, U+DFFF] (so a negated class never matches a surrogate encoding).
|
| |
| constexpr void | fold_ascii_case (char_class &klass) |
| | Closes klass under ASCII case folding.
|
| |
| constexpr bool | is_ascii_word_byte (std::uint8_t byte) |
| | Reports whether byte is an ASCII "word" byte ([0-9A-Za-z_]).
|
| |
| constexpr char_class | digit_set () |
| | The ASCII digit set behind \d (Python re.ASCII semantics).
|
| |
| constexpr char_class | word_set () |
| | The ASCII word set behind \w.
|
| |
| constexpr char_class | space_set () |
| | The ASCII whitespace set behind \s under flags::ascii / flags::bytes.
|
| |
| constexpr char_class | utf8_cont_set () |
| | The UTF-8 continuation-byte set 10xxxxxx.
|
| |
| constexpr char_class | utf8_lead2_set () |
| | The lead-byte set of a 2-byte UTF-8 sequence.
|
| |
| constexpr char_class | utf8_lead3_set () |
| | The lead-byte set of a 3-byte UTF-8 sequence.
|
| |
| constexpr char_class | utf8_lead4_set () |
| | The lead-byte set of a 4-byte UTF-8 sequence.
|
| |
| constexpr std::array< utf8_second_byte_bounds, 256 > | make_utf8_second_byte_bounds_table () |
| | Builds utf8_second_byte_bounds_table.
|
| |
| constexpr std::uint64_t | fingerprint_cp_class_content (const char_class &ascii, const code_range *ranges, std::uint32_t range_count) |
| | FNV-1a 64-bit content fingerprint of an ASCII bitmap and a range span, computed once at intern_cp_class (constexpr, for static_regex); match time reads cp_class::fingerprint.
|
| |
| dfa_nfa | dfa_flatten (std::span< const program_view > programs) |
| | Flattens programs into one union NFA, auditing DFA-ability.
|
| |
| void | dfa_set_bit (dfa_set &s, std::size_t i) |
| | Set bit i in s. Indices past the set's size are ignored (it is sized to fit).
|
| |
| bool | dfa_test_bit (const dfa_set &s, std::size_t i) |
| | Whether bit i is set in s.
|
| |
| dfa_set | dfa_closure (const dfa_nfa &nfa, const std::vector< std::uint32_t > &seeds, bool at_start) |
| | The epsilon-closure of seeds (a PC list), as a canonical PC bitset. at_start follows a text_start assertion (true only at offset 0).
|
| |
| dfa_set | dfa_move (const dfa_nfa &nfa, const dfa_set &set, std::uint8_t rep) |
| | The move on the byte rep: ε-closure of the successors of every PC in set that consumes rep.
|
| |
| std::int64_t | dfa_accept_of (const dfa_nfa &nfa, const dfa_set &set) |
| | The accepting rule of a state set: the SMALLEST rule index among its match PCs (the order tie-break), or -1 if none accept.
|
| |
| std::size_t | dfa_mask_words (std::size_t rule_count) noexcept |
| | Word count for a which-matched bitset over rule_count rules.
|
| |
| std::vector< std::uint64_t > | dfa_accept_mask_of (const dfa_nfa &nfa, const dfa_set &set) |
| | Bitset of ALL accepting rule indices in set (which-matched; word-packed). Empty vector when no rule accepts (or rule_count == 0).
|
| |
| std::int64_t | dfa_mask_min_rule (const std::vector< std::uint64_t > &mask) |
| | Smallest rule index set in mask, or -1 if empty (munch tag derivation).
|
| |
| dfa_byte_classes | dfa_compute_classes (const dfa_nfa &nfa) |
| | Partition 0..255 by the union NFA's consuming predicates.
|
| |
| void | dfa_seeds_all (const dfa_nfa &nfa, const dfa_set &set, const dfa_byte_classes &bc, const std::vector< std::vector< std::uint8_t > > &klass_members, std::vector< std::vector< std::uint32_t > > &seeds) |
| | Every class's seed list from one state in one pass over the state's PCs: the PCs after each consuming instruction that class passes, without rescanning the set once per class.
|
| |
| dfa_tables | dfa_build (std::span< const program_view > programs, std::size_t state_cap=max_dfa_states, bool unanchored=false) |
| | Subset construction over byte-classes, then Moore minimization.
|
| |
| bool | dfa_priority_closure (const dfa_nfa &nfa, std::uint32_t seed, bool at_start, std::vector< std::uint8_t > &seen, std::vector< std::uint32_t > &out) |
| | The priority-ordered epsilon closure of seed, as the Pike walk builds it: consuming pcs are appended to out in priority order, and reaching a match stops the walk, because a thread list is cut below its first accepting thread.
|
| |
| dfa_fidelity_raw | dfa_decide_fidelity (const program_view &prog, std::size_t budget) |
| | Decides whether prog's priority match equals its longest match on every input.
|
| |
| void | dfa_memo_misuse (const char *what) |
| | Throws the std::invalid_argument a misused dfa_munch_memo raises, out of line so the throw does not weigh on the per-token match that checks for it.
|
| |
| std::size_t & | ac_memory_budget () |
| | Bytes an automaton may hold (the layout rule is in the file header); ac_memory_budget_default unless a test shrinks it to reach the sparse form and the decline.
|
| |
| bool & | ac_dense_disabled () |
| | Test seam: search the sparse trie even where the dense table fits.
|
| |
| std::size_t & | ac_sparse_row_cap () |
| | Test seam: at most this many dense rows in the sparse form below what the budget allows. The root keeps its row whatever the cap: a miss there has no fail link to fall along.
|
| |
| REAL_BUILD_COLD std::optional< ac_automaton > | build_ac_automaton (std::span< const instr > code, std::span< const char_class > classes, std::size_t body_pc) |
| | Builds an ac_automaton from a fixed_alternation-shaped program's branch set, over the program's own byte classes (every class it tests is a union of them, so no position is split).
|
| |
| constexpr bool | word_before (std::string_view text, std::size_t pos, bool ascii_word) |
| | Word-ness of the code point ending at pos — the left side of a boundary. False at the text start or on a malformed sequence; ASCII / bytes / re.A (ascii_word) stay byte-level.
|
| |
| constexpr bool | word_after (std::string_view text, std::size_t pos, bool ascii_word) |
| | Word-ness of the code point starting at pos — the right side of a boundary. False at the text end or on a malformed sequence; ASCII / bytes / re.A stay byte-level.
|
| |
| constexpr bool | assertion_holds (assert_kind kind, std::string_view text, std::size_t pos, bool ascii_word) |
| | Evaluates a zero-width assertion at pos in text.
|
| |
| std::atomic< std::uint64_t > & | tally (counter c) noexcept |
| | The value of a test counter, to read or to reset.
|
| |
| constexpr void | note (counter c, std::uint64_t n=1) noexcept |
| | Bills n to a test counter. A no-op unless the test binary defines REAL_TEST_INSTRUMENT.
|
| |
| bool & | alternation_avx2_disabled () |
| | Test seam: keep the alternation fingerprint on 16-byte blocks where the CPU has AVX2, so a differential can compare both widths in one binary. Not for production use.
|
| |
| bool & | literal_avx2_disabled () |
| | Test seam: keep the literal filter on 16-byte blocks where the CPU has AVX2, so a differential can compare both widths in one binary. Not for production use.
|
| |
| constexpr bool | is_word_boundary_kind (assert_kind kind) noexcept |
| | True if kind is \b or \B (the only position asserts a fast path wraps).
|
| |
| constexpr std::uint8_t | wb_hint_of (assert_kind kind) noexcept |
| | Encodes kind as a wb_lead/wb_trail hint value (1 = \b, 2 = \B); 0 if not a word boundary.
|
| |
| constexpr bool | peel_optional_wb (std::span< const instr > code, std::size_t &p, std::uint8_t &hint) noexcept |
| | Peels an optional \b/\B assertion at p, lead or trail alike.
|
| |
| constexpr shape_lead | parse_shape_lead (std::span< const instr > code) noexcept |
| | Peels a fixed shape's save 0 and its optional lead \b/\B.
|
| |
| constexpr shape_close | parse_shape_close (std::span< const instr > code, std::size_t from) noexcept |
| | Peels a fixed shape's optional trail \b/\B, then its save 1 and match.
|
| |
| constexpr bool | is_full_ascii_word_class (const char_class &cls) noexcept |
| | True if cls is exactly the ASCII word set [0-9A-Za-z_] (\w under bytes/re.A).
|
| |
| constexpr bool | is_ascii_word_subset_class (const char_class &cls) noexcept |
| | True if every member of cls is an ASCII word byte (subset of \w under bytes/re.A).
|
| |
| constexpr bool | is_full_unicode_word_cp_class (const cp_class &cc, std::span< const code_range > all_ranges) noexcept |
| | True if cc is exactly the canonical Unicode \w class (not a user superset).
|
| |
| constexpr bool | wb_redundant_for_full_word (std::uint8_t lead, std::uint8_t trail) noexcept |
| | The DROP rule: \b next to a full-\w maximal run is redundant (\B never is).
|
| |
| constexpr bool | resolve_class_wb_hints (bool full_word, bool word_sub, bool maximal_run, std::uint8_t lead, std::uint8_t trail, std::uint8_t &out_lead, std::uint8_t &out_trail) noexcept |
| | DROP / WRAP policy for class / cp-class loops under optional \b/\B wraps.
|
| |
| constexpr bool | word_ranges_cover_interval_from (char32_t lo, char32_t hi, std::size_t &cursor) noexcept |
| | True if every code point in [lo, hi] is a Unicode word char (word_ranges), resuming the scan at cursor and leaving it past the last range consulted.
|
| |
| constexpr bool | word_ranges_cover_interval (char32_t lo, char32_t hi) noexcept |
| | True if every code point in [lo, hi] is a Unicode word char (covered by word_ranges). Standalone form of word_ranges_cover_interval_from.
|
| |
| constexpr bool | is_unicode_word_subset_cp_class (const cp_class &cc, std::span< const code_range > all_ranges) noexcept |
| | True if cc is a non-empty subset of Unicode \w (safe for maximal-run + \b wrap).
|
| |
| constexpr bool | cp_class_may_contain_ascii_byte (const cp_class &cc, std::uint8_t b) noexcept |
| | Whether byte b could be a member of cc; a delimiter that could hide in a possessive code-point loop makes the delimited fast path decline (pattern_hints::possessive_prefix).
|
| |
| constexpr bool | is_fixed_alternation (std::span< const instr > code, std::uint8_t *out_wb_lead=nullptr, std::uint8_t *out_wb_trail=nullptr, std::uint8_t *out_body_pc=nullptr, std::int32_t *out_branch_count=nullptr) |
| | Alternation of straight-line byte/klass branches, optionally wrapped in \b/\B.
|
| |
| constexpr void | extract_anchoring (std::span< const instr > code, pattern_hints &hints) |
| | Records start anchoring: the first non-save instruction tells whether every match must begin at position 0 (\A/^ non-multiline) or at a line start.
|
| |
| constexpr void | extract_prefix (std::span< const instr > code, pattern_hints &hints) |
| | Collects the required literal prefix and the exact-literal fast-path length.
|
| |
| constexpr void | compute_first_bytes (std::span< const instr > code, std::span< const char_class > classes, std::span< const cp_class > cp_classes, pattern_hints &hints) |
| | Computes the possible first-byte set by a DFS over the epsilon closure of pc 0.
|
| |
| constexpr int | class_range_count (const char_class &klass, std::uint8_t &lo0, std::uint8_t &hi0, std::uint8_t &lo1, std::uint8_t &hi1) |
| | Reports klass as up to two contiguous byte ranges.
|
| |
| REAL_BUILD_COLD constexpr void | detect_fast_shapes (std::span< const instr > code, std::span< const char_class > classes, std::span< const cp_class > cp_classes, std::span< const code_range > cp_ranges, std::int32_t cp_mark_ascii, std::int32_t cp_mark_offset, std::int32_t cp_mark_end, std::span< const lookaround_sub > lookarounds, pattern_hints &hints) |
| | Detects the whole-pattern fast-path shapes and sets their hint flags: class+, fixed-shape straight runs, a single codepoint class (./negated, optional +), an alternation of straight-line branches, and trailing-lookaround class+.
|
| |
| constexpr std::uint16_t | byte_frequency (std::uint8_t b) |
| | Approximate static frequency of a byte in mixed English and source text, per 10000, used only to rank candidate prefilter bytes: a rare required byte (-, @) is a far more selective memchr target than a common first-byte class.
|
| |
| constexpr std::uint8_t | literal_rarest_offset (std::string_view literal) |
| | Offset of the rarest byte of literal by byte_frequency (the first of equals).
|
| |
| constexpr void | extract_rare_byte (std::span< const instr > code, pattern_hints &hints) |
| | Records a required literal byte at a FIXED offset far rarer than the first-byte set (pattern_hints::rare_byte / rare_offset), so the search can memchr that one byte.
|
| |
| constexpr void | extract_rare_discriminant (std::span< const instr > code, pattern_hints &hints) |
| | Arms the rare-discriminant prefilter for shapes like https?://…: fixed prefix (http) + optional mono-byte (s?) + fixed mid with a rare disc (://).
|
| |
| constexpr bool | capture_free_walk_structural (std::span< const instr > code) noexcept |
| | The structural half of pattern_hints::capture_free_walk, that save 0 is the program's first instruction.
|
| |
| constexpr pattern_hints | analyze_program (std::span< const instr > code, std::span< const char_class > classes, std::span< const cp_class > cp_classes, std::span< const code_range > cp_ranges, std::int32_t cp_mark_ascii, std::int32_t cp_mark_offset, std::int32_t cp_mark_end, std::span< const lookaround_sub > lookarounds={}) |
| | Walks a compiled program once to derive its search hints.
|
| |
| constexpr bool | fixed_shape_walk_pays (std::span< const instr > code, const pattern_hints &hints) noexcept |
| | Whether a search over this fixed shape (pattern_hints::fixed_shape) should walk each candidate start rather than run the lazy DFA.
|
| |
| constexpr std::size_t | find_byte (std::string_view text, std::size_t pos, char byte) |
| | Index of byte in text[pos..), or real::npos.
|
| |
| constexpr std::size_t | find_rare_disc_candidate (std::string_view text, std::size_t pos, const pattern_hints &hints, bool *density_abandon=nullptr) |
| | Next candidate start for the rare-discriminant prefilter, or real::npos.
|
| |
| constexpr void | store_literal_density (literal_density &density, std::uint32_t cands, std::size_t origin, std::size_t next) noexcept |
| | Writes the adaptive search's local density back (find_literal_adaptive_rest keeps it in locals while it scans).
|
| |
| std::size_t | find_literal_adaptive_rest (std::string_view text, std::size_t pos, std::string_view literal, std::size_t rare, literal_density &density) |
| | The body of find_literal_adaptive past its first stop: the pair filter for a dense subject, else the rarest-byte scan that counts its stops and judges their density.
|
| |
| std::size_t | find_literal_adaptive (std::string_view text, std::size_t pos, std::string_view literal, std::size_t rare, literal_density &density) |
| | Index of the first occurrence of literal in text[pos..), or real::npos, by its rarest byte while that byte is rare in the subject and by the two-byte block filter once it is not.
|
| |
| bool | alternation_nibbles_supported () |
| | Whether the nibble fingerprint can run here: AArch64 always, x86 when the build enables SSSE3 or, with gcc or clang, when the running CPU has it. A plan built elsewhere never claims one: the fingerprint's reach is shorter than the pairs', and a scan that bounded its blocks by it while masking by the pairs would read past the subject.
|
| |
| constexpr std::size_t | find_literal (std::string_view text, std::size_t pos, std::string_view literal) |
| | Index of the first occurrence of literal in text[pos..), or real::npos.
|
| |
| constexpr std::size_t | find_prefix (std::string_view text, std::size_t pos, std::string_view prefix) |
| | First position >= pos where prefix occurs in text, or npos; the dispatch of find_literal.
|
| |
| std::size_t | find_members (std::string_view text, std::size_t pos, const std::array< std::uint8_t, 8 > &mem, std::uint8_t n) |
| | Least index at or after pos whose byte is one of n members, in ONE pass.
|
| |
| std::size_t | find_folded_literal (std::string_view text, std::size_t pos, std::string_view lit, std::uint16_t folded, std::size_t rare) |
| | The first occurrence at or after pos of a literal some of whose letters match in either case.
|
| |
| constexpr std::size_t | find_line_end_cr (std::string_view text, std::size_t pos) |
| | Where an ECMAScript line ends: the index of the first \n or \r in text[pos..), or real::npos.
|
| |
| constexpr std::size_t | find_bytes_cascade (std::string_view text, std::size_t pos, const char *set, std::uint8_t n) |
| | Index of the first byte in text[pos..) that belongs to a small first-byte set.
|
| |
| constexpr std::size_t | first_high_byte (std::string_view text, std::size_t pos, std::size_t end) |
| | Index of the first byte >= 0x80 in text[pos, end), or end if the range is pure ASCII.
|
| |
| constexpr std::vector< code_range > | coalesce_ranges (std::vector< code_range > ranges) |
| | Sorts ranges and merges overlapping or adjacent ones: the same code points in the fewest ranges.
|
| |
| constexpr std::vector< code_range > | complement_code_ranges (std::vector< code_range > ranges) |
| | Complements a set of code-point ranges within [0x80, 0x10FFFF] (negated classes, in-class \W/\D/\S). Input may be unsorted or overlapping; the gaps come sorted.
|
| |
| constexpr std::uint32_t | single_codepoint_atom (const ast &tree, std::int32_t index) |
| | The code point a node spells, when it is exactly one non-ASCII literal character.
|
| |
| constexpr digit_escape_result | decode_digit_escape (std::string_view text, std::size_t first) |
| | Decodes a \<digit> escape per CPython's rule; shared by the pattern and the replacement-template parsers so the two never drift.
|
| |
| constexpr ast | parse (std::string_view pattern, flags initial_flags=flags::none) |
| | Parses pattern into an ast (convenience over parser).
|
| |
| constexpr bool | is_any_non_ascii (const std::vector< code_range > &ranges) |
| | Whether ranges is exactly the whole non-ASCII space [U+0080, U+10FFFF] — the "any non-ASCII code point" shape emitted by compiler::emit_any_codepoint_class.
|
| |
| constexpr bool | cp_ranges_are_normalised (const std::vector< code_range > &ranges) |
| | Whether ranges is what every consumer of a cp_class requires: each range non-empty, the sequence strictly ascending and disjoint.
|
| |
| constexpr class_def | unicode_casefold (const class_def &in) |
| | Expands a character class to its Unicode simple case-fold closure (text-mode icase).
|
| |
| constexpr bool | node_nullable (const ast &tree, std::int32_t idx) |
| | True if the AST subtree rooted at idx can match the empty string. empty, anchor and lookaround are zero-width, so exactly nullable; byte/klass/any never are.
|
| |
| constexpr bool | subtree_has_nullable_capturing_group (const ast &tree, std::int32_t idx) |
| | True if the AST subtree rooted at idx contains, at any depth, a capturing group (group >= 0) whose body is nullable (node_nullable).
|
| |
| constexpr bool | ast_has_nullable_captured_repeat (const ast &tree, std::int32_t idx) |
| | True if a capturing group with a nullable body sits anywhere under a quantifier (? included): the source of pattern_hints::nullable_captured_repeat. A safe over-approximation: it flags the shape ((\b|x)+ counts), not a proven divergent capture.
|
| |
| constexpr dynamic_program | compile (const ast &tree, flags compile_flags) |
| | Compiles tree to an NFA program (convenience over compiler).
|
| |
| constexpr inner_literal | extract_inner_literal (const ast &tree) |
| | Extract the best required inner literal from a pattern's AST (a pure function on the node pool).
|
| |
| ast | build_prefix_ast (const ast &tree, std::int32_t count, std::int32_t skip=0) |
| | Build the prefix sub-AST: count top-level concat children starting after skip lead children.
|
| |
| std::size_t | prefix_reverse_start (const ast &tree, std::int32_t count, flags compile_flags, std::string_view text, std::size_t h, std::size_t min_start) |
| | The match start for a literal candidate at h: reverse-match the prefix (the first count top-level children) ending at h, bounded below by min_start. Runtime only (the reverse walk is not constexpr).
|
| |
| constexpr std::string_view | c_string_subject (const char *text) noexcept |
| | A C string as a subject. A null pointer is a caller's bug: a debug build stops on it, and a release build reads it as the empty subject, as the C API does, where constructing a std::string_view from it is undefined.
|
| |
| template<typename State , bool Bound, typename Slots > |
| constexpr bool | run_attempt (pike_vm< State, Bound > &vm, const program_view &prog, std::string_view subject, std::size_t pos, run_mode mode, Slots &slots, match_semantics sem=match_semantics::first) |
| | One attempt over a region, as every single search makes it: the trailing-lookaround walk where the pattern has one, else pike_vm::run with its memchr-cascade variant chosen once, here.
|
| |
| constexpr std::size_t | scratch_code_tier (std::size_t code_size) |
| | Rounds a program length up to the scratch capacity tier it shares with its neighbours.
|
| |
| constexpr bool | is_binprop_cp (binprop prop, char32_t cp) |
| | Whether cp has the binary property prop (== the UCD).
|
| |
| constexpr binprop | resolve_binprop (std::string_view loose) |
| | Resolve a loose-normalized binary-property name to its value, or count if unknown.
|
| |
| constexpr std::size_t | find_fold_lower_bound (std::uint32_t cp) |
| | Index of the first entry whose code point is at or after cp, or unicode_fold_table_size if none is. The seek half of find_fold_index, exposed on its own so a caller holding a RANGE enters the table once and walks forward instead of scanning it whole – see real::detail::unicode_casefold.
|
| |
| constexpr std::size_t | find_fold_index (std::uint32_t cp) |
| | Binary-searches unicode_fold_table for cp; returns its index, or unicode_fold_table_size if cp is not cased. An index (not a pointer into the table) keeps this usable in a constant expression on every compiler — g++ rejects a &table[i] != nullptr comparison inside a static_regex. Shared by the parser (is a literal cased?) and the compiler (its fold partners).
|
| |
| constexpr bool | is_gc_cp (gc_property prop, char32_t cp) |
| | Whether cp is in the General_Category property prop (== the UCD).
|
| |
| constexpr gc_property | resolve_gc (std::string_view loose) |
| | Resolve a loose-normalized General_Category name to its property, or count if unknown.
|
| |
| constexpr bool | cp_in_ranges (std::span< const code_range > ranges, char32_t cp) |
| | Binary-searches a sorted, non-overlapping range table for cp. Returns a bool (not a pointer into the table) so it stays constant-evaluable on every compiler.
|
| |
| constexpr bool | is_word_cp (char32_t cp) |
| | Whether cp is a Unicode word code point (== re \w).
|
| |
| constexpr bool | is_digit_cp (char32_t cp) |
| | Whether cp is a Unicode digit code point (== re \d).
|
| |
| constexpr bool | is_space_cp (char32_t cp) |
| | Whether cp is a Unicode whitespace code point (== re \s).
|
| |
| constexpr script | script_of (char32_t cp) |
| | The Script of cp (binary search; Unknown when no range covers it).
|
| |
| constexpr bool | is_script_cp (script sc, char32_t cp) |
| | Whether cp belongs to Script sc (== the UCD).
|
| |
| constexpr script | resolve_script (std::string_view loose) |
| | Resolve a loose-normalized Script name to its value, or count if unknown.
|
| |
| constexpr bool | is_scx_cp (script sc, char32_t cp) |
| | Whether cp is in the Script_Extensions of sc (== the UCD). NOT exclusive: a code point can satisfy this for several script values at once.
|
| |
| constexpr decoded_codepoint | decode_codepoint_strict (std::string_view text, std::size_t pos) |
| | Strictly decodes and validates the UTF-8 sequence at text[pos].
|
| |
| constexpr std::size_t | codepoint_advance (std::string_view text, std::size_t pos) |
| | Number of bytes from pos to the next code-point boundary, for advancing past an empty match during iteration.
|
| |
| constexpr std::size_t | codepoint_retreat (std::string_view text, std::size_t end, std::size_t floor) |
| | Width of the last code point before end, the mirror of codepoint_advance.
|
| |
|
|
constexpr int | max_loop_hops {8} |
| | Cap on how far a jump chain is followed to a loop head (empty-iteration exit routing), shared by every closure walk; a loop join reaches its split in one hop, so eight is headroom, not a knob.
|
| |
| constexpr std::size_t | lazy_dfa_default_byte_budget {std::size_t {64} << 20U} |
| |
| constexpr std::size_t | alternation_trie_min_branches {64} |
| |
| constexpr std::size_t | max_byte_program_size {20000} |
| |
|
constexpr std::size_t | il_warm_floor {4UL * 1024} |
| | Warm-regime IL minimum haystack, in bytes: below it the candidate scan can cost more than the route saves, even with the reverse DFA already built. A cold first scan uses the higher regex_immutables::il_min_haystack.
|
| |
|
constexpr std::size_t | il_short_scan_budget {64UL * 1024} |
| | Subject bytes a regex lets the inner-literal route decline under its floor before it builds what the route needs: the cold floor's least amortization. Short subjects that add up to it have paid for the build as one long subject would, and the build lifts both floors for good. Without it a regex only ever searched on short subjects stays on the bounded backtracker, many times dearer per search than the built route.
|
| |
|
constexpr std::uint32_t | onepass_prefix_warm_calls {8192} |
| | Anchored matches a regex runs before it builds its one-pass table for them. Rent before buying: the build (byte program, table, minimization) costs about as much as 5 000 to 13 000 of these calls made without it, so a regex matched a few times never pays it, and one matched in a loop pays at most about twice what the best choice made in hindsight would have.
|
| |
| constexpr std::array< utf8_second_byte_bounds, 256 > | utf8_second_byte_bounds_table |
| | First-continuation-byte bounds indexed by lead byte (only 0xC2–0xF4 are consulted).
|
| |
|
constexpr std::size_t | default_max_program_size {262144} |
| | max_program_size's default.
|
| |
|
constexpr std::int32_t | default_max_repeat_count {1000} |
| | max_repeat_count's default.
|
| |
|
constexpr std::int32_t | default_max_group_count {32766} |
| | max_group_count's default.
|
| |
|
constexpr std::int32_t | default_max_nesting_depth {200} |
| | max_nesting_depth's default.
|
| |
|
constexpr std::int32_t | default_max_lookaround_length {255} |
| | max_lookaround_length's default.
|
| |
|
constexpr std::size_t | default_max_dfa_states {65536} |
| | max_dfa_states's default.
|
| |
| constexpr std::size_t | max_program_size {default_max_program_size} |
| | Maximum number of NFA instructions in a compiled program — 256 Ki by default (REAL_MAX_PROGRAM_SIZE).
|
| |
|
constexpr std::string_view | program_too_large {"program too large"} |
| | The cause a program past max_program_size is rejected with (real::regex_error::cause).
|
| |
|
constexpr std::int32_t | max_repeat_count {default_max_repeat_count} |
| | Per-quantifier bounded-repeat cap, enforced at parse time (REAL_MAX_REPEAT_COUNT).
|
| |
|
constexpr std::int32_t | max_group_count {default_max_group_count} |
| | Maximum capture groups; bounds slot_count = 2 * (groups + 1) (REAL_MAX_GROUP_COUNT).
|
| |
|
constexpr std::int32_t | max_nesting_depth {default_max_nesting_depth} |
| | Maximum parser recursion depth; prevents stack overflow on deep nesting (REAL_MAX_NESTING_DEPTH).
|
| |
|
constexpr std::int32_t | max_lookaround_length {default_max_lookaround_length} |
| | Maximum bytes a bounded lookaround sub-pattern may consume (its L_max); bounding it keeps per-position evaluation linear (REAL_MAX_LOOKAROUND_LENGTH).
|
| |
| constexpr std::size_t | max_dfa_states {default_max_dfa_states} |
| | Maximum DFA states (opt-in real::dfa; REAL_MAX_DFA_STATES).
|
| |
|
constexpr std::uint64_t | fnv1a_offset_basis {14695981039346656037ULL} |
| | FNV-1a 64-bit offset basis. Reference it, never re-type it: a wrong basis still hashes, just not as FNV-1a, and nothing downstream notices.
|
| |
|
constexpr std::uint64_t | fnv1a_prime {1099511628211ULL} |
| | FNV-1a 64-bit prime, paired with fnv1a_offset_basis.
|
| |
|
constexpr std::size_t | bounded_backtrack_max_slots {34} |
| | Capture slots the bounded backtracker carries in a fixed array (16 groups and group 0).
|
| |
| constexpr std::size_t | bounded_backtrack_bits {8192} |
| | The bounded backtracker's budget: one bit per (instruction, position), (n + 1) x m bits for an n-byte subject and an m-instruction program.
|
| |
| constexpr std::size_t | max_dfa_byte_program {512} |
| | Cap on a pattern's expanded byte program before subset construction runs on it.
|
| |
|
constexpr std::uint32_t | dfa_no_rule {std::numeric_limits<std::uint32_t>::max()} |
| | dfa_tables::accept's "this state does not accept" marker.
|
| |
| constexpr std::size_t | ac_memory_budget_default {std::size_t {32} << 20U} |
| |
|
constexpr std::size_t | ac_max_branch_expansion = 64 |
| | Maximum class sequences one branch may expand into (one per combination of the classes its positions span); past this the WHOLE pattern takes the ordinary pattern_hints::fixed_alternation route. A case-folded letter is one class unless another branch tells its cases apart.
|
| |
|
constexpr std::size_t | fixed_shape_walk_max_width {8} |
| | Widest unfiltered shape kept on the walk.
|
| |
|
constexpr std::uint32_t | rare_disc_fail_abandon {32} |
| | Consecutive disc hits that fail back-verify before the density gate trips. Dense : filler (e.g. a:b:c:d…) makes memchr+verify lose to a selective http prefix.
|
| |
|
constexpr std::uint32_t | literal_dense_min_cands {8} |
| | Stops the rarest-byte scan makes before its density is judged: fewer say nothing.
|
| |
| constexpr std::size_t | literal_dense_gap {64} |
| | Mean bytes between stops below which the rarest byte counts as common: under it, a stop costs more than the pair filter spends crossing that many bytes.
|
| |
|
constexpr std::size_t | alternation_nibbles_min_branches {3} |
| | Fewest branches for which the fingerprint replaces the pairs. Two pairs are two compares a block, a fingerprint six table lookups: on x86 (SSSE3) two branches lose by the fingerprint and three break even; AArch64's lookups are cheap enough that two branches already gain.
|
| |
|
constexpr std::size_t | alternation_sample_min {4096} |
| | Shorter rests are scanned by the first bytes, unsampled (at least the sample and its reach).
|
| |
|
constexpr std::size_t | alternation_sample_bytes {512} |
| | Bytes sampled for the first bytes' density.
|
| |
|
constexpr std::size_t | alternation_wide_min_branches {2} |
| | Fewest branches the wide route considers: branches that open on classes outgrow the small set with a few.
|
| |
|
constexpr std::size_t | alternation_wide_max_branches {16} |
| | Most branches the fingerprint's plan holds.
|
| |
|
constexpr std::size_t | alternation_wide_false_budget {256} |
| | Past this many false candidates times branches in a sample, the automaton scans cheaper than the fingerprint and its verifier.
|
| |
|
constexpr std::size_t | alternation_dense_gap {32} |
| | Mean bytes between first-byte hits in the sample below which the pair filter takes over (both ISAs: where the first-byte scan and the pair filter crossed on 500 KB of log lines).
|
| |
|
constexpr std::size_t | alternation_dense_gap_nibbles {128} |
| | The same threshold for a plan masking by the nibble fingerprint, whose cost per block is fixed: it crossed the first-byte loop at 256-512 bytes between false stops (arm64, x86-64 AVX2, 3 to 10 branches); at 128 its worst ratio was 0.77 of the loop's time.
|
| |
| constexpr bool | have_members_scan {false} |
| | Whether a consumer should ARM a filter on find_members (one ISA only, measured).
|
| |
|
constexpr std::uint32_t | not_a_single_codepoint {0xFFFFFFFFU} |
| | Returned by single_codepoint_atom when the node is not one code point's bytes.
|
| |
| constexpr std::uint32_t | optional_literal_min_score {1800} |
| | The least inner_literal::score an inner run needs to be kept when an optional precedes or follows it.
|
| |
|
constexpr std::size_t | inner_literal_max {16} |
| | The most bytes an inner literal keeps; past this a longer needle costs storage without shrinking the candidate set much.
|
| |
|
constexpr const char * | unicode_binprop_unidata_version {"16.0.0"} |
| | The Unicode data version these tables were generated from.
|
| |
| constexpr code_range | binprop_ASCII_Hex_Digit_ranges [] |
| | \p{ASCII_Hex_Digit} — 3 ranges, 22 code points.
|
| |
|
constexpr code_range | binprop_Alphabetic_ranges [] |
| | \p{Alphabetic} — 757 ranges, 142759 code points.
|
| |
| constexpr code_range | binprop_Bidi_Control_ranges [] |
| | \p{Bidi_Control} — 4 ranges, 12 code points.
|
| |
|
constexpr code_range | binprop_Case_Ignorable_ranges [] |
| | \p{Case_Ignorable} — 452 ranges, 2749 code points.
|
| |
|
constexpr code_range | binprop_Cased_ranges [] |
| | \p{Cased} — 159 ranges, 4578 code points.
|
| |
|
constexpr code_range | binprop_Changes_When_Casefolded_ranges [] |
| | \p{Changes_When_Casefolded} — 626 ranges, 1533 code points.
|
| |
|
constexpr code_range | binprop_Changes_When_Casemapped_ranges [] |
| | \p{Changes_When_Casemapped} — 131 ranges, 2981 code points.
|
| |
|
constexpr code_range | binprop_Changes_When_Lowercased_ranges [] |
| | \p{Changes_When_Lowercased} — 614 ranges, 1460 code points.
|
| |
|
constexpr code_range | binprop_Changes_When_Titlecased_ranges [] |
| | \p{Changes_When_Titlecased} — 629 ranges, 1479 code points.
|
| |
|
constexpr code_range | binprop_Changes_When_Uppercased_ranges [] |
| | \p{Changes_When_Uppercased} — 630 ranges, 1552 code points.
|
| |
| constexpr code_range | binprop_Dash_ranges [] |
| | \p{Dash} — 24 ranges, 31 code points.
|
| |
| constexpr code_range | binprop_Default_Ignorable_Code_Point_ranges [] |
| | \p{Default_Ignorable_Code_Point} — 17 ranges, 4174 code points.
|
| |
| constexpr code_range | binprop_Deprecated_ranges [] |
| | \p{Deprecated} — 8 ranges, 15 code points.
|
| |
|
constexpr code_range | binprop_Diacritic_ranges [] |
| | \p{Diacritic} — 214 ranges, 1178 code points.
|
| |
|
constexpr code_range | binprop_Emoji_ranges [] |
| | \p{Emoji} — 150 ranges, 1431 code points.
|
| |
| constexpr code_range | binprop_Emoji_Component_ranges [] |
| | \p{Emoji_Component} — 10 ranges, 146 code points.
|
| |
| constexpr code_range | binprop_Emoji_Modifier_ranges [] |
| | \p{Emoji_Modifier} — 1 ranges, 5 code points.
|
| |
|
constexpr code_range | binprop_Emoji_Modifier_Base_ranges [] |
| | \p{Emoji_Modifier_Base} — 40 ranges, 134 code points.
|
| |
|
constexpr code_range | binprop_Emoji_Presentation_ranges [] |
| | \p{Emoji_Presentation} — 80 ranges, 1212 code points.
|
| |
|
constexpr code_range | binprop_Extended_Pictographic_ranges [] |
| | \p{Extended_Pictographic} — 78 ranges, 3537 code points.
|
| |
|
constexpr code_range | binprop_Extender_ranges [] |
| | \p{Extender} — 41 ranges, 59 code points.
|
| |
|
constexpr code_range | binprop_Grapheme_Base_ranges [] |
| | \p{Grapheme_Base} — 894 ranges, 152730 code points.
|
| |
|
constexpr code_range | binprop_Grapheme_Extend_ranges [] |
| | \p{Grapheme_Extend} — 375 ranges, 2193 code points.
|
| |
|
constexpr code_range | binprop_Grapheme_Link_ranges [] |
| | \p{Grapheme_Link} — 58 ranges, 69 code points.
|
| |
| constexpr code_range | binprop_Hex_Digit_ranges [] |
| | \p{Hex_Digit} — 6 ranges, 44 code points.
|
| |
| constexpr code_range | binprop_Hyphen_ranges [] |
| | \p{Hyphen} — 10 ranges, 11 code points.
|
| |
| constexpr code_range | binprop_IDS_Binary_Operator_ranges [] |
| | \p{IDS_Binary_Operator} — 3 ranges, 13 code points.
|
| |
| constexpr code_range | binprop_IDS_Trinary_Operator_ranges [] |
| | \p{IDS_Trinary_Operator} — 1 ranges, 2 code points.
|
| |
| constexpr code_range | binprop_IDS_Unary_Operator_ranges [] |
| | \p{IDS_Unary_Operator} — 1 ranges, 2 code points.
|
| |
| constexpr code_range | binprop_ID_Compat_Math_Continue_ranges [] |
| | \p{ID_Compat_Math_Continue} — 18 ranges, 43 code points.
|
| |
| constexpr code_range | binprop_ID_Compat_Math_Start_ranges [] |
| | \p{ID_Compat_Math_Start} — 13 ranges, 13 code points.
|
| |
|
constexpr code_range | binprop_ID_Continue_ranges [] |
| | \p{ID_Continue} — 793 ranges, 144541 code points.
|
| |
|
constexpr code_range | binprop_ID_Start_ranges [] |
| | \p{ID_Start} — 677 ranges, 141269 code points.
|
| |
| constexpr code_range | binprop_Ideographic_ranges [] |
| | \p{Ideographic} — 21 ranges, 106477 code points.
|
| |
| constexpr code_range | binprop_Join_Control_ranges [] |
| | \p{Join_Control} — 1 ranges, 2 code points.
|
| |
| constexpr code_range | binprop_Logical_Order_Exception_ranges [] |
| | \p{Logical_Order_Exception} — 7 ranges, 19 code points.
|
| |
|
constexpr code_range | binprop_Lowercase_ranges [] |
| | \p{Lowercase} — 675 ranges, 2569 code points.
|
| |
|
constexpr code_range | binprop_Math_ranges [] |
| | \p{Math} — 139 ranges, 2312 code points.
|
| |
| constexpr code_range | binprop_Modifier_Combining_Mark_ranges [] |
| | \p{Modifier_Combining_Mark} — 9 ranges, 14 code points.
|
| |
| constexpr code_range | binprop_Noncharacter_Code_Point_ranges [] |
| | \p{Noncharacter_Code_Point} — 18 ranges, 66 code points.
|
| |
|
constexpr code_range | binprop_Other_Alphabetic_ranges [] |
| | \p{Other_Alphabetic} — 250 ranges, 1495 code points.
|
| |
| constexpr code_range | binprop_Other_Default_Ignorable_Code_Point_ranges [] |
| | \p{Other_Default_Ignorable_Code_Point} — 11 ranges, 3776 code points.
|
| |
|
constexpr code_range | binprop_Other_Grapheme_Extend_ranges [] |
| | \p{Other_Grapheme_Extend} — 49 ranges, 160 code points.
|
| |
| constexpr code_range | binprop_Other_ID_Continue_ranges [] |
| | \p{Other_ID_Continue} — 7 ranges, 16 code points.
|
| |
| constexpr code_range | binprop_Other_ID_Start_ranges [] |
| | \p{Other_ID_Start} — 4 ranges, 6 code points.
|
| |
| constexpr code_range | binprop_Other_Lowercase_ranges [] |
| | \p{Other_Lowercase} — 28 ranges, 311 code points.
|
| |
|
constexpr code_range | binprop_Other_Math_ranges [] |
| | \p{Other_Math} — 134 ranges, 1362 code points.
|
| |
| constexpr code_range | binprop_Other_Uppercase_ranges [] |
| | \p{Other_Uppercase} — 5 ranges, 120 code points.
|
| |
| constexpr code_range | binprop_Pattern_Syntax_ranges [] |
| | \p{Pattern_Syntax} — 28 ranges, 2760 code points.
|
| |
| constexpr code_range | binprop_Pattern_White_Space_ranges [] |
| | \p{Pattern_White_Space} — 5 ranges, 11 code points.
|
| |
| constexpr code_range | binprop_Prepended_Concatenation_Mark_ranges [] |
| | \p{Prepended_Concatenation_Mark} — 7 ranges, 13 code points.
|
| |
| constexpr code_range | binprop_Quotation_Mark_ranges [] |
| | \p{Quotation_Mark} — 13 ranges, 30 code points.
|
| |
| constexpr code_range | binprop_Radical_ranges [] |
| | \p{Radical} — 3 ranges, 329 code points.
|
| |
| constexpr code_range | binprop_Regional_Indicator_ranges [] |
| | \p{Regional_Indicator} — 1 ranges, 26 code points.
|
| |
|
constexpr code_range | binprop_Sentence_Terminal_ranges [] |
| | \p{Sentence_Terminal} — 88 ranges, 170 code points.
|
| |
|
constexpr code_range | binprop_Soft_Dotted_ranges [] |
| | \p{Soft_Dotted} — 34 ranges, 50 code points.
|
| |
|
constexpr code_range | binprop_Terminal_Punctuation_ranges [] |
| | \p{Terminal_Punctuation} — 116 ranges, 291 code points.
|
| |
| constexpr code_range | binprop_Unified_Ideograph_ranges [] |
| | \p{Unified_Ideograph} — 17 ranges, 97680 code points.
|
| |
|
constexpr code_range | binprop_Uppercase_ranges [] |
| | \p{Uppercase} — 656 ranges, 1978 code points.
|
| |
| constexpr code_range | binprop_Variation_Selector_ranges [] |
| | \p{Variation_Selector} — 4 ranges, 260 code points.
|
| |
| constexpr code_range | binprop_White_Space_ranges [] |
| | \p{White_Space} — 10 ranges, 25 code points.
|
| |
|
constexpr code_range | binprop_XID_Continue_ranges [] |
| | \p{XID_Continue} — 800 ranges, 144522 code points.
|
| |
|
constexpr code_range | binprop_XID_Start_ranges [] |
| | \p{XID_Start} — 684 ranges, 141246 code points.
|
| |
|
constexpr std::span< const code_range > | binprop_ranges [] |
| | Range table indexed by binprop (parallel to the enum order).
|
| |
|
constexpr binprop_alias_entry | binprop_aliases [] |
| | Binary-property names, loose-keyed; for the \p{...} parser (no namespace prefix, same as PCRE2: \p{Alphabetic}, not \p{bp=Alphabetic}).
|
| |
|
constexpr const char * | unicode_fold_unidata_version {"16.0.0"} |
| | The Unicode data version these orbits were generated from.
|
| |
|
constexpr fold_entry | unicode_fold_table [] |
| | Fold orbits, sorted by fold_entry::cp for binary search.
|
| |
|
constexpr std::size_t | unicode_fold_table_size {2940} |
| | Number of entries in unicode_fold_table.
|
| |
|
constexpr const char * | unicode_property_unidata_version {"16.0.0"} |
| | The Unicode data version these tables were generated from.
|
| |
|
constexpr code_range | gc_Lu_ranges [] |
| | \p{Lu} — 651 ranges, 1858 code points.
|
| |
|
constexpr code_range | gc_Ll_ranges [] |
| | \p{Ll} — 662 ranges, 2258 code points.
|
| |
| constexpr code_range | gc_Lt_ranges [] |
| | \p{Lt} — 10 ranges, 31 code points.
|
| |
|
constexpr code_range | gc_Lm_ranges [] |
| | \p{Lm} — 75 ranges, 404 code points.
|
| |
|
constexpr code_range | gc_Lo_ranges [] |
| | \p{Lo} — 528 ranges, 136477 code points.
|
| |
|
constexpr code_range | gc_Mn_ranges [] |
| | \p{Mn} — 357 ranges, 2020 code points.
|
| |
|
constexpr code_range | gc_Mc_ranges [] |
| | \p{Mc} — 190 ranges, 468 code points.
|
| |
| constexpr code_range | gc_Me_ranges [] |
| | \p{Me} — 5 ranges, 13 code points.
|
| |
|
constexpr code_range | gc_Nd_ranges [] |
| | \p{Nd} — 71 ranges, 760 code points.
|
| |
| constexpr code_range | gc_Nl_ranges [] |
| | \p{Nl} — 12 ranges, 236 code points.
|
| |
|
constexpr code_range | gc_No_ranges [] |
| | \p{No} — 72 ranges, 915 code points.
|
| |
| constexpr code_range | gc_Pc_ranges [] |
| | \p{Pc} — 6 ranges, 10 code points.
|
| |
| constexpr code_range | gc_Pd_ranges [] |
| | \p{Pd} — 20 ranges, 27 code points.
|
| |
|
constexpr code_range | gc_Ps_ranges [] |
| | \p{Ps} — 79 ranges, 79 code points.
|
| |
|
constexpr code_range | gc_Pe_ranges [] |
| | \p{Pe} — 76 ranges, 77 code points.
|
| |
| constexpr code_range | gc_Pi_ranges [] |
| | \p{Pi} — 11 ranges, 12 code points.
|
| |
| constexpr code_range | gc_Pf_ranges [] |
| | \p{Pf} — 10 ranges, 10 code points.
|
| |
|
constexpr code_range | gc_Po_ranges [] |
| | \p{Po} — 193 ranges, 640 code points.
|
| |
|
constexpr code_range | gc_Sm_ranges [] |
| | \p{Sm} — 65 ranges, 950 code points.
|
| |
| constexpr code_range | gc_Sc_ranges [] |
| | \p{Sc} — 21 ranges, 63 code points.
|
| |
|
constexpr code_range | gc_Sk_ranges [] |
| | \p{Sk} — 31 ranges, 125 code points.
|
| |
|
constexpr code_range | gc_So_ranges [] |
| | \p{So} — 187 ranges, 7376 code points.
|
| |
| constexpr code_range | gc_Zs_ranges [] |
| | \p{Zs} — 7 ranges, 17 code points.
|
| |
| constexpr code_range | gc_Zl_ranges [] |
| | \p{Zl} — 1 ranges, 1 code points.
|
| |
| constexpr code_range | gc_Zp_ranges [] |
| | \p{Zp} — 1 ranges, 1 code points.
|
| |
| constexpr code_range | gc_Cc_ranges [] |
| | \p{Cc} — 2 ranges, 65 code points.
|
| |
| constexpr code_range | gc_Cf_ranges [] |
| | \p{Cf} — 21 ranges, 170 code points.
|
| |
| constexpr code_range | gc_Co_ranges [] |
| | \p{Co} — 3 ranges, 137468 code points.
|
| |
|
constexpr code_range | gc_Cn_ranges [] |
| | \p{Cn} — 731 ranges, 819533 code points.
|
| |
|
constexpr code_range | gc_L_ranges [] |
| | \p{L} — 677 ranges, 141028 code points.
|
| |
|
constexpr code_range | gc_M_ranges [] |
| | \p{M} — 321 ranges, 2501 code points.
|
| |
|
constexpr code_range | gc_N_ranges [] |
| | \p{N} — 144 ranges, 1911 code points.
|
| |
|
constexpr code_range | gc_P_ranges [] |
| | \p{P} — 198 ranges, 855 code points.
|
| |
|
constexpr code_range | gc_S_ranges [] |
| | \p{S} — 236 ranges, 8514 code points.
|
| |
| constexpr code_range | gc_Z_ranges [] |
| | \p{Z} — 8 ranges, 19 code points.
|
| |
|
constexpr code_range | gc_C_ranges [] |
| | \p{C} — 737 ranges, 957236 code points.
|
| |
|
constexpr std::span< const code_range > | gc_property_ranges [] |
| | Range table indexed by gc_property (parallel to the enum order).
|
| |
|
constexpr gc_alias_entry | gc_aliases [] |
| | Short codes (Lu) and long names (Uppercase_Letter), loose-keyed; for the \p{...} parser.
|
| |
|
constexpr const char * | unicode_props_unidata_version {"16.0.0"} |
| | The Unicode data version these tables were generated from.
|
| |
|
constexpr code_range | word_ranges [] |
| | Code-point ranges matched by \w (771 ranges, 142940 code points).
|
| |
|
constexpr std::size_t | word_ranges_size {771} |
| | Number of ranges in word_ranges.
|
| |
|
constexpr code_range | digit_ranges [] |
| | Code-point ranges matched by \d (71 ranges, 760 code points).
|
| |
|
constexpr std::size_t | digit_ranges_size {71} |
| | Number of ranges in digit_ranges.
|
| |
| constexpr code_range | space_ranges [] |
| | Code-point ranges matched by \s (10 ranges, 29 code points).
|
| |
|
constexpr std::size_t | space_ranges_size {10} |
| | Number of ranges in space_ranges.
|
| |
|
constexpr const char * | unicode_script_unidata_version {"16.0.0"} |
| | The Unicode data version these tables were generated from.
|
| |
|
constexpr script_range | script_ranges [] |
| | Script partition — 979 ranges, sorted and disjoint.
|
| |
|
constexpr script_alias_entry | script_aliases [] |
| | Script names, loose-keyed; for the \p{sc=...} / \p{scx=...} parsers. Both the long name (Latin) and the short UAX24/ISO 15924 code (Latn) resolve to the same value – \p{scx=...} states its overrides in short codes only, and since this table is shared, \p{sc=...} gains the short form too.
|
| |
|
constexpr const char * | unicode_scx_unidata_version {"16.0.0"} |
| | The Unicode data version these tables were generated from.
|
| |
| constexpr code_range | scx_Adlam_ranges [] |
| | \p{scx=Adlam} — 7 ranges, 92 code points.
|
| |
| constexpr code_range | scx_Ahom_ranges [] |
| | \p{scx=Ahom} — 3 ranges, 65 code points.
|
| |
| constexpr code_range | scx_Anatolian_Hieroglyphs_ranges [] |
| | \p{scx=Anatolian_Hieroglyphs} — 1 ranges, 583 code points.
|
| |
|
constexpr code_range | scx_Arabic_ranges [] |
| | \p{scx=Arabic} — 55 ranges, 1421 code points.
|
| |
| constexpr code_range | scx_Armenian_ranges [] |
| | \p{scx=Armenian} — 5 ranges, 97 code points.
|
| |
| constexpr code_range | scx_Avestan_ranges [] |
| | \p{scx=Avestan} — 4 ranges, 64 code points.
|
| |
| constexpr code_range | scx_Balinese_ranges [] |
| | \p{scx=Balinese} — 2 ranges, 127 code points.
|
| |
| constexpr code_range | scx_Bamum_ranges [] |
| | \p{scx=Bamum} — 2 ranges, 657 code points.
|
| |
| constexpr code_range | scx_Bassa_Vah_ranges [] |
| | \p{scx=Bassa_Vah} — 2 ranges, 36 code points.
|
| |
| constexpr code_range | scx_Batak_ranges [] |
| | \p{scx=Batak} — 2 ranges, 56 code points.
|
| |
| constexpr code_range | scx_Bengali_ranges [] |
| | \p{scx=Bengali} — 27 ranges, 114 code points.
|
| |
| constexpr code_range | scx_Bhaiksuki_ranges [] |
| | \p{scx=Bhaiksuki} — 4 ranges, 97 code points.
|
| |
| constexpr code_range | scx_Bopomofo_ranges [] |
| | \p{scx=Bopomofo} — 15 ranges, 122 code points.
|
| |
| constexpr code_range | scx_Brahmi_ranges [] |
| | \p{scx=Brahmi} — 3 ranges, 115 code points.
|
| |
| constexpr code_range | scx_Braille_ranges [] |
| | \p{scx=Braille} — 1 ranges, 256 code points.
|
| |
| constexpr code_range | scx_Buginese_ranges [] |
| | \p{scx=Buginese} — 3 ranges, 31 code points.
|
| |
| constexpr code_range | scx_Buhid_ranges [] |
| | \p{scx=Buhid} — 2 ranges, 22 code points.
|
| |
| constexpr code_range | scx_Canadian_Aboriginal_ranges [] |
| | \p{scx=Canadian_Aboriginal} — 3 ranges, 726 code points.
|
| |
| constexpr code_range | scx_Carian_ranges [] |
| | \p{scx=Carian} — 5 ranges, 53 code points.
|
| |
| constexpr code_range | scx_Caucasian_Albanian_ranges [] |
| | \p{scx=Caucasian_Albanian} — 5 ranges, 56 code points.
|
| |
| constexpr code_range | scx_Chakma_ranges [] |
| | \p{scx=Chakma} — 4 ranges, 91 code points.
|
| |
| constexpr code_range | scx_Cham_ranges [] |
| | \p{scx=Cham} — 4 ranges, 83 code points.
|
| |
| constexpr code_range | scx_Cherokee_ranges [] |
| | \p{scx=Cherokee} — 8 ranges, 182 code points.
|
| |
| constexpr code_range | scx_Chorasmian_ranges [] |
| | \p{scx=Chorasmian} — 1 ranges, 28 code points.
|
| |
|
constexpr code_range | scx_Common_ranges [] |
| | \p{scx=Common} — 159 ranges, 8585 code points.
|
| |
| constexpr code_range | scx_Coptic_ranges [] |
| | \p{scx=Coptic} — 10 ranges, 173 code points.
|
| |
| constexpr code_range | scx_Cuneiform_ranges [] |
| | \p{scx=Cuneiform} — 4 ranges, 1234 code points.
|
| |
| constexpr code_range | scx_Cypriot_ranges [] |
| | \p{scx=Cypriot} — 9 ranges, 112 code points.
|
| |
| constexpr code_range | scx_Cypro_Minoan_ranges [] |
| | \p{scx=Cypro_Minoan} — 2 ranges, 101 code points.
|
| |
| constexpr code_range | scx_Cyrillic_ranges [] |
| | \p{scx=Cyrillic} — 18 ranges, 521 code points.
|
| |
| constexpr code_range | scx_Deseret_ranges [] |
| | \p{scx=Deseret} — 1 ranges, 80 code points.
|
| |
| constexpr code_range | scx_Devanagari_ranges [] |
| | \p{scx=Devanagari} — 9 ranges, 221 code points.
|
| |
| constexpr code_range | scx_Dives_Akuru_ranges [] |
| | \p{scx=Dives_Akuru} — 8 ranges, 72 code points.
|
| |
| constexpr code_range | scx_Dogra_ranges [] |
| | \p{scx=Dogra} — 3 ranges, 82 code points.
|
| |
| constexpr code_range | scx_Duployan_ranges [] |
| | \p{scx=Duployan} — 10 ranges, 154 code points.
|
| |
| constexpr code_range | scx_Egyptian_Hieroglyphs_ranges [] |
| | \p{scx=Egyptian_Hieroglyphs} — 2 ranges, 5105 code points.
|
| |
| constexpr code_range | scx_Elbasan_ranges [] |
| | \p{scx=Elbasan} — 3 ranges, 42 code points.
|
| |
| constexpr code_range | scx_Elymaic_ranges [] |
| | \p{scx=Elymaic} — 1 ranges, 23 code points.
|
| |
|
constexpr code_range | scx_Ethiopic_ranges [] |
| | \p{scx=Ethiopic} — 37 ranges, 524 code points.
|
| |
| constexpr code_range | scx_Garay_ranges [] |
| | \p{scx=Garay} — 6 ranges, 72 code points.
|
| |
| constexpr code_range | scx_Georgian_ranges [] |
| | \p{scx=Georgian} — 13 ranges, 178 code points.
|
| |
| constexpr code_range | scx_Glagolitic_ranges [] |
| | \p{scx=Glagolitic} — 16 ranges, 144 code points.
|
| |
| constexpr code_range | scx_Gothic_ranges [] |
| | \p{scx=Gothic} — 5 ranges, 32 code points.
|
| |
| constexpr code_range | scx_Grantha_ranges [] |
| | \p{scx=Grantha} — 25 ranges, 116 code points.
|
| |
|
constexpr code_range | scx_Greek_ranges [] |
| | \p{scx=Greek} — 44 ranges, 531 code points.
|
| |
| constexpr code_range | scx_Gujarati_ranges [] |
| | \p{scx=Gujarati} — 17 ranges, 105 code points.
|
| |
| constexpr code_range | scx_Gunjala_Gondi_ranges [] |
| | \p{scx=Gunjala_Gondi} — 8 ranges, 66 code points.
|
| |
| constexpr code_range | scx_Gurmukhi_ranges [] |
| | \p{scx=Gurmukhi} — 19 ranges, 94 code points.
|
| |
| constexpr code_range | scx_Gurung_Khema_ranges [] |
| | \p{scx=Gurung_Khema} — 2 ranges, 59 code points.
|
| |
|
constexpr code_range | scx_Han_ranges [] |
| | \p{scx=Han} — 42 ranges, 99338 code points.
|
| |
| constexpr code_range | scx_Hangul_ranges [] |
| | \p{scx=Hangul} — 21 ranges, 11775 code points.
|
| |
| constexpr code_range | scx_Hanifi_Rohingya_ranges [] |
| | \p{scx=Hanifi_Rohingya} — 7 ranges, 55 code points.
|
| |
| constexpr code_range | scx_Hanunoo_ranges [] |
| | \p{scx=Hanunoo} — 1 ranges, 23 code points.
|
| |
| constexpr code_range | scx_Hatran_ranges [] |
| | \p{scx=Hatran} — 3 ranges, 26 code points.
|
| |
| constexpr code_range | scx_Hebrew_ranges [] |
| | \p{scx=Hebrew} — 10 ranges, 136 code points.
|
| |
| constexpr code_range | scx_Hiragana_ranges [] |
| | \p{scx=Hiragana} — 17 ranges, 433 code points.
|
| |
| constexpr code_range | scx_Imperial_Aramaic_ranges [] |
| | \p{scx=Imperial_Aramaic} — 2 ranges, 31 code points.
|
| |
| constexpr code_range | scx_Inherited_ranges [] |
| | \p{scx=Inherited} — 28 ranges, 558 code points.
|
| |
| constexpr code_range | scx_Inscriptional_Pahlavi_ranges [] |
| | \p{scx=Inscriptional_Pahlavi} — 2 ranges, 27 code points.
|
| |
| constexpr code_range | scx_Inscriptional_Parthian_ranges [] |
| | \p{scx=Inscriptional_Parthian} — 2 ranges, 30 code points.
|
| |
| constexpr code_range | scx_Javanese_ranges [] |
| | \p{scx=Javanese} — 3 ranges, 91 code points.
|
| |
| constexpr code_range | scx_Kaithi_ranges [] |
| | \p{scx=Kaithi} — 5 ranges, 89 code points.
|
| |
| constexpr code_range | scx_Kannada_ranges [] |
| | \p{scx=Kannada} — 21 ranges, 107 code points.
|
| |
| constexpr code_range | scx_Katakana_ranges [] |
| | \p{scx=Katakana} — 22 ranges, 375 code points.
|
| |
| constexpr code_range | scx_Kawi_ranges [] |
| | \p{scx=Kawi} — 3 ranges, 87 code points.
|
| |
| constexpr code_range | scx_Kayah_Li_ranges [] |
| | \p{scx=Kayah_Li} — 1 ranges, 48 code points.
|
| |
| constexpr code_range | scx_Kharoshthi_ranges [] |
| | \p{scx=Kharoshthi} — 8 ranges, 68 code points.
|
| |
| constexpr code_range | scx_Khitan_Small_Script_ranges [] |
| | \p{scx=Khitan_Small_Script} — 3 ranges, 472 code points.
|
| |
| constexpr code_range | scx_Khmer_ranges [] |
| | \p{scx=Khmer} — 4 ranges, 146 code points.
|
| |
| constexpr code_range | scx_Khojki_ranges [] |
| | \p{scx=Khojki} — 4 ranges, 85 code points.
|
| |
| constexpr code_range | scx_Khudawadi_ranges [] |
| | \p{scx=Khudawadi} — 4 ranges, 81 code points.
|
| |
| constexpr code_range | scx_Kirat_Rai_ranges [] |
| | \p{scx=Kirat_Rai} — 1 ranges, 58 code points.
|
| |
| constexpr code_range | scx_Lao_ranges [] |
| | \p{scx=Lao} — 11 ranges, 83 code points.
|
| |
|
constexpr code_range | scx_Latin_ranges [] |
| | \p{scx=Latin} — 65 ranges, 1555 code points.
|
| |
| constexpr code_range | scx_Lepcha_ranges [] |
| | \p{scx=Lepcha} — 3 ranges, 74 code points.
|
| |
| constexpr code_range | scx_Limbu_ranges [] |
| | \p{scx=Limbu} — 6 ranges, 69 code points.
|
| |
| constexpr code_range | scx_Linear_A_ranges [] |
| | \p{scx=Linear_A} — 4 ranges, 386 code points.
|
| |
| constexpr code_range | scx_Linear_B_ranges [] |
| | \p{scx=Linear_B} — 10 ranges, 268 code points.
|
| |
| constexpr code_range | scx_Lisu_ranges [] |
| | \p{scx=Lisu} — 5 ranges, 53 code points.
|
| |
| constexpr code_range | scx_Lycian_ranges [] |
| | \p{scx=Lycian} — 2 ranges, 30 code points.
|
| |
| constexpr code_range | scx_Lydian_ranges [] |
| | \p{scx=Lydian} — 4 ranges, 29 code points.
|
| |
| constexpr code_range | scx_Mahajani_ranges [] |
| | \p{scx=Mahajani} — 4 ranges, 62 code points.
|
| |
| constexpr code_range | scx_Makasar_ranges [] |
| | \p{scx=Makasar} — 1 ranges, 25 code points.
|
| |
| constexpr code_range | scx_Malayalam_ranges [] |
| | \p{scx=Malayalam} — 12 ranges, 127 code points.
|
| |
| constexpr code_range | scx_Mandaic_ranges [] |
| | \p{scx=Mandaic} — 3 ranges, 30 code points.
|
| |
| constexpr code_range | scx_Manichaean_ranges [] |
| | \p{scx=Manichaean} — 3 ranges, 52 code points.
|
| |
| constexpr code_range | scx_Marchen_ranges [] |
| | \p{scx=Marchen} — 3 ranges, 68 code points.
|
| |
| constexpr code_range | scx_Masaram_Gondi_ranges [] |
| | \p{scx=Masaram_Gondi} — 8 ranges, 77 code points.
|
| |
| constexpr code_range | scx_Medefaidrin_ranges [] |
| | \p{scx=Medefaidrin} — 1 ranges, 91 code points.
|
| |
| constexpr code_range | scx_Meetei_Mayek_ranges [] |
| | \p{scx=Meetei_Mayek} — 3 ranges, 79 code points.
|
| |
| constexpr code_range | scx_Mende_Kikakui_ranges [] |
| | \p{scx=Mende_Kikakui} — 2 ranges, 213 code points.
|
| |
| constexpr code_range | scx_Meroitic_Cursive_ranges [] |
| | \p{scx=Meroitic_Cursive} — 3 ranges, 90 code points.
|
| |
| constexpr code_range | scx_Meroitic_Hieroglyphs_ranges [] |
| | \p{scx=Meroitic_Hieroglyphs} — 2 ranges, 33 code points.
|
| |
| constexpr code_range | scx_Miao_ranges [] |
| | \p{scx=Miao} — 3 ranges, 149 code points.
|
| |
| constexpr code_range | scx_Modi_ranges [] |
| | \p{scx=Modi} — 3 ranges, 89 code points.
|
| |
| constexpr code_range | scx_Mongolian_ranges [] |
| | \p{scx=Mongolian} — 7 ranges, 178 code points.
|
| |
| constexpr code_range | scx_Mro_ranges [] |
| | \p{scx=Mro} — 3 ranges, 43 code points.
|
| |
| constexpr code_range | scx_Multani_ranges [] |
| | \p{scx=Multani} — 6 ranges, 48 code points.
|
| |
| constexpr code_range | scx_Myanmar_ranges [] |
| | \p{scx=Myanmar} — 5 ranges, 244 code points.
|
| |
| constexpr code_range | scx_Nabataean_ranges [] |
| | \p{scx=Nabataean} — 2 ranges, 40 code points.
|
| |
| constexpr code_range | scx_Nag_Mundari_ranges [] |
| | \p{scx=Nag_Mundari} — 1 ranges, 42 code points.
|
| |
| constexpr code_range | scx_Nandinagari_ranges [] |
| | \p{scx=Nandinagari} — 9 ranges, 86 code points.
|
| |
| constexpr code_range | scx_New_Tai_Lue_ranges [] |
| | \p{scx=New_Tai_Lue} — 4 ranges, 83 code points.
|
| |
| constexpr code_range | scx_Newa_ranges [] |
| | \p{scx=Newa} — 2 ranges, 97 code points.
|
| |
| constexpr code_range | scx_Nko_ranges [] |
| | \p{scx=Nko} — 6 ranges, 67 code points.
|
| |
| constexpr code_range | scx_Nushu_ranges [] |
| | \p{scx=Nushu} — 2 ranges, 397 code points.
|
| |
| constexpr code_range | scx_Nyiakeng_Puachue_Hmong_ranges [] |
| | \p{scx=Nyiakeng_Puachue_Hmong} — 4 ranges, 71 code points.
|
| |
| constexpr code_range | scx_Ogham_ranges [] |
| | \p{scx=Ogham} — 1 ranges, 29 code points.
|
| |
| constexpr code_range | scx_Ol_Chiki_ranges [] |
| | \p{scx=Ol_Chiki} — 1 ranges, 48 code points.
|
| |
| constexpr code_range | scx_Ol_Onal_ranges [] |
| | \p{scx=Ol_Onal} — 3 ranges, 46 code points.
|
| |
| constexpr code_range | scx_Old_Hungarian_ranges [] |
| | \p{scx=Old_Hungarian} — 7 ranges, 112 code points.
|
| |
| constexpr code_range | scx_Old_Italic_ranges [] |
| | \p{scx=Old_Italic} — 2 ranges, 39 code points.
|
| |
| constexpr code_range | scx_Old_North_Arabian_ranges [] |
| | \p{scx=Old_North_Arabian} — 1 ranges, 32 code points.
|
| |
| constexpr code_range | scx_Old_Permic_ranges [] |
| | \p{scx=Old_Permic} — 6 ranges, 50 code points.
|
| |
| constexpr code_range | scx_Old_Persian_ranges [] |
| | \p{scx=Old_Persian} — 2 ranges, 50 code points.
|
| |
| constexpr code_range | scx_Old_Sogdian_ranges [] |
| | \p{scx=Old_Sogdian} — 1 ranges, 40 code points.
|
| |
| constexpr code_range | scx_Old_South_Arabian_ranges [] |
| | \p{scx=Old_South_Arabian} — 1 ranges, 32 code points.
|
| |
| constexpr code_range | scx_Old_Turkic_ranges [] |
| | \p{scx=Old_Turkic} — 3 ranges, 75 code points.
|
| |
| constexpr code_range | scx_Old_Uyghur_ranges [] |
| | \p{scx=Old_Uyghur} — 3 ranges, 28 code points.
|
| |
| constexpr code_range | scx_Oriya_ranges [] |
| | \p{scx=Oriya} — 18 ranges, 97 code points.
|
| |
| constexpr code_range | scx_Osage_ranges [] |
| | \p{scx=Osage} — 6 ranges, 76 code points.
|
| |
| constexpr code_range | scx_Osmanya_ranges [] |
| | \p{scx=Osmanya} — 2 ranges, 40 code points.
|
| |
| constexpr code_range | scx_Pahawh_Hmong_ranges [] |
| | \p{scx=Pahawh_Hmong} — 5 ranges, 127 code points.
|
| |
| constexpr code_range | scx_Palmyrene_ranges [] |
| | \p{scx=Palmyrene} — 1 ranges, 32 code points.
|
| |
| constexpr code_range | scx_Pau_Cin_Hau_ranges [] |
| | \p{scx=Pau_Cin_Hau} — 1 ranges, 57 code points.
|
| |
| constexpr code_range | scx_Phags_Pa_ranges [] |
| | \p{scx=Phags_Pa} — 5 ranges, 61 code points.
|
| |
| constexpr code_range | scx_Phoenician_ranges [] |
| | \p{scx=Phoenician} — 2 ranges, 29 code points.
|
| |
| constexpr code_range | scx_Psalter_Pahlavi_ranges [] |
| | \p{scx=Psalter_Pahlavi} — 4 ranges, 30 code points.
|
| |
| constexpr code_range | scx_Rejang_ranges [] |
| | \p{scx=Rejang} — 2 ranges, 37 code points.
|
| |
| constexpr code_range | scx_Runic_ranges [] |
| | \p{scx=Runic} — 1 ranges, 89 code points.
|
| |
| constexpr code_range | scx_Samaritan_ranges [] |
| | \p{scx=Samaritan} — 3 ranges, 62 code points.
|
| |
| constexpr code_range | scx_Saurashtra_ranges [] |
| | \p{scx=Saurashtra} — 2 ranges, 82 code points.
|
| |
| constexpr code_range | scx_Sharada_ranges [] |
| | \p{scx=Sharada} — 8 ranges, 109 code points.
|
| |
| constexpr code_range | scx_Shavian_ranges [] |
| | \p{scx=Shavian} — 2 ranges, 49 code points.
|
| |
| constexpr code_range | scx_Siddham_ranges [] |
| | \p{scx=Siddham} — 2 ranges, 92 code points.
|
| |
| constexpr code_range | scx_SignWriting_ranges [] |
| | \p{scx=SignWriting} — 3 ranges, 672 code points.
|
| |
| constexpr code_range | scx_Sinhala_ranges [] |
| | \p{scx=Sinhala} — 15 ranges, 114 code points.
|
| |
| constexpr code_range | scx_Sogdian_ranges [] |
| | \p{scx=Sogdian} — 2 ranges, 43 code points.
|
| |
| constexpr code_range | scx_Sora_Sompeng_ranges [] |
| | \p{scx=Sora_Sompeng} — 2 ranges, 35 code points.
|
| |
| constexpr code_range | scx_Soyombo_ranges [] |
| | \p{scx=Soyombo} — 1 ranges, 83 code points.
|
| |
| constexpr code_range | scx_Sundanese_ranges [] |
| | \p{scx=Sundanese} — 2 ranges, 72 code points.
|
| |
| constexpr code_range | scx_Sunuwar_ranges [] |
| | \p{scx=Sunuwar} — 8 ranges, 51 code points.
|
| |
| constexpr code_range | scx_Syloti_Nagri_ranges [] |
| | \p{scx=Syloti_Nagri} — 3 ranges, 57 code points.
|
| |
| constexpr code_range | scx_Syriac_ranges [] |
| | \p{scx=Syriac} — 19 ranges, 119 code points.
|
| |
| constexpr code_range | scx_Tagalog_ranges [] |
| | \p{scx=Tagalog} — 3 ranges, 25 code points.
|
| |
| constexpr code_range | scx_Tagbanwa_ranges [] |
| | \p{scx=Tagbanwa} — 4 ranges, 20 code points.
|
| |
| constexpr code_range | scx_Tai_Le_ranges [] |
| | \p{scx=Tai_Le} — 6 ranges, 50 code points.
|
| |
| constexpr code_range | scx_Tai_Tham_ranges [] |
| | \p{scx=Tai_Tham} — 5 ranges, 127 code points.
|
| |
| constexpr code_range | scx_Tai_Viet_ranges [] |
| | \p{scx=Tai_Viet} — 2 ranges, 72 code points.
|
| |
| constexpr code_range | scx_Takri_ranges [] |
| | \p{scx=Takri} — 4 ranges, 80 code points.
|
| |
| constexpr code_range | scx_Tamil_ranges [] |
| | \p{scx=Tamil} — 25 ranges, 133 code points.
|
| |
| constexpr code_range | scx_Tangsa_ranges [] |
| | \p{scx=Tangsa} — 2 ranges, 89 code points.
|
| |
| constexpr code_range | scx_Tangut_ranges [] |
| | \p{scx=Tangut} — 6 ranges, 6931 code points.
|
| |
| constexpr code_range | scx_Telugu_ranges [] |
| | \p{scx=Telugu} — 17 ranges, 106 code points.
|
| |
| constexpr code_range | scx_Thaana_ranges [] |
| | \p{scx=Thaana} — 7 ranges, 66 code points.
|
| |
| constexpr code_range | scx_Thai_ranges [] |
| | \p{scx=Thai} — 6 ranges, 90 code points.
|
| |
| constexpr code_range | scx_Tibetan_ranges [] |
| | \p{scx=Tibetan} — 8 ranges, 211 code points.
|
| |
| constexpr code_range | scx_Tifinagh_ranges [] |
| | \p{scx=Tifinagh} — 7 ranges, 63 code points.
|
| |
| constexpr code_range | scx_Tirhuta_ranges [] |
| | \p{scx=Tirhuta} — 6 ranges, 97 code points.
|
| |
| constexpr code_range | scx_Todhri_ranges [] |
| | \p{scx=Todhri} — 7 ranges, 58 code points.
|
| |
| constexpr code_range | scx_Toto_ranges [] |
| | \p{scx=Toto} — 2 ranges, 32 code points.
|
| |
| constexpr code_range | scx_Tulu_Tigalari_ranges [] |
| | \p{scx=Tulu_Tigalari} — 16 ranges, 99 code points.
|
| |
| constexpr code_range | scx_Ugaritic_ranges [] |
| | \p{scx=Ugaritic} — 2 ranges, 31 code points.
|
| |
| constexpr code_range | scx_Vai_ranges [] |
| | \p{scx=Vai} — 1 ranges, 300 code points.
|
| |
| constexpr code_range | scx_Vithkuqi_ranges [] |
| | \p{scx=Vithkuqi} — 8 ranges, 70 code points.
|
| |
| constexpr code_range | scx_Wancho_ranges [] |
| | \p{scx=Wancho} — 2 ranges, 59 code points.
|
| |
| constexpr code_range | scx_Warang_Citi_ranges [] |
| | \p{scx=Warang_Citi} — 2 ranges, 84 code points.
|
| |
| constexpr code_range | scx_Yezidi_ranges [] |
| | \p{scx=Yezidi} — 7 ranges, 60 code points.
|
| |
| constexpr code_range | scx_Yi_ranges [] |
| | \p{scx=Yi} — 7 ranges, 1246 code points.
|
| |
| constexpr code_range | scx_Zanabazar_Square_ranges [] |
| | \p{scx=Zanabazar_Square} — 1 ranges, 72 code points.
|
| |
|
constexpr std::span< const code_range > | scx_ranges [] |
| | Range table indexed by script (parallel to the enum order, Unknown empty – no code point's scx is ever explicitly Unknown, see the generator).
|
| |