|
SciLex
A header-only C++20 lexer built on REAL
|
The .lex grammar format: rules written as text, parsed into scilex::rule lists.
More...
#include <algorithm>#include <cstddef>#include <fstream>#include <sstream>#include <stdexcept>#include <string>#include <string_view>#include <utility>#include <vector>#include <real/real.hpp>#include "lexer.hpp"Classes | |
| class | scilex::grammar_error |
Thrown for a malformed .lex grammar: where (origin, 1-based line, and 1-based byte column in the line when the cause has one) and why. More... | |
| struct | scilex::grammar |
| A parsed grammar: the rules in order, and each rule's name (index = the rule's kind). More... | |
| struct | scilex::detail::grammar_fields |
| The fields of one rule line, split on tabs, with the byte column each starts at. More... | |
| struct | scilex::detail::transition_site |
Where a rule's push= or set= option is written. More... | |
Namespaces | |
| namespace | scilex |
| The SciLex public API (scilex::lexer, scilex::rule, scilex::token). | |
| namespace | scilex::detail |
Functions | |
| std::string | scilex::detail::quoting (std::string_view before, std::string_view word, std::string_view after) |
before, then word in quotes, then after: a cause naming the offending word. | |
| grammar_fields | scilex::detail::split_grammar_line (std::string_view line, std::size_t first) |
Splits line on tabs, from first (0-based), keeping each field's column. | |
| std::string_view | scilex::detail::checked_mode (std::string_view name, std::string_view word, const std::string &origin, std::size_t line, std::size_t column) |
name, refused when it is empty or holds a comma (which separates the modes of in=). | |
| void | scilex::detail::set_transition (rule &out, mode_action::op operation, std::string_view target, const std::string &origin, std::size_t line, std::size_t column) |
Gives out its transition, refusing a second one. | |
| std::size_t | scilex::detail::apply_grammar_options (std::string_view options, std::size_t column, const std::string &origin, std::size_t line, rule &out) |
Applies the space-separated options of one rule to out. | |
| void | scilex::detail::check_transition_targets (const grammar &parsed, const std::vector< transition_site > &sites, const std::string &origin) |
Refuses a push= or set= whose mode no rule is active in, at the option. | |
| grammar | scilex::parse_grammar (std::string_view text, const std::string &origin="<string>") |
Parses a .lex grammar from text (see the file documentation for the format). | |
| grammar | scilex::load_grammar (const std::string &path) |
Reads and parses the .lex grammar at path. | |
The .lex grammar format: rules written as text, parsed into scilex::rule lists.
Optional, and not included by scilex.hpp: the lexer itself takes C++ rule lists, and a program that builds its rules in code needs none of this. It exists so that a grammar kept in a file — the CLI's input, a user-supplied grammar, a Python caller's text — goes through one tested parser.
One rule per line: a name, a tab, the pattern, then optionally a tab and space-separated options. Blank lines and lines whose first non-blank character is # are ignored; a UTF-8 byte order mark that starts the text is dropped, a trailing \r too, and so are the spaces and tabs that end a line, so a pattern never ends in one: write a final space [ ] or \x20. The options are:
skip — matches are consumed but not emitted;in=m1,m2 — the modes the rule is active in (default: default only);push=m, set=m, pop — the mode transition fired when the rule wins (at most one); a mode that push= or set= enters must have a rule active in it.A rule's kind is its 0-based position among the rules; scilex::grammar::names maps it back.