SciLex
A header-only C++20 lexer built on REAL
Loading...
Searching...
No Matches
Classes | Namespaces | Functions
grammar.hpp File Reference

The .lex grammar format: rules written as text, parsed into scilex::rule lists. More...

#include <algorithm>
#include <cstddef>
#include <fstream>
#include <sstream>
#include <stdexcept>
#include <string>
#include <string_view>
#include <utility>
#include <vector>
#include <real/real.hpp>
#include "lexer.hpp"
Include dependency graph for grammar.hpp:

Classes

class  scilex::grammar_error
 Thrown for a malformed .lex grammar: where (origin, 1-based line, and 1-based byte column in the line when the cause has one) and why. More...
 
struct  scilex::grammar
 A parsed grammar: the rules in order, and each rule's name (index = the rule's kind). More...
 
struct  scilex::detail::grammar_fields
 The fields of one rule line, split on tabs, with the byte column each starts at. More...
 
struct  scilex::detail::transition_site
 Where a rule's push= or set= option is written. More...
 

Namespaces

namespace  scilex
 The SciLex public API (scilex::lexer, scilex::rule, scilex::token).
 
namespace  scilex::detail
 

Functions

std::string scilex::detail::quoting (std::string_view before, std::string_view word, std::string_view after)
 before, then word in quotes, then after: a cause naming the offending word.
 
grammar_fields scilex::detail::split_grammar_line (std::string_view line, std::size_t first)
 Splits line on tabs, from first (0-based), keeping each field's column.
 
std::string_view scilex::detail::checked_mode (std::string_view name, std::string_view word, const std::string &origin, std::size_t line, std::size_t column)
 name, refused when it is empty or holds a comma (which separates the modes of in=).
 
void scilex::detail::set_transition (rule &out, mode_action::op operation, std::string_view target, const std::string &origin, std::size_t line, std::size_t column)
 Gives out its transition, refusing a second one.
 
std::size_t scilex::detail::apply_grammar_options (std::string_view options, std::size_t column, const std::string &origin, std::size_t line, rule &out)
 Applies the space-separated options of one rule to out.
 
void scilex::detail::check_transition_targets (const grammar &parsed, const std::vector< transition_site > &sites, const std::string &origin)
 Refuses a push= or set= whose mode no rule is active in, at the option.
 
grammar scilex::parse_grammar (std::string_view text, const std::string &origin="<string>")
 Parses a .lex grammar from text (see the file documentation for the format).
 
grammar scilex::load_grammar (const std::string &path)
 Reads and parses the .lex grammar at path.
 

Detailed Description

The .lex grammar format: rules written as text, parsed into scilex::rule lists.

Optional, and not included by scilex.hpp: the lexer itself takes C++ rule lists, and a program that builds its rules in code needs none of this. It exists so that a grammar kept in a file — the CLI's input, a user-supplied grammar, a Python caller's text — goes through one tested parser.

One rule per line: a name, a tab, the pattern, then optionally a tab and space-separated options. Blank lines and lines whose first non-blank character is # are ignored; a UTF-8 byte order mark that starts the text is dropped, a trailing \r too, and so are the spaces and tabs that end a line, so a pattern never ends in one: write a final space [ ] or \x20. The options are:

A rule's kind is its 0-based position among the rules; scilex::grammar::names maps it back.

WS \s+ skip
STRING " push=str
TEXT [^"\\]+ in=str
ESCAPE \\. in=str
END " in=str pop
NUMBER [0-9]+