SciLex
A header-only C++20 lexer built on REAL
Loading...
Searching...
No Matches
Classes | Public Member Functions | List of all members
scilex::token_stream Class Reference

Text fed in pieces, lexed as it arrives, with exactly the tokens lexer::tokenize gives the whole text: the same munches, modes, positions and errors. More...

#include <lexer.hpp>

Public Member Functions

 token_stream (const lexer &owner)
 A stream over owner, at the start of the text.
 
std::vector< token > feed (std::string_view chunk)
 Appends chunk and returns the tokens no text still to come can change.
 
std::vector< token > finish ()
 Ends the text and returns the tokens that remain.
 
std::size_t buffered () const noexcept
 Bytes the stream holds: from the first token not yet returned to the end of what was fed.
 

Detailed Description

Text fed in pieces, lexed as it arrives, with exactly the tokens lexer::tokenize gives the whole text: the same munches, modes, positions and errors.

A token is returned only once no text still to come can change it, which is not "every token but the last": a rule a[^z]*z gives way to one token the moment a z arrives, however many tokens other rules made of the text before it. So before each munch the stream asks every rule of the active mode whether more text could change its match there (real::regex::can_extend), and stops at the first munch where one could. finish lexes the rest with the end of the text as the end.

The stream keeps the text from the first token it has not returned, and before it the few bytes a rule may still read there (a lookbehind's width, one code point for \b; real::regex::left_context); the rest of what it returned is dropped at the next feed. So memory follows the longest token, not the input – and a token's lexeme views that buffer: it is valid until the next feed or finish. Positions are in the whole text.

scilex::token_stream in {lex.stream()};
while (read(chunk)) {
for (const scilex::token& t : in.feed(chunk)) { use(t); }
}
for (const scilex::token& t : in.finish()) { use(t); }
Text fed in pieces, lexed as it arrives, with exactly the tokens lexer::tokenize gives the whole text...
Definition lexer.hpp:1437
std::vector< token > feed(std::string_view chunk)
Appends chunk and returns the tokens no text still to come can change.
Definition lexer.hpp:1457
std::vector< token > finish()
Ends the text and returns the tokens that remain.
Definition lexer.hpp:1477
One lexical token: a typed slice of the source.
Definition token.hpp:59

Constructor & Destructor Documentation

◆ token_stream()

scilex::token_stream::token_stream ( const lexer &  owner)
inlineexplicit

A stream over owner, at the start of the text.

Parameters
[in]ownerThe lexer (must outlive the stream).

Member Function Documentation

◆ buffered()

std::size_t scilex::token_stream::buffered ( ) const
inlinenoexcept

Bytes the stream holds: from the first token not yet returned to the end of what was fed.

Returns
The count.

◆ feed()

std::vector< token > scilex::token_stream::feed ( std::string_view  chunk)
inline

Appends chunk and returns the tokens no text still to come can change.

Parameters
[in]chunkThe next piece of the text.
Returns
The tokens, in order; their lexemes are valid until the next call.
Exceptions
lex_errorWhere lexer::tokenize would throw on the whole text, once the text decides it, carrying the tokens this call decided before it (lex_error::decided).
std::logic_errorAfter finish.

◆ finish()

std::vector< token > scilex::token_stream::finish ( )
inline

Ends the text and returns the tokens that remain.

Returns
The tokens, in order; their lexemes are valid until the stream is destroyed.
Exceptions
lex_errorWhere lexer::tokenize would throw on the whole text, carrying the tokens this call decided before it (lex_error::decided).
std::logic_errorWhen called twice.

The documentation for this class was generated from the following file: