Text fed in pieces, lexed as it arrives, with exactly the tokens lexer::tokenize gives the whole text: the same munches, modes, positions and errors.
A token is returned only once no text still to come can change it, which is not "every token but the
last": a rule a[^z]*z gives way to one token the moment a z arrives, however many tokens other rules made of the text before it. So before each munch the stream asks every rule of the active mode whether more text could change its match there (real::regex::can_extend), and stops at the first munch where one could. finish lexes the rest with the end of the text as the end.
The stream keeps the text from the first token it has not returned, and before it the few bytes a rule may still read there (a lookbehind's width, one code point for \b; real::regex::left_context); the rest of what it returned is dropped at the next feed. So memory follows the longest token, not the input – and a token's lexeme views that buffer: it is valid until the next feed or finish. Positions are in the whole text.
while (read(chunk)) {
}
Text fed in pieces, lexed as it arrives, with exactly the tokens lexer::tokenize gives the whole text...
Definition lexer.hpp:1437
std::vector< token > feed(std::string_view chunk)
Appends chunk and returns the tokens no text still to come can change.
Definition lexer.hpp:1457
std::vector< token > finish()
Ends the text and returns the tokens that remain.
Definition lexer.hpp:1477
One lexical token: a typed slice of the source.
Definition token.hpp:59