Tokenizer

Parse a Schematika character stream into lexical tokens

Context

_images/ditaa-7cd4e5b7fa108bc2cd4d5779f08805a01bbb6ed8.png
#include <xo/tokenizer/tokenizer.hpp>

allowmixing

object tkz1<<tokenizer>>
tkz1 : input_state = ins1

object ins1<<input_state>>
ins1 : current_line = (9 * 8)

tkz1 o-- ins1

  • Assemble a stream of lexical tokens from a text stream.

  • Lexical errors reported via scan_result instance; errors reported with detailed context

Class

template<typename CharT>
class tokenizer

Parse a Schematika character stream into lexical tokens.

Use:

// see xo-tokenizer/example/tokenrepl/tokenrepl.cpp
// for exact working code

using tokenizer_type = tokenizer<char>;
using span_type = tokenizer_type::span_type;

tokenizer_type tkz;
span_type input = ...;

while (!input.empty()) {
    auto [tk, consumed, error] = tkz.scan(input);

    if (tk.is_valid()) {
        // do something with tk
    } else if (error.is_error()) {
        error.report(cout);
        break;
    }

    input = tkz.consume(consumed, input);
}

if endofinput {
    auto [tk, consumed, error] = tzk.notify_eof()

    // do something with (final) tk if tk.is_valid()
}

See tokentype.hpp for token types

Instance Variables

group tokenizer instance variables

Variables

input_state_type input_state_

track input state (line#,pos,..) for error messages. There’s an ordering problem here:

  1. input_state_.skip_leading_whitespace() advances current line automagically when it sees

  2. need to capture value of input_state_ before newline

  3. but neeed newline to end token Also recall input_state_type needed for reporting errors.

std::string prefix_

Accumulate partial token here. This will happen if input sent to tokenizer::scan ends without whitespace such that last available token’s extent is not determined

Constructors

group tokenizer constructors

Functions

tokenizer(bool debug_flag = false)

Methods

group tokenizer methods

Functions

static bool is_1char_punctuation(CharT ch)

identifies punctuation chars. These are chars that are not permitted to appear within a symbol token. Instead they force completion of a preceding token, and start a new token with themselves

static bool is_2char_punctuation(CharT ch)

more-relazed version of is_1char_punctuation. Chars that are not permitted to appear within a symbol token, but may form token combined with next character

static result_type assemble_token(std::size_t initial_whitespace, const span_type &token_text, input_state_type &input_state)

assemble token from text token_text. initial_whitespace Amount of whitespace input being consumed from input. token_text subset of input_line representing a single token. input_state input state containing input_line

retval.consumed will represent some possibly-empty prefix of input

static result_type assemble_final_token(const span_type &token_text, const input_state_type &input_state)

degenerate version of assemble_token() on reaching end-of-file

inline bool has_prefix() const

true if tokenizer contains stored prefix of possibly-incomplete token

result_type scan(const span_type &input, bool eof_flag)

scan for next input token, given input. Note:

  • tokenizer can consume input (e.g. whitespace) without completing a token

  • input will remember the extent of the last line of input for which parsing has begun, but not completed. It’s required that at least that portion of the input span remain valid across scan(), scan2() calls

Returns:

{parsed token, consumed span}

void discard_current_line()

discard current line after error. Just cleans up error-reporting state