Input State
Track detailed state of input stream to collect information useful for detailed error reporting
Context
#include <xo/tokenizer/input_state.hpp>
Class
-
template<typename CharT>
class input_state Track detailed input position for use in error messages.
input characters fall into two categories:
consumed: memory can be reclaimed/recycled
buffered: memory will be retained unaltered until consumed
remarks:
always in one of two states:
empty
contains exactly one line of input
also record current input position. Use this for example to identify where tokenizer rejected input.
.current_pos advances by one token
buffered characters always form a single contiguous range.
input_state does not own any storage; storage is owned elsewhere
@text
<——————.current_line——————> > <— .whitespace cccccccccccccccccccccccccccccccc__TTTTTTTTxxxxxxxxx ^ ^ ^ .current_line.lo | .current_line.hi .current_pos
<-—prev_line—-> <-—current_line—-> > <—whitespace ppppppppppppppppppp cccccccccccc__TTTTTTTT ^
@endtext
Instance Variables
- group input_state instance variables
Variables
-
span<const CharT> current_line_ = span<const CharT>()
remember current input line. Used only to report errors
-
size_t tk_start_ = 0
start of last token within current_line_
-
size_t current_pos_ = 0
input position within current_line_
-
size_t whitespace_ = 0
number of whitespace chars since end of preceding token, or last newline, whichever is less
-
bool debug_flag_ = false
true to log input activity
-
span<const CharT> current_line_ = span<const CharT>()
Constructors
- group input_state constructors
Static Methods
- group input_state static methods
Functions
-
static bool is_newline(CharT ch)
recognize the newline character ‘
’
-
static bool is_whitespace(CharT ch)
identifies whitespace chars. These are chars that do not belong to any token. They are not permitted to appear within a symbol or string token. Appearance of a whitespace char forces completioon of preceding token.
-
static bool is_newline(CharT ch)
Access Methods
- group Input-state-access-methods
General Methods
- group Input-state-general-methods
Functions
-
input_state rewind(std::size_t n) const
Input state less
nchars. Use to recover input state before a complete but error-triggering token
-
std::pair<input_error, span_type> capture_current_line(const span_type &input, bool eof_flag)
Capture prefix of
inputup to first newline. Set read position to start of line.Alters: .current_line .current_pos
Return pair comprising error code and input span representing first line (including trailing newline) from
input.
-
inline void add_whitespace(std::size_t z)
Add
zto the count of whitespace preceding the next token.capture_current_line() resets .whitespace, since it is counted from the start of a line. When whitespace separating two tokens straddles a line boundary the caller advances the line mid-skip, and uses this to carry the earlier count across — .whitespace counts whitespace since the end of the preceding token, and a newline between two tokens is whitespace like any other.
Alters: .whitespace
-
span_type consume_current_line()
atomically return current line while discarding it from input state
Alters .current_line .current_pos .whitespace
-
void discard_current_line()
Reset input state for start of next line. Expression parser may use this to discard remainder of input line after a parsing error.
Alters: .current_line .current_pos .whitespace
-
void advance(size_t z)
Advance input position by
zAlters: .current_pos
-
void advance_until(const CharT *pos)
Advance .current_pos to pos. Require: pos in current_line_
-
const CharT *skip_leading_whitespace()
Skip prefix of input, starting at current read position, comprising only whitespace.
Presume input position is at end of token; on return whitespace_ counts number of whitespace characters skipped.
Return pointer to first non-whitespace character after current_pos_ or current_line_.hi if reached end of buffered line.
Alters: .whitespace
-
input_state rewind(std::size_t n) const