html5tokenizer - Fork of html5gum with code span support

Age	Commit message (Collapse)	Author
2023-09-28	break!: move offsets out of Token	Martin Fischer
	Previously the Token enum contained the offsets using the O generic type parameter, which could be a usize if you're tracking offsets or a zero-sized type if you didn't care about offsets. This commit moves all the byte offset and syntax information to a new Trace enum, which has several advantages: * Traces can now easily be stored separately, while the tokens are fed to the tree builder. (The tree builder only has to keep track of which tree nodes originate from which tokens.) * No needless generics for functions that take a token but don't care about offsets (a tree construction implementation is bound to have many of such functions). * The FromIterator<(String, String)> impl for AttributeMap no longer has to specify arbitrary values for the spans and the value_syntax). * The PartialEq implementation of Token is now much more useful (since it no longer includes all the offsets). * The Debug formatting of Token is now more readable (since it no longer includes all the offsets). * Function pointers to functions accepting tokens are possible. (Since function pointer types may not have generic parameters.)
2023-09-28	chore: add BasicEmitter stub	Martin Fischer

2023-09-28	break!: rename DefaultEmitter to TracingEmitter	Martin Fischer

2023-09-28	refactor: clean up DefaultEmitter code	Martin Fischer

2023-09-28	refactor: move utils module under tokenizer::machine	Martin Fischer

2023-09-28	refactor: move machine module under tokenizer	Martin Fischer

2023-09-11	chore: move DefaultEmitter to own module	Martin Fischer

2023-09-09	refactor: merge token types with attr to new token module	Martin Fischer

2023-09-09	chore: group public modules together	Martin Fischer

2023-09-03	docs: add spans example	Martin Fischer

2023-09-03	feat: make DefaultEmitter public again	Martin Fischer

2023-09-03	fix!: remove adjusted_current_node_present_and_not_in_html_namespace	Martin Fischer
	Conceptually the tokenizer emits tokens, which are then handled in the tree construction stage (which this crate doesn't yet implement). While the tokenizer can operate almost entirely based on its state (which may be changed via Tokenizer::set_state) and its internal state, there is the exception of the 'Markup declaration open state'[1], the third condition of which depends on the "adjusted current node", which in turn depends on the "stack of open elements" only known to the tree constructor. In 82898967320f90116bbc686ab7ffc2f61ff456c4 I tried to address this by adding the adjusted_current_node_present_and_not_in_html_namespace method to the Emitter trait. What I missed was that adding this method to the Emitter trait effectively crippled the composability of the API. You should be able to do the following: struct TreeConstructor<R, O> { tokenizer: Tokenizer<R, O, SomeEmitter<O>>, stack_of_open_elements: Vec<NodeId>, // ... } However this doesn't work if the implementation of SomeEmitter depends on the stack_of_open_elements field. This commits remedies this oversight by removing this method and instead making the Tokenizer yield values of a new Event enum: enum Event<T> { Token(T), CdataOpen } Event::CdataOpen signals that the new Tokenizer::handle_cdata_open method has to be called, which accepts a CdataAction: enum CdataAction { Cdata, BogusComment } the variants of which correspond exactly to the possible outcomes of the third condition of the 'Markup declaration open state'. Removing this method also has the added benefit that the DefaultEmitter is now again spec-compliant, which lets us expose it again in the next commit in good conscience (previously it just hard-coded the method implementation to return false, which is why I had removed the DefaultEmitter from the public API in the last release). [1]: https://html.spec.whatwg.org/multipage/parsing.html#markup-declaration-open-state
2023-09-03	fix: BufReadReader skips line on invalid UTF-8	Martin Fischer

2023-09-03	docs: add changelog	Martin Fischer

2023-08-19	feat: introduce NaiveParser	Martin Fischer

2023-08-19	break!: remove DefaultEmitter from public API	Martin Fischer

2023-08-19	chore: move internal re-export after public API	Martin Fischer

2023-08-19	fix(docs): fix broken relative link in rustdoc	Martin Fischer

2023-08-19	break!: introduce AttributeMap	Martin Fischer
	This has a number of benefits: * it hides the implementation of the map * it hides the type used for the map values (which lets us e.g. change name_span to name_offset while still being able to provide a convenient `Attribute::name_span` method.) * it lets us provide convenience impls for the map such as `FromIterator<(String, String)>`
2023-08-19	chore: move Attribute to attr module	Martin Fischer
	This is done separately so that the following commit has a cleaner diff.
2023-08-19	feat!: add offset to comments	Martin Fischer

2023-08-19	refactor!: remove Span trait, just use Range	Martin Fischer
	`std::mem::size_of::<Range<NoopOffset>>()` is 0 so there's no need to abstract over Range.
2023-08-19	chore: demote missing_docs lint to warn	Martin Fischer
	`#![deny(missing_docs)]` makes `cargo test` abort immediately if any public API member is missing a doc comment ... which is quite annoying when experimenting with API designs. Also sometimes refactor commits (such as the very next commit) introduce new types that are then immediately removed afterwards, this should be possible without having to add a `/// TODO``` (which contrary to a compiler warning is easy to miss).
2023-08-19	break!: stop re-exporting reader traits & types	Martin Fischer
	This is primarily done to make the rustdoc more readable (by grouping Reader, IntoReader, StringReader and BufReadReader in the reader module). Ideally IntoReader is already implemented for your input type and you don't have to concern yourself with these traits / types at all.
2023-08-19	break!: remove Never in favor of std::convert::Infallible	Martin Fischer
	This change is a backport of 04e6cbe[1] from html5gum. [1]: https://github.com/untitaker/html5gum/commit/04e6cbe44bb7a388bd61d1c9cfe4c618eb3b0e29
2023-08-19	break!: remove InfallibleTokenizer in favor of Iterator::flatten	Martin Fischer

2023-08-19	break!: rename Readable to IntoReader	Martin Fischer
	The trait of the standard library is also called IntoIterator and not Iterable.
2021-12-05	spans: support attribute names	Martin Fischer

2021-12-05	spans: add span tests	Martin Fischer

2021-12-05	spans: copy DefaultEmitter to new span module	Martin Fischer

2021-12-05	allow setting the Tokenizer to Data, PlainText, RcData, RawText and ↵	Martin Fischer
	ScriptData states
2021-12-05	prepare for introduction of public State enum	Martin Fischer

2021-11-27	split up match-arms and tokenizer to isolate some tokenizer-internal state	Markus Unterwaditzer
	purpose: don't want to expose self.to_reconsume to the consume() method
2021-11-26	Read html from io::BufRead (#8)	Markus Unterwaditzer

2021-11-26	clean up reader interface	Markus Unterwaditzer

2021-11-24	hello world	Markus Unterwaditzer