Optionalbase: string
The retrieval URL for resolving references, overriding the URL of responses and in turn overridden by
base elements in the document; must be a hierarchical identifier, that is a scheme followed by a
root-relative path
A task converting a feed of HTML documents, given as text or as responses, into a feed of parsed trees
RangeError If base is not a hierarchical identifier, for instance a relative reference
or an opaque identifier such as urn:example:x, which would leave references
silently unresolved
Error While the feed is consumed, if the source feed fails or a response body can't be read
Creates an HTML parser.
The task converts a feed of HTML documents into a feed of parsed trees, one tree per document, so that downstream tasks work on document structure rather than on text.
Each document is given either as text or as a response carrying it in its body. Empty or whitespace-only documents produce no tree, and neither do responses without a body.
Trees have the same shape as the ones produced by xml, so the same path expressions work on both. Names are kept as written in the source,
xmlnsdeclarations are not resolved, and missinghtml,headorbodyelements are not added.Response bodies are decoded with the
charsetparameter of their content type. If no charset is stated, themetacharset declared in the first kilobyte of the document is used. If neither is stated, bodies are decoded as UTF-8 rather than as the windows-1252 legacy default of HTML. A leading byte order mark is stripped, both from text and from bodies in a Unicode charset.A response with a content type other than
text/htmlorapplication/xhtml+xmlis logged as a warning and parsed anyway. A response with a charset the platform can't decode is also logged, and its body is decoded as UTF-8. Since parsing never fails, these warnings are the only sign that a source is mis-declared.Trees carry the base URL needed to resolve the references they contain, recorded as an
xml:baseattribute on each root element. A root element that already declaresxml:basekeeps it, resolved against the recorded base URL.The base URL is the one stated by the first
baseelement of the document, resolved against the retrieval URL. The retrieval URL is thebaseargument, if given, or else the final URL of the response, after any redirects. No base URL is recorded if none of them yields an absolute URL, as for a document given as text with only a relativebaseelement. Unlike thebaseargument,baseelements are read leniently, since they come from the source.Parsing is forgiving and never fails. Emitted trees are always structurally sound, since unclosed elements are closed at the end of the input, but malformed input may be misrepresented rather than rejected. Trees also reflect the markup as written rather than as a browser would repair it: misnested elements stay misnested, and content a browser would relocate stays in place. Expressions should target the tree the parser actually produces.
Names are folded to lowercase, as HTML prescribes, except inside inline SVG and MathML, where the camelCase names defined by those languages are restored:
clipPathand@viewBoxare selected as written. HTML content insideforeignObjectand MathML text elements is folded like the rest of the document.