Metreeca MIME
    Preparing search index...

    Function html

    • Creates an HTML parser.

      The task converts a feed of HTML documents into a feed of parsed trees, one tree per document, so that downstream tasks work on document structure rather than on text.

      Each document is given either as text or as a response carrying it in its body. Empty or whitespace-only documents produce no tree, and neither do responses without a body.

      Trees have the same shape as the ones produced by xml, so the same path expressions work on both. Names are kept as written in the source, xmlns declarations are not resolved, and missing html, head or body elements are not added.

      Response bodies are decoded with the charset parameter of their content type. If no charset is stated, the meta charset declared in the first kilobyte of the document is used. If neither is stated, bodies are decoded as UTF-8 rather than as the windows-1252 legacy default of HTML. A leading byte order mark is stripped, both from text and from bodies in a Unicode charset.

      A response with a content type other than text/html or application/xhtml+xml is logged as a warning and parsed anyway. A response with a charset the platform can't decode is also logged, and its body is decoded as UTF-8. Since parsing never fails, these warnings are the only sign that a source is mis-declared.

      Trees carry the base URL needed to resolve the references they contain, recorded as an xml:base attribute on each root element. A root element that already declares xml:base keeps it, resolved against the recorded base URL.

      The base URL is the one stated by the first base element of the document, resolved against the retrieval URL. The retrieval URL is the base argument, if given, or else the final URL of the response, after any redirects. No base URL is recorded if none of them yields an absolute URL, as for a document given as text with only a relative base element. Unlike the base argument, base elements are read leniently, since they come from the source.

      Note

      • Incremental: each tree is emitted as soon as its document is parsed, so endless sources are processed for as long as the feed is consumed.
      • Materialising: each document is held in memory as a whole while it is parsed, so peak memory use is about twice the size of the largest document, regardless of the length of the feed.
      • Stateless: each document is parsed independently, so the result doesn't depend on how the feed is split across nested feeds or runs.
      Warning

      Parsing is forgiving and never fails. Emitted trees are always structurally sound, since unclosed elements are closed at the end of the input, but malformed input may be misrepresented rather than rejected. Trees also reflect the markup as written rather than as a browser would repair it: misnested elements stay misnested, and content a browser would relocate stays in place. Expressions should target the tree the parser actually produces.

      Note

      Names are folded to lowercase, as HTML prescribes, except inside inline SVG and MathML, where the camelCase names defined by those languages are restored: clipPath and @viewBox are selected as written. HTML content inside foreignObject and MathML text elements is folded like the rest of the document.

      Parameters

      • Optionalbase: string

        The retrieval URL for resolving references, overriding the URL of responses and in turn overridden by base elements in the document; must be a hierarchical identifier, that is a scheme followed by a root-relative path

      Returns Task<string | Response, Document>

      A task converting a feed of HTML documents, given as text or as responses, into a feed of parsed trees

      RangeError If base is not a hierarchical identifier, for instance a relative reference or an opaque identifier such as urn:example:x, which would leave references silently unresolved

      Error While the feed is consumed, if the source feed fails or a response body can't be read