The task converts a feed of parsed X/HTML trees into a feed of documents holding their main content, one document per
tree, so that downstream tasks work on the content of a page without its navigation, headers, footers, sidebars and
controls.
Trees without content produce no document, so the output feed is not aligned one to one with the input feed.
The main content of a tree is the first of the following regions the tree contains:
the first main element or, if there is none, the first element with role="main"
all article elements, taken together, so that listing pages are kept whole; articles inside page framing
(navigation, headers, footers, sidebars and the like), articles nested inside other articles, and articles
without any text are ignored
the element with the densest text, as described below
Regions are searched below the root of the tree, so a tree rooted at a main element is searched for a region
inside it. Element names are matched case-insensitively.
Text density favours long runs of text over the same amount of text split into short fragments, and penalises
containers holding a large share of framing. Scripts, styles, navigation, headers, footers, sidebars, controls and
embedded objects contribute no text, so pages with long menus are scored on their prose alone. Elements without text
of their own, such as line breaks, rules, images and metadata, are neutral. If two elements are equally dense, the
first one in document order is selected, which is the outermost of a chain of single children.
If the tree is a page, that is if it contains an html or body element or a title, the content is wrapped in the
body of a new html element. If the page has a title, the html element also includes a head with a copy of
it. The title is the first title element outside page framing, so that captions of embedded objects are not
mistaken for it. If the tree is a fragment, each content region becomes a document root of its own.
Each content region records as xml:base the base URL in scope at its original position, if any, so that its
relative references keep resolving correctly. In page output, the html element also records as xml:base the base
URL of the source tree, so that the page URL travels with its content. Source trees are not modified.
Note
Incremental: each document is emitted as soon as its tree is drawn, so endless sources are processed for as
long as the feed is consumed.
Streaming: trees are processed one at a time and released once their content is copied, so memory use
doesn't grow with the length of the feed.
Stateless: each tree is processed independently, so the result doesn't depend on how the feed is split
across nested feeds or runs.
Returns Task<AnyNode,Document>
A task converting a feed of parsed X/HTML trees into a feed of documents holding their main content
Throws
Error While the feed is consumed, if the source feed fails
Creates a content extractor.
The task converts a feed of parsed X/HTML trees into a feed of documents holding their main content, one document per tree, so that downstream tasks work on the content of a page without its navigation, headers, footers, sidebars and controls.
Trees without content produce no document, so the output feed is not aligned one to one with the input feed.
The main content of a tree is the first of the following regions the tree contains:
mainelement or, if there is none, the first element withrole="main"articleelements, taken together, so that listing pages are kept whole; articles inside page framing (navigation, headers, footers, sidebars and the like), articles nested inside other articles, and articles without any text are ignoredRegions are searched below the root of the tree, so a tree rooted at a
mainelement is searched for a region inside it. Element names are matched case-insensitively.Text density favours long runs of text over the same amount of text split into short fragments, and penalises containers holding a large share of framing. Scripts, styles, navigation, headers, footers, sidebars, controls and embedded objects contribute no text, so pages with long menus are scored on their prose alone. Elements without text of their own, such as line breaks, rules, images and metadata, are neutral. If two elements are equally dense, the first one in document order is selected, which is the outermost of a chain of single children.
If the tree is a page, that is if it contains an
htmlorbodyelement or a title, the content is wrapped in thebodyof a newhtmlelement. If the page has a title, thehtmlelement also includes aheadwith a copy of it. The title is the firsttitleelement outside page framing, so that captions of embedded objects are not mistaken for it. If the tree is a fragment, each content region becomes a document root of its own.Each content region records as
xml:basethe base URL in scope at its original position, if any, so that its relative references keep resolving correctly. In page output, thehtmlelement also records asxml:basethe base URL of the source tree, so that the page URL travels with its content. Source trees are not modified.