Contextractor 🧰
Extract clean, readable article text from any web page. Built on the Rust port of Trafilatura (extraction) and Crawlee (a TypeScript crawler that drives Playwright or fetches pages over plain HTTP with Cheerio) — strips navigation, ads, cookie banners, and boilerplate to leave the main content as plain text, Markdown, cleaned HTML, JSON, or the original raw page source, typically 80–90% fewer tokens than the raw HTML. Useful for feeding web content to LLMs or archiving articles. Free to try — no login required.
Evidence
Verified:
- www.contextractor.comcrawl-v1 · ce959cb91b88
Was this result useful?