Metascraper is a Node.js library designed to extract unified metadata from websites using various standards like Open Graph, Microdata, and JSON-LD. It utilizes a modular, plugin-based architecture to ensure high accuracy when scraping information such as authors, dates, and images from online articles.
Highlights
Supports multiple metadata protocols including Open Graph, Twitter Cards, and RDFa
Modular architecture allows users to include only the specific metadata plugins they need
Designed for high accuracy when extracting data from web articles
Compatible with headless browsers to handle rendered HTML content
GitHub - firecrawl/firecrawl: The context API to search, scrape, and interact with the web at scale. 🔥
The context API to search, scrape, and interact with the web at scale. 🔥 - firecrawl/firecrawl
github.com
GitHub - BrowserCash/teracrawl: High-performance web crawler API optimized for LLMs. Turn any search or website into clean Markdown using remote browsers.
High-performance web crawler API optimized for LLMs. Turn any search or website into clean Markdown using remote browsers. - BrowserCash/teracrawl
github.com
GitHub - ScrapeGraphAI/Scrapegraph-ai: Python scraper based on AI
Python scraper based on AI. Contribute to ScrapeGraphAI/Scrapegraph-ai development by creating an account on GitHub.