Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
Pythonadbar/trafilatura

trafilatura

Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML

90.7/100
6.8KForks: 427
View on GitHubHomepage →
Loading report...

Similar Projects

Scrapling

95

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

Python80.3K

Scrapegraph-ai

89

Python scraper based on AI

Python30.8K

Douyin_TikTok_Download_API

94

🚀 Self-hosted TikTok & Douyin scraper and no-watermark video downloader — async REST API, MCP server, CLI and web console for posts, profiles, comments and playlists. Self-healing identity pool, PostgreSQL archive, one docker compose up. 抖音、TikTok 数据采集与无水印视频下载 API,自托管,支持 MCP 调用与 Docker 一键部署。

Python20.1K

transformers

98

🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

Python165.1K
Back to List