Python scraper based on AI
-
Updated
Sep 25, 2026 - Python
Python scraper based on AI
🤖/👨🦰 Detect bots/crawlers/spiders using the user agent string
Amazon products scraper with using of rotating proxies and headless Chrome from ScrapingAnt
API definition, resources and reference implementation of URL Frontiers
一款强大的大模型微调数据集生成和管理工具。
Python course for 2nd year NLP students at NRU HSE, 2018-2019
GAMECHANGER Policy Analytics Site Crawlers
YiraBot: Simplifying Web Scraping for All. A user-friendly tool for developers and enthusiasts, offering command-line ease and Python integration. Ideal for research, SEO, and data collection.
Python course for 2nd year NLP students at NRU HSE, 2017-2018
Detect bots/crawlers/spiders via user-agent string
A Web Crawler developed in Python.
List of best web crawlers to extract data from the web. Find web crawling tools for different needs.
Web scraper for collecting product and review data from e-commerce sites using Scraping Bee, AWS, Selenium, and Pandas. Focuses on cost-effective solutions, user-friendly interfaces, and efficient data extraction and analysis.
Official documentation, crawler user agents, specifications, research, and tools for Answer Engine Optimization (AEO) and Generative Engine Optimization (GEO).
🤖 Generate optimized robots.txt files for AI search engine crawlers (GPTBot, PerplexityBot, ClaudeBot, and more)
Web crawlers based on travel information, crawler data from Ctrip, tuniu, qunar.
This is an advanced version of the previously released version of web-crawler
Template statique pour rendre un petit site plus lisible par les agents IA, crawlers et outils d’extraction.
Watchdog for AI crawlers at your own origin — access-log analyzer for Caddy/nginx. Zero dependencies.
Complete database of AI search engine crawler user-agents (GPTBot, PerplexityBot, ClaudeBot) with robots.txt configuration examples
To associate your repository with the web-crawlers topic, visit your repo's landing page and select "manage topics."