一个命令行网页抓取工具
项目描述
一个命令行网页抓取工具
scrape 是一个基于规则的网络爬虫和信息提取工具,能够操作和合并新的和现有的文档。XML 路径语言 (XPath) 和正则表达式用于定义过滤内容和 Web 遍历的规则。输出可以转换为文本、csv、pdf 和/或 HTML 格式。
安装
pip install scrape
或者
pip install git+https://github.com/huntrar/scrape.git#egg=scrape
或者
git clone https://github.com/huntrar/scrape cd scrape python setup.py install
您必须安装 wkhtmltopdf 才能将文件保存为 pdf。
用法
usage: scrape.py [-h] [-a [ATTRIBUTES [ATTRIBUTES ...]]] [-all]
[-c [CRAWL [CRAWL ...]]] [-C] [--csv] [-cs [CACHE_SIZE]]
[-f [FILTER [FILTER ...]]] [--html] [-i] [-m]
[-max MAX_CRAWLS] [-n] [-ni] [-no] [-o [OUT [OUT ...]]] [-ow]
[-p] [-pt] [-q] [-s] [-t] [-v] [-x [XPATH]]
[QUERY [QUERY ...]]
a command-line web scraping tool
positional arguments:
QUERY URLs/files to scrape
optional arguments:
-h, --help show this help message and exit
-a [ATTRIBUTES [ATTRIBUTES ...]], --attributes [ATTRIBUTES [ATTRIBUTES ...]]
extract text using tag attributes
-all, --crawl-all crawl all pages
-c [CRAWL [CRAWL ...]], --crawl [CRAWL [CRAWL ...]]
regexp rules for following new pages
-C, --clear-cache clear requests cache
--csv write files as csv
-cs [CACHE_SIZE], --cache-size [CACHE_SIZE]
size of page cache (default: 1000)
-f [FILTER [FILTER ...]], --filter [FILTER [FILTER ...]]
regexp rules for filtering text
--html write files as HTML
-i, --images save page images
-m, --multiple save to multiple files
-max MAX_CRAWLS, --max-crawls MAX_CRAWLS
max number of pages to crawl
-n, --nonstrict allow crawler to visit any domain
-ni, --no-images do not save page images
-no, --no-overwrite do not overwrite files if they exist
-o [OUT [OUT ...]], --out [OUT [OUT ...]]
specify outfile names
-ow, --overwrite overwrite a file if it exists
-p, --pdf write files as pdf
-pt, --print print text output
-q, --quiet suppress program output
-s, --single save to a single file
-t, --text write files as text
-v, --version display current version
-x [XPATH], --xpath [XPATH]
filter HTML using XPath
笔记
抓取的输入可以是链接、文件或两者的组合,允许您创建由现有内容和新抓取的内容构成的新文件。
默认情况下,多个输入文件/URL 保存到多个输出文件/目录。要合并它们,请使用 –single 标志。
保存为 pdf 或 HTML 时会自动包含图像;这涉及发出额外的 HTTP 请求,从而增加大量的处理时间。如果您希望放弃此功能,请使用 –no-images 标志,或设置环境变量 SCRAPE_DISABLE_IMGS。
请求缓存默认启用以缓存网页,可以通过设置环境变量 SCRAPE_DISABLE_CACHE 将其禁用。
页面在处理过程中临时保存为 PART.html 文件。除非将页面保存为 HTML,否则这些文件会在转换或退出时自动删除。
要无限制地抓取页面,请使用 –crawl-all 标志,或通过将一个或多个正则表达式传递给 –crawl 来过滤要按 URL 关键字抓取的页面。
如果您希望爬虫跟踪给定 URL 域之外的链接,请使用 –nonstrict。
可以通过 Ctrl-C 或通过使用 –maxpages 和 –maxlinks 设置要爬行的页面或链接的数量来停止爬行。一个页面可能包含零个或多个指向更多页面的链接。
抓取文件的文本输出可以打印到stdout,而不是通过输入-print来保存。
过滤 HTML 可以使用 –xpath 来完成,而过滤文本则是通过在 –filter 中输入一个或多个正则表达式来完成。
如果您只想指定要提取的特定标记属性而不是整个 XPath,请使用 –attributes。默认选择是仅提取文本属性,但您可以指定一个或多个不同的属性(例如 href、src、title 或任何可用的属性..)。