本教程假设你的系统已经安装Scrapy。若尚未安装,请参阅安装指南。
我们将抓取quotes.toscrape.com,一个展示著名作者名言的网站。
本教程将逐步完成以下任务:
-
创建新的Scrapy项目
-
编写爬虫(spider),抓取网站并提取数据
-
使用命令行导出抓取的数据
-
修改爬虫,使其递归跟随链接
-
使用爬虫参数
Scrapy使用Python编写。对Python了解越多,就越能充分利用Scrapy。
如果已经熟悉其他语言,希望快速学习Python,Python官方教程是很好的资源。
如果刚接触编程,并希望从Python开始,以下书籍可能有帮助:
还可以查看面向非程序员的Python学习资源列表,以及learnpython社区推荐的资源。
创建项目
开始抓取前,需要建立新的Scrapy项目。进入希望存储代码的目录,然后运行:
scrapy startproject tutorial
这会创建tutorial目录,内容如下:
tutorial/
scrapy.cfg # deploy configuration file
tutorial/ # project's Python module, you'll import your code from here
__init__.py
items.py # project items definition file
middlewares.py # project middlewares file
pipelines.py # project pipelines file
settings.py # project settings file
spiders/ # a directory where you'll later put your spiders
__init__.py
抓取任何内容前,打开settings.py,取消USER_AGENT所在行的注释,以标明自己的身份,例如项目名称加URL或电子邮件地址。这样,对爬虫有异议的网站所有者就可以联系你要求调整,而不是直接封禁。
第一个爬虫
爬虫是你定义、由Scrapy用来从一个或一组网站采集信息的类。它们必须继承Spider,定义初始请求,也可以进一步定义如何跟随页面中的链接,以及如何解析下载的页面内容来提取数据。
下面是第一个爬虫的代码。将它保存为项目tutorial/spiders目录中的quotes_spider.py:
from pathlib import Path
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
async def start(self):
urls = [
"https://quotes.toscrape.com/page/1/",
"https://quotes.toscrape.com/page/2/",
]
for url in urls:
yield scrapy.Request(url=url, callback=self.parse)
def parse(self, response):
page = response.url.split("/")[-2]
filename = f"quotes-{page}.html"
Path(filename).write_bytes(response.body)
self.log(f"Saved file {filename}")
可以看到,爬虫继承了scrapy.Spider,并定义了以下属性和方法:
-
name:标识爬虫。在同一个项目中必须唯一,也就是说,不同爬虫不能使用同一个名称。 -
start():必须是异步生成器,产出用于开始抓取的请求,也可以产出数据项。后续请求会从这些初始请求逐步生成。 -
parse():每个请求的响应下载完成后,用于处理响应的方法。参数response是TextResponse实例,包含页面内容,并提供进一步处理内容的实用方法。
如何运行爬虫
要让爬虫运行,进入项目顶层目录并执行:
scrapy crawl quotes
此命令运行刚添加的quotes爬虫,向quotes.toscrape.com域发送请求。输出会类似下面这样:
... (omitted for brevity)
2016-12-16 21:24:05 [scrapy.core.engine] INFO: Spider opened
2016-12-16 21:24:05 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min)
2016-12-16 21:24:05 [scrapy.extensions.telnet] DEBUG: Telnet console listening on 127.0.0.1:6023
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (404) <GET https://quotes.toscrape.com/robots.txt> (referer: None)
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/1/> (referer: None)
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/2/> (referer: None)
2016-12-16 21:24:05 [quotes] DEBUG: Saved file quotes-1.html
2016-12-16 21:24:05 [quotes] DEBUG: Saved file quotes-2.html
2016-12-16 21:24:05 [scrapy.core.engine] INFO: Closing spider (finished)
...
检查当前目录,应该可以发现新增的quotes-1.html和quotes-2.html文件。它们包含对应URL的内容,这是parse方法要求保存的。
提示
如果你想知道为什么还没有解析HTML,请稍等,下面就会介绍。
底层刚刚发生了什么?
Scrapy发送爬虫start()方法产出的第一批scrapy.Request对象。每收到一个响应,它便把Response对象传给与该请求关联的回调方法,本例为parse。
start方法的简便写法
除了实现start(),从URL产出Request对象,也可以定义start_urls类属性,保存URL列表。start()的默认实现会利用该列表创建爬虫的初始请求。
from pathlib import Path
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
start_urls = [
"https://quotes.toscrape.com/page/1/",
"https://quotes.toscrape.com/page/2/",
]
def parse(self, response):
page = response.url.split("/")[-2]
filename = f"quotes-{page}.html"
Path(filename).write_bytes(response.body)
即使没有显式告诉Scrapy,这些URL的请求仍由parse()处理,因为parse()是默认回调:没有显式指定回调的请求会调用它。
提取数据
学习如何用Scrapy提取数据,最好的方式是在Scrapy shell中试用选择器。运行:
scrapy shell 'https://quotes.toscrape.com/page/1/'
提示
从命令行运行Scrapy shell时,记得始终用引号包住URL,否则含参数,也就是含&字符的URL会无法正常工作。
Windows上改用双引号:
scrapy shell "https://quotes.toscrape.com/page/1/"
你会看到类似如下内容:
[ ... Scrapy log here ... ]
2016-09-19 12:09:27 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/1/> (referer: None)
[s] Available Scrapy objects:
[s] scrapy scrapy module (contains scrapy.Request, scrapy.Selector, etc)
[s] crawler <scrapy.crawler.Crawler object at 0x7fa91d888c90>
[s] item {}
[s] request <GET https://quotes.toscrape.com/page/1/>
[s] response <200 https://quotes.toscrape.com/page/1/>
[s] settings <scrapy.settings.Settings object at 0x7fa91d888c10>
[s] spider <DefaultSpider 'default' at 0x7fa91c8af990>
[s] Useful shortcuts:
[s] shelp() Shell help (print this help)
[s] fetch(req_or_url) Fetch request (or URL) and update local objects
[s] view(response) View response in a browser
在shell中,可以结合response对象,尝试用CSS选择元素:
>>> response.css("title")
[<Selector query='descendant-or-self::title' data='<title>Quotes to Scrape</title>'>]
response.css('title')返回一个类似列表的SelectorList对象。它表示一组Selector对象,每个对象封装一个XML或HTML元素,可以继续查询,以细化选择或提取数据。
要提取上面标题的文本,可以这样做:
>>> response.css("title::text").getall()
['Quotes to Scrape']
这里有两点要注意。首先,CSS查询添加了::text,表示只选择<title>元素内部的直接文本。如果没有::text,得到的是包含标签的完整标题元素:
>>> response.css("title").getall()
['<title>Quotes to Scrape</title>']
其次,.getall()返回列表。选择器可能得到多个结果,因此提取全部结果。如果像本例一样,只需要第一个结果,可以这样:
>>> response.css("title::text").get()
'Quotes to Scrape'
另一种写法是:
>>> response.css("title::text")[0].get()
'Quotes to Scrape'
当没有结果时,对SelectorList使用索引会抛出IndexError:
>>> response.css("noelement")[0].get()
Traceback (most recent call last):
...
IndexError: list index out of range
更适合的方式是直接对SelectorList调用.get();没有结果时,它会返回None:
>>> response.css("noelement").get()
这里的经验是:大多数抓取代码都应当能够容忍页面中找不到某些内容的情况。这样,即使部分内容提取失败,至少仍能得到一部分数据。
除了getall()和get(),还可以使用re()通过正则表达式提取内容:
>>> response.css("title::text").re(r"Quotes.*")
['Quotes to Scrape']
>>> response.css("title::text").re(r"Q\w+")
['Quotes']
>>> response.css("title::text").re(r"(\w+) to (\w+)")
['Quotes', 'Scrape']
为了找到合适的CSS选择器,可以在shell中使用view(response),通过浏览器打开响应页面,再用浏览器开发者工具检查HTML、确定选择器,参见使用浏览器开发者工具抓取数据。
Selector Gadget也很实用,能够为通过可视化方式选中的元素快速找到CSS选择器,支持多种浏览器。
XPath简介
>>> response.xpath("//title")
[<Selector query='//title' data='<title>Quotes to Scrape</title>'>]
>>> response.xpath("//title/text()").get()
'Quotes to Scrape'
XPath表达式十分强大,是Scrapy选择器的基础。实际上,CSS选择器在底层会转换为XPath。仔细查看shell中选择器对象的文本表示,就能看到这一点。
XPath也许没有CSS选择器那么流行,但功能更强,因为它除了遍历结构,还可以检查内容。例如,可以选择包含“Next Page”文字的链接。这使XPath非常适合抓取任务。即使已经会构建CSS选择器,也建议学习XPath,它会让抓取更容易。
本文不深入讲解XPath,可以在Scrapy选择器文档中了解其用法。进一步学习推荐通过示例学习XPath的教程以及介绍如何用XPath思考的教程。
在爬虫中提取数据
回到爬虫。目前它尚未提取具体数据,只是将整个HTML页面保存到本地文件。现在把上述提取逻辑整合进去。
Scrapy爬虫通常生成许多包含页面数据的字典。为此,在回调中使用Python的yield关键字,如下:
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
start_urls = [
"https://quotes.toscrape.com/page/1/",
"https://quotes.toscrape.com/page/2/",
]
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
"tags": quote.css("div.tags a.tag::text").getall(),
}
运行这个爬虫前,输入以下内容退出Scrapy shell:
quit()
然后运行:
scrapy crawl quotes
现在,提取的数据应该和日志一起输出:
2016-09-19 18:57:19 [scrapy.core.scraper] DEBUG: Scraped from <200 https://quotes.toscrape.com/page/1/>
{'tags': ['life', 'love'], 'author': 'André Gide', 'text': '“It is better to be hated for what you are than to be loved for what you are not.”'}
2016-09-19 18:57:19 [scrapy.core.scraper] DEBUG: Scraped from <200 https://quotes.toscrape.com/page/1/>
{'tags': ['edison', 'failure', 'inspirational', 'paraphrased'], 'author': 'Thomas A. Edison', 'text': "“I have not failed. I've just found 10,000 ways that won't work.”"}
保存抓取的数据
最简单的保存方式是使用Feed导出,命令如下:
scrapy crawl quotes -O quotes.json
这会生成quotes.json,包含全部抓取的数据项,并按JSON格式序列化。
命令行选项-O会覆盖已有文件;-o则向已有文件追加内容。不过,向JSON文件追加内容会使它成为无效JSON。需要追加时,可考虑JSON Lines等其他序列化格式:
scrapy crawl quotes -o quotes.jsonl
JSON Lines具有类似流的结构,便于追加新记录,运行两次也不会有JSON的上述问题。此外,每条记录各占一行,因此可以处理大文件,无需将所有内容放进内存。JQ等工具可以在命令行中帮助完成处理。
对于本教程这样的小项目,这些方式已足够。如果想对数据项做更复杂的处理,可以编写Item Pipeline。项目创建时,已经在tutorial/pipelines.py为流水线建立占位文件。不过,如果只想保存数据项,就无需实现任何Item Pipeline。
跟随链接
假设你希望抓取https://quotes.toscrape.com的所有页面,而不只是前两页的名言。
已经知道如何从页面提取数据,接下来看看如何跟随其中的链接。
首先,提取希望继续访问的页面链接。检查页面可以发现,下一页链接的HTML如下:
<ul class="pager">
<li class="next">
<a href="/page/2/">Next <span aria-hidden="true">→</span></a>
</li>
</ul>
可以在shell中尝试提取:
>>> response.css('li.next a').get()
'<a href="/page/2/">Next <span aria-hidden="true">→</span></a>'
这样得到的是锚点元素,但我们需要href属性。Scrapy支持一个CSS扩展,用于选择属性内容:
>>> response.css("li.next a::attr(href)").get()
'/page/2/'
还可以使用attrib属性,更多说明见选择元素属性:
>>> response.css("li.next a").attrib["href"]
'/page/2/'
下面修改爬虫,递归跟随下一页链接,并提取其中的数据:
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
start_urls = [
"https://quotes.toscrape.com/page/1/",
]
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
"tags": quote.css("div.tags a.tag::text").getall(),
}
next_page = response.css("li.next a::attr(href)").get()
if next_page is not None:
next_page = response.urljoin(next_page)
yield scrapy.Request(next_page, callback=self.parse)
提取数据之后,parse()寻找下一页链接,使用urljoin()构建完整绝对URL,因为链接可能是相对地址;再产出下一页的新请求,并把自己注册为回调,以继续提取数据、抓取所有页面。
这就是Scrapy跟随链接的机制:在回调方法中产出Request后,Scrapy会调度请求,并注册在请求完成后执行的回调。
利用这种方式,可以构建复杂爬虫,根据自定义规则跟随链接,并按访问页面的不同提取不同数据。
本例形成一种循环:不断跟随下一页链接,直到找不到为止。这很适合抓取博客、论坛等分页网站。
创建Request的简便方式
创建Request对象的简便方式是使用response.follow:
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
start_urls = [
"https://quotes.toscrape.com/page/1/",
]
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("span small::text").get(),
"tags": quote.css("div.tags a.tag::text").getall(),
}
next_page = response.css("li.next a::attr(href)").get()
if next_page is not None:
yield response.follow(next_page, callback=self.parse)
与scrapy.Request不同,response.follow直接支持相对URL,无需调用urljoin。注意,它只返回Request实例,仍需要将这个Request产出。
也可以向response.follow传入选择器而不是字符串,选择器应当提取所需属性:
for href in response.css("ul.pager a::attr(href)"):
yield response.follow(href, callback=self.parse)
对于<a>元素,还有一种更简便的方式:response.follow自动使用其href属性,因此代码可以进一步缩短:
for a in response.css("ul.pager a"):
yield response.follow(a, callback=self.parse)
要从可迭代对象创建多个请求,可以使用response.follow_all:
anchors = response.css("ul.pager a")
yield from response.follow_all(anchors, callback=self.parse)
也可以进一步简写:
yield from response.follow_all(css="ul.pager a", callback=self.parse)
更多示例与模式
下面的另一个爬虫演示回调和跟随链接,这次抓取作者信息:
import scrapy
class AuthorSpider(scrapy.Spider):
name = "author"
start_urls = ["https://quotes.toscrape.com/"]
def parse(self, response):
author_page_links = response.css(".author + a")
yield from response.follow_all(author_page_links, self.parse_author)
pagination_links = response.css("li.next a")
yield from response.follow_all(pagination_links, self.parse)
def parse_author(self, response):
def extract_with_css(query):
return response.css(query).get(default="").strip()
yield {
"name": extract_with_css("h3.author-title::text"),
"birthdate": extract_with_css(".author-born-date::text"),
"bio": extract_with_css(".author-description::text"),
}
该爬虫从主页开始,跟随全部作者页面链接,对每个作者页面调用parse_author回调;同时像前面一样,用parse回调跟随分页链接。
这里把回调作为位置参数传给response.follow_all,以缩短代码。Request也支持这种写法。
parse_author定义了辅助函数,用于提取和清理CSS查询结果,然后产出包含作者数据的Python字典。
此例还说明:即使同一作者有多条名言,也无需担心重复访问同一作者页面。Scrapy默认会过滤已访问URL的重复请求,避免编程错误造成对服务器的过量访问。可以通过DUPEFILTER_CLASS配置这一行为。
到这里,你应该已经理解如何在Scrapy中使用跟随链接和回调机制。
另一个利用跟随链接机制的例子是CrawlSpider:这个通用爬虫实现了一套小型规则引擎,可以在其基础上编写自己的爬虫。
另一个常见模式是组合多个页面的数据来构建数据项,这可以通过向回调传递额外数据的方法实现。
使用爬虫参数
运行爬虫时,可以用-a传入命令行参数:
scrapy crawl quotes -O quotes-humor.json -a tag=humor
这些参数会传给爬虫的__init__方法,并默认成为爬虫属性。
本例中,tag参数的值可通过self.tag访问。可以据此构建URL,让爬虫只抓取具有某个标签的名言:
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
async def start(self):
url = "https://quotes.toscrape.com/"
tag = getattr(self, "tag", None)
if tag is not None:
url = url + "tag/" + tag
yield scrapy.Request(url, self.parse)
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
}
next_page = response.css("li.next a::attr(href)").get()
if next_page is not None:
yield response.follow(next_page, self.parse)
如果传入tag=humor,爬虫就只会访问humor标签下的URL,例如https://quotes.toscrape.com/tag/humor。
更多处理方法见爬虫参数文档。
下一步
本教程只介绍了Scrapy基础,还有许多功能未提及。Scrapy概览中的“还有什么?”一节,可以快速了解其他重要功能。
可以继续阅读基本概念,了解命令行工具、爬虫、选择器,以及本教程未覆盖的数据建模等内容。如果更喜欢尝试示例项目,请查看示例。
如果使用编程代理,请参阅结合编程代理使用Scrapy。
原文:Scrapy完整入门教程;作者:Scrapy文档贡献者;日期:滚动latest版本。原文及源码权利归原作者和相应权利人所有。











暂无评论内容