Scrapy入门教程

本教程假设你的系统已经安装Scrapy。若尚未安装,请参阅安装指南。

我们将抓取quotes.toscrape.com,一个展示著名作者名言的网站。

本教程将逐步完成以下任务:

  1. 创建新的Scrapy项目

  2. 编写爬虫(spider),抓取网站并提取数据

  3. 使用命令行导出抓取的数据

  4. 修改爬虫,使其递归跟随链接

  5. 使用爬虫参数

Scrapy使用Python编写。对Python了解越多,就越能充分利用Scrapy。

如果已经熟悉其他语言,希望快速学习Python,Python官方教程是很好的资源。

如果刚接触编程,并希望从Python开始,以下书籍可能有帮助:

还可以查看面向非程序员的Python学习资源列表,以及learnpython社区推荐的资源。

创建项目

开始抓取前,需要建立新的Scrapy项目。进入希望存储代码的目录,然后运行:

scrapy startproject tutorial

这会创建tutorial目录,内容如下:

tutorial/
    scrapy.cfg            # deploy configuration file

    tutorial/             # project's Python module, you'll import your code from here
        __init__.py

        items.py          # project items definition file

        middlewares.py    # project middlewares file

        pipelines.py      # project pipelines file

        settings.py       # project settings file

        spiders/          # a directory where you'll later put your spiders
            __init__.py

抓取任何内容前,打开settings.py,取消USER_AGENT所在行的注释,以标明自己的身份,例如项目名称加URL或电子邮件地址。这样,对爬虫有异议的网站所有者就可以联系你要求调整,而不是直接封禁。

第一个爬虫

爬虫是你定义、由Scrapy用来从一个或一组网站采集信息的类。它们必须继承Spider,定义初始请求,也可以进一步定义如何跟随页面中的链接,以及如何解析下载的页面内容来提取数据。

下面是第一个爬虫的代码。将它保存为项目tutorial/spiders目录中的quotes_spider.py:

from pathlib import Path

import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"

    async def start(self):
        urls = [
            "https://quotes.toscrape.com/page/1/",
            "https://quotes.toscrape.com/page/2/",
        ]
        for url in urls:
            yield scrapy.Request(url=url, callback=self.parse)

    def parse(self, response):
        page = response.url.split("/")[-2]
        filename = f"quotes-{page}.html"
        Path(filename).write_bytes(response.body)
        self.log(f"Saved file {filename}")

可以看到,爬虫继承了scrapy.Spider,并定义了以下属性和方法:

  • name:标识爬虫。在同一个项目中必须唯一,也就是说,不同爬虫不能使用同一个名称。

  • start():必须是异步生成器,产出用于开始抓取的请求,也可以产出数据项。后续请求会从这些初始请求逐步生成。

  • parse():每个请求的响应下载完成后,用于处理响应的方法。参数response是TextResponse实例,包含页面内容,并提供进一步处理内容的实用方法。

    parse()通常解析响应,把抓取数据提取为字典,同时寻找要继续访问的新URL,并据此创建新的Request。

如何运行爬虫

要让爬虫运行,进入项目顶层目录并执行:

scrapy crawl quotes

此命令运行刚添加的quotes爬虫,向quotes.toscrape.com域发送请求。输出会类似下面这样:

... (omitted for brevity)
2016-12-16 21:24:05 [scrapy.core.engine] INFO: Spider opened
2016-12-16 21:24:05 [scrapy.extensions.logstats] INFO: Crawled 0 pages (at 0 pages/min), scraped 0 items (at 0 items/min)
2016-12-16 21:24:05 [scrapy.extensions.telnet] DEBUG: Telnet console listening on 127.0.0.1:6023
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (404) <GET https://quotes.toscrape.com/robots.txt> (referer: None)
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/1/> (referer: None)
2016-12-16 21:24:05 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/2/> (referer: None)
2016-12-16 21:24:05 [quotes] DEBUG: Saved file quotes-1.html
2016-12-16 21:24:05 [quotes] DEBUG: Saved file quotes-2.html
2016-12-16 21:24:05 [scrapy.core.engine] INFO: Closing spider (finished)
...

检查当前目录,应该可以发现新增的quotes-1.html和quotes-2.html文件。它们包含对应URL的内容,这是parse方法要求保存的。

提示

如果你想知道为什么还没有解析HTML,请稍等,下面就会介绍。

底层刚刚发生了什么?

Scrapy发送爬虫start()方法产出的第一批scrapy.Request对象。每收到一个响应,它便把Response对象传给与该请求关联的回调方法,本例为parse。

start方法的简便写法

除了实现start(),从URL产出Request对象,也可以定义start_urls类属性,保存URL列表。start()的默认实现会利用该列表创建爬虫的初始请求。

from pathlib import Path

import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"
    start_urls = [
        "https://quotes.toscrape.com/page/1/",
        "https://quotes.toscrape.com/page/2/",
    ]

    def parse(self, response):
        page = response.url.split("/")[-2]
        filename = f"quotes-{page}.html"
        Path(filename).write_bytes(response.body)

即使没有显式告诉Scrapy,这些URL的请求仍由parse()处理,因为parse()是默认回调:没有显式指定回调的请求会调用它。

提取数据

学习如何用Scrapy提取数据,最好的方式是在Scrapy shell中试用选择器。运行:

scrapy shell 'https://quotes.toscrape.com/page/1/'

提示

从命令行运行Scrapy shell时,记得始终用引号包住URL,否则含参数,也就是含&字符的URL会无法正常工作。

Windows上改用双引号:

scrapy shell "https://quotes.toscrape.com/page/1/"

你会看到类似如下内容:

[ ... Scrapy log here ... ]
2016-09-19 12:09:27 [scrapy.core.engine] DEBUG: Crawled (200) <GET https://quotes.toscrape.com/page/1/> (referer: None)
[s] Available Scrapy objects:
[s]   scrapy     scrapy module (contains scrapy.Request, scrapy.Selector, etc)
[s]   crawler    <scrapy.crawler.Crawler object at 0x7fa91d888c90>
[s]   item       {}
[s]   request    <GET https://quotes.toscrape.com/page/1/>
[s]   response   <200 https://quotes.toscrape.com/page/1/>
[s]   settings   <scrapy.settings.Settings object at 0x7fa91d888c10>
[s]   spider     <DefaultSpider 'default' at 0x7fa91c8af990>
[s] Useful shortcuts:
[s]   shelp()           Shell help (print this help)
[s]   fetch(req_or_url) Fetch request (or URL) and update local objects
[s]   view(response)    View response in a browser

在shell中,可以结合response对象,尝试用CSS选择元素:

>>> response.css("title")
[<Selector query='descendant-or-self::title' data='<title>Quotes to Scrape</title>'>]

response.css('title')返回一个类似列表的SelectorList对象。它表示一组Selector对象,每个对象封装一个XML或HTML元素,可以继续查询,以细化选择或提取数据。

要提取上面标题的文本,可以这样做:

>>> response.css("title::text").getall()
['Quotes to Scrape']

这里有两点要注意。首先,CSS查询添加了::text,表示只选择<title>元素内部的直接文本。如果没有::text,得到的是包含标签的完整标题元素:

>>> response.css("title").getall()
['<title>Quotes to Scrape</title>']

其次,.getall()返回列表。选择器可能得到多个结果,因此提取全部结果。如果像本例一样,只需要第一个结果,可以这样:

>>> response.css("title::text").get()
'Quotes to Scrape'

另一种写法是:

>>> response.css("title::text")[0].get()
'Quotes to Scrape'

当没有结果时,对SelectorList使用索引会抛出IndexError:

>>> response.css("noelement")[0].get()
Traceback (most recent call last):
...
IndexError: list index out of range

更适合的方式是直接对SelectorList调用.get();没有结果时,它会返回None:

>>> response.css("noelement").get()

这里的经验是:大多数抓取代码都应当能够容忍页面中找不到某些内容的情况。这样,即使部分内容提取失败,至少仍能得到一部分数据。

除了getall()和get(),还可以使用re()通过正则表达式提取内容:

>>> response.css("title::text").re(r"Quotes.*")
['Quotes to Scrape']
>>> response.css("title::text").re(r"Q\w+")
['Quotes']
>>> response.css("title::text").re(r"(\w+) to (\w+)")
['Quotes', 'Scrape']

为了找到合适的CSS选择器,可以在shell中使用view(response),通过浏览器打开响应页面,再用浏览器开发者工具检查HTML、确定选择器,参见使用浏览器开发者工具抓取数据。

Selector Gadget也很实用,能够为通过可视化方式选中的元素快速找到CSS选择器,支持多种浏览器。

XPath简介

除了CSS,Scrapy选择器也支持XPath表达式:

>>> response.xpath("//title")
[<Selector query='//title' data='<title>Quotes to Scrape</title>'>]
>>> response.xpath("//title/text()").get()
'Quotes to Scrape'

XPath表达式十分强大,是Scrapy选择器的基础。实际上,CSS选择器在底层会转换为XPath。仔细查看shell中选择器对象的文本表示,就能看到这一点。

XPath也许没有CSS选择器那么流行,但功能更强,因为它除了遍历结构,还可以检查内容。例如,可以选择包含“Next Page”文字的链接。这使XPath非常适合抓取任务。即使已经会构建CSS选择器,也建议学习XPath,它会让抓取更容易。

本文不深入讲解XPath,可以在Scrapy选择器文档中了解其用法。进一步学习推荐通过示例学习XPath的教程以及介绍如何用XPath思考的教程。

提取名言及作者

掌握选择与提取的基本知识后,接下来完善爬虫,编写从网页提取名言的代码。

https://quotes.toscrape.com中的每条名言都由类似下面的HTML元素表示:

<div class="quote">
    <span class="text">“The world as we have created it is a process of our
    thinking. It cannot be changed without changing our thinking.”</span>
    <span>
        by <small class="author">Albert Einstein</small>
        <a href="/author/Albert-Einstein">(about)</a>
    </span>
    <div class="tags">
        Tags:
        <a class="tag" href="/tag/change/page/1/">change</a>
        <a class="tag" href="/tag/deep-thoughts/page/1/">deep-thoughts</a>
        <a class="tag" href="/tag/thinking/page/1/">thinking</a>
        <a class="tag" href="/tag/world/page/1/">world</a>
    </div>
</div>

打开Scrapy shell,尝试确定如何提取需要的数据:

scrapy shell 'https://quotes.toscrape.com'

以下查询会得到名言HTML元素的选择器列表:

>>> response.css("div.quote")
[<Selector query="descendant-or-self::div[@class and contains(concat(' ', normalize-space(@class), ' '), ' quote ')]" data='<div class="quote" itemscope itemtype...'>,
<Selector query="descendant-or-self::div[@class and contains(concat(' ', normalize-space(@class), ' '), ' quote ')]" data='<div class="quote" itemscope itemtype...'>,
...]

每个返回的选择器都可以继续查询其子元素。先把第一个选择器赋给变量,以便直接对一条特定名言运行CSS选择器:

>>> quote = response.css("div.quote")[0]

使用刚创建的quote对象,提取该名言的text、author和tags:

>>> text = quote.css("span.text::text").get()
>>> text
'“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”'
>>> author = quote.css("small.author::text").get()
>>> author
'Albert Einstein'

标签是字符串列表,因此可以用.getall()提取全部标签:

>>> tags = quote.css("div.tags a.tag::text").getall()
>>> tags
['change', 'deep-thoughts', 'thinking', 'world']

确定各部分的提取方式后,就可以遍历所有名言元素,将它们整理成Python字典:

>>> for quote in response.css("div.quote"):
...     text = quote.css("span.text::text").get()
...     author = quote.css("small.author::text").get()
...     tags = quote.css("div.tags a.tag::text").getall()
...     print(dict(text=text, author=author, tags=tags))
...
{'text': '“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”', 'author': 'Albert Einstein', 'tags': ['change', 'deep-thoughts', 'thinking', 'world']}
{'text': '“It is our choices, Harry, that show what we truly are, far more than our abilities.”', 'author': 'J.K. Rowling', 'tags': ['abilities', 'choices']}
...

在爬虫中提取数据

回到爬虫。目前它尚未提取具体数据,只是将整个HTML页面保存到本地文件。现在把上述提取逻辑整合进去。

Scrapy爬虫通常生成许多包含页面数据的字典。为此,在回调中使用Python的yield关键字,如下:

import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"
    start_urls = [
        "https://quotes.toscrape.com/page/1/",
        "https://quotes.toscrape.com/page/2/",
    ]

    def parse(self, response):
        for quote in response.css("div.quote"):
            yield {
                "text": quote.css("span.text::text").get(),
                "author": quote.css("small.author::text").get(),
                "tags": quote.css("div.tags a.tag::text").getall(),
            }

运行这个爬虫前,输入以下内容退出Scrapy shell:

quit()

然后运行:

scrapy crawl quotes

现在,提取的数据应该和日志一起输出:

2016-09-19 18:57:19 [scrapy.core.scraper] DEBUG: Scraped from <200 https://quotes.toscrape.com/page/1/>
{'tags': ['life', 'love'], 'author': 'André Gide', 'text': '“It is better to be hated for what you are than to be loved for what you are not.”'}
2016-09-19 18:57:19 [scrapy.core.scraper] DEBUG: Scraped from <200 https://quotes.toscrape.com/page/1/>
{'tags': ['edison', 'failure', 'inspirational', 'paraphrased'], 'author': 'Thomas A. Edison', 'text': "“I have not failed. I've just found 10,000 ways that won't work.”"}

保存抓取的数据

最简单的保存方式是使用Feed导出,命令如下:

scrapy crawl quotes -O quotes.json

这会生成quotes.json,包含全部抓取的数据项,并按JSON格式序列化。

命令行选项-O会覆盖已有文件;-o则向已有文件追加内容。不过,向JSON文件追加内容会使它成为无效JSON。需要追加时,可考虑JSON Lines等其他序列化格式:

scrapy crawl quotes -o quotes.jsonl

JSON Lines具有类似流的结构,便于追加新记录,运行两次也不会有JSON的上述问题。此外,每条记录各占一行,因此可以处理大文件,无需将所有内容放进内存。JQ等工具可以在命令行中帮助完成处理。

对于本教程这样的小项目,这些方式已足够。如果想对数据项做更复杂的处理,可以编写Item Pipeline。项目创建时,已经在tutorial/pipelines.py为流水线建立占位文件。不过,如果只想保存数据项,就无需实现任何Item Pipeline。

使用爬虫参数

运行爬虫时,可以用-a传入命令行参数:

scrapy crawl quotes -O quotes-humor.json -a tag=humor

这些参数会传给爬虫的__init__方法,并默认成为爬虫属性。

本例中,tag参数的值可通过self.tag访问。可以据此构建URL,让爬虫只抓取具有某个标签的名言:

import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"

    async def start(self):
        url = "https://quotes.toscrape.com/"
        tag = getattr(self, "tag", None)
        if tag is not None:
            url = url + "tag/" + tag
        yield scrapy.Request(url, self.parse)

    def parse(self, response):
        for quote in response.css("div.quote"):
            yield {
                "text": quote.css("span.text::text").get(),
                "author": quote.css("small.author::text").get(),
            }

        next_page = response.css("li.next a::attr(href)").get()
        if next_page is not None:
            yield response.follow(next_page, self.parse)

如果传入tag=humor,爬虫就只会访问humor标签下的URL,例如https://quotes.toscrape.com/tag/humor。

更多处理方法见爬虫参数文档。

下一步

本教程只介绍了Scrapy基础,还有许多功能未提及。Scrapy概览中的“还有什么?”一节,可以快速了解其他重要功能。

可以继续阅读基本概念,了解命令行工具、爬虫、选择器,以及本教程未覆盖的数据建模等内容。如果更喜欢尝试示例项目,请查看示例。

如果使用编程代理,请参阅结合编程代理使用Scrapy。


原文:Scrapy完整入门教程;作者:Scrapy文档贡献者;日期:滚动latest版本。原文及源码权利归原作者和相应权利人所有。

© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容