scrapy 2.3 在蜘蛛中提取数据

2021-05-31 16:59 更新

让我们回到蜘蛛身边。到目前为止，它还没有提取任何数据，特别是将整个HTML页面保存到一个本地文件中。让我们把上面的提取逻辑集成到蜘蛛中。

剪贴蜘蛛通常会生成许多字典，其中包含从页面中提取的数据。为此，我们使用 yield 回调中的python关键字，如下所示：

import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"
    start_urls = [
        'http://quotes.toscrape.com/page/1/',
        'http://quotes.toscrape.com/page/2/',
    ]

    def parse(self, response):
        for quote in response.css('div.quote'):
            yield {
                'text': quote.css('span.text::text').get(),
                'author': quote.css('small.author::text').get(),
                'tags': quote.css('div.tags a.tag::text').getall(),
            }

如果运行这个spider，它将用日志输出提取的数据：

2016-09-19 18:57:19 [scrapy.core.scraper] DEBUG: Scraped from <200 http://quotes.toscrape.com/page/1/>
{'tags': ['life', 'love'], 'author': 'André Gide', 'text': '“It is better to be hated for what you are than to be loved for what you are not.”'}
2016-09-19 18:57:19 [scrapy.core.scraper] DEBUG: Scraped from <200 http://quotes.toscrape.com/page/1/>
{'tags': ['edison', 'failure', 'inspirational', 'paraphrased'], 'author': 'Thomas A. Edison', 'text': "“I have not failed. I've just found 10,000 ways that won't work.”"}

以上内容是否对您有帮助：

← scrapy 2.3 提取数据

scrapy 2.3 存储抓取的数据 →

写笔记

我要补充

scrapy 2.3 在蜘蛛中提取数据

推荐文章

推荐教程

推荐课程