python爬蟲（3）——python爬取大規模資料的的方法和步驟

阿新 • • 發佈：2019-01-06

python爬取大規模資料的的方法和步驟：

一、爬取我們所需要的一線連結

channel_extract.py
這裡的一線連結也就是我們所說的大類連結：

from bs4 import BeautifulSoup
import requests

start_url = 'http://lz.ganji.com/wu/'
host_url = 'http://lz.ganji.com/'

def get_channel_urls(url):
    wb_data = requests.get(url)
    soup = BeautifulSoup(wb_data.text, 'lxml' 
)
    links = soup.select('.fenlei > dt > a')
    #print(links)
    for link in links:
        page_url = host_url + link.get('href')
        print(page_url)

#get_channel_urls(start_url)


channel_urls = '''
http://lz.ganji.com/jiaju/
http://lz.ganji.com/rirongbaihuo/
http://lz.ganji.com/shouji/
http://lz.ganji.com/bangong/
http://lz.ganji.com/nongyongpin/
http://lz.ganji.com/jiadian/
http://lz.ganji.com/ershoubijibendiannao/
http://lz.ganji.com/ruanjiantushu/
http://lz.ganji.com/yingyouyunfu/
http://lz.ganji.com/diannao/
http://lz.ganji.com/xianzhilipin/
http://lz.ganji.com/fushixiaobaxuemao/
http://lz.ganji.com/meironghuazhuang/
http://lz.ganji.com/shuma/
http://lz.ganji.com/laonianyongpin/
http://lz.ganji.com/xuniwupin/
'''

那麼拿我爬取的58同城為例就是爬取了二手市場所有品類的連結，也就是我說的大類連結；
找到這些連結的共同特徵，用函式將其輸出，並作為多行文字儲存起來。

二、獲取我們所需要的詳情頁面的連結和詳情資訊

page_parsing.py

1、說說我們的資料庫：

先看程式碼：

#引入庫檔案
from bs4 import BeautifulSoup
import requests
import pymongo #python操作MongoDB的庫
import re
import time

#連結和建立資料庫
client = pymongo.MongoClient('localhost' 
, 27017)
ceshi = client['ceshi'] #建ceshi資料庫
ganji_url_list = ceshi['ganji_url_list'] #建立表文件
ganji_url_info = ceshi['ganji_url_info']

2、判斷頁面結構是否和我們想要的頁面結構相匹配，比如有時候會有404頁面；

3、從頁面中提取我們想要的連結，也就是每個詳情頁面的連結；

這裡我們要說的是一個方法就是:
item_link = link.get('href').split('?')[0]

這裡的這個link什麼型別的，這個get方法又是什麼鬼？
後來我發現了這個型別是

<class 'bs4.element.Tab>

如果我們想要單獨獲取某個屬性，可以這樣，例如我們獲取它的 class 叫什麼

print soup.p['class']
#['title']

還可以這樣，利用get方法，傳入屬性的名稱，二者是等價的

print soup.p.get('class')
#['title']

下面我來貼上程式碼：

#爬取所有商品的詳情頁面連結：
def get_type_links(channel, num):
    list_view = '{0}o{1}/'.format(channel, str(num))
    #print(list_view)
    wb_data = requests.get(list_view)
    soup = BeautifulSoup(wb_data.text, 'lxml')
    linkOn = soup.select('.pageBox') #判斷是否為我們所需頁面的標誌；
    #如果爬下來的select連結為這樣：div.pageBox > ul > li:nth-child(1) > a > span  這裡的:nth-child(1)要刪掉
    #print(linkOn)
    if linkOn:
        link = soup.select('.zz > .zz-til > a')
        link_2 = soup.select('.js-item > a')
        link = link + link_2
        #print(len(link))
        for linkc in link:
            linkc = linkc.get('href')
            ganji_url_list.insert_one({'url': linkc})
            print(linkc)
    else:
        pass

4、爬取詳情頁中我們所需要的資訊

我來貼一段程式碼：

#爬取趕集網詳情頁連結：
def get_url_info_ganji(url):
    time.sleep(1)
    wb_data = requests.get(url)
    soup = BeautifulSoup(wb_data.text, 'lxml')
    try:
        title = soup.select('head > title')[0].text
        timec = soup.select('.pr-5')[0].text.strip()
        type = soup.select('.det-infor > li > span > a')[0].text
        price = soup.select('.det-infor > li > i')[0].text
        place = soup.select('.det-infor > li > a')[1:]
        placeb = []
        for placec in place:
            placeb.append(placec.text)
        tag = soup.select('.second-dt-bewrite > ul > li')[0].text
        tag = ''.join(tag.split())
        #print(time.split())
        data = {
            'url' : url,
            'title' : title,
            'time' : timec.split(),
            'type' : type,
            'price' : price,
            'place' : placeb,
            'new' : tag
        }
        ganji_url_info.insert_one(data) #向資料庫中插入一條資料；
        print(data)
    except IndexError:
        pass

四、我們的主函式怎麼寫？

main.py
看程式碼：

#先從別的檔案中引入函式和資料：
from multiprocessing import Pool
from page_parsing import get_type_links,get_url_info_ganji,ganji_url_list
from channel_extract import channel_urls

#爬取所有連結的函式：
def get_all_links_from(channel):
    for i in range(1,100):
        get_type_links(channel,i)

#後執行這個函式用來爬取所有詳情頁的檔案：
if __name__ == '__main__':
#     pool = Pool()
#     # pool = Pool()
#     pool.map(get_url_info_ganji, [url['url'] for url in ganji_url_list.find()])
#     pool.close()
#     pool.join()


#先執行下面的這個函式，用來爬取所有的連結：
if __name__ == '__main__':
    pool = Pool()
    pool = Pool()
    pool.map(get_all_links_from,channel_urls.split())
    pool.close()
    pool.join()

五、計數程式

count.py
用來顯示爬取資料的數目；

import time
from page_parsing import ganji_url_list,ganji_url_info
while True:
    # print(ganji_url_list.find().count())
    # time.sleep(5)
    print(ganji_url_info.find().count())
    time.sleep(5)

python爬蟲（3）——python爬取大規模資料的的方法和步驟

python爬取大規模資料的的方法和步驟：一、爬取我們所需要的一線連結 channel_extract.py 這裡的一線連結也就是我們所說的大類連結： from bs4 import BeautifulSoup import requests

小白學 Python 爬蟲（25）：爬取股票資訊

人生苦短，我用 Python 前文傳送門：小白學 Python 爬蟲（1）：開篇小白學 Python 爬蟲（2）：前置準備（一）基本類庫的安裝小白學 Python 爬蟲（3）：前置準備（二）Linux基礎入門小白學 Python 爬蟲（4）：前置準備（三）Docker基礎入門小白學 Pyth

Python網路爬蟲（九）：爬取頂點小說網站全部小說，並存入MongoDB

前言：本篇部落格將爬取頂點小說網站全部小說、涉及到的問題有：Scrapy架構、斷點續傳問題、Mongodb資料庫相關操作。背景： Python版本：Anaconda3 執行平臺：Windows IDE：PyCharm 資料庫：MongoDB 瀏

Python爬蟲（一）--城市公交網路站點資料的爬取

作者：WenWu_Both 出處：http://blog.csdn.net/wenwu_both/article/ 版權：本文版權歸作者和CSDN部落格共有轉載：歡迎轉載，但未經作者同意，必須保留此段聲必須在文章中給出原文連結；否則必究法律責任

54. Python 爬蟲（3）

你是需要理解 match 網站 for 3.2 rst e30 【基於python3的版本】rllib下載：當不知道urlretrieve方法，寫法如下：from urllib import request url = "http://inews.gtimg.

python爬蟲（3）——SSL證書與Handler處理器

pan 高級訪問網站 size cos 中文名 ssl 內核 pos 一、SSL證書問題　　　　　　　　　　　　上一篇文章，我們創建了一個小爬蟲，下載了上海鏈家房產的幾個網頁。實際上我們在使用urllib聯網的過程中，會遇到證書訪問受限的問題。　　　　處理HTTPS

[Python]網路爬蟲（一）：抓取網頁的含義和URL基本構成

一、網路爬蟲的定義網路爬蟲，即Web Spider，是一個很形象的名字。把網際網路比喻成一個蜘蛛網，那麼Spider就是在網上爬來爬去的蜘蛛。網路蜘蛛是通過網頁的連結地址來尋找網頁的。從網站某一個頁面（通常是首頁）開始，讀取網頁的內容，找到在網頁中的其它連結地址

Python爬蟲入門——3.6 Selenium 爬取淘寶資訊

上一節我們介紹了Selenium工具的使用，本節我們就利用Selenium跟Chrome瀏覽器結合來爬取淘寶相關男士羽絨服商品的資訊，當然你可以用相同的方法來爬取淘寶其他商品的資訊。我們要爬取羽絨服的價格、圖片連線、賣家、賣家地址、收貨人數等資訊，並將其儲存在csv中 fr

小白學 Python 爬蟲（3）：前置準備（二）Linux基礎入門

人生苦短，我用 Python 前文傳送門：小白學 Python 爬蟲（1）：開篇小白學 Python 爬蟲（2）：前置準備（一）基本類庫的安裝 Linux 基礎 CentOS 官網： https://www.centos.org/ 。 CentOS 官方下載連結： https://www.cent

爬蟲（七）：爬取貓眼電影top100

all for rip pattern 分享爬取 values findall proc 一：分析網站目標站和目標數據目標地址：http://maoyan.com/board/4?offset=20目標數據：目標地址頁面的電影列表，包括電影名，電影圖片，主演，上映日期以

webmagic是個神奇的爬蟲（二）-- webmagic爬取流程細講

webmagic流程圖鎮樓：第一篇筆記講到了如何建立webmagic專案，這一講來說一說webmagic爬取的主要流程。 webmagic主要由Downloader（下載器）、PageProcesser（解析器）、Schedule（排程器）和Pi

爬蟲（4）：抓取ajax資料

import urllib.request import json # 請求頭 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHT

pyhton微博爬蟲（3）——獲取微博評論資料

本文的主要目標是獲取微博評論資料，具體包括微博評論連結、總評論數、使用者ID、使用者暱稱、評論時間、評論內容、使用者詳情連結等。實現程式碼如下所示： # -*- coding: utf-8 -*- """ Created on Tue Aug 8 16:

一個鹹魚的Python爬蟲之路（三）：爬取網頁圖片

you os.path odin 路徑生成存在 parent lose exist 學完Requests庫與Beautifulsoup庫我們今天來實戰一波，爬取網頁圖片。依照現在所學只能爬取圖片在html頁面的而不能爬取由JavaScript生成的圖。所以我找了這個網站

python 爬蟲（一） requests+BeautifulSoup 爬取簡單網頁代碼示例

utf-8 bs4 rom 文章都是 Coding man header 文本以前搞偷偷摸摸的事，不對，是搞爬蟲都是用urllib，不過真的是很麻煩，下面就使用requests + BeautifulSoup 爬爬簡單的網頁。詳細介紹都在代碼中註釋了，大家可以參閱。

自學python爬蟲（四）Requests+正則表示式爬取貓眼電影

前言學了requests庫和正則表示式之後我們可以做個簡單的專案來練練手咯！先附上專案GitHub地址，歡迎star和fork，也可以pull request哦~ 地址：https://github.com/zhangyanwei233/Maoyan100.git 正文開始哈哈哈

python爬蟲（五）：實戰【3. 使用正則來爬創客實驗室】

依然爬取創科實驗室網站中講座的資訊（只爬標題，其它同）但技術上採用requests+正則表示式思想： #通過正則表示式，獲取講座標題規則：<h3>中文字元出現4次任意字元</h3> m = str(re.findall('<h3

python 爬蟲（五）爬取多頁內容

import urllib.request import ssl import re def ajaxCrawler(url): headers = {"User-Agent":"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/5

python 爬蟲（三）模擬post請求，爬取資料

import urllib.request import urllib.parse url =r"http://www.baidu.com" #將要傳送的資料合成一個字典 #字典的鍵值在網頁裡找 data = { "username":"1507", "password":"230

python爬蟲（1）——簡單的爬取網頁的資訊

獲取網上真實的語料資料，本身對Py的掌握不是很好，記錄下自己學習的過程，希望對你有幫助。 #python3 獲得taoeba的語料（不知道從哪翻到的這個網站，有各國語言的句子，訪問速度較慢 # -

python爬蟲（3）——python爬取大規模資料的的方法和步驟

python爬取大規模資料的的方法和步驟：

一、爬取我們所需要的一線連結

二、獲取我們所需要的詳情頁面的連結和詳情資訊

1、說說我們的資料庫：

2、判斷頁面結構是否和我們想要的頁面結構相匹配，比如有時候會有404頁面；

3、從頁面中提取我們想要的連結，也就是每個詳情頁面的連結；

4、爬取詳情頁中我們所需要的資訊

四、我們的主函式怎麼寫？

五、計數程式

相關推薦