同时下载多张宝可梦图片
这篇文章会使用 Python 的 Requests 和 Beautiful Soup 函数库,搭配 threading 内建函数库进行多执行绪处理,实作爬取宝可梦图鉴里的宝可梦网页,并自动将每只宝可梦的图片下载到电脑中。
本篇使用的 Python 版本为 3.7.12,所有范例可使用 Google Colab 实作,不用安装任何软件 ( 参考:使用 Google Colab )
观察网页结构
前往宝可梦图鉴官方网站,点击任何一只宝可梦进入该宝可梦的页面,本篇教学范例使用杰尼龟的页面。
在画面中按下鼠标右键,检视网页的源代码。
从源代码中可以看见 meta 标签里,包含了一张杰尼龟的缩图,待会就会透过爬虫抓取这张缩图。
爬取 meta 标签里的图片
参考“Requests 函数库”和“Beautiful Soup 函数库”安装 Requests 和 Beautiful Soup 函数库 ( 如果是 Colab 或 Anaconda Jupyter 应该已经内建 ),执行下方的程序码,使用 select 方法爬取具有 property 属性为 og:image 的 meta 标签,就能爬取到 meta 标签里的图片网址。
import requests
from bs4 import BeautifulSoup
web = requests.get(f'https://tw.portal-pokemon.com/play/pokedex/0007')
soup = BeautifulSoup(web.text, "html.parser")
img = soup.select('meta[property="og:image"]') # 爬取具有 property 屬性為 og:image 的 meta 標籤
imgUrl = img[0]['content'] # 取出該標籤的 content 內容
print(imgUrl)
参考“内建函数 ( 文件读写 open )”教学搭配 Requests 函数库,就能读取网址并将该图片下载到 pokemon 的文件夹中 ( 请先建立该文件夹 )。
如果是 Colab 环境需要先连动 Google Drive,并透过 os.chdir 切换根目录。
import requests
from bs4 import BeautifulSoup
web = requests.get(f'https://tw.portal-pokemon.com/play/pokedex/0007')
soup = BeautifulSoup(web.text, "html.parser")
img = soup.select('meta[property="og:image"]') # 爬取具有 property 屬性為 og:image 的 meta 標籤
imgUrl = img[0]['content'] # 取出該標籤的 content 內容
print(imgUrl)
import os
os.chdir('/content/drive/MyDrive/Colab Notebooks') # Colab 換路徑使用
imgFile = requests.get(imgUrl) # 讀取圖片資訊
f = open(f'pokemon/0007.png', 'wb') # 建立 0007.png 圖片
f.write(imgFile.content) # 寫入圖片
f.close()
print('ok')
搭配 threading 同时下载多张宝可梦图片
使用 threading 函数库,将原本的程序,从“同步”改为“异步”执行,因为有多张图片要处理,所以采用网址的编号 ( 宝可梦编号 ) 作为图片的名称,透过循环的方式,就能同时一次下载上百张图片。
import os
os.chdir('/content/drive/MyDrive/Colab Notebooks') # Colab 換路徑使用
import requests
from bs4 import BeautifulSoup
import threading
def download(num):
# 加入 try 保護避免遇到無法下載的狀況而發生錯誤
try:
web = requests.get(f'https://tw.portal-pokemon.com/play/pokedex/{num}') # 使用變數替換網址
soup = BeautifulSoup(web.text, "html.parser")
img = soup.select('meta[property="og:image"]')
imgUrl = img[0]['content']
imgFile = requests.get(imgUrl)
f = open(f'pokemon/{num}.png', 'wb')
f.write(imgFile.content)
f.close()
print(num)
except:
print('error')
pass
# 使用迴圈,一次可以下載 1~99 張圖片
for i in range(1,100):
n = f'{i:04d}'
threading.Thread(target=download, args=(n,)).start()
程序执行后,就会看见图片“几乎同时”下载到电脑里。
使用 concurrent.futures
如果遇到 Colab 无法使用 threading 的情形,可以改用 concurrent.futures 处理异步下载。
import os
os.chdir('/content/drive/MyDrive/Colab Notebooks') # Colab 換路徑使用
import requests
from bs4 import BeautifulSoup
from concurrent.futures import ThreadPoolExecutor
def download(num):
try:
web = requests.get(f'https://tw.portal-pokemon.com/play/pokedex/{num}')
soup = BeautifulSoup(web.text, "html.parser")
img = soup.select('meta[property="og:image"]')
imgUrl = img[0]['content']
imgFile = requests.get(imgUrl)
f = open(f'pokemon/{num}.png', 'wb')
f.write(imgFile.content)
f.close()
print(num)
except:
print('error')
pass
numArr = [f'{j:04d}' for j in range(1,10)] # 建立圖片檔名清單
executor = ThreadPoolExecutor() # 建立非同步的多執行緒的啟動器
with ThreadPoolExecutor() as executor:
executor.map(download, numArr) # 同時下載圖片
程序执行后,就会看见图片“几乎同时”下载到电脑里。
小结
如果要用爬虫下载大量图片,使用“异步”( 多执行绪平行处理 ) 的方式是相当方便的做法,不仅不用一张张的等待,更能发挥闲置 CPU 最大的效益 ( 同一时间里可以做很多事情 )。
微信扫码关注
抖音扫码关注