辨识剪刀、石头、布
这篇教学会使用 Teachable Machine 训练“剪刀、石头、布”的图像模型,再透过 OpenCV 搭配 tensorflow 读取摄影镜头图像进行辨识。
快速导览:
因为程序使用 Jupyter 搭配 Tensorflow 进行开发,所以请先阅读“使用 Anaconda”和“Jupyter 安装 Tensorflow”,安装对应的软件包。
安装相关软件包
因 Tensorflow 和 Python 版本更新的缘故 ( 原本 Jupyter 方式可能无法操作 ),建议参考“使用 Python 虚拟环境”,建立虚拟环境进行实作 ( 目前测试可运行的 Python 版本为 3.9.6 ),并使用下列指令安装软件包:
安装 OpenCV,测试可运行版本为 4.9.0.80。
pip install opencv-python
安装 Tensorflow,测试可运行版本为 2.15.0。
pip install tensorflow
Teachable Machine 建立分类,训练模型
打开 Teachable Machine 网站后,点击“开始使用”建立新专案 ( 最下方可以切换语系为繁体中文 ),选择 “图片专案 > 标准图片模型”,进入图片模型训练流程。
参考“使用 Teachable Machine”文章,训练“剪刀”、“石头”和“布”以及“背景”共四个分类,范例使用 a 作为剪刀分类的名称,b 作为石头分类名称,c 作为布分类名称,bg 作为背景分类名称 ( 建议左手和右手都要训练 )。
为什么要训练“背景”呢?因为如果没有背景分类,会在没有出拳的时候自动判断最接近的分类,导致判断出错,所以建议一定要训练一个背景的分类 ( 背景的分类除了单纯的背景,也可以加入许多杂乱的图片,增加背景的准确度 )。
分类建立后点击“训练模型”,等待训练完成,可以从预览区域测试模型。
确认辨识结果没问题,就可以将模型导出为 OpenCV Keras 类型,解压缩后将 .h5 的模型文件放到指定的文件夹里。
搭配 OpenCV 进行辨识,并加入文字
下方的程序码使用 OpenCV 读取摄影镜头的图像,即时判断现在出现的图像是什么分类,并透过 putText() 方法,在图像中加入分类的名称。
from keras.models import load_model # TensorFlow is required for Keras to work
import cv2 # Install opencv-python
import numpy as np
# Disable scientific notation for clarity
np.set_printoptions(suppress=True)
# Load the model
model = load_model("keras_Model.h5", compile=False)
def text(text): # 建立顯示文字的函式
global show_img # 設定 img 為全域變數
org = (0,50) # 文字位置
fontFace = cv2.FONT_HERSHEY_SIMPLEX # 文字字型
fontScale = 2.5 # 文字尺寸
color = (255,255,255) # 顏色
thickness = 5 # 文字外框線條粗細
lineType = cv2.LINE_AA # 外框線條樣式
cv2.putText(show_img, text, org, fontFace, fontScale, color, thickness, lineType) # 放入文字
cap = cv2.VideoCapture(0)
if not cap.isOpened():
print("Cannot open camera")
exit()
while True:
ret, frame = cap.read()
if not ret:
print("Cannot receive frame")
break
img = cv2.resize(frame , (398, 224))
show_img = img[0:224, 80:304]
img = np.asarray(show_img, dtype=np.float32).reshape(1, 224, 224, 3)
img = (img / 127.5) - 1
prediction = model.predict(img)
index = np.argmax(prediction)
print(index)
if index == 0:
text('a') # 使用 text() 函式,顯示文字
if index == 1:
text('b')
if index == 2:
text('c')
cv2.imshow("Webcam Image", show_img)
if cv2.waitKey(1) == ord('q'):
break # 按下 q 鍵停止
cap.release()
cv2.destroyAllWindows()
如果要加入中文,必须要使用 pillow 函数库,打开命令提示字元或终端机,启动 tensorflow 虚拟环境安装 pillow 函数库,再使用下方的程序码,就能够加入中文的文字。
from keras.models import load_model # TensorFlow is required for Keras to work
import cv2 # Install opencv-python
import numpy as np
from PIL import ImageFont, ImageDraw, Image # 載入 PIL 相關函式庫
fontpath = 'NotoSansTC-Regular.ttf' # 設定字型路徑
# Disable scientific notation for clarity
np.set_printoptions(suppress=True)
# Load the model
model = load_model("keras_Model.h5", compile=False)
def text(text): # 建立顯示文字的函式
global show_img # 設定 img 為全域變數
org = (0,50) # 文字位置
font = ImageFont.truetype(fontpath, 50) # 設定字型與文字大小
imgPil = Image.fromarray(show_img) # 將 img 轉換成 PIL 影像
draw = ImageDraw.Draw(imgPil) # 準備開始畫畫
draw.text((0, 0), text, fill=(255, 255, 255), font=font) # 寫入文字
show_img = np.array(imgPil)
cap = cv2.VideoCapture(0)
if not cap.isOpened():
print("Cannot open camera")
exit()
while True:
ret, frame = cap.read()
if not ret:
print("Cannot receive frame")
break
img = cv2.resize(frame , (398, 224))
show_img = img[0:224, 80:304]
img = np.asarray(show_img, dtype=np.float32).reshape(1, 224, 224, 3)
img = (img / 127.5) - 1
prediction = model.predict(img)
index = np.argmax(prediction)
print(index)
if index == 0:
text('布') # 使用 text() 函式,顯示文字
if index == 1:
text('剪刀')
if index == 2:
text('石頭')
cv2.imshow("Webcam Image", show_img)
if cv2.waitKey(1) == ord('q'):
break # 按下 q 鍵停止
cap.release()
cv2.destroyAllWindows()
微信扫码关注
抖音扫码关注