树莓派实战指南:调用 Custom Vision 目标检测完成货架库存盘点
【免费下载链接】IoT-For-Beginners12 Weeks, 24 Lessons, IoT for All!项目地址: https://gitcode.com/GitHub_Trending/io/IoT-For-Beginners
在 IoT-For-Beginners 的零售项目中,stock-detector目标检测模型已经训练并发布,接下来要让树莓派或虚拟 IoT 设备把摄像头对准货架,调用 Custom Vision 目标检测接口,给每罐番茄酱标出边界框,再按重叠面积去重、输出商品数量,完成一次库存盘点。这份实战指南从 3 行代码的改动讲起,覆盖树莓派调用目标检测、边界框相对坐标换算,以及用 Shapely 剔除重复预测的完整库存统计流程。
一张表看懂:图像分类 vs 目标检测
stock-counter程序的骨架直接复用自制造项目"从设备检查水果"的流程:拍照 → 调模型 → 解析输出。分类与检测的区别集中在下面这张表:
| 对比维度 | 图像分类器 | 目标检测器 |
|---|---|---|
| SDK 调用方法 | predictor.classify_image(project_id, iteration_name, image) | predictor.detect_image(project_id, iteration_name, image) |
| 返回结构 | 每个标签一条:tag_name+probability | 一张图返回多条:每条含tag_name、probability和bounding_box |
| 是否必须阈值 | 通常不用(每标签返回最高置信度) | 必须(同一物体可能被重复框出,要过滤低概率结果) |
| 典型用途 | 判断图里"是什么" | 判断物体"在哪、有几个",进而盘点库存 |
环境与设备准备:Raspberry Pi 与虚拟设备对照
| 项目 | Raspberry Pi | 虚拟 IoT 设备 |
|---|---|---|
| 项目位置 | 在 Pi 上创建stock-counter文件夹 | 在电脑上创建,并先建好虚拟环境 |
| 摄像头 | 安装 PiCamera,固定对准货架(排线搭在盒子上方或双面胶固定均可) | 安装 CounterFit 与counterfit_shims_picamera;用静态图就准备检测器没见过的货架照片,用网络摄像头则保证视野拍全库存 |
| 依赖安装 | 直接执行pip3 install | 在激活的虚拟环境内执行所有pip3 install |
💡 拍照、调用模型这两段的代码与制造项目完全一致,直接复用即可;本指南只讲"改"的部分。
三行代码:从 classify_image 切到 detect_image
拍照和创建predictor客户端的代码不动,把分类调用整段替换掉即可。
改动前:
results = predictor.classify_image(project_id, iteration_name, image) for prediction in results.predictions: print(f'{prediction.tag_name}:\t{prediction.probability * 100:.2f}%')改动后:
results = predictor.detect_image(project_id, iteration_name, image) threshold = 0.3 predictions = list(prediction for prediction in results.predictions if prediction.probability > threshold) for prediction in predictions: print(f'{prediction.tag_name}:\t{prediction.probability * 100:.2f}%')这里做了三件事:换成detect_image运行目标检测器;用列表推导式收集所有概率大于threshold的预测;遍历打印标签名和百分比概率。虚拟设备版本只在开头多一行CounterFitConnection.init('127.0.0.1', 5000)连接 CounterFit 服务、并把摄像头换成counterfit_shims_picamera里的PiCamera,其余完全一致。完整实现见 Pi 版 app.py 和 虚拟设备版 app.py。
Custom Vision 阈值设置:哪些检测概率可信
为什么检测必须过阈值?分类器每个标签只给一条结果,而检测器一张图能返回多个框,同一罐货可能被重复框出。threshold = 0.3就是把低于 30% 的"不太像"预测全部丢掉,只留下可信的。
跑起来后,树莓派上的典型输出是这样的:
pi@raspberrypi:~/stock-counter $ python3 app.py tomato paste: 34.13% tomato paste: 33.95% tomato paste: 35.05% tomato paste: 32.80%💁 阈值过低会带入大量误检,过高又会漏检,请按自己的货架照片调整
threshold。
设备拍的照片和这些概率都能在 Custom Vision 门户的Predictions页逐条复核:
边界框相对坐标:top、left、height、width 换算
每个检测结果除了标签和概率,还带一个bounding_box,由top、left、height、width四个 0~1 的相对值定义,分别表示框距图片上边、左边的距离及框本身高宽占图像对应维度的比例。原点(0, 0)在图片左上角,框的下边缘等于top + height。
拿一张宽 600 像素、高 800 像素的图来算账:框从上边 320 像素处开始,top = 0.4(800 × 0.4 = 320);距左边 240 像素,left = 0.4(600 × 0.4 = 240);框高 240 像素,height = 0.3(800 × 0.3 = 240);框宽 120 像素,width = 0.2(600 × 0.2 = 120)。
| 坐标 | 值 |
|---|---|
| Top | 0.4 |
| Left | 0.4 |
| Height | 0.3 |
| Width | 0.2 |
用相对值而不是像素值,好处是同一组数字无论图片缩放到 640×480 还是 1280×960 都能按原比例定位,不用跟着分辨率重算。
把边界框画出来看效果
先打印核对:把for循环里的输出改成带框信息——
print(f'{prediction.tag_name}:\t{prediction.probability * 100:.2f}%\t{prediction.bounding_box}')控制台就会多出left、top、width、height四个 0~1 的值。想直观看到框在哪,用 Pillow 把相对坐标换算成像素再画上去。安装依赖(虚拟设备请在激活的虚拟环境里执行):
pip3 install pillow顶部加导入,文件末尾追加绘制代码:
from PIL import Image, ImageDraw, ImageColorwith Image.open('image.jpg') as im: draw = ImageDraw.Draw(im) for prediction in predictions: scale_left = prediction.bounding_box.left scale_top = prediction.bounding_box.top scale_right = prediction.bounding_box.left + prediction.bounding_box.width scale_bottom = prediction.bounding_box.top + prediction.bounding_box.height left = scale_left * im.width top = scale_top * im.height right = scale_right * im.width bottom = scale_bottom * im.height draw.rectangle([left, top, right, bottom], outline=ImageColor.getrgb('red'), width=2) im.save('image.jpg')换算原理就是"相对值 × 图像尺寸":left为 0.5、图宽 600 像素时得到 300 像素(0.5 × 600 = 300);右下两个角由left + width、top + height先算出相对值再乘尺寸。保存时直接覆盖image.jpg,在 VS Code 资源管理器里点开就能看到红色矩形:
Shapely 重叠面积去重:数出货架上的商品数
货架上罐头前后排摆放时,框会互相重叠;重叠越大,越可能两个框指向同一罐货。统计前先安装依赖(树莓派要先装系统库):
sudo apt install libgeos-dev pip3 install shapely顶部导入from shapely.geometry import Polygon,然后在画框代码之前加两段。第一段定义重叠阈值和框转多边形的辅助函数:
overlap_threshold = 0.20 def create_polygon(prediction): scale_left = prediction.bounding_box.left scale_top = prediction.bounding_box.top scale_right = prediction.bounding_box.left + prediction.bounding_box.width scale_bottom = prediction.bounding_box.top + prediction.bounding_box.height return Polygon([(scale_left, scale_top), (scale_right, scale_top), (scale_right, scale_bottom), (scale_left, scale_bottom)])第二段做两两比较并输出库存数:
to_delete = [] for i in range(0, len(predictions)): polygon_1 = create_polygon(predictions[i]) for j in range(i+1, len(predictions)): polygon_2 = create_polygon(predictions[j]) overlap = polygon_1.intersection(polygon_2).area smallest_area = min(polygon_1.area, polygon_2.area) if overlap > (overlap_threshold * smallest_area): to_delete.append(predictions[i]) break for d in to_delete: predictions.remove(d) print(f'Counted {len(predictions)} stock items')要点逐条说清:
- overlap_threshold = 0.20允许 20% 的重叠比例;阈值不是绝对面积,而是相对两框中较小者
smallest_area的比例; Polygon.intersection返回两框的交集多边形,取.area就是重叠面积;- 比较顺序是"第 1 个 vs 第 2、3、4…个,第 2 个 vs 第 3、4…个",即
i与j = i+1的嵌套循环,判定重复后立即break; - Python 不允许遍历列表时删元素,所以先收集进to_delete,循环结束后统一
remove; - 这段逻辑放在画框之前,生成的图片只画出去重后幸存的框;
Counted N stock items输出的数量可以直接发给 IoT 服务,库存偏低时触发补货告警。
💡 这是刻意简化的方案——遇到重叠就删掉配对中的第一个。生产环境需要更细的策略,比如多个物体之间的复杂重叠、一个框完全包含在另一个框里的情形。仓库里 code-count 虚拟设备版 app.py 把
overlap_threshold设为 0.002,可以对比不同阈值对剔除结果的影响。
接下来:完整代码与 IoT Edge 迁移作业
- 计数版完整代码:Pi 版 code-count app.py、虚拟设备版 code-count app.py
- 分步文档:单板机目标检测调用、单板机库存计数
- 进阶作业:把目标检测器导出为紧凑模型并部署到 Azure IoT Edge,再从 IoT 设备访问边缘版本,见 assignment.md
零售项目到这里就完整了:从训练stock-detector到树莓派上的库存盘点。作业做完后,记得按 clean-up.md 清理云端资源。
【免费下载链接】IoT-For-Beginners12 Weeks, 24 Lessons, IoT for All!项目地址: https://gitcode.com/GitHub_Trending/io/IoT-For-Beginners
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考