Welcome back, everyone! In today’s lesson, AI on the Edge LESSON 52: Gesture Recognition on the Raspberry Pi 5, we are building a custom, real-time hand gesture recognition system from scratch. Up until now, we have been using MediaPipe to track keypoints, identify individual hand joints, and extract raw 3D coordinate data. But raw coordinates alone don’t tell the computer what your hand is actually doing. Today, we bridge that gap by teaching our Raspberry Pi 5 how to train on custom hand positions and dynamically classify live hand gestures using pure mathematics and vector geometry—no heavy neural network retraining required.
The core secret to making gesture recognition work reliably across different camera distances is scale-invariant distance normalization. If you simply measure the pixel distance between your fingertips and your wrist, bringing your hand closer to the camera lens will blow out those numbers and break your classifier. To solve this, our script establishes a baseline unit of measurement for every single frame: the distance between the wrist (landmark 0) and the index finger base joint (landmark 5). By dividing all measured finger-tip-to-wrist and tip-to-tip distances by this dynamic scale factor, the resulting feature vector remains identical whether your hand is two feet away or five feet away from the lens.
Once we extract this normalized 15-element feature vector—representing the spatial ratios between all five fingertips relative to the wrist and each other—we construct an interactive training phase directly inside the live video loop. The program prompts you in the terminal for the number of custom gestures you wish to record, along with their labels (such as “Peace”, “Thumbs Up”, or “Fist”). As you present each gesture to your camera feed and press the spacebar, the algorithm captures the instantaneous geometric signature of your hand and stores it in a runtime training dictionary.
During live inference, our classifier calculates the Sum of Absolute Errors (SAE), or Manhattan distance, between the live incoming feature vector and every stored profile in our training dataset. The function loops through all saved gestures to find the candidate with the absolute lowest cumulative error score. To prevent false positives when an untrained or messy hand pose is shown, we compare that lowest error score against a strict threshold value (maxErrorThreshold = 3.5). If the error falls within the allowable boundary, the matched gesture label is instantly displayed on the live OpenCV HUD; otherwise, the system safely defaults to “Unknown”.
We manage our camera feed using high-speed frame capture with Picamera2 running at 60 frames per second at 1280×720 resolution, processing hand detection with MediaPipe Hands, and rendering real-time performance diagnostics—including an exponential moving average FPS counter—directly onto the video output. Go ahead, study the methodology, load the concept onto your Raspberry Pi 5, grab a hot cup of coffee, and let’s get building!
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 |
# ==================================================================== # DISCLAIMER: # This code is provided as-is for educational and experimental # purposes only. The author makes no representations or warranties of # any kind concerning the safety, suitability, or accuracy of this # code. Use at your own risk. The author assumes no liability for any # damages, system failures, security breaches, or network issues # resulting from the use or implementation of this script. # ==================================================================== import cv2 import time from picamera2 import Picamera2 import mediapipe as mp import math def findDistance(pt1, pt2): dX = pt1.x - pt2.x dY = pt1.y - pt2.y dZ = pt1.z - pt2.z dist = math.sqrt(dX*dX + dY*dY + dZ*dZ) return dist def getTipDistances(handLm): lm = handLm.landmark wrist = lm[0] indexBase = lm[5] scaleFactor = findDistance(wrist, indexBase) tips = [lm[4],lm[8],lm[12],lm[16],lm[20]] tipDistances =[] for tip in tips: dist = findDistance( wrist, tip) / scaleFactor tipDistances.append(dist) for startTip in range(len(tips)): for endTip in range(startTip+1, len(tips)): dist = findDistance(tips[startTip],tips[endTip]) / scaleFactor tipDistances.append(dist) return tipDistances def calculateError(tipDist, trainDist): totalError = 0 for idx in range(len(tipDist)): diff = tipDist[idx] -trainDist[idx] totalError = totalError + abs(diff) return totalError def findGesture(tipLive, trainedData, maxAllowedError): bestMatch = "Unknown" lowestError = 999999999 for gestureName, trainedDist in trainedData.items(): error = calculateError(tipLive, trainedDist) if error < lowestError: lowestError = error bestMatch = gestureName if lowestError > maxAllowedError: return "Unknown" if lowestError <= maxAllowedError: return bestMatch print("HAND GESTURE TRAINING SETUP") numGestures = int(input("Enter Number of Gestures to Train On: ")) gestureNames =[] for idx in range(numGestures): name = input("Enter Name for Gesture #"+ str(idx+1)+ ": ") gestureNames.append(name) W=1280 H=720 tStart = time.time() fps = 15 RES = (W,H) piCam = Picamera2(1) piCam.preview_configuration.main.size = RES piCam.preview_configuration.main.format = "RGB888" piCam.preview_configuration.controls.FrameRate=60 piCam.preview_configuration.align() piCam.configure("preview") piCam.start() textLowerLeft = (int(W*.01),int(H*.07)) fontFace = cv2.FONT_HERSHEY_SIMPLEX fontThickness = int(W/425) fontScale = H*.002 fontColor = (0,0,255) textFps = (int(W*.01),int(H*.07)) textGesture= (int(W*.01),int(H*.14)) textGesturePrompt = (int(W*.01),int(H*.21)) maxErrorThreshold = 3.5 hands = mp.solutions.hands.Hands( model_complexity = 0, min_detection_confidence = .5, min_tracking_confidence = .5, max_num_hands =2) trainedData ={} trainIdx = 0 while True: deltaT = time.time() - tStart tStart=time.time() fps = fps*.95 + (1/deltaT)*.05 frame= piCam.capture_array() frame=cv2.flip(frame,-1) rgbFrame= cv2.cvtColor(frame, cv2.COLOR_BGR2RGB) results = hands.process(rgbFrame) if results.multi_hand_landmarks: frameHeight, frameWidth, _ = frame.shape for handLm in results.multi_hand_landmarks: # Draw landmark numbers on hand joints for idx, landmark in enumerate(handLm.landmark): pixelX = int(landmark.x * frameWidth) pixelY = int(landmark.y * frameHeight) cv2.putText( frame, str(idx), (pixelX, pixelY), cv2.FONT_HERSHEY_SIMPLEX, .6*fontScale, (255, 0, 0), fontThickness ) if trainIdx < len(gestureNames): targetName = gestureNames[trainIdx] promptText = "show " + targetName + " & Press Space" cv2.putText(frame,promptText, textGesturePrompt, fontFace,fontScale, (0,255,255), fontThickness) key = cv2.waitKey(1) if key == ord(' '): if results.multi_hand_landmarks: targetHand = results.multi_hand_landmarks[0] trainedData[targetName] = getTipDistances(targetHand) trainIdx = trainIdx + 1 if key == ord('q'): break detectGesture = "Unknown" if results.multi_hand_landmarks: liveDistances = getTipDistances(results.multi_hand_landmarks[0]) detectGesture = findGesture(liveDistances, trainedData, maxErrorThreshold) labelText = "Gesture: " + detectGesture cv2.putText(frame, labelText, textGesture,fontFace,fontScale, (0,255,0),fontThickness) myText = "FPS: "+str(round(fps,1)) cv2.putText(frame,myText,textLowerLeft,fontFace,fontScale,fontColor,fontThickness) cv2.imshow("Camera", frame) cv2.moveWindow("Camera",0,60) if cv2.waitKey(1)==ord('q'): break cv2.destroyAllWindows() print('Program Terminated') |