Tag Archives: Fusion AI Hat

AI on the Edge LESSON 32: Facial Recognition and Eye Tracking in OpenCV

Hey guys, Paul McWhorter here from TopTechBoy.com. Welcome back to our AI on the Edge series. If you’ve been following along, you already know how to pull high-frame-rate video off your Raspberry Pi 5 using the new picamera2 library, and you know how to use OpenCV to hunt down faces in a crowded frame.

But today, we are taking things a massive step forward. We aren’t just looking for faces anymore—we are looking inside the face to track the eyes.

This lesson highlights one of the most vital concepts in all of computer vision: The Region of Interest (ROI). If you try to scan an entire 1280×720 frame for tiny features like eyes, your frame rate will absolutely tank. Instead, we are going to act like real engineers. We will use a cascading logic approach: find the face first, isolate that exact box, and search only inside that small window for the eyes.

Go ahead and pour yourself a nice, cold glass of iced coffee or a hot cup of black coffee, get your code ready, and let’s break down exactly how this program works.

This is the code we developed in the video:

Code Architecture & Codex Breakdown

Since you already have the script loaded up in your IDE, let’s dissect the critical logic gates that make this tracking script fast and accurate.

1. Setting Up the High-Performance Pipeline

We configure the Picamera2 frontend to grab a crisp RGB888 array at a resolution of 1280×720 targeting 60 FPS. By using .capture_array(), we bypass slow formatting overhead and feed raw pixel data directly into OpenCV. Because the camera orientation might be flipped depending on your desktop mounting rig, we use cv2.flip(frame, -1) to keep the spatial coordinates intuitive.

2. The Cascading Filter Matrix

Notice how we initialize two distinct classifiers using pre-trained Haar Cascades:

  • haarcascade_frontalface_default.xml (To grab the macro features of the face)

  • haarcascade_eye.xml (To grab the micro features of the eyes)

We pass a minSize parameter of 100×100 pixels for the face detector. Why? Because we don’t care about background noise or tiny false positives across the room. We want to find you, sitting right in front of the workstation.

3. The Magic of the Region of Interest (ROI)

This is where the real engineering happens. Look closely at this inner loop:

Instead of passing the massive gray frame to the eye finder, we slice the array: gray[y:y+h, x:x+w]. This isolates a tiny sub-matrix containing nothing but your face. The search area drops exponentially, keeping our frame rates close to maximum velocity.

4. Re-Mapping Local to Global Coordinates

When the eye detector finds a match inside the sliced face frame, it returns local coordinates (i, j, w, h) relative to the top-left corner of that face box, not the whole screen. If you tried to draw a rectangle directly at (i, j), your eye boxes would be floating erratically in the top-left corner of your monitor!

To fix this spatial offset, we map them back to global coordinate space by adding the face’s original offsets:

  • Global X Position: x + i

  • Global Y Position: y + j

General Knowledge: How Haar Cascades and ROIs Work Under the Hood

Now that you understand the mechanics of the script, let’s dive into the fundamental computer vision theory that makes legacy Edge AI tracking so efficient.

The Viola-Jones Framework

Haar Cascade classifiers are based on the Viola-Jones object detection framework. Instead of using massive, compute-heavy deep learning neural networks that require powerful discrete GPUs, Haar Cascades utilize simple, binary pixel-intensity features called Haar-like features.

These features act like digital templates looking for specific shifts in brightness:

  • Edge Features: Detects boundaries where a dark zone transitions into a light zone (like the bridge of your nose versus your cheek).

  • Line Features: Useful for identifying long, horizontal elements like eyebrows or the line of the mouth.

  • Center-Surround Features: Excellent for finding eyes, where the dark pupil is surrounded by lighter skin and sclera.

Why Slicing the Array Saves Your Processor

Every time you invoke .detectMultiScale(), OpenCV has to pass a sliding window across the image matrix at multiple scales, performing thousands of additions and subtractions per frame.

Mathematically, if an entire frame has a pixel area, scanning it scales linearly with that total area. By filtering for the face first and establishing a tight Region of Interest (ROI), you reduce the eye tracking search space down to a fractional area.

On resource-constrained hardware like an edge microcontroller or a single-board computer, isolating the matrix dimensions before calling nested lookups is the difference between a sluggish, unusable slideshow and a silky smooth tracking experience.

AI on the Edge LESSON 30: Tune Object Tracker with Mouse Selected ROI

AI on the Edge LESSON 30: Tune Object Tracker with Mouse Selected ROI

Welcome, Makers!

Well, hello there! It is absolutely fantastic to have you back. I’m Paul McWhorter, and today, we are taking a massive step forward in our AI on the Edge journey.

Up until now, we’ve been hard-coding our color thresholds (those pesky Lower Color and Upper Color values) to tell our camera what to look for. That’s fine for a science experiment, but it’s not exactly “smart,” is it? If the lighting changes, or if we want to track a different colored object, we have to go back into the code and manually edit those numbers.

Not anymore!

In today’s lesson, we are building a tool that lets us teach the AI. We’re going to use the mouse to draw a Region of Interest (ROI) right on our camera feed. The system will look at the pixels inside that box, calculate the average Hue, Saturation, and Value, and automatically set our tracking range for us.

This is the kind of professional-level functionality that turns a hobby project into a true, intelligent machine.

The Concept: From Hard-Coding to Dynamic Learning

The magic happens in our mouseAction function. Instead of just reading pixel values, we are now implementing a “click-and-drag” system:

  1. Click and Hold: We capture the startX and startY coordinates.

  2. Drag: We draw a rectangle in real-time so we can see exactly what area we are selecting.

  3. Release: We take that specific slice of the image, convert it to the HSV color space, and use the cv2.mean() function to find the average color properties.

  4. Auto-Tune: We set our LC (Lower Color) and UC (Upper Color) based on that average.

By doing this, the system learns what “object” we want to track on the fly. It’s elegant, it’s powerful, and it feels like real magic when you see those servos snap onto your target after a quick mouse drag.

What We’ve Accomplished

By the end of this lesson, you will have a system that:

  • Visually selects an object using the mouse.

  • Automatically calculates the optimal HSV thresholds for that specific object.

  • Updates the tracking behavior immediately without needing to stop or re-run the code.

  • Maintains that professional “Edge” feel, giving you real-time feedback on your FPS and mouse position data.

A Note on the “Edge”

Remember, we aren’t just running code; we are running on hardware. When we calculate the mean of the ROI, we are doing real image processing on the fly. You’ll notice the Composite and Mask windows updated immediately, giving you a visual confirmation that your “teacher” (you!) has successfully guided the “student” (the AI).

This is the power of working with OpenCV and the Raspberry Pi. You are building a system that observes, thinks, and reacts—all in real-time.

Get Ready to Build

Grab your Pi, make sure your servos are ready to go, and let’s get that camera calibrated. You’ve put in the work to get this far, and today is where all that effort starts to feel really rewarding.

I’m incredibly proud of how far you’ve come. Let’s dive in and start building!

Are you ready to see how accurately your Pi can “see” once you’ve given it the ability to learn from your selections?

In the video lesson we developed the following code.

 

AI on the Edge LESSON 28: Use Pan Tilt Camera to Track Object of Interest in OpenCV

Hey everyone, Paul McWhorter here from TopTechBoy.com. Welcome back to our channel, where we don’t just write abstract software—we build real, physical, intelligent machines. Go ahead and grab yourself a nice hot cup of coffee or a big glass of iced tea, because today we are closing the loop between the digital world of computer vision and the physical world of robotics.

In our last lesson, we successfully taught OpenCV how to find a specific color, isolate the largest shape, and draw a beautiful green bounding box around it. That was great, but it had a massive limitation: if your object moved off the edge of the frame, it was gone forever. The camera just sat there, blind and helpless.

Today, we change that. We are taking that tracking data from our code and using it to command a physical pan-tilt mechanism powered by two servos on our Fusion Hat. By the time we finish today, your camera will physically turn, tilt, and hunt down your object, keeping it locked dead center in the middle of your video feed.

The Big Leap: Closing the Loop

What we are building today is a foundational concept in automation engineering known as a feedback control loop.

Up until now, your camera was an open-loop observer. It saw things, but it couldn’t react physically. To make an autonomous tracking system, we need to implement a simple pipeline:

  1. Sense: The camera captures the image frame.

  2. Think: OpenCV finds the target object and calculates its position.

  3. Act: The script commands the hardware servos to move the camera mount to correct any positioning errors.

The Mathematics of the Target Error

To make a camera track an object, we have to define what “perfect tracking” looks like to a computer. Perfect tracking means the center of our tracked object is sitting exactly at the center of our video frame.

Because we are running our camera at a crisp resolution of 1280×720, the mathematical center of our universe is fixed. We calculate our frame’s horizontal and vertical centers by dividing our dimensions in half. This gives us a permanent anchor point right in the middle of our grid.

When an object appears on screen, our contour detection gives us its bounding box. We calculate the exact center of that box by taking its starting coordinate and adding half of its width and height. Now we have two sets of coordinates:

  • Where we want the object to be (The Frame Center).

  • Where the object actually is (The Box Center).

The difference between where the object is and where it belongs is called the Error Signal. We calculate an X Error and a Y Error by simply subtracting the frame center from the box center.

Managing Jitter with a Control Deadband

If our object is perfectly centered, our Error is zero. If the object moves to the right, the X Error becomes a positive number. If it moves to the left, it becomes a negative number. The same logic applies vertically to our Y Error.

Now, you might think we should tell the servos to move every single time the Error is anything other than zero. But remember what we learned about camera sensors: pixels dance, light fluctuates, and your calculations will always have a tiny amount of natural mathematical noise. If you try to correct for every single fractional pixel change, your servos will constantly buzz, twitch, and jitter themselves to death.

To fix this, we implement an engineering safety margin called a Deadband. In this lesson, we establish a 40-pixel safety zone around the center of the frame.

  • If the object is within 40 pixels of the center, the error is too small to care about, and we tell the servos to sit perfectly still.

  • The moment the object drifts outside that 40-pixel window, our control logic triggers.

If the X Error is greater than 40, we decrement our pan angle by one degree to turn the camera toward the target, pass that new angle to our servo handler, and pause for a tiny fraction of a second (20 milliseconds) to give the mechanical gears time to physically move. If it’s negative, we increment the angle. We apply the exact same behavioral logic to our tilt servo using the Y Error.

Visualizing the System Matrix

To help us calibrate and troubleshoot this system, we overlay clear visual indicators directly onto our live video feed:

  • The Reticle: We draw a solid blue dot directly at our fixed frame center. This acts as our tracking crosshair.

  • The Target: We draw a large red circle directly over the center of our moving object’s bounding box.

When your system is working properly, you can physically watch the machine think. As you move an object around, the red circle moves away from the blue dot, the error threshold trips, the servos kick in, and the camera moves until the red circle swallows the blue dot once again.

Your Homework Assignment

You guys know the drill: watching me build a tracking rig doesn’t make you an automation engineer. You have to write the logic, feel the hardware move, and tune it yourself.

Here is your homework challenge for Lesson 28: Right now, our tracking logic uses what is called an incremental step controller. No matter how far away the object is from the center, the camera always moves at the exact same speed—one lazy degree at a time. If you move your target slowly, the camera keeps up. If you snap your target quickly across the room, the camera falls behind and loses it because it can’t accelerate.

Your assignment is to upgrade this control loop. Instead of stepping by a hardcoded value of 1, I want you to make the servo adjustment step proportional to the size of the error. If the object is close to the center, it should move gently by a fraction of a degree. If the object takes off like a rocket and creates a massive error signal, the camera should aggressively throw the servos open to catch up instantly.

Get your proportional tracking loops tuned, shoot a video showing your camera tracking a fast-moving object smoothly, upload it to YouTube, and share your link down in the comments below. See you guys in the next lesson!

AI on the Edge LESSON 1: Introduction and Class Overview

Welcome to our all new AI on the Edge class! I will need you to buckle up, get your hardware together, and get ready to teach AI who is boss! We will be using a Pi 5, and the Fusion AI Lab kit. I will show links to the hardware below. In today’s lesson I describe the Class Introduction, and will show you some demos of the types of projects we will be doing. You will either Drive AI or your will be Destroyed by AI. Don’t be one of the ones who will be eaten by it

The Future will Belong to Those Who Can Drive AI

Guys, get your gear, and make sure you end up on the right side of the Dystopian future that awaits the world.

I have provided Amazon links, so you can order everything in the same place

You Will need a Raspberry Pi 5
Order Pi 5

You will need a heat sink and fan
Order Heat Sink and Fan

You Will Need the Fusion AI Lab Kit
Order Fusion AI Lab Kit

You Will Need a 25 Watt Power Supply
Order Power Supply

You Will Need a Micro HDMI Cable
Order Micro HDMI Cable

You Will Need a Keyboard and Mouse
Order Wireless Keyboard and Mouse

This isn’t just another Raspberry Pi class. This is a complete journey where we’re going to take the powerful Raspberry Pi 5, combine it with the SunFounder Fusion AI Lab kit, and build real, practical, intelligent systems that run completely on the edge — no cloud, no internet required.

In this class, you’re not going to just learn how to blink an LED or run someone else’s pre-made script. You’re going to learn how to build smart machines that can see, listen, speak, think, and act in the real world. We’re going to combine computer vision, voice recognition, speech synthesis, sensor reading, motor control, and modern AI techniques — all running locally on your Raspberry Pi 5.

Over the course of this series, you will learn how to:

  • Capture and process live video from the Raspberry Pi Camera
  • Detect faces and track objects in real time using MediaPipe and OpenCV
  • Control hardware with voice commands
  • Make your Raspberry Pi speak with natural-sounding Text-to-Speech
  • Build smooth, responsive control systems using threading
  • Use displays like the SSD1306 OLED to show live information
  • Combine everything into impressive AI-powered projects

This class is designed for makers, students, hobbyists, and engineers who want to move beyond basic tutorials and start building real intelligent edge devices. Whether you dream of building smart robots, autonomous monitoring systems, interactive AI companions, or just want to gain serious skills in modern embedded AI, this class is for you.

I’m going to teach this the way I always do — step by step, clearly, and with lots of hands-on projects. We’ll start with the fundamentals and gradually build up to more advanced and exciting projects as the class progresses.

If you’ve ever wanted to move from “playing with the Raspberry Pi” to “building truly intelligent systems,” then you’re in the right place. This is going to be a fun, challenging, and incredibly rewarding journey.

So if you’re ready to stop just watching AI videos and start building your own AI on the edge… then buckle up, because we’re about to do exactly that.

Welcome to the class! I’m really glad you’re here. Let’s get started!