Tag Archives: Fusion HAT+

AI on the Edge LESSON 48: Hand Detection in OpenCV and MediaPipe on the Raspberry Pi

In Lesson 48 we’re taking a huge step forward — we’re bringing real-time hand detection to the Raspberry Pi using the incredible power of MediaPipe combined with OpenCV and the Pi Camera. This is one of those projects that feels like pure magic when you see it working: the camera picks up your hands instantly, tracks all 21 landmarks on each hand, draws beautiful connections in real time, and does it all right on the edge with no cloud required.
You’ll learn how to set up MediaPipe’s Hands solution for reliable multi-hand tracking, how to efficiently process frames from the Picamera2 library, flip and convert images for proper display, and draw professional-looking landmarks and connections with custom styling. We also keep a smooth FPS counter running so you can see exactly how well your Pi is performing.
This lesson is exciting because hand tracking is the foundation for so many advanced gesture-control projects — think touchless interfaces, sign language recognition, robotic control, virtual instruments, and interactive art installations. Once you have solid hand detection running smoothly on your Raspberry Pi, the creative possibilities are almost endless.
So grab your Raspberry Pi, fire up the camera, and let’s get those hands dancing on the screen! As always, the full code is waiting for you below. Watch the video, follow along, and then start experimenting — I can’t wait to see what you build with this!
Let’s make something awesome!

 

AI on the Edge LESSON 46: Ultimate Dazzleing Running Rainbow On a NeoPixel Ring

In this lesson we create one of the most beautiful and satisfying NeoPixel effects — a smooth, continuous, running rainbow on a 12-LED ring. We affectionately have named this pattern the “Runbow”. This is the “ultimate” version because it is both visually stunning and highly educational, as we explore two fundamentally different programming approaches the Runbow.
The first method is the most intuitive: we place a fixed rainbow across the ring (each LED gets a different hue equally spaced around the color wheel) and then simply rotate the entire pattern one position at a time. This approach is easy to understand because it mimics physically moving a colorful wheel. We save the color of the first pixel, shift all the other pixels one place, and then place the saved color at the end. It feels like a conveyor belt of color circling around the ring.
The second method is more elegant and mathematically pure. Instead of storing and shifting colors, we recalculate the color of every LED on every frame using a moving offset. For each LED we compute its hue as (i / LED_COUNT + offset) % 1.0, where i is the LED’s position and offset is a value that slowly increases over time. This creates a perfectly smooth rainbow that flows around the ring without ever shifting raw RGB values. Because we regenerate the pattern fresh each cycle, there is no risk of color corruption or accumulated errors.
Both techniques produce a gorgeous running rainbow, but they teach very different programming mindsets. The first method helps you deeply understand array manipulation and data movement. The second method introduces the powerful concept of using mathematics and offsets to create motion — a technique used frequently in advanced LED animations, games, and visual effects. In the end, the offset method tends to look smoother and is easier to extend with additional effects (such as brightness pulsing), but both approaches are valuable skills for any embedded AI or IoT developer working with addressable LEDs.You can adjust the speed of the rainbow by changing how much you increment the offset each loop or by modifying the delay. Once you master these two methods, you will be able to create almost any animated pattern you can imagine on your NeoPixel ring. Lets get this party started!
Here is the code we developed in this video:

This is the schematic we are using to connect the NeoPixel ring:

NeoPixel
NeoPixel Schematic

AI on the Edge LESSON 45: Adding a NeoPixel Ring To Your Raspberry Pi Project

In Lesson 45 of our AI on the Edge series, we take our Fusion AI Lab Kit to the next level by adding a 12-pixel NeoPixel ring. This lesson bridges the gap between pure AI processing and vibrant physical output, showing how your edge AI projects can communicate visually with the real world in a beautiful and engaging way.
We begin by setting up the NeoPixel ring using the SPI interface on the SunFounder Fusion Hat. After importing the necessary libraries (time, board, neopixel_spi, colorsys, and math), we create a simple and reusable hsv2rgb() function that converts Hue values (0.0 to 1.0) into RGB colors that the NeoPixels can understand. This function becomes the foundation for smooth rainbow effects later in the lesson.The lesson starts with basic pixel control. We manually light up each of the 12 pixels one at a time using different colors (red, green, blue, cyan, magenta, yellow, etc.). This slow, deliberate approach lets you clearly see how individual pixel addressing works and helps students understand the coordinate system of the ring.Next, we explore full-ring control by making the entire ring blink between bright red and blue. We then move into motion with a running green pixel moving across a red background — a great introduction to animation techniques. This is followed by a more advanced chasing effect where a blue pixel chases a green pixel around a dim red background.One of the highlights of this lesson is the gentle pulsating aqua effect. Using a sine wave (math.sin), we create a smooth breathing/pulsing animation where the brightness of the color rises and falls naturally. This technique produces a very professional and visually pleasing result that students can easily adapt for future projects.
Finally, we create a rainbow effect. First, we display a uniform rainbow where all 12 pixels show the same color that cycles smoothly through the entire spectrum.
We finish the lesson by assigning the homework. The homework is for you to create the classic running rainbow (what I like to call a “Runbow”), where the colors flow continuously around the ring — one of the most popular and impressive NeoPixel animations.
This lesson reinforces important programming concepts including loops, functions, color theory (HSV vs RGB), timing control, and animation techniques, while giving students an exciting visual payoff. The skills learned here open the door to creating stunning visual feedback for future AI projects — whether it’s status indicators, emotional displays, or attention-grabbing outputs from your edge AI models.

In this class we are using this as our standard components. You should already have the core circuit built and should already have the OLED connected. Today you will add the NeoPixel array.

Our core circuit is:

Fusion Hat Circuit Diagram
This is the circuit we will use moving forward in the class

Last week we also added the SSD1306 OLED Display.

OLED
SSD1306 OLED Connected to the Fusion AI Hat

And finally today we add the NeoPixel ring from the Fusion AI Lab Kit.

NeoPixel
NeoPixel Schematic

AI on the Edge LESSON 40: Active Face Tracker with Pan Tilt Camera and MediaPipe on Pi 5

Boys and girls, welcome back! In today’s lesson, we are going to tie together everything we’ve been building in the AI on the Edge series and construct something truly interactive: a fully autonomous, voice-controlled, pan-tilt face tracking robot running locally right on your Raspberry Pi 5!

In our previous lessons, we learned how to detect faces using MediaPipe and how to drive physical servos to point a camera. Today, we step up our game. We are bringing in multithreading, Speech-to-Text (STT) using the Fusion Hat, and Text-to-Speech (TTS) with Piper to give our Pi a voice, a personality, and the physical ability to track down humanoids in real time.

What We Are Building in This Lesson

Imagine setting up a camera system that constantly scans its environment. The moment a human face enters the frame, the system locks on and speaks up: “Humanoid Detected, Shall I track?”

Using real-time voice commands, you can issue directions straight to the Pi without touching a keyboard:

  • “Track” — Activates proportional control on the pan-tilt kit. The servos will calculate pixel error relative to the center of the frame and smoothly adjust their angles to keep your face dead center.

  • “Release” — Disables active tracking, letting the servos hold their position while the vision loop continues monitoring.

  • “Blind” — Isolates the facial keypoints for the subject’s eyes and draws solid black circles over them in real time, causing the robot to announce: “Subject Has Been Blinded, Shall I Vaporize?”

  • “Restore” — Removes the eye overlay and brings vision back to normal.

  • “Quit” — Safely terminates all background threads, announces shutdown, and closes down the application gracefully.

Key Technical Concepts Covered

1. Multi-Threaded Architecture & Thread-Safe Queues

Audio processing—both listening for voice input and generating spoken speech—is computationally heavy and blocking by nature. If you run speech recognition directly inside your primary video processing loop, your frame rate will plummet from a smooth 60 FPS down to a complete crawl.

To solve this, we spin up two independent background threads using Python’s threading module:

  • Speech Thread: Monitors a thread-safe speakQ (Queue) and handles text-to-speech output using Piper without stalling the main loop.

  • Command Thread: Continuously listens to the microphone via Speech-to-Text, strips and parses incoming voice triggers, and pushes valid commands into a commandQ.

2. MediaPipe Facial Landmark Detection

We leverage MediaPipe’s high-speed face detection solution running at 1280×720 resolution on the Raspberry Pi 5. By calculating relative bounding boxes and keypoint coordinate matrices (x, y), the system identifies both face centroids and precise feature locations like eye coordinates.

3. Proportional Servo Error Correction

To keep the camera centered on a moving subject, the script computes positional error delta values between the center of the bounding box and the exact midpoint of the camera frame:

xError = xBoxCenter – xFrameCenter

yError = yBoxCenter – yFrameCenter

These error values are scaled down and applied directly to update the current pan and tilt servo angles, ensuring smooth, continuous tracking movement without jarring overshoots.

Your Homework Assignment

Get your Raspberry Pi 5, mount your pan-tilt camera assembly with the Fusion Hat, and implement the multithreaded architecture outlined in this lesson. Tune your servo scaling factors to ensure your tracking motion is fluid and responsive at 60 FPS. Have fun!


 

AI on the Edge LESSON 36: Select Active Camera in OpenCV With Voice Commands

Welcome back, makers, engineers, and AI enthusiasts! In our previous lessons, we built out robust multi-camera setups, streaming feeds from both Raspberry Pi cameras and high-definition USB cameras. But watching four feeds in separate tiles is only half the battle. What happens when you want to interact with your system hands-free?

In this lesson, we are taking our edge vision projects to the next level by integrating voice commands. You will learn how to run a dedicated speech-to-text thread in the background, catch voice triggers safely using a thread-safe queue, and dynamically switch your active main camera feed in OpenCV on the fly—just by speaking.

What You Will Learn in This Lesson

  • Multi-Threaded Speech Recognition: How to run the stt listener in a separate daemon thread so it never blocks or stutters your high-FPS video processing loop.

  • Thread-Safe Communication: Using Python’s Queue module to pass voice commands seamlessly from the background listening thread into your main application loop.

  • Handling STT Variations: Accounting for common speech-to-text homophone variations (like “two”, “to”, and “too”, or “four” and “for”) to make your voice control robust and reliable.

  • Dynamic Frame Routing: Mapping spoken commands to specific camera streams (piCam1, piCam2, usbCam1, and usbCam2) and updating your primary display window instantly.

  • Organized Window Layouts: Positioning and sizing multiple OpenCV windows on your desktop workspace for a clean, professional multi-camera dashboard.

Step-by-Step Breakdown of the Script

1. Setting Up the Background Voice Thread

When working with real-time video processing in OpenCV, blocking functions are your worst enemy. If you call a speech recognition listening function directly inside your main while loop, your video frames will freeze while waiting for audio input.

To prevent this, we initialize a background worker function (getCamera) and launch it as a daemon thread:

  • The thread continuously listens for your voice commands using the Fusion Hat STT module.

  • Once a command is captured, it strips whitespace, checks for exit triggers (like saying 'quit'), and pushes valid commands directly into our commandQ.

2. Initializing Multiple Cameras

Our setup harnesses the full power of the hardware by combining native Raspberry Pi camera interfaces with standard USB webcams:

  • Pi Cameras: Configured via Picamera2 at a crisp 1280x720 resolution with RGB888 formatting running smoothly at 60 frames per second.

  • USB Cameras: Initialized through OpenCV’s VideoCapture class, set to target resolution and optimized for 30 FPS.

3. Processing and Routing Commands in the Main Loop

Inside the primary application loop, our script handles three core tasks simultaneously:

  1. Calculate Performance: Continuously tracks and smooths out the frames-per-second (FPS) metric so you can monitor system load.

  2. Check the Command Queue: Non-blockingly checks if the background thread has dropped a new camera selection into commandQ. If a new command is waiting, it updates the mainCam variable.

  3. Route the Main Frame: Evaluates the active camera string—including clever fallback checks for common voice misinterpretations like 'camera to' or 'camera for'—and assigns the corresponding video stream to mainFrame.

4. Managing Your OpenCV Windows

To give you a complete command-center experience, the script generates a large primary display window for your active camera view, accompanied by a clean row of smaller preview tiles across the bottom of your screen for all four connected feeds.