Tag Archives: AI

AI on the Edge LESSON 42: Create Composite Images Using Masks in OpenCV and MediaPipe

In this exciting project, we combine a Raspberry Pi 5, the Fusion AI Lab Kit, a Pi Camera, and a remote IP camera to generate a stunning real-time composite video. Watch as a glowing, translucent MediaPipe face mesh of my face hovers magically over live video of the Mighty River Nice scenery captured by an IP camera. The effect looks futuristic and professional — perfect for creative video effects, interactive installations, or just blowing your mind with computer vision! Using Picamera2 for high-frame-rate local capture and OpenCV with an RTSP stream from the river camera, we process everything in real time. MediaPipe’s Face Mesh detects and tracks facial landmarks, which we draw as beautiful cyan/teal contours with glowing irises. Then we create a clean mask, separate the mesh foreground from the river background, and blend them seamlessly into one composite frame. You’ll see every debugging layer live on screen too — meshLayer, mask, inverted mask, riverBG, and meshFG — so you can understand exactly how the magic happens.This tutorial is beginner-to-intermediate friendly and packed with practical OpenCV + MediaPipe techniques you can adapt for your own augmented reality projects. Whether you’re a longtime follower of the Paul McWhorter channel or new to the Fusion AI Kit, you’ll walk away inspired and ready to build your own hovering effects, overlays, or interactive displays.Grab the full code from the video description, fire up your Pi 5, and start creating jaw-dropping computer vision projects today. Drop a comment and let me know what you’d like to overlay next — another face mesh, hand tracking, or something completely different? Let’s keep pushing the limits of what we can do with affordable AI hardware!

 

AI on the Edge 38: Using MediaPipe for Face Recognition on the Raspberry Pi 5

In this video lesson we introduce you to MediaPipe. The OS we had you flash in LESSON 1 already has the MediaPipe framework installed, and all the needed and working dependencies. If you have installed that OS and not modified it, this and future lessons will work. If you find dependency errors, you might need to reflash the original OS.

MediaPipe is a free, open-source framework developed by Google that makes it much easier to add advanced computer vision and AI features to your Python programs. It is especially popular among developers who use OpenCV because it works seamlessly with it and delivers excellent real-time performance, even on devices like the Raspberry Pi 5.

With MediaPipe, you can quickly add powerful capabilities such as face detection, face mesh (detailed facial landmarks), hand tracking, body pose estimation, and more — all without having to write complex deep learning code from scratch. It comes with pre-trained machine learning models that are optimized for speed, allowing your programs to run smoothly at 30 frames per second or higher.

The biggest advantage for Python + OpenCV users is its simplicity. You capture video frames using OpenCV or picamera2, pass them to MediaPipe for processing, and then draw the results (such as bounding boxes or landmarks) back onto your image using normal OpenCV functions. This combination gives developers a fast and straightforward way to build interactive computer vision projects like face trackers, gesture-controlled robots, or smart camera applications.

In short, MediaPipe acts as a powerful, easy-to-use toolkit that bridges the gap between OpenCV and modern AI vision technology.

We introduce you to MediaPipe using a simple example where we create a faceFinder on our Raspberry Pi 5.

 

AI on the Edge LESSON 37: Using RTSP and IP Cameras in OpenCV on Raspberry Pi 5

The code below shows the work we did in this lesson.

AI on the Edge Lesson 37: Using RTSP and IP Cameras in OpenCV on Raspberry Pi 5

Hey guys! Welcome back to our AI on the Edge series. In our previous lessons, we’ve had a blast working with standard USB webcams, but if you are building a real-world computer vision application, an automation rig, or a security monitoring setup around your home or farm, USB cables just aren’t going to cut it. You need to pull video feeds from remote IP cameras using the Real-Time Streaming Protocol (RTSP).

Today, we are taking that exact step on the Raspberry Pi 5, connecting to an IP camera, streaming the feed smoothly into OpenCV, and—most importantly—solving the dreaded latency problem that plagues RTSP feeds.

The Big Challenge: Conquering RTSP Latency

If you’ve ever tried pulling an RTSP stream into OpenCV straight out of the box, you’ve probably noticed something frustrating: the video lags behind real-time, sometimes by several seconds or even tens of seconds.

Why does that happen? Because by default, FFmpeg and OpenCV buffer incoming frames to ensure smooth playback. But when you are doing computer vision, AI inferencing, or real-time tracking on the edge, you don’t want old history—you want right now.

To fix that, we pass the cv2.CAP_FFMPEG backend flag and immediately flush the buffer by setting the property to 0. This forces OpenCV to drop the backlog and grab the absolute newest frame available from the camera stream, keeping your Pi 5 processing live data in real-time.

Understanding the Script Structure

Let’s break down the key parts of today’s implementation:

  • Credentials & Resolution: We import a separate secret file to keep our camera IP addresses, usernames, and passwords safe and out of public repositories. We lock our resolution at 1280×720 to balance crisp detail with the Pi 5’s processing overhead.
  • Smooth FPS Calculation: Instead of a jittery raw frame-rate readout, we use an exponential moving average to give us a stable, readable performance metric on screen.
  • The Display Window: We configure a GUI window using OpenCV’s window flags so we can easily position and resize our output feed on the desktop.

Drop Your Questions Below

Working with network streams can sometimes be tricky depending on your specific camera’s firmware, codec settings, and network stability. If you run into any connection drops or lag spikes on your Raspberry Pi 5, drop a comment on the video!

Keep building, stay creative, and I will see you guys in Lesson 38!

Here is the code developed in the video

 

Local Voice Control of NVIDIA Jetson Orin Nano with STT: Getting Started with Vosk

Engineering Your Own Local Voice Assistant: No Cloud, No Compromise

Most “smart” voice assistants are just glorified remote controls for someone else’s server. Today, we’re changing that. We are going to build a local, offline voice command pipeline. This isn’t just about saving data; it’s about ownership. When you can control your hardware—like opening and closing a farm gate—without an internet connection, you have built a system that is robust, private, and yours to control forever. Today you are going to get Speech to Text up and running on your NVIDIA Jetson Orin Nano running under Jetpack 7.2.

The “Why” Behind the Setup

You might ask, “Why not just use a cloud API?” Because cloud APIs are fragile. They rely on internet stability, external servers, and privacy-invasive data logging. By running Vosk locally, we keep the processing on your hardware (like the NVIDIA Jetson). It’s faster, it works in the middle of a power-isolated homestead, and it’s 100% secure.

Part 1: Preparing the Environment

Before we can make the machine listen, we have to prepare the battlefield. We aren’t just downloading files; we are setting up a stable environment where your dependencies won’t conflict with your OS.

This is IMPORTANT!

Now you can post the code below. You also have to point Thonny to run in the virtual environment. Open Thonny, and under run –  select interpreter. Then you must point it to /home/yourUserName/STT/ttsVenv/bin/python3. For me, my username is pjm, but you put in your user name in path above. Here is what mine looked like:

Part 2: Solving the PipeWire Challenge

The biggest headache in modern Linux audio is PipeWire. If you try to open a microphone stream using a hardcoded sample rate that doesn’t match your hardware, your program won’t just fail—it will segfault. We use the validation script below to programmatically query the hardware, asking it: “What sample rate are you running at?” before we even try to open a stream.

Homework: Your Gate Controller

You now have a system that identifies audio input, resamples it to 16kHz, and outputs text. Your assignment: Transform this text output into an action.

I want you to add a conditional statement to the main loop. If the recognized text is “open”, print an ASCII art representation of an open gate. If it’s “close”, print the closed version. This is the first step in closing the loop between your AI and the physical world. Go get ’em, and don’t just copy the code—understand how the data flows from the microphone to your decision logic!

AI on the Edge LESSON 34 SUPPLEMENT: Simple Improvement to FPS

In this video I show you how to dramatically improve the FPS of our work in lesson 34. I also show the solution to the issue of the Fusion Hat microphone not working with the project.