AI camera system
A full-stack system that watches live camera feeds, spots people, vehicles, and animals, recognizes familiar faces, and raises alerts, with a desktop app, a web dashboard, and a phone app.
- LanguagePython 3.11
- DetectionYOLOv8 nano
- APIFastAPI, 19 endpoints
- Phone appFlutter
- Runs onA laptop CPU, no GPU
Follow a frame
Every frame from every camera takes the same path. Most frames stop early: if nothing moved, there's no reason to run the expensive AI. Select a stage to see what it does.
Camera
This is the real loop from the camera worker, lightly trimmed. The motion gate decides whether the AI runs at all. Face recognition runs on every frame until a person is identified, then only every fifth frame.
src/camera/manager.py
has_motion = (
not self._app_cfg.motion.enabled
or self._motion.detect(frame)
or frames_since_detection >= max_skip
)
if has_motion:
detections = self._detector.detect(frame)
tracks = self._tracker.update(detections)
# Eager mode: run every frame until a face is confirmed,
# then throttle to every N frames for updates.
for track in tracks:
if track.class_id != 0:
continue
unconfirmed = track.face_id is None
due = cnt % interval == 0
if unconfirmed or due:
self._run_face_recognition(frame, track)
event_ids = self._persist_tracks(frame, tracks)
By the numbers
The per-frame speed is from the notes in Leo's detector code, measured on an Intel i5-1135G7.
The API
Everything the apps can do goes through a FastAPI server on port 8000. FastAPI also generates interactive documentation at /docs, so every endpoint can be tried in a browser.
| Method | Path | What it does |
|---|---|---|
| GET | /api/cameras | List cameras |
| GET | /api/cameras/{id}/status | FPS, frame count, active tracks, errors |
| GET | /api/cameras/{id}/stream | Live MJPEG video |
| GET | /api/events | Search event history |
| GET | /api/events/{id} | One event |
| GET | /api/events/stream | Live event feed (server-sent events) |
| GET | /api/alerts | List alerts |
| POST | /api/alerts/{id}/acknowledge | Mark an alert as seen |
| GET | /api/stats | Totals for the dashboard |
| GET | /api/faces | List enrolled people |
| GET | /api/faces/{id} | One person |
| POST | /api/faces | Enroll a new person from a photo |
| DELETE | /api/faces/{id} | Remove a person |
| GET | /api/faces/{id}/image | Enrollment photo |
| GET | /api/export/events.csv | Download events as CSV |
| GET | /api/export/alerts.csv | Download alerts as CSV |
| GET | /api/system/status | Is the AI running? |
| POST | /api/system/start | Start AI detection |
| POST | /api/system/stop | Stop AI detection |
Design decisions
Speed
Skip frames where nothing moves
Running YOLO on an empty hallway 15 times a second wastes the CPU. A cheap motion check runs first, and only frames with movement reach the AI. It's tuned to err toward running YOLO, because missing a real event is worse than wasting one inference.
Accuracy
Look for faces inside people
Face recognition only searches inside a detected person's bounding box, scaled down to about 300 pixels. That's faster, and it avoids false matches on posters or screens in the background.
Hardware
No GPU required
Everything is chosen to run on a normal laptop: the nano YOLO model, the CPU-friendly HOG face detector, and CPU-only PyTorch.
Control
AI starts off
Cameras connect at launch, but detection stays off until someone presses Start in the desktop app or the phone app calls POST /api/system/start.
Build your own
This guide builds the same pipeline one piece at a time, using the same libraries. Each step works on its own, so you'll have something running early and can stop wherever you like. You need a Linux or macOS computer, a webcam, and some Python experience. Your progress is saved in this browser.
Only point cameras where you have permission, and tell people they're on camera. Laws about recording and especially facial recognition differ by state and country. Only enroll faces of people who have agreed to it.
-
1
Set up the environment
Face recognition depends on
dlib, which is slow and fussy to compile. Conda has a pre-built version, so start there. PyTorch (used by YOLO) should be installed as a matched CPU-only pair.environment.yml
name: surveillance channels: - conda-forge - pytorch dependencies: - python=3.11 - dlib>=19.24 # pre-compiled binary from conda-forge - pytorch::pytorch=2.6.0=*cpu* - pytorch::torchvision=0.21.0=*cpu* - pip - pip: - opencv-python-headless>=4.8 - ultralytics - supervision>=0.19 - face-recognition>=1.3 - fastapi>=0.109 - uvicorn[standard]>=0.27 - sqlalchemy>=2.0conda env create -f environment.yml conda activate surveillance
-
2
Read frames from a camera
OpenCV reads webcams, IP cameras, and video files through the same call. The source is
0for the first webcam, anrtsp://URL for an IP camera, or a file path. Start here and make sure you can see your feed.import cv2 cap = cv2.VideoCapture(0) # webcam, "rtsp://...", or "video.mp4" while True: ok, frame = cap.read() if not ok: break cv2.imshow("camera", frame) if cv2.waitKey(1) == ord("q"): break cap.release()For this window version, install the regular
opencv-pythonpackage. The headless version in the environment has no windows, which is what you want later when the API serves the video. -
3
Add a motion gate
Background subtraction learns what the empty scene looks like and highlights anything new. Leo's version shrinks and blurs the frame first so it's fast, cleans up the mask, and adds up the moving area.
src/detection/motion.py
bg = cv2.createBackgroundSubtractorMOG2(history=500, varThreshold=16, detectShadows=False) kernel = cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (5, 5)) def has_motion(frame, min_area=800): small = cv2.resize(frame, (320, 240)) gray = cv2.cvtColor(cv2.GaussianBlur(small, (21, 21), 0), cv2.COLOR_BGR2GRAY) mask = bg.apply(gray) mask = cv2.morphologyEx(mask, cv2.MORPH_OPEN, kernel) # remove noise mask = cv2.morphologyEx(mask, cv2.MORPH_CLOSE, kernel) # fill holes contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE) return sum(cv2.contourArea(c) for c in contours) >= min_area -
4
Detect objects with YOLOv8
The
ultralyticspackage downloads the model the first time you use it. Filtering to only the classes you care about makes it faster and cuts noise. The numbers are COCO class IDs: 0 is a person, 2 is a car, 16 is a dog.src/detection/detector.py
from ultralytics import YOLO model = YOLO("yolov8n.pt") # nano: fastest, fine on a CPU results = model( frame, imgsz=640, conf=0.62, classes=[0, 1, 2, 3, 5, 7, 14, 15, 16], device="cpu", verbose=False, ) for box in results[0].boxes: x1, y1, x2, y2 = box.xyxy[0].tolist() label = model.names[int(box.cls)] print(label, float(box.conf), (x1, y1, x2, y2)) -
5
Track objects across frames
YOLO looks at each frame separately, so it can't tell you that the person in this frame is the same one from the last frame. A tracker can. Leo's system uses ByteTrack from the
supervisionlibrary:import supervision as sv tracker = sv.ByteTrack() detections = sv.Detections.from_ultralytics(results[0]) detections = tracker.update_with_detections(detections) for tracker_id, class_id in zip(detections.tracker_id, detections.class_id): print(f"track {tracker_id}: {model.names[class_id]}")Now "a person appeared" becomes one event per person, not one event per frame.
-
6
Recognize known faces
The
face_recognitionlibrary turns a face into a list of 128 numbers, called an encoding. Two photos of the same person produce encodings that are close together. Enroll someone once from a clear photo, then compare every new face against your list.import face_recognition as fr # Enroll once, from a clear photo of someone who agreed to it known = [fr.face_encodings(fr.load_image_file("alex.jpg"))[0]] names = ["Alex"] # Then, for each person crop from YOLO (converted from BGR to RGB): rgb = cv2.cvtColor(person_crop, cv2.COLOR_BGR2RGB) for encoding in fr.face_encodings(rgb, fr.face_locations(rgb, model="hog")): distances = fr.face_distance(known, encoding) best = distances.argmin() name = names[best] if distances[best] <= 0.55 else "unknown"Lower tolerance means stricter matching. Leo uses 0.55, and the usual range is 0.4 to 0.65.
-
7
Save events to a database
SQLite needs no server; it's just a file. SQLAlchemy lets you describe tables as Python classes. This is a trimmed version of Leo's event table:
src/database/models.py
class Event(Base): __tablename__ = "events" id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True) timestamp: Mapped[datetime.datetime] = mapped_column(DateTime, index=True) event_type: Mapped[str] = mapped_column(String(64), index=True) label: Mapped[str | None] = mapped_column(String(64)) # "person", "car" confidence: Mapped[float | None] = mapped_column(Float) track_id: Mapped[int | None] = mapped_column(Integer, index=True) face_name: Mapped[str | None] = mapped_column(String(128)) screenshot_path: Mapped[str | None] = mapped_column(String(512)) is_alert: Mapped[bool] = mapped_column(Boolean, default=False) -
8
Add alert rules
Keep the rules in a config file so they can change without touching code. The cooldown matters most: without it, one visitor standing still produces an alert every frame.
config.yaml
alerts: unknown_person: enabled: true min_confidence: 0.55 cooldown_seconds: 60 # don't repeat for the same track after_hours: enabled: false start: "22:00" end: "06:00" restricted_zones: enabled: false zones: ["server_room"] -
9
Serve an API and live video
MJPEG is the simplest way to stream video to a browser: the server sends JPEG after JPEG, and an ordinary
<img>tag plays them. This is Leo's stream endpoint:src/api/routes/cameras.py
@router.get("/{camera_id}/stream") async def mjpeg_stream(camera_id: int): async def frame_generator(): while True: jpeg = _manager.get_frame(camera_id) if jpeg: yield ( b"--frame\r\n" b"Content-Type: image/jpeg\r\n\r\n" + jpeg + b"\r\n" ) await asyncio.sleep(0.033) # ~30 fps ceiling return StreamingResponse( frame_generator(), media_type="multipart/x-mixed-replace; boundary=frame", )uvicorn app:app --host 0.0.0.0 --port 8000
Then open
http://localhost:8000/docsto try every endpoint. -
10
Build the dashboard
With the API in place, any front end works. The quickest is one HTML page:
<img src="http://localhost:8000/api/cameras/1/stream" width="640">
Leo went further with a Flutter phone app that connects over Wi-Fi, with screens for cameras, events, alerts, and enrolled faces. Its live view is the same MJPEG stream. Flutter builds for Android, iOS, and desktop from one Dart codebase.
Keep the API on your home network. Exposing a camera feed to the internet without proper authentication lets anyone watch it.