LT
Back to home

AI camera system

A full-stack system that watches live camera feeds, spots people, vehicles, and animals, recognizes familiar faces, and raises alerts, with a desktop app, a web dashboard, and a phone app.

Follow a frame

Every frame from every camera takes the same path. Most frames stop early: if nothing moved, there's no reason to run the expensive AI. Select a stage to see what it does.

Camera

This is the real loop from the camera worker, lightly trimmed. The motion gate decides whether the AI runs at all. Face recognition runs on every frame until a person is identified, then only every fifth frame.

src/camera/manager.py

has_motion = (
    not self._app_cfg.motion.enabled
    or self._motion.detect(frame)
    or frames_since_detection >= max_skip
)
if has_motion:
    detections = self._detector.detect(frame)
    tracks = self._tracker.update(detections)

    # Eager mode: run every frame until a face is confirmed,
    # then throttle to every N frames for updates.
    for track in tracks:
        if track.class_id != 0:
            continue
        unconfirmed = track.face_id is None
        due = cnt % interval == 0
        if unconfirmed or due:
            self._run_face_recognition(frame, track)

    event_ids = self._persist_tracks(frame, tracks)

By the numbers

~3,900lines of Python
~1,800lines of Dart in the Flutter app
19REST API endpoints
9object classes it watches for
~3 msper frame for YOLOv8n on the laptop's CPU
0.62minimum confidence to count a detection
0.55face-match tolerance (lower is stricter)
60 scooldown before re-alerting on the same person

The per-frame speed is from the notes in Leo's detector code, measured on an Intel i5-1135G7.

The API

Everything the apps can do goes through a FastAPI server on port 8000. FastAPI also generates interactive documentation at /docs, so every endpoint can be tried in a browser.

MethodPathWhat it does
GET/api/camerasList cameras
GET/api/cameras/{id}/statusFPS, frame count, active tracks, errors
GET/api/cameras/{id}/streamLive MJPEG video
GET/api/eventsSearch event history
GET/api/events/{id}One event
GET/api/events/streamLive event feed (server-sent events)
GET/api/alertsList alerts
POST/api/alerts/{id}/acknowledgeMark an alert as seen
GET/api/statsTotals for the dashboard
GET/api/facesList enrolled people
GET/api/faces/{id}One person
POST/api/facesEnroll a new person from a photo
DELETE/api/faces/{id}Remove a person
GET/api/faces/{id}/imageEnrollment photo
GET/api/export/events.csvDownload events as CSV
GET/api/export/alerts.csvDownload alerts as CSV
GET/api/system/statusIs the AI running?
POST/api/system/startStart AI detection
POST/api/system/stopStop AI detection

Design decisions

Speed

Skip frames where nothing moves

Running YOLO on an empty hallway 15 times a second wastes the CPU. A cheap motion check runs first, and only frames with movement reach the AI. It's tuned to err toward running YOLO, because missing a real event is worse than wasting one inference.

Accuracy

Look for faces inside people

Face recognition only searches inside a detected person's bounding box, scaled down to about 300 pixels. That's faster, and it avoids false matches on posters or screens in the background.

Hardware

No GPU required

Everything is chosen to run on a normal laptop: the nano YOLO model, the CPU-friendly HOG face detector, and CPU-only PyTorch.

Control

AI starts off

Cameras connect at launch, but detection stays off until someone presses Start in the desktop app or the phone app calls POST /api/system/start.

Build your own

This guide builds the same pipeline one piece at a time, using the same libraries. Each step works on its own, so you'll have something running early and can stop wherever you like. You need a Linux or macOS computer, a webcam, and some Python experience. Your progress is saved in this browser.

Only point cameras where you have permission, and tell people they're on camera. Laws about recording and especially facial recognition differ by state and country. Only enroll faces of people who have agreed to it.

  1. 1

    Set up the environment

    Face recognition depends on dlib, which is slow and fussy to compile. Conda has a pre-built version, so start there. PyTorch (used by YOLO) should be installed as a matched CPU-only pair.

    environment.yml

    name: surveillance
    channels:
      - conda-forge
      - pytorch
    dependencies:
      - python=3.11
      - dlib>=19.24              # pre-compiled binary from conda-forge
      - pytorch::pytorch=2.6.0=*cpu*
      - pytorch::torchvision=0.21.0=*cpu*
      - pip
      - pip:
        - opencv-python-headless>=4.8
        - ultralytics
        - supervision>=0.19
        - face-recognition>=1.3
        - fastapi>=0.109
        - uvicorn[standard]>=0.27
        - sqlalchemy>=2.0
    conda env create -f environment.yml
    conda activate surveillance
  2. 2

    Read frames from a camera

    OpenCV reads webcams, IP cameras, and video files through the same call. The source is 0 for the first webcam, an rtsp:// URL for an IP camera, or a file path. Start here and make sure you can see your feed.

    import cv2
    
    cap = cv2.VideoCapture(0)  # webcam, "rtsp://...", or "video.mp4"
    while True:
        ok, frame = cap.read()
        if not ok:
            break
        cv2.imshow("camera", frame)
        if cv2.waitKey(1) == ord("q"):
            break
    cap.release()

    For this window version, install the regular opencv-python package. The headless version in the environment has no windows, which is what you want later when the API serves the video.

  3. 3

    Add a motion gate

    Background subtraction learns what the empty scene looks like and highlights anything new. Leo's version shrinks and blurs the frame first so it's fast, cleans up the mask, and adds up the moving area.

    src/detection/motion.py

    bg = cv2.createBackgroundSubtractorMOG2(history=500, varThreshold=16, detectShadows=False)
    kernel = cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (5, 5))
    
    def has_motion(frame, min_area=800):
        small = cv2.resize(frame, (320, 240))
        gray = cv2.cvtColor(cv2.GaussianBlur(small, (21, 21), 0), cv2.COLOR_BGR2GRAY)
        mask = bg.apply(gray)
        mask = cv2.morphologyEx(mask, cv2.MORPH_OPEN, kernel)   # remove noise
        mask = cv2.morphologyEx(mask, cv2.MORPH_CLOSE, kernel)  # fill holes
        contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
        return sum(cv2.contourArea(c) for c in contours) >= min_area
  4. 4

    Detect objects with YOLOv8

    The ultralytics package downloads the model the first time you use it. Filtering to only the classes you care about makes it faster and cuts noise. The numbers are COCO class IDs: 0 is a person, 2 is a car, 16 is a dog.

    src/detection/detector.py

    from ultralytics import YOLO
    
    model = YOLO("yolov8n.pt")  # nano: fastest, fine on a CPU
    
    results = model(
        frame,
        imgsz=640,
        conf=0.62,
        classes=[0, 1, 2, 3, 5, 7, 14, 15, 16],
        device="cpu",
        verbose=False,
    )
    for box in results[0].boxes:
        x1, y1, x2, y2 = box.xyxy[0].tolist()
        label = model.names[int(box.cls)]
        print(label, float(box.conf), (x1, y1, x2, y2))
  5. 5

    Track objects across frames

    YOLO looks at each frame separately, so it can't tell you that the person in this frame is the same one from the last frame. A tracker can. Leo's system uses ByteTrack from the supervision library:

    import supervision as sv
    
    tracker = sv.ByteTrack()
    
    detections = sv.Detections.from_ultralytics(results[0])
    detections = tracker.update_with_detections(detections)
    for tracker_id, class_id in zip(detections.tracker_id, detections.class_id):
        print(f"track {tracker_id}: {model.names[class_id]}")

    Now "a person appeared" becomes one event per person, not one event per frame.

  6. 6

    Recognize known faces

    The face_recognition library turns a face into a list of 128 numbers, called an encoding. Two photos of the same person produce encodings that are close together. Enroll someone once from a clear photo, then compare every new face against your list.

    import face_recognition as fr
    
    # Enroll once, from a clear photo of someone who agreed to it
    known = [fr.face_encodings(fr.load_image_file("alex.jpg"))[0]]
    names = ["Alex"]
    
    # Then, for each person crop from YOLO (converted from BGR to RGB):
    rgb = cv2.cvtColor(person_crop, cv2.COLOR_BGR2RGB)
    for encoding in fr.face_encodings(rgb, fr.face_locations(rgb, model="hog")):
        distances = fr.face_distance(known, encoding)
        best = distances.argmin()
        name = names[best] if distances[best] <= 0.55 else "unknown"

    Lower tolerance means stricter matching. Leo uses 0.55, and the usual range is 0.4 to 0.65.

  7. 7

    Save events to a database

    SQLite needs no server; it's just a file. SQLAlchemy lets you describe tables as Python classes. This is a trimmed version of Leo's event table:

    src/database/models.py

    class Event(Base):
        __tablename__ = "events"
    
        id: Mapped[int] = mapped_column(Integer, primary_key=True, autoincrement=True)
        timestamp: Mapped[datetime.datetime] = mapped_column(DateTime, index=True)
        event_type: Mapped[str] = mapped_column(String(64), index=True)
        label: Mapped[str | None] = mapped_column(String(64))       # "person", "car"
        confidence: Mapped[float | None] = mapped_column(Float)
        track_id: Mapped[int | None] = mapped_column(Integer, index=True)
        face_name: Mapped[str | None] = mapped_column(String(128))
        screenshot_path: Mapped[str | None] = mapped_column(String(512))
        is_alert: Mapped[bool] = mapped_column(Boolean, default=False)
  8. 8

    Add alert rules

    Keep the rules in a config file so they can change without touching code. The cooldown matters most: without it, one visitor standing still produces an alert every frame.

    config.yaml

    alerts:
      unknown_person:
        enabled: true
        min_confidence: 0.55
        cooldown_seconds: 60     # don't repeat for the same track
      after_hours:
        enabled: false
        start: "22:00"
        end: "06:00"
      restricted_zones:
        enabled: false
        zones: ["server_room"]
  9. 9

    Serve an API and live video

    MJPEG is the simplest way to stream video to a browser: the server sends JPEG after JPEG, and an ordinary <img> tag plays them. This is Leo's stream endpoint:

    src/api/routes/cameras.py

    @router.get("/{camera_id}/stream")
    async def mjpeg_stream(camera_id: int):
        async def frame_generator():
            while True:
                jpeg = _manager.get_frame(camera_id)
                if jpeg:
                    yield (
                        b"--frame\r\n"
                        b"Content-Type: image/jpeg\r\n\r\n" + jpeg + b"\r\n"
                    )
                await asyncio.sleep(0.033)  # ~30 fps ceiling
    
        return StreamingResponse(
            frame_generator(),
            media_type="multipart/x-mixed-replace; boundary=frame",
        )
    uvicorn app:app --host 0.0.0.0 --port 8000

    Then open http://localhost:8000/docs to try every endpoint.

  10. 10

    Build the dashboard

    With the API in place, any front end works. The quickest is one HTML page:

    <img src="http://localhost:8000/api/cameras/1/stream" width="640">

    Leo went further with a Flutter phone app that connects over Wi-Fi, with screens for cameras, events, alerts, and enrolled faces. Its live view is the same MJPEG stream. Flutter builds for Android, iOS, and desktop from one Dart codebase.

    Keep the API on your home network. Exposing a camera feed to the internet without proper authentication lets anyone watch it.

↑↓ to move Enter to select Esc to close