Skip to main content
Harmanpreet Singh
All work

Independent project: Platform engineering and machine learning

HomeSense: a smart-home intelligence platform

My thesis model turned into a running platform: live Kafka streaming, anomaly detection with human review, a citation-checked assistant, a resident simulator, gated model releases, web and phone apps, and a Kubernetes deployment.

The HomeSense web app showing the digital-twin viewer: an apartment floor plan with sensors and their detection radius, a green marker for the simulated resident, playback controls, and side panels for what really happened and what the platform concluded.
The digital-twin viewer. A simulated resident walks the floor plan, the sensors react, and the panels compare what really happened with what the platform concluded.

What it is

A smart home's motion and door sensors produce a stream of events. HomeSense turns that stream into what the resident is doing, flags unusual patterns for a person to review, answers questions about the home in plain language, and lets you test all of it against a simulated resident.

It began as the model from my M.S. thesis. HomeSense is everything around that model that a real deployment would need: versioned contracts, an API and database, a live streaming pipeline, web and phone apps, a model-promotion process, and Kubernetes packaging. I built it in about four months, in ten documented phases.

My part

Sole author. I designed and built every layer described here: the ML core library, API, streaming services, worker, web and mobile apps, simulator, anomaly detectors, assistant, model lifecycle and Kubernetes packaging. All 53 commits in the repository are mine.

June to October 2026. Independent project in a single repository. The model it serves is the one from my thesis, and the datasets are the public CASAS smart-home datasets.

Stack

  • Python
  • FastAPI
  • PostgreSQL
  • Kafka
  • Redis
  • React 19
  • TypeScript
  • Expo
  • Kubernetes
  • Helm
  • PyTorch
  • MLflow

What is inside

Events enter through HTTP ingest, Kafka or the simulator and pass through four stream stages. PostgreSQL is the system of record: the API, web app, phone app and assistant all read from it.

  1. 01

    Contracts and a framework-free ML core

    Versioned pydantic contracts export JSON Schemas and generated TypeScript types. The ML core has no web, database or queue imports, which a test enforces. According to the project's golden tests, window files built by the original thesis code and the new library are bit-identical for all five houses.

  2. 02

    API, database and worker

    FastAPI with OpenID Connect (PKCE), PostgreSQL with Alembic migrations, optimistic locking and immutable published house snapshots, S3-compatible object storage, and a worker that claims jobs from PostgreSQL. Redis only wakes workers sooner.

  3. 03

    Production inference

    A small service serves exactly one approved, checksummed model. It reports ready only after download, a SHA-256 match, load and a real warm-up inference, uses micro-batching, and keeps the old model serving if a reload fails.

  4. 04

    Real-time streaming

    Kafka topics run from sensor events to windows, predictions and anomalies, plus a dead-letter topic. Windowing uses event time with a watermark, delivery is at-least-once with idempotent writes, and clients receive updates over server-sent events.

  5. 05

    Web and phone apps

    A React 19 web app and an Expo Android app share a typed API client generated from the OpenAPI schema and shared design tokens. The phone app has OIDC sign-in, an offline cache, push notifications by design and an offline review queue with idempotency keys.

  6. 06

    A controlled assistant

    A fixed LangGraph workflow classifies the question, calls read-only, workspace-scoped, bounded tools, retrieves documents from pgvector and then verifies the draft. Any statement without supporting evidence is removed. An offline deterministic provider runs in CI.

  7. 07

    Digital-twin simulator

    A statistical engine fitted per house, and a physical engine in which one resident walks the real floor plan (A* navigation) while motion and door sensors respond. Runs publish through Kafka like real sensors and store ground truth for evaluation.

  8. 08

    Anomaly detection and review

    A rule detector and a learned per-house detector with a threshold tuned to about one false alarm a day. Four verdicts (confirmed, false positive, expected, insufficient evidence) each create a new label version.

  9. 09

    Model lifecycle

    Models move through candidate, challenger, champion and rollback aliases. Promotion needs measured checks, a shadow run next to the current model and a researcher's review. Drift checks only ask for an evaluation, and lineage answers which model and checks produced a prediction.

  10. 10

    Kubernetes

    A Helm chart renders 15 workloads, with autoscaling on the right signals, restricted pod security, default-deny network policies and resources derived from a load test. Worker pools are split by job kind.

  1. 1. Sensors

    • Real events
    • Replay
    • Simulated resident
  2. 2. Kafka

    • Keyed by house
    • Dead-letter topic
  3. 3. Window builder

    • Event-time windows
    • Watermark
  4. 4. Predictor

    • Inference service
    • One pinned model
  5. 5. Anomaly detectors

    • Rules
    • Learned per-house model
  6. 6. PostgreSQL

    • System of record
    • Idempotent writes
  7. 7. Web and phone

    • Live updates
    • Human review
The live path. Events become windows, predictions and anomalies, and are stored in PostgreSQL.
The HomeSense anomaly review page: a list of flagged windows on the left, and on the right the evidence, an explanation of why it was raised, and four verdict buttons: confirmed anomaly, false positive, expected behaviour and insufficient evidence.
Anomaly review. The evidence is never rewritten, and each human verdict is saved as a new label version.
The HomeSense model lifecycle page with four cards for candidate, challenger, champion and rollback, a model table showing the dataset and commit it was trained from, promotion checks, drift and an append-only deployment history.
The model lifecycle. A model reaches a house through checks, a shadow run and a review, and every change is recorded.
Phone screen showing the assistant answering the question: were there anomalies in the last 24 hours, and what did the model think the resident was doing. The answer cites numbered evidence.
The assistant answers from tool results and cites its evidence.
Phone screen showing one anomaly: its category and severity, why it was raised with a score against a threshold, buttons to show the prediction behind it and to explain it, and a Mark as reviewed button.
Anomaly detail with the reasons it was raised and a review action.
Phone screen showing a finished simulation run, Digital twin night wandering, with progress, a Replay button, what the platform thinks against what the simulator is doing, and the injected nocturnal wandering anomaly marked as noticed.
A simulation run with its injected anomaly and whether the platform noticed it.