Independent project: Platform engineering and machine learning
HomeSense: a smart-home intelligence platform
My thesis model turned into a running platform: live Kafka streaming, anomaly detection with human review, a citation-checked assistant, a resident simulator, gated model releases, web and phone apps, and a Kubernetes deployment.

What it is
A smart home's motion and door sensors produce a stream of events. HomeSense turns that stream into what the resident is doing, flags unusual patterns for a person to review, answers questions about the home in plain language, and lets you test all of it against a simulated resident.
It began as the model from my M.S. thesis. HomeSense is everything around that model that a real deployment would need: versioned contracts, an API and database, a live streaming pipeline, web and phone apps, a model-promotion process, and Kubernetes packaging. I built it in about four months, in ten documented phases.
My part
Sole author. I designed and built every layer described here: the ML core library, API, streaming services, worker, web and mobile apps, simulator, anomaly detectors, assistant, model lifecycle and Kubernetes packaging. All 53 commits in the repository are mine.
June to October 2026. Independent project in a single repository. The model it serves is the one from my thesis, and the datasets are the public CASAS smart-home datasets.
Stack
- Python
- FastAPI
- PostgreSQL
- Kafka
- Redis
- React 19
- TypeScript
- Expo
- Kubernetes
- Helm
- PyTorch
- MLflow
What is inside
Events enter through HTTP ingest, Kafka or the simulator and pass through four stream stages. PostgreSQL is the system of record: the API, web app, phone app and assistant all read from it.
01
Contracts and a framework-free ML core
Versioned pydantic contracts export JSON Schemas and generated TypeScript types. The ML core has no web, database or queue imports, which a test enforces. According to the project's golden tests, window files built by the original thesis code and the new library are bit-identical for all five houses.
02
API, database and worker
FastAPI with OpenID Connect (PKCE), PostgreSQL with Alembic migrations, optimistic locking and immutable published house snapshots, S3-compatible object storage, and a worker that claims jobs from PostgreSQL. Redis only wakes workers sooner.
03
Production inference
A small service serves exactly one approved, checksummed model. It reports ready only after download, a SHA-256 match, load and a real warm-up inference, uses micro-batching, and keeps the old model serving if a reload fails.
04
Real-time streaming
Kafka topics run from sensor events to windows, predictions and anomalies, plus a dead-letter topic. Windowing uses event time with a watermark, delivery is at-least-once with idempotent writes, and clients receive updates over server-sent events.
05
Web and phone apps
A React 19 web app and an Expo Android app share a typed API client generated from the OpenAPI schema and shared design tokens. The phone app has OIDC sign-in, an offline cache, push notifications by design and an offline review queue with idempotency keys.
06
A controlled assistant
A fixed LangGraph workflow classifies the question, calls read-only, workspace-scoped, bounded tools, retrieves documents from pgvector and then verifies the draft. Any statement without supporting evidence is removed. An offline deterministic provider runs in CI.
07
Digital-twin simulator
A statistical engine fitted per house, and a physical engine in which one resident walks the real floor plan (A* navigation) while motion and door sensors respond. Runs publish through Kafka like real sensors and store ground truth for evaluation.
08
Anomaly detection and review
A rule detector and a learned per-house detector with a threshold tuned to about one false alarm a day. Four verdicts (confirmed, false positive, expected, insufficient evidence) each create a new label version.
09
Model lifecycle
Models move through candidate, challenger, champion and rollback aliases. Promotion needs measured checks, a shadow run next to the current model and a researcher's review. Drift checks only ask for an evaluation, and lineage answers which model and checks produced a prediction.
10
Kubernetes
A Helm chart renders 15 workloads, with autoscaling on the right signals, restricted pod security, default-deny network policies and resources derived from a load test. Worker pools are split by job kind.
1. Sensors
- Real events
- Replay
- Simulated resident
2. Kafka
- Keyed by house
- Dead-letter topic
3. Window builder
- Event-time windows
- Watermark
4. Predictor
- Inference service
- One pinned model
5. Anomaly detectors
- Rules
- Learned per-house model
6. PostgreSQL
- System of record
- Idempotent writes
7. Web and phone
- Live updates
- Human review




