Skip to main content
Harmanpreet Singh
Research

M.S. Computer Science (thesis track), University of Georgia, 2026

My thesis, explained

Multimodal Self-Supervised Activity Recognition on CASAS Smart-Home Sensor Streams

A smart home can tell what a resident is doing without a camera, just from motion and door sensors clicking on and off. My thesis asks how best to read those sparse event logs, and finds that the answer depends on the house.

PDF, 58 pages, 2.2 MB. Advisor: Fei Dou.

01

The short version

The data are timestamped sensor events from four public CASAS homes: a kitchen sensor turns on, a hallway sensor fires a moment later, a door opens. There is no video and nothing worn, which is good for privacy but makes the signal thin and irregular.

I cut each stream into windows of 30 events and turned every window into four different views: an ordered sequence, a short text summary, an image on the floor plan, and a graph. Each view is read by its own neural network, and a small Transformer combines them.

The headline result is that no single view wins everywhere. The best model for each home used a different combination of views, and the thesis shows why a single score would have hidden that.

  • 4

    CASAS homes

    Aruba, Cairo, Kyoto and Milan

  • 4

    views of each window

    sequence, text, image and graph

  • 30

    events per window

    stride of 15, so windows overlap

  • 3

    random seeds

    every result is a mean over seeds 10, 30 and 50

02

Why it is hard

Recognising what someone is doing from ambient sensors is harder than it sounds.

  • Sparse, irregular events

    A window of 30 events can span seconds during cooking or many minutes at night. The data are discrete events, not a steady signal.

  • Ambiguity

    Kitchen sensors can mean cooking, eating, taking medicine or just walking through. Evidence for different activities overlaps.

  • Class imbalance

    Common background behaviour vastly outnumbers rare but important events such as taking medicine or getting up at night, so a model can score well and still miss them.

  • Every house is different

    Sensor sets, floor plans and residents differ: Aruba has one resident and 34 sensors, Cairo has two residents and no door sensors, Kyoto has the richest set, Milan has ten activity classes.

03

Four views of the same 30 events

The idea is that the views are not copies of each other. Each one makes a different kind of evidence easy to see.

  1. Sensor log

    • Timestamped motion and door events
    • One house, one metadata file
  2. Event windows

    • 30 events, stride 15
    • Windows never mix activity labels
  3. Four derived views

    • Sequence of event features
    • Deterministic text summary
    • Floor-plan raster image
    • Event graph
  4. Encoders

    • Transformer
    • Transformer
    • CNN
    • Graph network
  5. Fusion and classifier

    • Fusion transformer
    • Auxiliary heads per view
    • Activity label

The pipeline for one window. Simplified from the thesis.

  • Sequence

    The 30 events in order, each with its sensor, state, time since the window started, the gap before it and the time of day.

    Good at: Order and timing, for example leaving home versus coming back.

  • Text

    A deterministic written summary of the window: time of day, weekday or weekend, dominant room, pace, room-to-room moves and per-sensor counts. No language model writes it.

    Good at: Context that survives small changes in event order.

  • Image

    A 96 by 96 picture on the house floor plan showing where sensors fired, how often, and whether early or late in the window, blurred into a heat map.

    Good at: Where things happened, which separates many activities better than counts do.

  • Graph

    Events as nodes, linked to the physical sensors they came from, with edges for order, repeated sensors and events close in time.

    Good at: Repeated sensors and neighbourhood structure, even in a house with no door sensors.

04

How it works

  1. Shared windows and splits

    All models, mine and the baselines, start from the same windows and the same train and test split. Differences in the results are therefore differences in how a window is represented and learned, not in how it was cut.

  2. Four views, four encoders

    A Transformer reads the sequence, a small Transformer reads the text, a convolutional network reads the image and a graph network reads the graph. Each produces a 256-number summary.

  3. Self-supervised pretraining

    Without using any labels, the encoders learn that the four views of one window belong together and that views of other windows do not (contrastive learning). The sequence encoder also learns to guess hidden sensor names.

  4. Fusion

    A two-layer Transformer combines the active views. Each view also keeps its own small classifier, and learned per-class weights decide how much each class trusts each view. Confident views are given more say.

  5. Fine-tuning for rare activities

    Training handles imbalance with weighted sampling, class weights and focal loss, and keeps the checkpoint with the best validation macro-F1 so rare activities are not ignored.

05

What I found

Weighted F1 is the main score. It follows the test set's real mix of activities. All numbers are means over three seeds.

Best proposed model for each house by weighted F1, with the strongest baseline for comparison.
HouseBest proposed modelWeighted F1Macro F1Best baseline (weighted F1)
ArubaGraph, image and sequence0.8500.735SICA 0.847
CairoGraph, text and image0.8420.524SICA 0.755
KyotoImage, text and sequence0.8430.766SICA 0.821

Baselines are SICA, CPC, DeepCASAS and DCNN, from the literature. Only the best baseline for each house is shown here.

Milan

Milan is the hardest home, with ten activity classes that mostly happen in the same few rooms. Its strongest proposed model used all four views, and the thesis reports that model as the best balanced one by macro F1, ahead of the strongest baseline. The thesis tables for Milan are in the PDF.

  • No one view wins

    Aruba did best with graph, image and sequence together. Cairo, with no door sensors, did best with graph, text and image. Kyoto, with rich floor-plan zones, did best with image, text and sequence. Milan needed all four. The best set of views differs by house, and adding a view does not always help.

  • One score can hide the failure

    Weighted F1 and accuracy reward getting the common activities right. Macro F1 gives every activity equal weight. Cairo scores 0.842 weighted but only 0.524 macro, because rare activities are still hard. For monitoring, the rare activities can be the ones that matter.

  • Some confusions are real

    The worst errors cluster where activities share the same room, sensors and rhythm, such as eating, cooking and taking medicine in a kitchen. There, intent is not visible in the sensor evidence.

06

The tools around the model

Part of the thesis is the research software that made the experiments repeatable.

  • One-file experiment configuration

    A single short intent file sets house, window size, stride and seeds, and the rest of the configuration is derived from it, so changing the window size changes everything consistently.

  • Sensor annotation tool

    A graphical tool for placing sensors on a floor plan and drawing zones as polygons, which writes the metadata the image and graph views need. A new house can be added with a visual pass instead of custom code.

  • Smart-home playback simulator

    Replays the raw event log over the floor plan and shows the four views of the current window, used to check that the pipeline is doing what it should.

The sensor annotation tool showing the Aruba floor plan with each of the 34 sensors placed and the office zone drawn as a polygon.
The sensor annotation tool (thesis Figure 3.1). Each of the 34 Aruba sensors is placed on the floor plan and assigned to a zone; the result drives the image and graph views.
The smart-home simulator replaying the Aruba event log over the floor plan with active sensors highlighted and tabs for the graph, image, text and sequence views.
The playback simulator (thesis Figure 3.2). It replays raw events over the floor plan and shows the four views of the current window.

Try an interactive version of the playback idea, on synthetic data

07

Limits and next steps

Limitations

  • Activity labels are not identical across houses (nine, seven and ten classes), so macro F1 values should be compared across houses with care.
  • Rare activities have little support, so their scores shift with the split and the seed.
  • Some figures show validation-style diagnostics while the tables show test-set means.
  • The annotation and simulator tools were not evaluated in a user study.

Future work

  • Study window lengths beyond 10 and 30 events, and how each activity changes as context grows.
  • Analyse the learned per-class weights to see which activities trust which view.
  • Add objectives aimed at rare activities such as taking medicine and getting up at night.
  • Test on more CASAS homes and on moving to a new resident or house.

The full thesis

Read it in full

This page is a plain-language summary. The PDF has the full method, every table and the references. It is the thesis as submitted to the University of Georgia Graduate School.

58 pages, 2.2 MB