Project information

  • Category: PhD Research, Civil Engineering — University of Michigan
  • Lab: Infrastructure for All (INFRALL) Lab
  • Data: Multi-year national smartphone GPS panel data; parcel-level land use data; American Community Survey demographics
TRB 2026 ASCE ICTD 2026

The Problem

Raw smartphone location data has to be filtered before it's usable—but the thresholds that decide who makes the cut are often set by convention, not evidence. With no ground truth to check them against, it's unclear whether they're removing noise or quietly removing entire populations.

Diagram showing the two-part problem: arbitrary sample retention thresholds and bias assessment limited to a single retained sample, and the two-part solution: data-driven sample selection and bias evaluation across thresholds

Approach

I built a classifier to predict when a device's inferred home location is unreliable, using a non-residential label as a proxy for high risk in place of ground truth. The high-risk class is rare, representing roughly 10% of cases, so the model has to stay precise about that class specifically rather than just optimizing overall accuracy. From there, instead of checking representativeness at one fixed threshold, I tracked how the retained sample's demographics shift as filtering tightens, making the equity tradeoff visible instead of hidden.

This builds on the Total Survey Error (TSE) framework from survey methodology, extending it to the specific ways bias enters an LBS pipeline—from who owns a smartphone in the first place, down to who gets excluded by data-quality filtering at the end.

Total Survey Error framework showing coverage error, selection error, nonresponse error, and inference/imputation error as a population is filtered down from target population to structured dataset

Key Results

  • Replaced arbitrary filtering conventions with a data-driven classifier
  • Addressed severe class imbalance—the high-risk class is only ~10% of cases—while keeping precision high on that class specifically
  • Validated across two years of data to confirm results hold over time
  • Checked across different seasonal windows to rule out artifacts
  • Showed representativeness shifts continuously with filtering—not just at one cutoff

Why It Matters

This is a data-quality problem with no ground truth to check against—the same challenge that shows up anywhere a pipeline has to filter noisy, passively-collected data, handle a rare but critical class, and decide who or what gets kept.