Population-aware hierarchical Bayesian domain adaptation via multi-component invariant learning

Vishwali Mhasawade, Nabeel Abdur Rehman, Rumi Chunara

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

While machine learning is rapidly being developed and deployed in health settings such as influenza prediction, there are critical challenges in using data from one environment to predict in another due to variability in features. Even within disease labels there can be differences (e.g. "fever" may mean something different reported in a doctor's office versus in an online app). Moreover, models are often built on passive, observational data which contain different distributions of population subgroups (e.g. men or women). Thus, there are two forms of instability between environments in this observational transport problem. We first harness substantive knowledge from health research to conceptualize the underlying causal structure of this problem in a health outcome prediction task. Based on sources of stability in the model and the task, we posit that we can combine environment and population information in a novel population-aware hierarchical Bayesian domain adaptation framework that harnesses multiple invariant components through population attributes when needed. We study the conditions under which invariant learning fails, leading to reliance on the environment-specific attributes. Experimental results for an influenza prediction task on four datasets gathered from different contexts show the model can improve prediction in the case of largely unlabelled target data from a new environment and different constituent population, by harnessing both environment and population invariant information. This work represents a novel, principled way to address a critical challenge by blending domain (health) knowledge and algorithmic innovation. The proposed approach will have significant impact in many social settings wherein who the data comes from and how it was generated, matters.

Original languageEnglish (US)
Title of host publicationACM CHIL 2020 - Proceedings of the 2020 ACM Conference on Health, Inference, and Learning
PublisherAssociation for Computing Machinery, Inc
Pages182-192
Number of pages11
ISBN (Electronic)9781450370462
DOIs
StatePublished - Feb 4 2020
Event2020 ACM Conference on Health, Inference, and Learning, CHIL 2020 - Toronto, Canada
Duration: Apr 2 2020Apr 4 2020

Publication series

NameACM CHIL 2020 - Proceedings of the 2020 ACM Conference on Health, Inference, and Learning

Conference

Conference2020 ACM Conference on Health, Inference, and Learning, CHIL 2020
Country/TerritoryCanada
CityToronto
Period4/2/204/4/20

Keywords

  • data generating process
  • influenza prediction

ASJC Scopus subject areas

  • Public Health, Environmental and Occupational Health
  • Education
  • Health(social science)

Fingerprint

Dive into the research topics of 'Population-aware hierarchical Bayesian domain adaptation via multi-component invariant learning'. Together they form a unique fingerprint.

Cite this