Computer Science > Machine Learning

arXiv:2306.15083 (cs)

[Submitted on 26 Jun 2023 (v1), last revised 17 Jun 2024 (this version, v3)]

Title:Balanced Filtering via Disclosure-Controlled Proxies

Authors:Siqi Deng, Emily Diana, Michael Kearns, Aaron Roth

View PDF

Abstract:We study the problem of collecting a cohort or set that is balanced with respect to sensitive groups when group membership is unavailable or prohibited from use at deployment time. Specifically, our deployment-time collection mechanism does not reveal significantly more about the group membership of any individual sample than can be ascertained from base rates alone. To do this, we study a learner that can use a small set of labeled data to train a proxy function that can later be used for this filtering or selection task. We then associate the range of the proxy function with sampling probabilities; given a new example, we classify it using our proxy function and then select it with probability corresponding to its proxy classification. Importantly, we require that the proxy classification does not reveal significantly more information about the sensitive group membership of any individual example compared to population base rates alone (i.e., the level of disclosure should be controlled) and show that we can find such a proxy in a sample- and oracle-efficient manner. Finally, we experimentally evaluate our algorithm and analyze its generalization properties.

Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:2306.15083 [cs.LG]
	(or arXiv:2306.15083v3 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2306.15083
Journal reference:	5th Symposium on Foundations of Responsible Computing (FORC 2024)
Related DOI:	https://doi.org/10.4230/LIPIcs.FORC.2024.4

Submission history

From: Emily Diana [view email]
[v1] Mon, 26 Jun 2023 21:55:24 UTC (760 KB)
[v2] Wed, 5 Jul 2023 20:12:37 UTC (487 KB)
[v3] Mon, 17 Jun 2024 19:21:28 UTC (1,249 KB)

Computer Science > Machine Learning

Title:Balanced Filtering via Disclosure-Controlled Proxies

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Balanced Filtering via Disclosure-Controlled Proxies

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators