文章基本信息

标题：Taming Technical Bias in Machine Learning Pipelines
本地全文：下载
作者：Sebastian Schelter ; Julia Stoyanovich
期刊名称：Bulletin of the Technical Committee on Data Engineering
出版年度：2020
卷号：43
期号：4
页码：39-50
出版社：IEEE Computer Society
摘要：Machine Learning (ML) is commonly used to automate decisions in domains as varied as credit and lending, medical diagnosis, and hiring. These decisions are consequential, imploring us to carefully balance the benefits of efficiency with the potential risks. Much of the conversation about the risks centers around bias — a term that is used by the technical community ever more frequently but that is still poorly understood. In this paper we focus on technical bias — a type of bias that has so far received limited attention and that the data engineering community is well-equipped to address. We discuss dimensions of technical bias that can arise through the ML lifecycle, particularly when it’s due to preprocessing decisions or post-deployment issues. We present results of our recent work, and discuss future research directions. Our over-all goal is to support the development of systems that expose the knobs of responsibility to data scientists, allowing them to detect instances of technical bias and to mitigate it when possible.