Research Project
AI Voices of America
Building the first nationally representative dataset of American English voices — to study privacy, fairness, and trust in AI voice technology.
About the project
AI voice technology — smart speakers, phone assistants, customer service agents — has quietly become part of everyday life. Yet the datasets used to train these systems disproportionately reflect internet-derived speech from a narrow slice of the country, and consumers rarely realize how much personal information their voice can reveal. AI Voices of America is building the first nationally representative dataset of American English voices — roughly 3,000 participants reflecting the geographic, demographic, and political diversity of the United States — to enable rigorous research on three connected questions.
- What can voice reveal about a person? Voice carries information about age, gender, region, socioeconomic background, personality, mood, and political identity. We aim to map the actual scope of consumer privacy risk in voice-based AI systems and inform data-governance and disclosure policy.
- Do AI voice systems perform equitably? Commercial speech recognition has been shown to err nearly twice as often for Black speakers as for white speakers, and similar gaps likely exist along regional, socioeconomic, and political lines — driven by training data that underrepresents rural, non-coastal, and politically conservative communities. The dataset will support equity audits and benchmarking across these dimensions.
- How can we design more representative and trustworthy AI voices? Listeners trust AI voices more when they sound like their own community, even when they know the voice is synthetic. Building voice systems on representative data is a precondition for honest and fair human–AI interaction.
Team
-
Scott Schanke
Assistant Professor of Management Information Systems · University of Georgia
scott.schanke@uga.edu