Article
Zoom has announced improvements in how it transcribes the speech of people with disabilities, thanks to data from the Speech Accessibility Project.
“Internally, we evaluated the production Zoom AI Services Scribe API on this benchmark exactly as it is available today,” it announced in a blog post. “There was no benchmark-specific tuning, and no separate model created for the evaluation. The result was a word error rate of 5.72%.”
That’s 52% lower than the next-closest system, it said.
“The true measure of speech AI is not how well it recognizes the average voice, but how faithfully it recognizes every voice,” wrote Zoom Chief Technology Officer Xuedong Huang in the blog post. “That is why the Speech Accessibility Project is so important.
“Led by the University of Illinois Urbana-Champaign, with participation from organizations across academia and industry, the project was created to address a persistent gap in speech technology,” Huang wrote.
Zoom used a Speech Accessibility Project dataset of about 1,000 hours of contributed speech from people with diverse and non-standard speech patterns.
“This is not a laboratory prototype created for a benchmark,” Huang added in a LinkedIn post. “It is a real-world production API operating at the leading edge of one of speech AI’s most demanding accessibility challenges.”
Speech Accessibility Project
405 N Mathews Ave., Urbana, IL 61801