Modulate raises $25M to sell AI that listens to voice calls, not just transcripts

The Velma platform reads emotion, tone and synthetic speech from audio directly, and is used for fraud detection, voice agent supervision and trust and safety work


Modulate co-founders Mike Pappas, left, and Carter Huffman laugh on a tan leather sofa beside a window, with Modulate logo cushions.

Modulate co-founders Mike Pappas, chairman, left, and Carter Huffman, chief executive.

Image Credits Credit: Modulate

Modulate, a Boston company that builds AI models to analyze speech directly rather than working from text transcripts, has raised $25M in a round led by Future Ventures. Hyperplane and Lakestar also took part, bringing the company’s total funding to $60M.

The company was founded by Mike Pappas and Carter Huffman, who first met at MIT when Mike worked out a physics problem that Carter was solving in a hallway. They soon became friends because of their shared interest in how computers affect people’s ability to connect.

Modulate first became known in the gaming world. Since 2023, its ToxMod system has helped moderate voice chat in Call of Duty, and Lakestar led a $30 million Series A funding round in 2022. Now, the company plans to use its technology for fraud prevention, customer service, and a new challenge: checking if AI voice agents are working as intended.

“Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can’t be solved from a transcript,” said chief executive and co-founder Carter Huffman, who took over the top job from co-founder Mike Pappas, now chairman.

The main product, Velma, can analyze audio signals, picking up on emotion, tone, intent, emphasis, and whether the voice is synthetic. The various signals can then be combined to identify higher-level events such as a fraud attempt, a harassment case, or a frustrated customer, and this can be done in real time during an ongoing call.

Instead of using one large model, Velma employs what Modulate calls an Ensemble Listening Model, which selects and combines more than 100 smaller, specialized audio models for each task.

The company states that this method is up to 1,000 times more efficient than using a single large model and that Velma is twice as accurate as general-purpose large language models at identifying actual problems while generating seven times fewer false alarms; those comparisons are made by Modulate itself.

Its public benchmark results are easier to check. In July, Modulate’s transcription models took first place out of 88 entries on Hugging Face’s Open ASR Leaderboard, and its deepfake detector has ranked first on Hugging Face’s Speech Deepfake Arena since March, with an equal error rate of 1.1%. Batch transcription costs $0.03 an hour.

The deepfake work lands in a market where cloned voices have become an established tool for fraud. Modulate says hospitals use its models to defend against deepfake callers, and that its systems now process more than 10mn hours of audio a month, with more than 600mn hours analyzed in total. It also offers voice masking to protect staff in high-risk roles and detects child grooming in voice conversations.

Steve Jurvetson, co-founder of Future Ventures, said the company had “gained a significant technical lead in audio-native AI” and that demand was spreading well beyond gaming into AI agents, security, and customer experience.

The money will go into research, engineering, and a push to court developers with new SDKs, APIs, and partner integrations, so that companies building voice products can buy audio understanding rather than train their own models.

Get the TNW newsletter

Get the most important tech news in your inbox each week.

Published
Back to top