Investment Updates

Meet Rime: the linguistics-first voice AI designing the new interface for computing

Rime, led by Stanford-trained linguist Lily Clifford, builds the speech models behind nearly 100 million enterprise phone calls a month. The company is building the world’s first enterprise-ready speech-to-speech model.

By
Katie Boysen
,
Jul 28, 2026

7 MIN

TLDR 

  • Rime raised a $24 million Series A led by M13 to fund a 10x expansion of its proprietary conversational dataset and strategic hires across engineering and research so voice agents can think and respond at the same time.
  • Rime powers nearly 100 million monthly phone calls for enterprises like Mayo Clinic, Dialpad and Upstart.
  • Rime has statistically significantly lower "Hung Up During Intro" (HUDI) than major competitors across 100,000 calls.

Why is M13 investing in voice AI now?

Voice is the last interface computers are still bad at. Rime sits at the point where enterprise demand for voice AI is exploding and the quality bar for a live, high-stakes conversation is still far out of reach for most of the market. 

Rime is an enterprise voice AI company that builds the speech models powering high-stakes, regulated phone calls, including appointment scheduling, loan servicing, and patient care, for companies like Mayo Clinic, Dialpad and Upstart.

"There are many application companies in the market, and we began to feel like the problems they were solving would be resolved at the foundational layer," said M13 Partner Morgan Blumberg. "The question was whether there was a challenger to other voice AI solutions thinking about the problem differently. That's when we found Lily.

"The design patterns that will define the speech-to-speech era of AI haven't been built yet," said Blumberg. "Rime is doing the foundational work by combining frontier AI with deep linguistic expertise to create the speech infrastructure the next generation of voice builders will rely on. Lily, Brooke Larson and Ares Geovanis are building Rime with the rare combination of world-class AI research and deep linguistic expertise to deploy real-time voice people can trust in complex, regulated environments."

The $24 million round will go toward a roughly 10x increase in Rime's proprietary conversational dataset and strategic hires across engineering and research. As part of the round, Blumberg joins Rime's board, and Rafael Valle—who previously led audio research at Meta's Superintelligence Lab—joins as Chief Science Officer.

Proof and early validation

Rime powers nearly 100 million phone calls a month for enterprises including Mayo Clinic, Dialpad and Upstart, with particular traction in healthcare and financial services, where pronunciation accuracy and compliance are non-negotiable. 

In an independent study by Miravoice, which evaluated 12 voices from three vendors across 100,000 calls, Rime's voices produced statistically significantly lower "Hung Up During Intro" (HUDI) rates than other major voice companies and the fastest median time to completion of any provider tested. 

How does voice AI quality affect business outcomes?

Rime builds the speech models behind the AI voices enterprises put in front of customers including appointment scheduling, outbound calls, loan servicing, customer care. The premise is linguistics-first. Sounding human is table stakes; but sounding right for the specific conversation is what makes an application viable. A narration model can sound nearly indistinguishable from a person, but that same voice is wrong for a live exchange where tone and appropriateness carry the meaning.

In regulated, high-stakes industries like financial services, healthcare, hospitality, that distinction decides whether a product ships at all. A health system won't put a voice agent in front of patients for genetic counseling if it can't reliably pronounce prednisone or cystic fibrosis. When the stakes are high and an experience sounds flat or robotic, enterprises would rather wait than risk it.

The payoff is measurable. Enterprises running tens of millions of calls a year track how willing people are to stay on the line, and small gains compound fast. As Clifford puts it, "increasing willingness to engage even by a single percentage point can mean huge ROI."

What makes Rime's approach to voice AI different?

Nearly every AI company says that they have better models. Rime focuses on making better design choices. By blending design with data, Rime is able to create three wedges in the voice AI industry that rejects the concept of benchmark maxing, and focuses on the nuances in human speech that drive connection and meaning. 

  • Proprietary conversational data. As a Stanford PhD student, Clifford proved that a model trained only on clean audiobook narration scored well on narration but collapsed on real conversation. Today, almost every voice application is still built on that easy narration data, because conversation is far harder to collect and annotate. So Rime built its own recording studio and now holds what may be the largest proprietary corpus of studio-quality conversational speech.
  • A single speech-to-speech model. Today's voice agents run a cascaded pipeline—transcription, then reasoning, then synthesis—three steps in sequence, and a big reason they still sound robotic. Rime collapses them into a single speech-to-speech model, so an agent can think and respond at once. LLMs made voice applications easier to build; they didn't change how it feels to talk to one. Speech-to-speech is what closes that gap.
  • Design primitives for the last mile. LLMs made it possible to spin up a voice demo in seconds. Getting from that demo to production is another story. "It's a lot easier to go from zero to one now," Clifford says, "but the last mile, from one to 10, is really long." Today developers hand-script every branch and every stall message. Rime's aim is a world where they observe a conversation in production and improve it with a prompt—designing conversations instead of building forms.

Rime's origin story

Rime was founded in 2022 by Clifford, Larson and Geovanis. Clifford was a Stanford PhD student working on data selection for speech models at a time when "no one was saying voice AI, and no one was saying AI. People were saying deep learning." Clifford saw that no one was focused on the conversational element because the data was too hard to get. So the founders built a recording studio in a Mid-Market basement in San Francisco and started collecting it themselves.

From day one, the team became the engine. Larson is a linguist who taught in Harvard's linguistics department before joining Amazon Alexa; Geovanis is a Stanford engineer; and Clifford considers herself a linguist first, even before a machine learning researcher—someone drawn to how different ways of speaking carry different meaning, not just the words themselves.

That level of focus is also a hiring advantage because Rime attracts people obsessed with the craft of language. There's an old adage in natural language processing that your product gets better the day you fire the linguist; Rime is built on the opposite conviction. Its bar isn't the engineering problems the frontier labs optimize for, but the ones that are hard to verify—whether a conversation felt empathetic, whether a customer felt heard. 

For Blumberg, that gravitational pull was a signal in itself. "Lily and the team have a relentless obsession with the problem they're solving," she said. "That breeds resilience and dedication, and it attracts other people who are relentlessly obsessed with the same problem."

What is the future of voice AI in the enterprise? 

Enterprise adoption is still in its early innings, and it tends to move one use case at a time before the floodgates open. In practice, this looks like a health system that starts with appointment scheduling, then adds outbound research calls, then, eventually, a voice agent becomes the front door—and for many large institutions, more people already reach them by phone than through a web portal.

Clifford sees voice becoming a dominant computing interface as more of our interaction with software moves into the background. Every era of computing has deserved its own interface.  The constraint isn't demand; it's that enterprise teams still lack the modeling primitives to build these applications easily. Making the models more capable and defining the design tools around them is how that adoption gets unlocked.

FAQ

  • What is Rime? Rime is an enterprise conversational voice AI company building speech models for high stakes conversations. Its technology powers nearly 100 million monthly calls for enterprises including Mayo Clinic, Dialpad and Upstart.
  • What is a speech-to-speech model? A speech-to-speech model lets AI listen and respond directly with speech instead of separate transcription, language, and speech generation steps, making conversations faster and more natural.
  • Rime vs. ElevenLabs: what's the difference? ElevenLabs is known for consumer voice generation. Rime is built for live enterprise conversations, combining proprietary conversational data with speech-to-speech models for regulated industries.
  • Who uses Rime for voice AI? Enterprises including Mayo Clinic, Dialpad and Upstart use Rime for patient communication, appointment scheduling, loan servicing and customer support. Rime powers financial services, healthcare, travel and hospitality use cases, among others.
  • What is a cascaded voice AI pipeline? A cascaded voice AI pipeline separates speech recognition, language generation, and speech synthesis. Speech-to-speech models combine them into a single system for lower latency and a more natural feeling that customers stay on the line for.
  • What industries use voice AI? Voice AI is used across healthcare, financial services, insurance, hospitality, retail and customer support, especially for high volume, customer-facing conversations.

Read more about Rime

Learn more at rime.ai

Follow Rime

Follow Rime on LinkedIn and follow Lily Clifford for updates on the company's research and product.

The views expressed here are those of the individual M13 personnel quoted and are not the views of M13 Holdings Company, LLC (“M13”) or its affiliates. This content is for general informational purposes only and does not and is not intended to constitute legal, business, investment, tax or other advice. You should consult your own advisers as to those matters and should not act or refrain from acting on the basis of this content. This content is not directed to any investors or potential investors, is not an offer or solicitation and may not be used or relied upon in connection with any offer or solicitation with respect to any current or future M13 investment partnership. Past performance is not indicative of future results. Unless otherwise noted, this content is intended to be current only as of the date indicated. Any projections, estimates, forecasts, targets, prospects, and/or opinions expressed in these materials are subject to change without notice and may differ or be contrary to opinions expressed by others. Any investments or portfolio companies mentioned, referred to, or described are not representative of all investments in funds managed by M13, and there can be no assurance that the investments will be profitable or that other investments made in the future will have similar characteristics or results. A list of investments made by funds managed by M13 is available at m13.co/portfolio.

Media & appearances
No items found.

At a glance

There are no Cinderella stories. Blockbuster venture successes are always preceded by years of sleepless nights.”
There are no Cinderella stories. Blockbuster venture successes are always preceded by years of sleepless nights.”
—Sarah Tomolonius, M13 Partner & Head of Investor Relations