Founder Stack

Wispr raises $280M to push AI voice technology bey

Wispr raises $280M to push AI voice technology beyond dictation

· Funding · YourStory

US-based artificial intelligence startup Wispr has raised $280 million in a Series B funding round at a $2 billion valuation, as the company behind the voice-based writing platform Flow seeks to improve speech recognition and expand its broader human and AI interaction technology. The funding round was led by Menlo Ventures, an early investor known for backing technology companies such as Uber, Roku, and Anthropic. Existing investors Notable Capital, NEA, Neo Ventures, 8VC, and MVP Ventures also participated, alongside new backers Acrew, Activate, Forerunner, Goodwater, Peak XV, Together Fund and PLUS Capital. The latest round brings Wispr’s total funding to $361 million. Founded to make voice a practical alternative to typing, Wispr develops AI tools that convert spoken language into written text. Its flagship product, Flow, is designed to help users write emails, messages and documents by speaking naturally. The company recently introduced Notetaker, a meeting transcription tool that automatically captures conversations and generates notes. The newly-raised capital will be directed largely towards research and development, particularly in improving speech recognition accuracy and expanding the deployment of Wispr’s technology across the places where people already communicate and work. Chief executive and co-founder Tanay Kothari said accuracy would determine whether voice technology becomes a genuine replacement for typing rather than remaining a novelty. He noted that voice systems fail when users must repeatedly stop to correct mistakes, because interruptions break concentration and reduce trust in the technology. Alongside the funding announcement, Wispr unveiled a preview of Canto, its first proprietary speech model. Speech models are AI systems trained to recognise spoken language and convert it into text. According to the company, many existing systems are evaluated using clean recordings produced in quiet environments, despite the fact that most users interact with voice tools in cars, on busy streets or in open offices. Wispr said Canto was developed specifically for those real-world conditions. The company claims that, in particularly noisy environments, the model can reduce word error rates from more than 30% to between 5 and 10%. In everyday use, it expects the number of voice-generated texts requiring manual editing to fall by 30 to 35%. One of Canto’s distinguishing features is its ability to handle multilingual speech. This is especially relevant in countries such as India, where many speakers routinely switch between languages within a single conversation, a behaviour known as code-switching. Wispr used Hinglish, a blend of Hindi and English, to illustrate how traditional speech systems often struggle to recognise mixed languages and convert them into the correct written form. The company said its model also incorporates personal dictionaries and frequently used names to improve accuracy. The emphasis on multilingual AI reflects

Original source: YourStory
Read more on Founder Stack