Exclusive: Wispr Raises $280M to Make Voice a Serious Computing Interface
Image credit : Wispr Flow
Wispr wants speaking to a computer to feel as natural as speaking to another person. Its first product, Flow, turns everyday speech into polished writing across apps, but the larger ambition reaches beyond faster dictation. The company is building toward an interface where people can communicate, create, and eventually act through voice without repeatedly returning to a keyboard or screen.
That ambition now has considerably more capital behind it. Wispr has raised $280 million in Series B funding at a $2 billion valuation, led by Menlo Ventures. Existing investors Notable Capital, NEA, Neo Ventures, 8VC and MVP Ventures participated alongside new backers including Acrew, Activate, Forerunner, Goodwater, Peak XV, Together Fund and Capital. The round takes total funding to $361 million.
Wispr was founded by Tanay Kothari and Sahaj Garg, two Stanford-trained technologists whose backgrounds span artificial intelligence, machine learning, and product engineering. Their long-term thesis is unusually ambitious: make voice a primary way people interact with computing rather than another feature buried inside software.
The biggest obstacle is trust. Voice loses its advantage the moment a user has to stop, find a mistake, and correct it manually. Wispr is directing much of the new funding toward accuracy and has previewed Canto, its first proprietary speech model. The company says that under difficult conditions involving noise, wind, accents, or music, Canto can reduce word-error rates from above 30% to roughly 5%–10%.
Wispr says users have already written more than 60 billion words with Flow, while people at almost all Fortune 500 companies and more than 10,000 enterprises use the product. Its next challenge is bigger than transcription: proving that voice can become dependable enough to carry meaningful work, understand context and eventually turn what someone says into action. If that happens, the keyboard may remain important, but it will no longer own every conversation between humans and computers.