Voice is fast becoming the primary way people and businesses interact with artificial intelligence. As speech models grow more natural and more affordable, enterprises across banking, insurance, healthcare, customer support and commerce are moving from typed interfaces and rigid menus to conversational voice, in the languages their customers actually speak. Industry watchers expect voice to be one of the biggest drivers of enterprise AI adoption over the next few years.

That shift is being powered by a sharp fall in the cost of building and running AI. The Stanford AI Index 2026 notes that the cost of running a model at GPT-3.5 level of quality fell roughly 280 times in about eighteen months, while open-source model development has scaled rapidly. 

McKinsey’s latest State of AI report finds that more than 80 percent of companies now use AI in at least one function, with attention shifting from simply scaling compute toward deploying models more efficiently. Together, these trends have opened the door for smaller, focused teams to build advanced AI that until recently was the preserve of billion-dollar labs.

This is a global pattern. Efficient, open-weight players have repeatedly shown that lean teams can reach the frontier, and a growing number of them are emerging outside the traditional hubs of the United States and China. India’s Maya Research is one such example. The Bengaluru-based startup says it has built frontier speech AI models with a team of about five people and roughly one million dollars in total spend.

“The cost of building at the frontier has collapsed, and the advantage has shifted from capital to conviction and craft. The moat is no longer the size of your cheque. It is the quality of your decisions,” said Dheemanth Reddy, co-founder and CEO of Maya Research.

Rather than relying on expensive infrastructure, Maya says it has developed an architecture designed to run on consumer-grade GPUs and to keep improving through conversations collected on its own voice application. According to Reddy, the company now ranks first globally for Hindi speech models on public leaderboards and among the top five for English.

For India and other emerging markets, the opportunity is tied to two things: the multilingual needs of the next billion users coming online across India, Southeast Asia and West Asia, and access to affordable computing power. Government-backed compute initiatives and cheaper hardware are gradually easing the constraint, but founders say it remains the biggest hurdle.

“The honest bottleneck today is not model quality. It is access to affordable GPUs. A country that wants to build AI infrastructure has to treat compute as national infrastructure,” Reddy said.

Industry demand is expected to be strongest where a multilingual, human-sounding voice makes a measurable difference: banking and insurance, customer support and contact centres, healthcare and commerce. In these sectors, the ability to hold a natural conversation in a customer’s own language is increasingly seen as a driver of both cost savings and better experience.

The next frontier for the field is real-time, full-duplex speech-to-speech AI, systems that let people converse naturally without converting speech into text first. Several global labs are racing toward it, and Maya counts it as its next milestone.

“If we ship speech-to-speech, make voice drastically cheaper, and get it on-device, Maya can be the voice interface for the next billion users,” Reddy said.