google.com, pub-8701563775261122, DIRECT, f08c47fec0942fa0
USA

Model routing on AI is a problem for OpenAI and Anthropic

A new spending discipline is beginning to take hold in corporate America as CFOs and boards begin to guard against inefficient AI spending. This change has the potential to reshape the AI ​​business.

For the past two years, the tactic has been to use the strongest AI model by default and route all queries through that model, regardless of complexity. Now that AI bills are far beyond budgets, companies are starting to ask whether every task really needs limits. Two leaders at the center of the AI ​​framework told CNBC this week that a solution has emerged: model steering.

What is model routing?

Routing is a tool that matches work to model, sending hard problems to expensive frontier models and easy problems to cheaper, faster alternatives.

Scott Wu, CEO of Cognition, which makes the coding agent Devin, said the gains in routine work are huge. For most standard jobs, companies can achieve five to 10 times better cost efficiencies by using models that are still good enough for the task, he said.

Nowadays, most companies do not provide referrals at all. Glean CEO Arvind Jain estimates that roughly 95% of enterprise AI use is still running on the most expensive frontier models, even for tasks that cheaper alternatives can easily handle. Wu gave the example of asking a model to name the third US president. Every single one of them, no matter how expensive, will tell you it was Thomas Jefferson.

Glean CEO Arvind Jain on the SaaS Monster stage during day one of Web Summit 2022 at the Altice Arena in Lisbon, Portugal, on November 2, 2022.

Harry Murphy | Sports file | Getty Images

Sellers are under pressure

AI companies are aware of the concern.

Cognition has announced what it calls its AI productivity guarantee. If Devin delivers less engineering value than what the customer paid, Cognition will fund the usage for up to $10 million until it reaches parity. Wu framed this as a way to cut through the noise on return on investment, a metric that has dogged the industry.

Rather than measuring activities like tokens or lines of code consumed, Wu said, Cognition estimates the number of human engineering hours the agent actually saved and backs up that estimate with a refund. “You can spend billions of coins and not do anything with it,” he said. Companies should strive for output, not efficiency.

If companies start redirecting easy, high-volume work to cheaper open-source models in China or elsewhere, OpenAI and Anthropic will stop charging for every task. They only take on more complex jobs. Both companies established their businesses and IPO prospects based on the assumption of massive demand at premium prices.

Patel doesn’t think this will bankrupt frontier labs and says cutting-edge technology will remain valuable. But he sees the pricing model changing. Labs will need to become more efficient in how they use models rather than charging more, and Patel predicts this will lead to a concerted effort across the industry.

The question was whether companies would continue to spend as AI bills increased. It seems like many people will now find a way to spend wisely. Pricing power is shifting from companies selling world-class AI to companies buying it.

Frontier labs will still charge a premium for the most demanding work. So how much of the market do other things make up? The answer could go a long way in determining valuations of leading AI companies.

Select CNBC as your preferred source on Google and never miss a beat from the most trusted name in business news.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button