Glean CEO explains how model routing and agentic workflows cut enterprise AI costs
Glean’s Arvind Jain describes how dynamic model selection, human feedback loops, and agentic search (Waldo) reduce latency and token usage while controlling costs across thousands of employees.
1 source · cross-referenced
- Glean’s model routing lets enterprises choose or restrict models per task or user, with most customers opting for automatic selection to reduce costs.
- Glean reports its routing and agentic search model Waldo reduces latency by 50% and tokens by 25% by filtering queries before handing off to frontier models.
- Glean’s ARR reached $300M in 2026, up threefold in 15 months, as enterprises adopt its AI coworker and agent platform.
- Open-weight models are now a key part of enterprise AI strategies due to cost concerns, Jain says, with usage rising sharply in the last three months.
Glean, an enterprise AI platform led by ex-Google Distinguished Engineer Arvind Jain, uses model routing to help organizations control costs as frontier model prices climb. The company’s system lets employees explicitly select models, administrators restrict usage, or Glean automatically route tasks to the most cost-effective model for the job. Most customers choose automatic routing for economic reasons, Jain told Latent Space.
Glean’s agentic search model, Waldo, introduced in April 2026, reduces latency by 50% and tokens by 25% by filtering queries and assembling “raw materials” before handing off to a frontier model. Waldo decides how to break down questions, which tools to use, and when there’s enough evidence to escalate—reserving advanced models for work that demands them.
Glean’s business momentum reflects enterprise demand for cost control. The company reached $300M in annual recurring revenue in 2026, a threefold increase over 15 months, and was last valued at $7.2B after a $150M Series F in June 2025. Customers like Zillow report 80% adoption across 7,000 employees, while Booking.com adopted Glean as its first company-wide AI platform.
Jain highlighted the role of human feedback loops in improving model routing. With Glean deployed to thousands of employees and used to build agents across departments, the company observes which models users select first, when they upgrade, and why. This data feeds continuous updates to the routing system, helping Glean’s AI judges assess route quality on real-world workloads.
Cost pressures are driving interest in open-weight models, Jain said. Enterprise usage of open-source LLMs was “minuscule” in 2025 but surged in mid‑2026 as businesses sought cheaper alternatives. “Nobody is willing anymore to rely on only one or two providers,” Jain said, noting that open-weight models are now considered a key part of enterprise AI strategies. Glean’s internal evals compare real-world workloads across query classes, testing cheaper and more expensive routes in parallel to refine routing decisions.
- Aug 18, 2026 · Hugging Face
IBM Research study finds agentic memory dose must be calibrated to model capability
Trust79 - Aug 17, 2026 · arXiv cs.CL
Researchers propose InflationAgent to cut agentic LLM costs by routing based on token inflation
Trust79 - Aug 15, 2026 · Latent Space — swyx
Flue 2 introduces React-style hooks for building dynamic agent harnesses
Trust78