Model Routing Made Our Customers' AI Work 76% Cheaper

Most AI requests in regulatory and quality work are small. Naming a chat, finding the sentence a citation came from, suggesting a follow-up question and pulling a product name out of a document come out the same on a small, fast model as on the largest one. A gap analysis across a dozen SOPs, or a submission question that spans several guidance documents, needs the strongest model available.
Arca sends each request to the smallest model that handles it well. Over the last 30 days, our customers' AI work cost 76% less than it would have with every request on our top tier, Max. The answers held up. On 50 everyday questions, Auto scored 4 points below Arca running Claude Opus 5.5 on every turn, at 29% of the cost. Workspace admins can now see their own savings in Analytics, and on our usage-based plans the savings come off the bill.
- Saved against Max
- 76%
- Across 56 customer workspaces, last 30 days
- Cost on Max
- 4.2x
- The same work with every call on the Max tier
- Tier chosen by Auto
- 91%
- Chat turns where Arca picked the tier
Chat has Fast, Standard and Max tiers
In chat, people pick Fast, Standard or Max, or leave the tier on Auto and Arca picks one for each turn. Background work, such as citation checks, document processing and grid cells, runs on a model chosen for that task. When a provider releases a stronger model, we move the matching tier to it.
On our regulatory benchmark, Standard answers about as well as Max. We ran Arca on Reg Affairs Bench once with every request on Standard and once on Max. On the 28 tasks both runs completed, Standard scored 84.6 out of 100 and Max 84.3, and each fully passed 22 of them. Max cost 3.4 times as much per task.2
Auto picks Standard when unsure
For each turn on Auto, a frontier decision model answers a few questions about the request. Does it need several dependent reasoning steps? Does it combine three or more sources, or search a whole library? How costly would a wrong answer be? The model returns a probability for each answer instead of generated text. A fixed rule in our code turns the answers into a tier. When an answer is close to 50/50 and flipping it would change the tier, a language-model classifier answers the same questions instead.
| Tier | When Auto picks it |
|---|---|
| Fast | A bounded, one-step request with no documents attached and nothing high-stakes. |
| Standard | Everything else, including anything ambiguous, and whenever the classifier fails. |
| Max | The request needs both synthesis across many sources and several dependent steps. |
Requests with documents or high stakes never run on Fast. Anyone can pick a tier by hand, including Max, and people did on 9% of turns.
Chat saves 45%, background work over 80%
We priced every AI call our customers made in the last 30 days twice. The first price is for the model that ran the call. The second is for the same tokens on the Max-tier model. Both use the providers' public list rates.1
| Area | Share of cost | Saved against Max |
|---|---|---|
| Chat | 69% | 45% |
| Citations and review | 20% | 82% |
| Document processing | 6% | 96% |
| Agents and automations | 3% | 71% |
| Titles and suggestions | under 1% | 93% |
| Grids | under 1% | 97% |
In chat, 21% of turns ran on Max, 72% on Standard and 7% on Fast. Citation checks, document tagging and grid cells save more because they run in large numbers on small models.
Savings depend on the team. Among our eight most active customers they ranged from 37% to 87%. The median was 72%. Teams that do more research across documents use Max more.
Auto on 50 everyday questions
Cheaper routing is only worth it if the answers stay good. To check, we wrote 50 everyday questions our customers ask. Some are quick lookups and others are full regulatory analyses. We ran each through Arca three times. One run was on Auto, one had every turn on Anthropic's Claude Opus 5.5, and one had every turn on OpenAI's GPT-6 Astra. All three runs used Arca's tools and workflow, so only the model choice differed. We also sent each question to the Opus 5.5 and GPT-6 Astra APIs directly, as in Reg Affairs Bench. Those runs could use the provider's own web search but none of Arca's tools. A jury of three judge models scored each answer against a written rubric.3
Auto cost 71% less per task than Arca on Opus 5.5 and 88% less than Arca on GPT-6 Astra. It scored 89.0 out of 100. That is 4 points below Arca on Opus 5.5 and within a point of Arca on GPT-6 Astra. Auto fully passed 47 of the 50 questions, against 49 and 44. It ran 19 questions on Fast, 28 on Standard and 3 on Max.
Auto also cost less than both provider APIs and scored at least as well. It cost 27% less per task than the Opus 5.5 API and scored 5.7 points higher. It cost 55% less than the GPT-6 Astra API at about the same score (87.7). Arca on a fixed model costs more than that model's API because it reads far more source material before it answers, about 20 times as many input tokens per question.
Admins can see their savings in Analytics
Savings differ from team to team, so each workspace gets its own numbers. We are rolling out an AI usage tab in Analytics for workspace admins. It shows the savings against Max per day or per week, the parts of Arca the usage went to, the tier each chat turn ran on and each person's share of usage.
Our first usage-based plans start this month. On those plans, each request is billed at the rate of the model that ran it, not the Max tier, so the savings come off the customer's bill.

References
- Figures cover 56 customer workspaces from 5 September to 5 October 2026. Internal and demo workspaces are excluded. Each call is compared with today's Max-tier model at public list prices. Embeddings for search indexing, 1.4% of cost, have no Max equivalent and are left out of the savings rate. The Auto share counts the chat turns that recorded how their tier was chosen.↩
- Run on 5 October 2026 over the 34 headline tasks of Reg Affairs Bench, with today's Standard and Max models. Six tasks failed to return an answer in one or both runs and are left out of the comparison. A full pass means every critical rubric line was met and no penalty fired. Max finished tasks faster, in 6.6 minutes on average against 8.9 on Standard.↩
- Run on 7 October 2026, and the two provider API runs on 8 October. The questions can be answered from public sources and are not published. Costs are at provider list prices for the tokens and searches each run used. Provider web searches cost $0.01 each. In 23 of the 150 Arca runs the model asked a clarifying question instead of answering. It then got one fixed reply asking for a general answer with its assumptions stated, and that answer was scored. A full pass means every critical rubric line was met and no penalty fired. Bars in the chart are 95% bootstrap intervals over the 50 questions.↩
