Mobile Reality logoMobile Reality logo

Prompt Caching: A Practical Guide for Claude, OpenAI, Gemini

Prompt caching for Claude, OpenAI, and Gemini visualized as a light beam passing through clear and dark blocks.

Introduction

Prompt caching lets a model provider reuse the work it already did on the start of your prompt, so repeated content such as a system prompt, tool definitions or a long conversation history is billed at a fraction of the normal input price and processed faster. In 2026 every major provider supports it, but each one implements it differently, and the differences decide whether your cache hits or silently misses.

This practical guide is for engineers and CTOs running agents, chat products or document pipelines who want lower cost and latency without changing models. I will explain how prompt caching works, how Claude, OpenAI, Google Gemini, Amazon Bedrock and self-hosted vLLM each handle it, how Claude Code uses it, and how to structure prompts so the cache hits.

You will also find the results of our own experiment: the same 20-turn agent conversation run under different caching strategies, with the cached tokens, cost and latency each one produced. Where we have not measured something ourselves, I say so and point to the provider's documentation.

What Is Prompt Caching?

Prompt caching stores the processed state of a prompt prefix so that later requests starting with the same tokens can skip reprocessing it. The provider bills those cached input tokens at a discounted rate and starts generating sooner, because the expensive first pass over the prompt is mostly done.

How Prompt Caching Works Under the Hood

When a model reads a prompt, every layer computes attention keys and values for each token; this stored state is called the KV cache. Normally the KV cache is thrown away after the request. With prompt caching, the provider keeps it for a while, indexed by the exact token sequence it was built from.

The next request that begins with the same tokens can reuse that state and only compute attention for the new part. Self-hosted engines make the mechanism visible: vLLM's automatic prefix caching documentation describes caching the KV cache of existing queries so a new query sharing the same prefix can reuse it directly. vLLM stores the KV cache in fixed-size blocks and identifies each block by a hash of its tokens and everything before it, which is why a single changed token early in the prompt invalidates every block after it.

Why Only the Prefix Can Be Cached

Attention makes every token depend on all the tokens before it. That means the cached state for token 5,000 is only valid if tokens 1 to 4,999 are identical. There is no way to cache "the middle" of a prompt: the cache always covers a prefix, and it stops at the first difference.

A line of engineered crystalline blocks linked sequentially, showing a sudden structural break where an incompatible link is inserted.
A line of engineered crystalline blocks linked sequentially, showing a sudden structural break where an incompatible link is inserted.

This one rule explains almost every cache miss. A timestamp in the system prompt, a user name inserted near the top, tool definitions in a different order or a re-serialised JSON blob with keys in a new order will each break caching for everything that follows.

Prompt Caching vs Semantic Caching

Semantic caching is a different technique with a similar name. It stores whole answers and returns a stored answer when a new question is similar enough to an old one, usually by comparing embeddings. Prompt caching never skips the model; it only skips reprocessing an identical prefix, and the model still generates a fresh answer.

We have not added semantic caching to the agents we build. For agents that call tools and act on live data, returning a stored answer to a near-match question is a correctness risk, while exact-prefix caching is safe by construction.

What Caching Saves: Cost and Latency

Prompt caching saves money on input tokens and time on the prefill phase. It does not make output tokens cheaper or faster, so its effect depends on how much of each request is a repeated prefix. Long system prompts, many tool definitions, large documents asked about repeatedly and long multi-turn conversations benefit most; short prompts with long answers benefit least.

Prompt Caching Across Providers: How They Differ

The concept is the same everywhere, but the controls, prices and lifetimes differ. This table summarises the main providers as documented in October 2026:

Provider / How you enable it / Cache lifetime / Price of a cache read / Price of a cache write
ProviderHow you enable itCache lifetimePrice of a cache readPrice of a cache write
Anthropic ClaudeExplicit cache_control breakpoints, up to 45 minutes default, 1 hour optional0.1x input on most models1.25x (5 min) or 2x (1 hour)
OpenAI (GPT-5.6 and later)Automatic, optional explicit breakpoints and prompt_cache_keyAt least 30 minutes0.1x input on most models1.25x input
Google GeminiImplicit by default, explicit caches optionalExplicit caches have a TTL you setDiscounted (0.1x on Gemini 3.8 Flash via OpenRouter)Depends on cache type
Amazon BedrockImplicit, or explicit cachePoint checkpointsModel-specific, often 5 minutes or 1 hourModel's cache-read rateModel-specific
Self-hosted vLLMenable_prefix_caching=TrueUntil evicted from GPU memoryFree (your hardware)Free (your hardware)

Two practical consequences follow. First, on Claude you must design breakpoints yourself, and a missing or misplaced breakpoint means no discount at all. Second, cache writes are not free on Claude and on recent OpenAI models, so caching a prefix that is never read again costs more than not caching it.

Claude Prompt Caching: cache_control, Breakpoints and Cache Lifetime

Claude prompt caching is the most explicit of the major providers, which makes it the most controllable and the easiest to get wrong. According to Anthropic's prompt caching documentation, you mark up to four breakpoints per request, and everything before each breakpoint becomes a cacheable prefix.

Explicit cache_control Breakpoints

A breakpoint is a cache_control field on a content block, and Claude allows up to four cache breakpoints per request. Claude processes a request in the order tools, then system prompt, then messages, so a breakpoint on the last system block caches the tool definitions and the system prompt together:

json
{
  "system": [
    {
      "type": "text",
      "text": "You are a document authoring agent...",
      "cache_control": { "type": "ephemeral" }
    }
  ],
  "messages": [
    { "role": "user", "content": "Create a claim intake form." }
  ]
}

On the next request, Claude looks for the longest cached prefix that matches up to each breakpoint. If the tool definitions and system prompt are byte-identical, they are read from cache; the new user message is processed normally.

5-Minute vs 1-Hour Cache Lifetime

The default cache lifetime is five minutes, and every hit refreshes it at no extra charge. You can request a one-hour cache with "ttl": "1h". A 5-minute write costs 1.25 times the base input price, a 1-hour write costs 2 times, and a read costs 0.1 times on most models, with lower read multipliers on Anthropic's newest flagship models.

The right TTL depends on the gap between requests, not on the length of the work. Agent loops that fire calls seconds apart only need five minutes. A chat where a user may come back after a coffee break, or a background agent whose steps take longer than five minutes, needs the hour cache, or every return trip pays the full input price again.

Minimum Prompt Length and Silent Misses

Prompts shorter than a model-specific minimum are not cached, and Anthropic does not return an error when that happens. The minimum ranges from 512 tokens on the newest models to 4,096 tokens on some older ones and on Claude Haiku 4.5. A short system prompt on a model with a high minimum will never cache, no matter how many breakpoints you add.

The only reliable check is the response usage fields. Anthropic reports cachecreationinputtokens and cachereadinputtokens on every response; if both stay at zero, caching is not happening.

Our Claude Cache Policy: Three Breakpoints, Two TTLs

The shared inference library we use across our AI workspace and a client document portal places at most three Claude breakpoints, each with its own TTL:

  1. System prompt, off by default, with a 1-hour TTL when enabled.
  2. Long prefix, on by default, with a 1-hour TTL. It sits ten messages behind the latest one and only moves every six messages, so the marker advances in steps and one long prefix is reused across many turns instead of being rewritten each time.
  3. Recent turn, on by default, with a 5-minute TTL, placed on the last block that can be cached.

During a single agent run, continuation calls happen seconds apart, so the runtime forces the recent-turn marker down to the cheaper 5-minute tier and keeps the hour cache for the first call, whose prefix a later user message may reuse. The library also enforces the rule that a longer-TTL breakpoint must come before a shorter one, never marks thinking blocks, and lets a team cap all markers at five minutes when an hour cache would never pay back. The policy is covered by 41 unit tests, because a breakpoint bug does not throw an error; it only shows up on the bill.

What One Cache Hit Is Worth

Before our larger experiment, we ran a two-call check against Claude Sonnet 5.5 through OpenRouter with our MDMA authoring system prompt, about 5,900 tokens, marked with a one-hour breakpoint. The first call wrote 5,862 tokens to the cache and its prompt cost $0.02347. The second, identical call read the same 5,862 tokens from cache and its prompt cost $0.00119, about 95% less.

The first call is also the warning. A one-hour write costs twice the uncached price, so a prefix written once and never read again doubles its input cost instead of cutting it.

Prompt Caching in Claude Code

Claude Code, Anthropic's coding agent, is one of the heaviest users of prompt caching there is, and we use it for engineering and content work. Each turn re-sends the system prompt, tool definitions, project context and the whole conversation, so without caching a long coding session would reprocess hundreds of thousands of tokens on every message.

How Claude Code Organises the Cache

Claude Code's prompt caching documentation describes three layers ordered from most to least stable: the system prompt with tool definitions, then project context such as CLAUDE.md and memory, then the conversation itself. New content is appended at the end, so on a normal turn the entire previous request is the cached prefix and only the latest exchange is new.

That ordering is the same advice we give for any agent. Put what never changes first, what changes per session next, and what changes every turn last.

A monolithic crystalline bedrock supporting fluid, shifting geometric structures that represent volatile conversational turns.
A monolithic crystalline bedrock supporting fluid, shifting geometric structures that represent volatile conversational turns.

Commands and Actions That Invalidate the Cache

Some actions force a slower, more expensive turn because they change the prefix:

  • Switching models with /model, because each model has its own cache.
  • Changing effort level on most models, and turning on fast mode.
  • Connecting or removing an MCP server when its tools load into the system prompt.
  • Compacting the conversation with /compact, which replaces the history with a summary.
  • Upgrading Claude Code, which usually changes the system prompt for the next new conversation.

Editing CLAUDE.md mid-session does not break the cache, but it also does not apply until /clear, /compact or a restart, because Claude Code loads the file once at session start. Invoking skills and changing permission modes append to the conversation and keep the cache intact.

Session Cache Lifetime and TTL Settings

Claude Code picks the TTL per request. On a Claude subscription within plan usage, the main conversation gets a one-hour cache; with an API key, a cloud provider or usage credits, the default is five minutes. You can choose yourself with the promptCacheTtl setting or the CLAUDECODEPROMPTCACHETTL environment variable, force five minutes with FORCEPROMPTCACHING5M=1, or turn caching off entirely with DISABLEPROMPT_CACHING=1 when debugging.

Checking Cache Performance in a Session

Run /usage to see a "Prompt cache (main)" line with the request count, the share of input tokens served from cache, the number of misses and whether the cache is still warm, plus a likely cause for the last miss when Claude Code can identify one. A status line script can read the same numbers live. A high read-to-write ratio means caching is working; repeated writes turn after turn mean something in your prefix keeps changing.

OpenAI Prompt Caching: Automatic Caching, promptcachekey and Breakpoints

OpenAI prompt caching works without any code changes. OpenAI's prompt caching guide states that caching is enabled by default on supported models, with a minimum cacheable prompt of 1,024 tokens on GPT-5.6 and later, and that a cached prefix stays eligible for reuse for 30 minutes after its last write or reuse.

Automatic Caching on Every Request

Automatic caching places a breakpoint at the end of the latest message for you, so a multi-turn conversation caches its growing history turn after turn. The response usage reports the result in prompttokensdetails.cachedtokens on Chat Completions (or inputtokensdetails.cachedtokens on the Responses API). On GPT-5.6 and later, cache writes are billed at 1.25 times the input rate and reads at a fraction of it.

Routing Requests With promptcachekey

Caches live on specific machines, so two requests with the same prefix only share a cache if they land in the same place. The promptcachekey parameter is a routing hint: requests with the same key are steered towards the same cache. In our shared inference library, we send a per-conversation promptcachekey to the OpenAI API, and the equivalent conversation header to xAI.

Explicit Breakpoints With promptcacheoptions

GPT-5.6 and later models also accept explicit breakpoints. Amazon Bedrock's prompt caching documentation describes the controls: a promptcachebreakpoint on a content block marks the end of a reusable prefix, and promptcacheoptions.mode switches between implicit (an automatic breakpoint on the latest message plus any explicit ones) and explicit (only your breakpoints). Explicit mode is useful in agent loops where you want system instructions and tool definitions cached and nothing else to consume write slots.

Google Gemini: Implicit Caching vs Explicit Caches

Google Gemini offers both models of caching. Google's context caching documentation describes implicit caching as on by default for Gemini 2.5 and newer, with cost savings passed on automatically when a request hits a cache, and sets minimum cacheable sizes of 4,096 tokens for the current Flash models.

Why Implicit Caching Is Not Enough for Us

Implicit caching is best effort: repeating an identical prompt does not guarantee a hit. Our shared inference library therefore uses explicit cached content for Gemini, because it is the only Gemini caching with a guaranteed discount. One shared cache is created per combination of model, system prompt and tools, lives for ten minutes and is recreated on expiry, and any cache failure falls back to a plain uncached request.

The library can also fold the first messages of a conversation into the cache when they are stable, and it skips caching for prompts too short to meet Gemini's minimum. Our experiment below shows how large the gap between implicit and explicit caching can be.

Amazon Bedrock Prompt Caching

Amazon Bedrock supports both implicit caching and explicit cache checkpoints, depending on the model and API. The Bedrock documentation is unusually detailed about the mechanics, and three rules matter most.

Cache Checkpoints in the Converse API

In the Converse API, you add a cachePoint block after the content you want cached, in the tools, system or messages fields. Checkpoints are processed in the order tools, system, messages, and the minimum token count applies to the whole prefix before each checkpoint. Changing an earlier section, such as the tool list, invalidates every cache after it.

Reading the Usage Fields Correctly

With caching on, Bedrock's inputTokens field counts only uncached input. Total input is inputTokens + cacheReadInputTokens + cacheWriteInputTokens, so a dashboard that only reads inputTokens will under-report usage and over-report savings.

Our Experience: Caching on a Document Pipeline

On an invoice and contract processing platform we built for a fintech client, the extraction service adds a Bedrock cache point after the system prompt, because it costs nothing to add. But our own assessment for that client was blunt: each document is processed once, the fixed instructions are a few hundred tokens against several thousand tokens of document text, so caching would save around 5 to 10 percent of input cost at best, and they should expect a few percent. We also found that the open-weight model we tested there rejected cache points, so its runs were always uncached.

Self-Hosted Models: vLLM Automatic Prefix Caching

If you serve your own model, prefix caching is a server setting rather than a pricing feature. vLLM's automatic prefix caching is switched on with enableprefixcaching=True and reuses KV cache blocks for any request that shares a prefix with a recent one.

When Prefix Caching Helps on Your Own Hardware

The vLLM documentation names two cases: long documents queried repeatedly with different questions, and multi-turn conversations whose history is reused across rounds. It also names the limit: prefix caching only speeds up the prefill phase, so it does not help when requests share no prefix or when most of the time is spent generating long answers.

On your own GPUs there is no cache-read discount, because you already pay for the hardware. The benefit is throughput and memory: requests that reuse cached blocks finish prefill faster and leave more GPU capacity for other requests. Our guide to running a self-hosted LLM for business covers the serving side.

Our Experiment: Prompt Caching on a Real Agent Conversation

Provider documentation tells you what caching can do. We wanted to know what it does on a conversation that looks like our own agents, so on 7 October 2026 we measured it.

How We Set It Up

We scripted a 20-turn session with a document authoring agent: a user building a home insurance claim workflow step by step, adding forms, an approval gate, a table and a webhook, then editing and finally printing the whole document. Every request carried our real MDMA authoring system prompt (about 4,000 to 6,000 tokens depending on the tokenizer) plus six tool definitions, and the conversation history grew to over 10,000 tokens by the last turn.

  • Models: openai/gpt-6-sol and google/gemini-3.8-flash, the same models that run the analyst and writer roles in our own AI editor, called through OpenRouter and pinned to one provider each.
  • Strategies: OpenAI with automatic caching only, OpenAI with a per-conversation promptcachekey, Gemini with implicit caching only, Gemini with an explicit cache breakpoint on the system prompt, and Gemini with an explicit breakpoint on the latest message.
  • Runs: every strategy ran the full conversation twice, one conversation at a time, with a unique session marker at the top of the system prompt so no run could reuse another run's cache. That is 200 model calls in total.
  • Measures: cached_tokens and cache-write tokens from the response usage fields, cost as reported by the gateway, and time to first token.

The Results

Input cost below is the measured cost minus output tokens at list price, compared with the same prompt tokens at the full input price, averaged over both runs:

Conceptual rendering of server nodes distributing repetitive prompts across AI models.
Conceptual rendering of server nodes distributing repetitive prompts across AI models.
Strategy / Input served from cache / Turns with a cache hit / Input cost per conversation / Same input at full price / Input saving
StrategyInput served from cacheTurns with a cache hitInput cost per conversationSame input at full priceInput saving
OpenAI, automatic caching91%38 of 40$0.058$0.29380%
OpenAI, with promptcachekey90%38 of 40$0.067$0.32179%
Gemini, implicit caching0%0 of 40$0.119$0.1190%
Gemini, explicit cache on system prompt62%40 of 40$0.051$0.11556%
Gemini, explicit breakpoint on latest message0%0 of 40$0.119$0.1190%

What We Learned

Automatic caching on OpenAI worked out of the box. From the second turn on, every request read the previous turn's prefix from cache, and 91% of all input tokens were served from cache. Because GPT-6 Sol bills cache writes at a premium, the saving on input cost was 80%, not 91%.

promptcachekey made no measurable difference here. In a single conversation sent from one machine, requests already reached the same cache. The key is a routing hint for traffic spread across many users and servers, which this test did not simulate, so we still send it in production.

Gemini's implicit caching never hit. Across 40 turns with an identical, growing prefix, implicit caching served zero cached tokens. This matches what led us to use explicit caches for Gemini in production: implicit caching is best effort, and in our test it delivered nothing.

An explicit cache on the stable prefix cut Gemini input cost by more than half. Caching the system prompt and tools hit on every turn, and 62% of input came from cache. The conversation history was not cached in this setup, which is why the share is lower than on OpenAI.

Our breakpoint on the latest message did not cache on Gemini through the gateway. We do not read this as a rule of the Gemini API; through our gateway, only the cache on the stable leading block took effect. It is a good example of why you verify cached tokens per provider and per gateway rather than trusting that a marker was honoured.

Latency did not improve measurably. With prefixes of 4,000 to 11,000 tokens, time to first token was dominated by generation and, on Gemini, by the model's reasoning before it answers. Caching should matter more for latency on much longer prompts, but at these sizes we saw a cost effect, not a speed effect.

We also planned the same runs on Claude; they did not complete, so the Claude figures in this guide are the two-call check described above.

Prompt Caching Use Cases: Where It Pays Off

Caching pays off wherever many requests share a long, stable beginning. These are the use cases where we see it matter most.

Conversational Agents

Conversational agents re-send the whole history on every turn, so the prefix grows with each message and most of each request has been seen before. This is the case our experiment measured, and it is where automatic caching on OpenAI and well-placed cache breakpoints on Claude save the most.

Coding and Knowledge Work

Coding and knowledge work agents carry large system prompts, many tool definitions and long sessions of file reads and tool results. Claude Code is the reference example: its cache keeps a long session affordable, and its documentation is the clearest public description of which actions break the cache.

Repeated Questions Over the Same Documents

Asking many questions about one contract, one codebase or one knowledge base is the textbook case: put the document first, cache it once, and every later question pays the cache-read rate for it. Large document processing where each document is read only once is the opposite case, covered at the end of this guide.

An intricate three-dimensional matrix of software code bathed in soft laser light during rapid and repeated algorithmic analysis.
An intricate three-dimensional matrix of software code bathed in soft laser light during rapid and repeated algorithmic analysis.

Batch Processing With a Shared Prompt

Batch jobs that run thousands of items through the same long instructions benefit when the items are processed close together in time, inside the cache lifetime. Keep the shared instructions in the prefix, the item at the end, and run the batch in one burst rather than spread across hours.

How to Design Prompts for Cache Hits

Every provider rewards the same prompt structure. These are the practices we apply to every agent we build.

Put Stable Content First

Order every request from least to most volatile: tool definitions, then the system prompt, then stable reference documents, then the conversation, then the new message. Anything that changes per request belongs at the end, never in the middle.

Keep the Prefix Byte-Identical

Treat the prefix as a build artefact. Serialise tool definitions in a fixed order, avoid timestamps, request IDs and user names in the system prompt, and pin the formatting of any JSON you embed. A prefix that is semantically identical but differs by one character is a cache miss.

Keep Tool Definitions Stable for the Whole Session

Adding or removing a tool mid-session changes the start of the prompt and invalidates everything after it. If an agent needs tools on demand, load them as messages later in the conversation, or keep the full tool list from the first request for the whole session, as Claude Code does with deferred tool loading.

Match the Cache Lifetime to the Gaps Between Requests

Use the short TTL for agent loops and batch processing where calls follow each other within seconds, and the long TTL only where users or slow steps create gaps longer than five minutes. Paying for hour-long writes on a burst of calls that finishes in two minutes is wasted money.

Keep a Conversation on One Model and One Provider

Caches are per model and usually per provider. Switching models mid-conversation, or letting a router send consecutive turns to different providers, throws the cache away. Our article on how an LLM router decides which model handles each request covers how to route without breaking caching.

Monitor Cache Hits and Misses

Log cached_tokens and cache-write tokens on every call, and track the share of input served from cache per conversation. Our pricing layer reports how much each cache read saved compared with the base rate, so a sudden drop in that number points straight at a prompt change that broke the prefix.

When Prompt Caching Saves Almost Nothing

Caching is not a universal discount. It saves little or nothing in these cases:

  • Single-use large document processing, where each document is read once and the fixed instructions are small.
  • Short prompts below the minimum, which are never cached and never report an error.
  • Workloads dominated by output, where most of the cost and time is generation, not reading the prompt.
  • Long gaps between requests, where the cache expires before it is read and every write is paid for twice.

Before you tune caching, check what each call sends. On the fintech document project, the biggest saving came from cutting an early version of the pipeline from around 95,000 tokens per document to roughly 3,000, which no cache policy could have matched.

CEO of Mobile Reality

Matt Sadowski

CEO of Mobile Reality

Cut What Your AI Agent Costs per Task

We instrument AI agents to show what every task costs, then lower the bill without lowering quality.

  • Per-call USD pricing rolled up per task and per conversation.

  • Context trimming, so each call carries only what the step needs.

  • Role-based model routing and prompt caching designed for your providers.

  • Fallbacks that never retry errors that will fail again.

  • AI automation measured before and after every change.

Conclusion

Prompt caching is the cheapest performance win available for agents and chat products, but only when the prompt is designed for it and the cache lifetime matches the traffic. Our experiment and production code point to a short list of rules.

  • Prompt caching only reuses an identical prefix: put tools and system prompt first, the conversation last, and keep the prefix byte-identical.
  • On Claude, cache_control breakpoints and the 5-minute or 1-hour TTL are your job; a missing breakpoint means no discount at all.
  • OpenAI caches automatically; a stable promptcachekey per conversation helps requests reach the same cache.
  • Gemini's implicit caching is best effort; use explicit caches when you need a guaranteed discount.
  • Measure cached tokens on every response, because a broken cache never throws an error.

If your agent's input costs keep growing with every turn, we can instrument it, find where the cache misses and restructure the prompts around it, through our AI automation agency services.

Frequently Asked Questions

What is prompt caching?

Prompt caching lets a model provider reuse the processed state of a prompt prefix it has already seen, so repeated content such as a system prompt, tool definitions or conversation history is billed at a discounted rate and processed faster. It only works on an identical prefix: the cache stops at the first token that differs.

How long does the Claude prompt cache last?

By default five minutes, refreshed for free on every cache hit. You can request a one-hour cache lifetime with a ttl of 1h on the cache_control breakpoint, at a higher write price: 1.25 times the input price for a 5-minute write and 2 times for a 1-hour write, with reads at 0.1 times on most models.

Does OpenAI prompt caching work automatically?

Yes. Caching is on by default for supported models with prompts of at least 1,024 tokens. In our 20-turn agent test, 91% of input tokens were served from cache without any configuration, and input cost dropped by about 80%. A prompt_cache_key helps route related requests to the same cache when traffic is spread across many servers.

Why is my prompt cache not hitting?

The most common causes are a prefix that changes between requests (timestamps, user names or reordered tool definitions near the top), a prompt shorter than the provider's minimum, a gap longer than the cache lifetime, or switching models or providers mid-conversation. Check the cached token fields in every response, because a miss never returns an error.

More on AI Cost, Small Models and Evaluation

What an AI system costs, which model it runs on, and how you prove it still works are one decision, not three. These articles cover the build and run cost, model choice and the evaluation that makes switching safe:

Want to know whether a smaller or self-hosted model could handle part of your workload? Our custom AI model development team starts with an evaluation on your own cases.

Did you like the article?Find out how we can help you.

Matt Sadowski

CEO of Mobile Reality

CEO of Mobile Reality

Related articles

What LLM routers are, how a router differs from an AI gateway, and how we route model calls by role, classifier and failure to cut cost without losing quality.

07.10.2026

LLM Router: How We Route Models by Cost, Quality and Failure

What LLM routers are, how a router differs from an AI gateway, and how we route model calls by role, classifier and failure to cut cost without losing quality.

Read full article

A self-hosted LLM guide for business: hardware and VRAM needs, Ollama and Docker for prototypes, production serving, and when owning the model beats APIs.

07.10.2026

Self-Hosted LLM Guide: Hardware, Tools and When It Pays Off

A self-hosted LLM guide for business: hardware and VRAM needs, Ollama and Docker for prototypes, production serving, and when owning the model beats APIs.

Read full article

SLM vs LLM in 2026: key differences in speed, cost and hardware, real use cases, and our benchmark where small language models matched a flagship on one format.

07.10.2026

SLM vs LLM: Key Differences and When Small Models Win

SLM vs LLM in 2026: key differences in speed, cost and hardware, real use cases, and our benchmark where small language models matched a flagship on one format.

Read full article