Deliverable

On-Device LLM Gateway
Local retrieval and compression in front of your cloud LLM API

One appliance on the LAN. It indexes the internal corpus once, retrieves and compresses per request on a local model, and forwards a 4,000-token curated context instead of a 14,400-token pile. Modelled at 63.3% lower monthly cost than cloud-side RAG — with the appliance's own hardware inside that figure.

InquireAll solutions
Why it exists

Three things the gateway changes

The gateway does not make your cloud model smarter. It changes what the cloud model is asked to read.

You pay for what you send

Answer length is a few hundred tokens; prompt length is tens of thousands. At pro-tier list price, unmatched input runs 4.50 CNY per million tokens against 13.50 CNY for output, and cached input is 30× cheaper than uncached. Cutting the prompt is the only lever with this much travel.

Relevance and accuracy move together

If 2,000 tokens are what the task needs and the call carries 14,400, only 25.0% of the context is on point. The gateway lifts that to 50.0% — less spend and a lower chance of the model anchoring on the wrong passage.

The corpus stays inside the LAN

Retrieval, ranking and summarisation run on the appliance. The raw documents, mail and repositories never leave the network — only the curated context does. For teams that cannot ship a whole corpus to an API, this is the difference between using an LLM and not using one.

Internal knowledge baseSharePoint · wiki · repos · mailOn-device gatewayindex · retrieve · compress · freeze prefixCurated context4,000 tokCloud LLM APIindex built once, then amortisednetwork boundaryAnswered on the appliance25% of requestsAnswer

Placement of the gateway. Only the curated context leaves the LAN; the raw corpus never does.

Cost model

What the gateway does to the monthly bill

One team, one knowledge base, one set of questions. 19,800 requests per month at pro-tier pricing; the gateway scheme carries its own hardware and power.

SchemeContext / turnTurnsInput / callInput / monthTotal / month
No orchestration32,000 tok2.270,400 tok1,393.9 M tok6,078 CNY
Cloud-side RAG8,000 tok1.814,400 tok285.1 M tok1,372 CNY
On-device gateway4,000 tok1.45,600 tok83.2 M tok503 CNY

Cloud-side RAG is the baseline we quote against, because it is what most teams already run. Comparing a good design to no design would produce a larger, less honest number.

63.3%lower than cloud-side RAG
869 CNYsaved per month · 10,427 CNY per year
503 CNYtotal per month, appliance included
No orchestration6,078Cloud-side RAG1,372On-device gateway503CNY / month

Total monthly run cost — cloud plus the gateway appliance. The right-hand bar is the on-device scheme.

Where the saving comes from

Removing one mechanism at a time

The four mechanisms interact, so their effects cannot be summed. The meaningful figure is the penalty for removing one while the other three stay on.

MechanismCost if disabledShare of total
Stable prefixes — cache hit 10% → 65%+163 CNY/mo32%
Context compression — 8,000 → 4,000 tok+139 CNY/mo27%
Local offload — 25% resolved on the appliance+112 CNY/mo22%
Turn convergence — 1.8 → 1.4 turns+96 CNY/mo19%

The biggest single lever is not compression — it is cache hit rate. At 32% it outweighs context compression at 27%, because cached input is 30× cheaper than uncached. A prompt prefix that never changes is worth more than a shorter one, which is why prefix stability is specified as a deliverable rather than left as an implementation detail.

Sensitivity

How far the compression has to go

Compression ratio is the assumption most likely to differ in your environment. Here it is scanned rather than asserted.

Context / turnvs cloud-side RAGRelevant shareTotal / monthSaving
2,000 tok4.0×100.0%433 CNY68.4%
3,000 tok2.7×66.7%468 CNY65.9%
4,000 tok — specified2.0×50.0%503 CNY63.3%
6,000 tok1.3×33.3%572 CNY58.3%
8,000 tok — no compression1.0×25.0%642 CNY53.2%
16,000 tok0.5×12.5%920 CNY32.9%

Even with compression switched off entirely, local offload and turn convergence still carry roughly half the saving. The value of the appliance does not rest on aggressive compression alone.

Deliverable

What METO ships

The gateway is a productised configuration of our RK182X edge inference hardware, not a bare board.

Appliance and BSP

RK182X-class edge inference host sized for the workload, with the BSP, thermal validation and enclosure. Throughput reference: 102.01 tok/s decode on a 3B model, 215.86 tok/s on 0.5B for routing alone.

Gateway software

Corpus ingestion and incremental index build, retrieval and re-ranking, context compression, prefix freezing, per-request token accounting and a local-only fallback path.

Integration

An OpenAI-compatible endpoint that sits in front of whichever cloud provider you already use. Existing client code keeps its interface; only the base URL changes.

Measurement pack

The cost model above re-run with your token logs, plus a written baseline of signal-to-noise measured on a sample of your own requests.

Stated honestly: the appliance is not free. 150 CNY/month of amortised hardware and 17 CNY/month of power sit inside the 503 CNY figure. The first full pass over a large corpus is prefill-bound and must be amortised into the index — a configuration that re-embeds the whole corpus per request will lose more in latency than it saves in tokens.

Deployment

Five steps from purchase order to measured savings

#StepOutput
1Token accounting on the current setupSpend attributed per feature, split input / cached / output
2Corpus survey and index designWhat gets indexed, refresh cadence, exclusion list
3Appliance build and indexingGateway running on the LAN, first full index complete
4Retrieval and compression tuningCompressed context still answers the sampled questions correctly
5Cut over and measureBefore / after cost and signal-to-noise on your own traffic

Steps 1 and 2 are done on your side with our model; steps 3 to 5 are ours. We will not quote a saving before step 1 is complete — the number depends on your token logs, not on ours.

FAQ

Questions we get from engineering and procurement

Does the gateway replace our cloud provider?

No. It sits in front of the provider you already use and reduces what is sent to it. You keep the provider relationship, the model choice and the account.

What is the smallest useful deployment?

A single appliance covers a team of the size modelled here (30 people). Below roughly a dozen users the appliance's own amortisation outweighs the token saving, and we would say so.

How much of the work can run locally?

Modelled at 25% of requests answered without a cloud call. That share is a measurement, not a claim — step 1 of the deployment establishes it for your traffic.

What leaves the LAN?

Only the curated context. Documents, mail and repositories are read on the appliance; the raw corpus is never transmitted. The index itself also stays on the appliance.

Which models does it use on-device?

Class and task routing can run on a 215.86 tok/s 0.5B model; retrieval and summarisation on a 102.01 tok/s 3B model, which is the specified configuration. Larger on-device models are available at lower throughput.

What if the compression is worse than modelled?

Then the result is worse, and the sensitivity table above says by how much. At 8,000 tokens of context — no compression at all — the saving still reaches 53.2%.

Related

Related solutions

—

Basis of the figures

[measured] Byte-to-token conversion was measured with a 200k-vocabulary tokeniser on this study's own corpora: 3.0 bytes/token CJK-dominant, 3.2 English with markup, 3.7 English plain text. Our internal memory stack holds 284,222 bytes = 95,303 tokens and injects 8,766 bytes = 3,004 tokens per session — a measured 31.7× compression. [published] Pro-tier list prices, CNY per million tokens, quiet-hours band: unmatched input 4.50, cached input 0.15, output 13.50. On-device decode throughput for 0.5B / 3B / 4B / 8B: 215.86 / 102.01 / 90 / 61.11 tok/s. [assumed] Team size, turn counts, cache hit rates, output length, electricity price and hardware amortisation. [derived] Costs, savings, signal-to-noise and latency. The model claims a compression ratio of 2.0× where our own measurement is 31.7× — roughly a quarter of the observed value.

Want the model re-run on your own token logs?

Send us a month of usage records and the shape of your corpus. We will return the cost model with your numbers in place of our assumptions.

Inquire