ArcenArcen

Whitepaper · v0.1 · July 2026

The economics of machine-speed inference.

Two advantages, one stack: a serving plane we own end-to-end and can optimize at every layer, and a payment rail built for machines, not checkout pages.

01

The problem: the margin stack

A typical inference call doesn't buy compute. It buys a supply chain. The request passes through a gateway, a reseller, a model host, and a payment processor, and every hop prices in its own cut: token markups, platform fees, subscriptions, credit-pack float. None of those margins buy you a single FLOP. They pay for the fact that the stack is split across four companies.

Worse: none of those hops can make the inference itself any better. A reseller can't reschedule a GPU. Optimization lives in the serving plane, and you only control the serving plane if you run it.

02

One stack, one price sheet

Arcen collapses the supply chain: compute, gateway, metering, and payment are one integrated stack run by one operator. So the bill has exactly one line:

call cost = (prompt tokens × model input rate) + (output tokens × model output rate) // no per-call fee. no gas. no subscription. nothing else.

The token rates, what the catalog lists, are the price of the compute itself. They vary per model because a 400B-parameter model burns more silicon time than a 20B one. What they don't contain is anyone else's margin: no gateway cut, no reseller cut, no platform percentage, no per-call surcharge. A compute provider has one thing to sell, silicon time, and one number to publish for it.

03

The serving stack: why we can optimize inference

Owning the racks means owning the scheduler. Every technique below is standard, hard-won systems engineering. The point is that only the operator of the serving plane gets to apply them, and their effects compound:

A

Continuous batching

Requests join a running batch mid-flight instead of waiting for a free slot, so the GPUs never idle between calls. This is the single biggest utilization win in modern serving: more tokens per GPU-second, cheaper tokens for everyone in the batch.

B

Paged KV-cache

Attention memory is paged like virtual memory, so long contexts don't strand VRAM. Less memory waste means a bigger effective batch on the same card: throughput without new hardware.

C

Prefix caching

Agents resend the same system prompt and tool schemas thousands of times a day. The cluster computes that prefix once and reuses it across calls, so the most expensive part of an agent's request becomes nearly free on repeat.

D

Quantized + speculative serving

Open weights served quantized (FP8-class) with a small draft model speculating ahead of the big one. Same output, materially more tokens per second per card: headroom we can pass into the price.

E

Exact-response cache

Identical repeat calls (retries, evals, agent loops) are served from the response cache at half-price tokens and ~10ms, per-tenant isolated. This one is a billing rule, not just an optimization: the discount is posted.

F

Tier routing

The catalog is tiered High / Balanced / Eco so agents buy exactly the intelligence a task needs. The cheapest token is the frontier token you didn't need: routing classification traffic to Eco is a 10–50× saving before any hardware trick.

04

Elastic compute: capacity that follows demand

We don't keep every model's GPUs provisioned around the clock waiting for the next call, that's capital sitting idle for traffic that hasn't arrived yet. Instead, a routing layer decides, per request, where the call actually runs:

1

Request arrives

Every call hits the response cache first (Mechanism E). An exact repeat is served in ~10ms at half-price tokens, no routing decision needed.

cache miss
2

Is this model's GPU already warm?

A model getting steady traffic stays provisioned on our own racks. One that's been idle scales back down, we don't pay standby cost for capacity nobody is calling.

3a

Cold: passthrough covers the gap

The call routes through the same passthrough tier that serves our labeled 3rd Party Provider models, for the seconds it takes our GPU to boot. Nothing waits on us to spin up.

3b

Warm: served on our own compute

The call runs on Arcen's own GPU, the open weights served with NVIDIA NIM under our NVIDIA AI Enterprise subscription. Every following call for that model stays on our own compute until it goes idle again.

StateServed byWhat you see
Cache hitResponse cache, per-tenant~10ms · half tokens
Cold (GPU idle)Passthrough tierlabeled per call
Warm (GPU promoted)Arcen's own compute"Arcen Compute"

The label always matches who actually served the call. "Arcen Compute" means it ran on our own racks that specific time, not that the model is nominally in our catalog, same honesty rule as the 3rd Party Provider tag on the models page. We'd rather show a passthrough label than a false one.

05

The payment rail: one asset, zero gas

The second advantage is that paying Arcen was designed for machines. Your wallet holds one asset, USDG. That's the whole setup:

  • No gas, ever. Settlement rides a gasless payment protocol, you never hold a second token just to pay transaction fees, and no call ever fails because a gas balance ran dry.
  • Fair by construction. Calls are metered at their exact token count and billed against posted rates, no credits, no rounding into bundles, no subscription floor.
  • Pre-authorize, then refund. Accountless calls quote a ceiling up front and automatically refund the difference between the ceiling and the actual metered charge. You can cap any key's lifetime or rolling spend on top.
  • Machine-speed settlement. Each call settles itself on-chain in the background, an agent can transact with Arcen with no human, no card, and no account state to manage.
06

A worked example

Assumptions, stated plainly: 100 calls of 5k input + 20k output tokens on a Balanced-class model whose compute prices at $0.30 / $0.90 per 1M tokens; the comparison route accumulates 25% in intermediary margins on the same compute. Change the assumptions and the gap changes, that's the point of showing them.

Multi-hop route (+25% stacked margins)$2.44

$1.95 compute × 1.25 in accumulated cuts

Arcen: compute, nothing else$1.95

tokens × posted rates. done.

Arcen with 30% cache hit rate$1.66

cached calls bill half tokens

In this shape: 20% saved from pricing structure alone, ~32% once the cache participates, before counting prefix-cache and batching effects, which land in the token rates themselves. Cache-friendly agent loops push the effective gap toward 30% and beyond; unique, one-off calls see the structural 20%. Illustrative, live prices are always on the models page, and the math is yours to check.

07

The flywheel

Cheap inference is a consequence of utilization. Machine traffic is steady , agents don't sleep, and steady load is exactly what continuous batching converts into cheap tokens. Every call raises cluster utilization; utilization drops the unit cost of a token; the price sheet passes it on, which attracts more agents. Half of every call's revenue funds monthly distributions to $ARCEN holders (see Rewards); the other half buys more compute. The network gets cheaper because it gets used.

08

Roadmap

  • Response cache GA, half-price repeat calls, on by default, per-tenant isolated.
  • Prefix-cache-aware pricing, passing cached-prefix savings into the posted input rate for agent-shaped workloads.
  • Prepaid sessions, deposit once, spend from an off-chain balance at ~10ms per call; one settlement covers many calls.
  • Distribution engine, automated monthly holder payouts, pro-rata to supply.
  • Cluster expansion, more owned capacity, more models, same price-sheet discipline.

The thesis

We optimize the silicon.
You hold one asset. Nothing in between.

See the posted prices →