Under a millisecond: The Engineering behind Mesh


Mesh: A Lightweight Gateway for Attributing and Governing LLM Spend


By Aditya Raj, Engineer, Bold AI


Earlier this week Alan wrote about why we built Mesh and why we are giving it away. This is the other half of that story: how Mesh is built, and what it measured like before we let anyone else run it.


I wrote it up as a short paper rather than a changelog, on purpose. A paper forces you to state the setup, report the numbers you got rather than the ones you hoped for, and put the limitations in the same document as the results. That discipline is the point. Everything below comes from that paper, and the paper itself ships in the repo as a PDF.


The paper as written. Five pages: architecture, evaluation over about 13.2 million requests, discussion and limitations.



Results at a glance

  • Overhead per request: +0.79 ms at p50 and +0.77 ms at p99 against a direct call on the same wire shape

  • Streaming: passed through delta by delta; time to first byte 6.2 ms on a stream that takes 17.5 ms end to end

  • Raw capacity on a shared laptop: about 50k authenticated and 117k unauthenticated requests per second

  • Connections: 5,000 concurrent clients held with zero connection errors

  • Stability: 25,200 requests at 140 rps, 100% success, no leaks

  • Sample: about 13.2 million requests against emulated providers, all on one laptop



The problem with provider billing


Provider billing tells you cost per account and per provider. The questions an organisation actually asks are indexed by person, product and model, and they span providers. Sharing one provider key across teams destroys attribution, turns revocation into a coordinated outage, and leaves nowhere to enforce a budget or restrict a model.


Mesh moves identity, policy and accounting into a single hop in front of every provider. Callers authenticate with keys that Mesh issues (they start with mesh_live_), and provider credentials never leave the gateway. I set four constraints before writing code: under a millisecond of added latency, no database round trip on the request path, streaming passed through untouched, and one inexpensive process.



How Mesh is built: two planes, one process


Mesh is a single Go 1.23 binary, about 9.3k lines of code with 8 direct dependencies. Two planes share one process and one Postgres database.

  • Control plane (/admin/*): providers, models, users, keys, access grants, policies, logs and a kill switch. Postgres is the source of truth. It runs under an admin database role with full write access.

  • Data plane (/v1/*): proxy, embeddings, generate, models and usage. Its state is an immutable in memory snapshot plus atomic counters. It runs under a gateway role that can read configuration and append logs and counters, and nothing else. It does no database I/O per request; telemetry is asynchronous and batched.


Provider API keys are sealed with AES-256-GCM under a 32 byte master key and decrypted only while a snapshot is being built. Gateway keys are 256 bit random tokens stored only as a peppered SHA-256 hash with a six character display prefix. Administrator credentials use bcrypt, hashed session tokens and a 5 s in process session cache. Privilege is split across two database roles with row level security forced on: the gateway role can read configuration and append logs and counters but cannot change configuration, while log retention, executed in bounded chunks with a 30 day default, runs only under the administrative role.


Configuration reaches the data plane by rebuilding the snapshot. Every administrative write triggers a rebuild, and concurrent writes are coalesced through a monotonically increasing sequence number, so one rebuild absorbs all pending changes. A 45 s reconcile loop bounds staleness for edits made outside the API. Policy is scoped per principal and model: a daily USD cap (default $20), an hourly call limit and an allowed window over weekdays in an IANA timezone. Per provider: maximum outbound rps, maximum concurrent upstream calls, an hourly quota and a retry policy. A global kill switch suspends all traffic.


Providers are data, not code. Each is described by an endpoint, header and body JSON templates with placeholders, and dot paths that locate the streamed delta, token usage and completion signal. Five wire families (Anthropic, OpenAI Chat and Responses, Gemini, Ollama) plus a generic family translate typed content blocks (text, image, document, audio, video) into provider native shapes.


Attribution is one request_logs row per call, denials included: key, principal, provider, model, outcome, deny reason, status, latency, input and output tokens, cost at NUMERIC(16,9) precision, and client device. Composite indices on (user, model, created_at), (provider, created_at) and (key, created_at) turn the dashboard's per person, per product and per model aggregations into index range scans.



The nine steps of a request, and what keeps each one cheap


Every POST /v1/proxy goes through the same nine steps. Steps 1 to 4 touch only memory; the first I/O is the upstream call.

  1. Authenticate. SHA-256 of the bearer token, then a 32 byte lookup in the key index map. The snapshot is read through an atomic pointer: lock free, RCU style.

  2. Authorize. Per key, pre joined maps from model name to model access: provider config, prices, policy and parsed timezone.

  3. Policy. A sync.Map of atomic 64 bit integers: hourly calls and daily spend in integer nano USD, with the time window precomputed as minutes.

  4. Admission. A weighted semaphore of 200 units. Cost is content weight (text 1, image 5, document 8, audio 10, video 25) times a load factor of 1, 2 or 4 taken from goroutine and heap pressure polled every 3 s. A 4 s timeout returns 503 server_busy.

  5. Provider limiter. A token bucket (rps, with burst equal to rps) plus a slot channel for maximum concurrent upstream calls. A 4 s timeout returns 503 provider_busy.

  6. Render. Regex placeholder substitution with JSON escaped values; the pre built content JSON is spliced in raw.

  7. Dispatch. One shared HTTP transport: keep alive with 256 idle connections per host, HTTP/2, 64 KB buffers, and SSRF safe redirects.

  8. Stream. A 64 KB buffered SSE reader; dot paths pre split when the snapshot is built; a hand rolled JSON string escaper writing into a reused buffer; a flush per delta; a 300 s idle timer.

  9. Account. Cost equals tokens times price; an atomic add to the spend counter; a non blocking enqueue of the log event.


All persistence is deferred off the critical path. Log events enter two bounded channels of 4,096 entries, denials drained first, and two workers write them as pgx batches of 500 rows or once per second. Enqueueing never blocks, so database latency cannot reach callers. Spend counters are reconciled to Postgres every 30 s through batched upserts, hourly counters are checkpointed and reset at hour boundaries, and every counter is re seeded from Postgres at start up. Every background loop runs under a supervisor that recovers panics and restarts it.



How it was measured


One laptop: a Go load generator, Mesh (a dev build under air, no PGO), and testserve, a dummy Anthropic, OpenAI, Gemini and Ollama upstream on one uvicorn worker. Intel Core Ultra 7 255H, 16 threads, 22.9 GB RAM. Everything shared the CPU, so the capacity figures are lower bounds. About 13.2 million requests in total, with the configuration unchanged throughout: max_outbound_rps 50, max_concurrent_upstream 20, an admission queue of 200 units, a 4 s admit timeout, and request logging off on the test providers.



What Mesh adds per request


Figure 1. (a) All endpoints at 20 rps, n=300, 100% success; grey is the direct to provider baseline. (b) Split on a shared clock: phase 1 is auth, policy, admission, limiter, render and dispatch; phase 2 is the upstream call plus SSE to NDJSON translation.


A direct call on the same wire shape measures 1.40 / 2.60 / 3.20 ms at p50 / p95 / p99. Through /v1/proxy it measures 2.19 / 3.68 / 3.97 ms: an overhead of +0.79 / +1.08 / +0.77 ms. For embeddings the p50 overhead is +0.65 ms. The two phases are symmetric at about 1 ms each (p50 0.99 / 0.99 ms; p99 1.98 / 2.49 ms at 40 rps). Routes that never touch an upstream (/v1/models, /v1/usage) stay under 1.5 ms at p99. On ts-huge, a stream of about 200 deltas, time to first byte is 6.2 ms against 17.5 ms total, which is what pass through streaming should look like.



What happens under load


Figure 2. Closed loop (N users) and open loop (fixed arrival rate). Goodput counts 2xx only. Red dash dot line: the 8 s worst case wait (4 s admission plus 4 s limiter). Grey dotted line in (a): the theoretical N/50 s limiter latency.


Latency tracks N/50 seconds exactly (Figure 2a): all queueing happens in the provider limiter, before dispatch. At 50 users, phase 1 is 998.6 ms while phase 2 is 1.14 ms. Throughput is set by the provider limit, about 53 rps per provider; the mixed workload caps at 150 to 170 rps because 33% of it hits one provider (50 / 0.33 is about 150). The open loop knee is sharp: about 2 ms up to 50 rps, then backlog growth. Past the 8 s ceiling, Mesh sheds with 503 instead of queueing without bound.



Raw capacity and resources


Peak raw throughput and connection scale, measured:

  • GET /healthz, 100 users: 116,710 rps, p99 4.35 ms. The middleware floor.

  • GET /v1/models with key auth, 1,000 users: 50,395 rps, p99 139 ms. About 0.17 ms of CPU per request for hash, lookup and JSON.

  • Invalid key to 401, 500 users: 96,640 rps. Rejection p99 1.16 ms at 20 rps.

  • /healthz with held connections, 5,000 users: 80,507 rps, p99 306 ms. 7,363 sockets, 0 errors.

  • /v1/models with held connections, 5,000 users: 42,398 rps, p99 905 ms. 7,753 goroutines, 398 MB RSS.


Figure 3. Paths without a provider limiter; the load generator shares the CPU. 0 failures in every run.


Each held connection costs about 1 to 1.5 goroutines and about 50 KB of RSS. Idle footprint is 25 to 28 MB, 21 threads and about 18 goroutines, and Mesh never exceeded 28 OS threads. Payload latency grows about 17 to 20 ms per MB (1 MB of text, p50 18.9 ms; a 5 MB image, p50 100.1 ms). At 5,000 users goroutines crossed 90% of the goroutine ceiling, the load factor stepped to 4x, and admission throttled itself as designed.


Figure 4. (a, b) Peak goroutines and RSS per run against open client sockets. (c) Payload size, ts-claude, 5 rps.



What a dedicated host would do
(a projection, not a measurement)


Extrapolating the measured per connection cost (about 50 KB RSS and 1 to 1.5 goroutines per held connection) and per request CPU to a dedicated 24 GB RAM CPU host, one Mesh instance projects to about 16,000 users streaming at the same moment, or about 100,000 active chat users at one message per minute, and could keep about 130,000 to 250,000 connections open. These are projections. They assume provider limits and the queue ceilings are raised accordingly, because with the defaults the provider token bucket and the 5,000 goroutine ceiling bind first.



Stability and correctness


Figure 5. (a) Time to first byte against last byte per stream shape. (b) Rejection latency. (c) A new TCP connection costs +0.5 to 0.8 ms at p50.

  • Soak, mixed traffic at 140 rps for 180 s: 25,200 requests, 100% success, p50 1.65 ms, p99 3.75 ms, max 6.8 ms. RSS went 93.8 MB to 94.0 MB to 54.9 MB, goroutines peaked at 22, no leaks.

  • Soak at 200 rps for 120 s, over capacity: 84.1% success, the rest shed as 503. RSS went 64 MB to 109 MB to 94 MB.

  • Upstream correctness, n=1,150: 0 keys in any body, 0 unrendered placeholders, correct model rewrite, and every 2xx stream ended with done plus usage.

  • Telemetry: 132,611 rows written, 0 insert failures, database pool 3 of 16 in use.



Why the overhead is low


The roughly 1 ms per direction follows from four choices.


First, the request path is free of I/O until the upstream dial. Configuration is one immutable snapshot published through an atomic pointer, and all mutable state is atomic integers reconciled asynchronously.


Second, authentication and authorization reduce to one SHA-256 evaluation and two hash map lookups over pre joined structures, about 0.17 ms of CPU over a health check.


Third, back pressure is explicit and bounded at every stage: a weighted semaphore, a token bucket, a concurrency slot channel and fixed timeouts. Overload shows up as fast 503 responses and flat latency for admitted work, not as unbounded queues and memory growth.


Fourth, streaming is incremental: 64 KB buffered reads, dot paths pre split at snapshot build, an allocation light NDJSON encoder and a flush per delta keep time to first byte independent of response length, while pooled keep alive upstream connections avoid the measured 0.5 to 0.8 ms cost of a new connection.



Limitations


Mesh runs as a single instance against a single Postgres database, with no clustering, sharding or multi region support, so availability and capacity are bounded by one host.


All numbers come from one laptop where the load generator, Mesh, the dummy upstream and a dev frontend shared the CPU, using an unoptimised dev build and dummy providers rather than real ones. Production latency will be dominated by the provider, and absolute capacity may differ.


Retries cover transport errors only. Upstream HTTP 500 and 429 are passed through as 502 after one attempt, and Retry-After is not honoured.


The admission queue is global, so one saturated provider can cause 503s for providers that still have headroom, and the worst case wait before rejection is about 8 s because the same 4 s timeout applies twice.


Budgets are checked before a call and charged after it, so concurrent in flight calls can overshoot a daily cap. Spend counters live in memory between 30 s checkpoints, so a crash can lose up to 30 s of accounting.


Telemetry is lossy by design under floods. During bad key floods, 4.19 million denial events were dropped, so attacks are under counted in the logs.


Memory scales with open connections and in flight payload, and the goroutine ceiling is the first protection limit to engage.



Run it, break it, and tell us


The code, the docs, the SDKs and the test server that produced these numbers are all in the repo: github.com/Bold-AI-Inc/meshcore. The paper itself, Mesh: A Lightweight Gateway for Attributing and Governing LLM Spend, ships in the repo as a PDF.


If your numbers differ from mine, I want to know. And if you can make a step on the request path cheaper, one concern per pull request with a before and after comparison is how we take changes.


Aditya Raj is an engineer at Bold AI and the author of Mesh.



Book a 15 Min Meeting | Bold AI