VantEdge
Platform/02 · Compute

Your GPU fleet,
scheduled and served.

Serve inference across a multi-backend fleet — hosted providers and your own GPUs — and run training and fine-tuning workloads on infrastructure you control.

Compute · Fleet
5 nodes96% up
gcp-us-central-01H100 ×8Kimi K3
cloud · gcp
95%
onprem-lab-02H100 ×4Qwen3.5-27B
on-prem · lab-a
78%
vpc-aws-east-04A100 ×8Llama-3.3-70B
vpc · aws-us-east-1
62%
onprem-lab-03RTX 6000 ×2embed-Qwen3
on-prem · lab-a
12%
gcp-warm-pool-01H100 ×4warm pool
cloud · gcp · ready in 340ms
idle · ready
34,120tok/s0 queuedautoscaling
01The problem

GPUs are stranded or overpriced

Renting frontier APIs is expensive and sends your data off-site; self-hosting is hard to schedule and scale. Most teams end up locked into one and paying for the trade-off.

02Capabilities

What this layer gives you.

Self-hosted serving

Serve open models on your own GPUs with SGLang and vLLM — tensor-parallel, prefix caching, and tuned KV-cache.

Multi-backend fleet

Blend self-hosted models with hosted providers behind the gateway; route by cost, capability, or data residency.

Elastic scheduling

Workloads scheduled across Kubernetes, with warm pools kept ready for instant assignment.

Train on your accelerators

Bring fine-tuning and training jobs to your own GPUs — no data or weights leave your perimeter.

03How it works
  1. 1Connect your acceleratorsOn-prem, VPC, or cloud GPUs — including RTX PRO 6000-class hardware.
  2. 2Deploy modelsStand up self-hosted open models or point at hosted providers, all behind one gateway.
  3. 3Scale on demandWarm router pools and autoscaling keep latency low without idle spend.
Deployment · kimi-k3
running · v1.2
Config
servingSGLang · tensor-parallel × 8bf16
hardwareH100 ×8 · gcp-us-central-01sovereign
endpoint/kimi-k3 · openai-compatible
tuningprefix cache · KV cache 128 GB · ctx 128K
Performance
last 15 min
throughput4,120 tok/s
latencyp50 28 · p95 142 · p99 210 ms
cache hit
78%
active reqs
27 / 128
Elasticity
warm pool2 nodes idle · ready 340 msready
autoscale+1 @ 85% util (30s) · −1 after 5m idle
SLAburst-first · p95 target 500 ms
every request logged via gateway · exported to your telemetry

Raw compute, in your perimeter

Your GPUs, your data, your weights.

04Where it sits

One integrated stack, GPU to Agent. This layer, in context.

00 · COMPUTEGPU Fleet01 · ACCESSAI Gateway02 · COMPUTEInference & Training Manager03 · KNOWLEDGEContext Enginedet · verified04 · TRUSTDeterministic Execution05 · BUILDAgent Builder
COMPUTE
Serving
SGLang · vLLM
GPUs
on-prem · VPC · cloud
Scheduling
Kubernetes + warm pools
Models
open + hosted
0data leaves your infra·multicloud + on-prem·open+ hosted models
05The rest of the stack