Product C · Inference

A drop-in OpenAI-compatible endpoint.

Deploy any trained model or RAG pipeline as an OpenAI-compatible API — base URL, key, model name, usage dashboard, logs, latency and token metering included.

/v1/chat/completionsStreamingPer-token metering
Integration

Swap one line of code

Change the base URL and key. Your existing OpenAI client keeps working — no new SDK, no rewrite.

310ms
p50 latency
99.95%
uptime
OpenAI
compatible
quickstart.py
from openai import OpenAI
client = OpenAI(
    base_url="https://api.neva.ai/v1",
    api_key="sk-neva-prod-a91f…",
)
resp = client.chat.completions.create(
    model="acme-support-v3",
    messages=[{"role": "user",
               "content": "How do I reset my device?"}],
)
print(resp.choices[0].message.content)
Serving tiers

From dev workloads to private clusters

The inference layer is the recurring product — start shared, grow into dedicated and reserved capacity.

Shared

Small customers and development workloads. Pay per token.

Dedicated

Production and business users. Isolated endpoint, higher limits.

Reserved GPU

Heavy users with predictable demand. Committed capacity.

Enterprise private

Support, SLA, audit logs, and custom terms.

Everything metered

Keys, logs, latency and billing in one place

API keys & scopes

Create, rotate and scope keys per environment.

Token metering

Per-request token counts with included allowance and overage.

Request logs

Inspect every call with latency and status.

Rate limits

Per-key RPS and burst controls.

Audit logs

Enterprise-grade access trails and DPAs.

Billing controls

Monthly minimums, prepaid credits, invoicing.

Deploy once. Bill per token.

Recurring inference is the platform's cash cow — sticky, predictable, and the natural upgrade path from a first RAG project.