Learn Netra
Everything you need to understand AI inference, optimize performance, and bring your models to production.
STARTER
What is Netra Runtime?
Category, capabilities, pricing, and benchmarked performance in one place.
STARTER
What is LLM inference?
How language models generate output, and why the serving engine matters.
CORE CONCEPTS
Inference latency & throughput
TTFT, time-per-output-token, end-to-end latency, throughput, and concurrency explained.
CORE CONCEPTS
OpenAI-compatible inference API
What it is, why it matters, and which endpoints are supported.
CORE CONCEPTS
Continuous batching
The serving technique that keeps GPUs busy and raises throughput.
CORE CONCEPTS
KV cache
What it stores, the memory trade-off, and why it matters for long context.
CORE CONCEPTS
Model quantization
Lower-precision weights for smaller, faster, cheaper inference.
CORE CONCEPTS
Token streaming
Sending output token by token so responses feel instant.
CORE CONCEPTS
Knowledge distillation
Transferring a large model's knowledge into a smaller, faster one.
CORE CONCEPTS
Paged attention
Virtual memory for the KV cache, and why every modern engine uses it.
CORE CONCEPTS
Prompt caching
Reuse a repeated prompt prefix to cut cost and time-to-first-token.
CORE CONCEPTS
Speculative decoding
Draft with a small model, verify with the large one, same output.
CORE CONCEPTS
Sovereign AI
Control over data, model, compute, and governance, at company scale.
CORE CONCEPTS
GPU spot instances
Spot vs on-demand, preemption, and when cheap interruptible GPUs pay off.
GUIDES
Deploy a fine-tuned LLM
Steps and trade-offs for serving your customized model in production.
GUIDES
Reduce LLM inference cost
Practical levers to produce more tokens per GPU-second.
GUIDES
Serve small models in production
When small fine-tuned models win, and how to deploy them.
GUIDES
LLM inference for AI agents
Why agentic, long-context workloads stress inference, and how Netra serves them.
INTEGRATIONS
Any OpenAI-compatible tool
The three-values pattern behind every integration, plus a curl test.
INTEGRATIONS
OpenClaw
Make a Netra-served model the brain of your personal AI assistant.
INTEGRATIONS
n8n
Power AI Agent and chat model nodes with one credential change.
INTEGRATIONS
Dify
Register your model with the OpenAI-API-compatible provider.
INTEGRATIONS
Hermes Agent
Point Nous Research's agent at your endpoint with one wizard run.
INTEGRATIONS
LangChain
Two ChatOpenAI arguments in Python or JavaScript, nothing else.
INTEGRATIONS
Open WebUI
A private ChatGPT-style UI on top of a model you own.
INTEGRATIONS
LlamaIndex
RAG over your documents with OpenAILike and your endpoint.
INTEGRATIONS
Vercel AI SDK
generateText, streamText, and useChat on your own endpoint.
INTEGRATIONS
CrewAI
Run multi-agent crews on a model served by Netra.
COMPARISONS
Netra Runtime vs vLLM
An honest comparison anchored on a published long-context benchmark.
COMPARISONS
Hosted vs self-hosted inference
Control, cost, latency, and ops — and when each makes sense.
TOOLS & CALCULATORS
LLM memory requirements
How much VRAM an LLM needs: weights, KV cache, and quantization.
TOOLS & CALCULATORS
Fine-tuning memory
GPU memory to fine-tune: weights, gradients, optimizer, and activations.
TOOLS & CALCULATORS
LoRA & QLoRA
Fine-tune with small adapters, and a 4-bit base, to save memory.
TOOLS & CALCULATORS
Tokenization
What a token is, and why counts differ across GPT, Claude, and Llama.
TOOLS & CALCULATORS
OCR
How optical character recognition and AI OCR turn images into text.
TOOLS & CALCULATORS
Image segmentation
Pixel-level masks, Segment Anything, and running it locally.