Learn Netra

Everything you need to understand AI inference, optimize performance, and bring your models to production.

STARTER

What is Netra Runtime?

Category, capabilities, pricing, and benchmarked performance in one place.

STARTER

What is LLM inference?

How language models generate output, and why the serving engine matters.

CORE CONCEPTS

Inference latency & throughput

TTFT, time-per-output-token, end-to-end latency, throughput, and concurrency explained.

CORE CONCEPTS

OpenAI-compatible inference API

What it is, why it matters, and which endpoints are supported.

CORE CONCEPTS

Continuous batching

The serving technique that keeps GPUs busy and raises throughput.

CORE CONCEPTS

KV cache

What it stores, the memory trade-off, and why it matters for long context.

CORE CONCEPTS

Model quantization

Lower-precision weights for smaller, faster, cheaper inference.

CORE CONCEPTS

Token streaming

Sending output token by token so responses feel instant.

CORE CONCEPTS

Knowledge distillation

Transferring a large model's knowledge into a smaller, faster one.

CORE CONCEPTS

Paged attention

Virtual memory for the KV cache, and why every modern engine uses it.

CORE CONCEPTS

Prompt caching

Reuse a repeated prompt prefix to cut cost and time-to-first-token.

CORE CONCEPTS

Speculative decoding

Draft with a small model, verify with the large one, same output.

CORE CONCEPTS

Sovereign AI

Control over data, model, compute, and governance, at company scale.

CORE CONCEPTS

GPU spot instances

Spot vs on-demand, preemption, and when cheap interruptible GPUs pay off.

GUIDES

Deploy a fine-tuned LLM

Steps and trade-offs for serving your customized model in production.

GUIDES

Reduce LLM inference cost

Practical levers to produce more tokens per GPU-second.

GUIDES

Serve small models in production

When small fine-tuned models win, and how to deploy them.

GUIDES

LLM inference for AI agents

Why agentic, long-context workloads stress inference, and how Netra serves them.

INTEGRATIONS

Any OpenAI-compatible tool

The three-values pattern behind every integration, plus a curl test.

INTEGRATIONS

OpenClaw

Make a Netra-served model the brain of your personal AI assistant.

INTEGRATIONS

n8n

Power AI Agent and chat model nodes with one credential change.

INTEGRATIONS

Dify

Register your model with the OpenAI-API-compatible provider.

INTEGRATIONS

Hermes Agent

Point Nous Research's agent at your endpoint with one wizard run.

INTEGRATIONS

LangChain

Two ChatOpenAI arguments in Python or JavaScript, nothing else.

INTEGRATIONS

Open WebUI

A private ChatGPT-style UI on top of a model you own.

INTEGRATIONS

LlamaIndex

RAG over your documents with OpenAILike and your endpoint.

INTEGRATIONS

Vercel AI SDK

generateText, streamText, and useChat on your own endpoint.

INTEGRATIONS

CrewAI

Run multi-agent crews on a model served by Netra.

COMPARISONS

Netra Runtime vs vLLM

An honest comparison anchored on a published long-context benchmark.

COMPARISONS

Hosted vs self-hosted inference

Control, cost, latency, and ops — and when each makes sense.

TOOLS & CALCULATORS

LLM memory requirements

How much VRAM an LLM needs: weights, KV cache, and quantization.

TOOLS & CALCULATORS

Fine-tuning memory

GPU memory to fine-tune: weights, gradients, optimizer, and activations.

TOOLS & CALCULATORS

LoRA & QLoRA

Fine-tune with small adapters, and a 4-bit base, to save memory.

TOOLS & CALCULATORS

Tokenization

What a token is, and why counts differ across GPT, Claude, and Llama.

TOOLS & CALCULATORS

OCR

How optical character recognition and AI OCR turn images into text.

TOOLS & CALCULATORS

Image segmentation

Pixel-level masks, Segment Anything, and running it locally.