Fastest DeepSeek V4.1 Flash

DeepSeek

DeepSeek V4.1 Flash is available on Netra Runtime with separate rates for input, cached input, and output tokens. Compare its pricing with the other models in the catalog as you choose a model for your application and plan for usage at scale.

Provider benchmarks

Netra reports approximately 380 output tokens per second for this model. The chart also shows Artificial Analysis measurements for four selected providers. Output speed measures how quickly tokens arrive once generation has started; it does not include the wait for the first token.

0 100 200 300 400 Output speed (tokens/s) · Higher is better ≈380* Netra* 375 Fireworks 239 CoreWeave 233 DeepSeek 194 Baseten
Selected providers, not a market-wide ranking. Competitor measurements: reasoning, max effort · 10,000 input tokens. Output speed rounded to whole tokens/s. Source: Artificial Analysis, accessed . Prices and performance vary over time. *Netra: internally reported estimate. Workload and reasoning settings have not been matched to Artificial Analysis; these figures are not a controlled speed comparison.

Pricing comparison

Compare input, cached input, and output rates in USD per 1 million tokens.

Provider pricing in USD per 1 million tokens
Provider Input Cached input Output
Netra $0.20 $0.05 $0.50
Fireworks $0.22 $0.007 $0.66
CoreWeave $0.20 $0.03 $0.65
DeepSeek $0.30 $0.006 $1.20
Baseten $0.30 $0.03 $1.20

Competitor rates are as listed by Artificial Analysis, in USD per 1 million tokens. Provider prices and availability may change; cached input is priced separately. Total cost depends on your token mix.

Source: Artificial Analysis provider pricing. Checked . Prices may change.