Fastest DeepSeek V4 Flash 0731

DeepSeek

DeepSeek V4 Flash 0731 is available on Netra Runtime with usage-based token pricing. Input, cached input, and output are priced separately, so you can estimate costs around the prompts you send and the responses your application generates.

Provider benchmarks

Netra reports approximately 380 output tokens per second for this model. The chart also shows Artificial Analysis measurements for four selected providers. Output speed measures how quickly tokens arrive once generation has started; it does not include the wait for the first token.

0 100 200 300 400 Output speed (tokens/s) · Higher is better ≈380* Netra* 288 Baseten 275 Fireworks 221 DeepSeek 119 CoreWeave
Selected providers, not a market-wide ranking. Competitor measurements: reasoning, max effort · 10,000 input tokens. Output speed rounded to whole tokens/s. Source: Artificial Analysis, accessed . Prices and performance vary over time. *Netra: internally reported estimate. Workload and reasoning settings have not been matched to Artificial Analysis; these figures are not a controlled speed comparison.

Pricing comparison

Compare input, cached input, and output rates in USD per 1 million tokens.

Provider pricing in USD per 1 million tokens
Provider Input Cached input Output
Netra $0.20 $0.05 $0.50
Baseten $0.13 $0.028 $0.26
Fireworks $0.22 $0.007 $0.66
DeepSeek $0.44 $0.014 $1.32
CoreWeave $0.14 $0.07 $0.28

Competitor rates are as listed by Artificial Analysis, in USD per 1 million tokens. Provider prices and availability may change; cached input is priced separately. Total cost depends on your token mix. Artificial Analysis marks this model as deprecated and continues benchmarking only the default 10,000-input-token workload.

Source: Artificial Analysis provider pricing. Checked . Prices may change.