Fastest DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is available on Netra Runtime with usage-based token pricing. Input, cached input, and output are priced separately, so you can estimate costs around the prompts you send and the responses your application generates.
Provider benchmarks
Netra reports approximately 380 output tokens per second for this model. The chart also shows Artificial Analysis measurements for four selected providers. Output speed measures how quickly tokens arrive once generation has started; it does not include the wait for the first token.
Pricing comparison
Compare input, cached input, and output rates in USD per 1 million tokens.
| Provider | Input | Cached input | Output |
|---|---|---|---|
| Netra | $0.20 | $0.05 | $0.50 |
| Baseten | $0.13 | $0.028 | $0.26 |
| Fireworks | $0.22 | $0.007 | $0.66 |
| DeepSeek | $0.44 | $0.014 | $1.32 |
| CoreWeave | $0.14 | $0.07 | $0.28 |
Competitor rates are as listed by Artificial Analysis, in USD per 1 million tokens. Provider prices and availability may change; cached input is priced separately. Total cost depends on your token mix. Artificial Analysis marks this model as deprecated and continues benchmarking only the default 10,000-input-token workload.
Source: Artificial Analysis provider pricing. Checked . Prices may change.