DeepSeek: DeepSeek V4 Flash Vision Exp

by DeepSeek Intelligent Image Input Multimodal

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents, reasoning, and world knowledge. It is a sparse mixture-of-experts model with 13B active parameters out of 284B total. It is suited for document and chart understanding, visual question answering, and multimodal agent workflows that in

Choose a model to compare against DeepSeek V4 Flash Vision Exp

Specifications

Specifications for DeepSeek V4 Flash Vision Exp
AttributeValue
Lab DeepSeek
Tags Intelligent Image Input Multimodal
Release Date 2026-08
Context Window 1,048,576 tokens
Input Price / 1M $0.22
Output Price / 1M $0.66
Input Modalities Text, Image
Output Modalities Text

Strengths

  • Supports text and image input with text output
  • Million-token context window for very long tasks
  • Low input pricing for cost-sensitive production workloads
  • Tool-use support for agentic workflows

Weaknesses

  • Newer listing with limited independent benchmark coverage
  • Multimodal performance may vary by provider endpoint

Best For

  • Agentic workflows and tool-using applications
  • Multimodal document and screenshot analysis
  • Long-context document and codebase processing

In Depth: DeepSeek V4 Flash Vision Exp

Summary

DeepSeek V4 Flash Vision Exp is an AI model from DeepSeek.

Released 2026-08. It supports Text, Image input and produces Text output, with a context window of 1,048,576 tokens. Input pricing is $0.22 per 1M tokens and output is $0.66 per 1M tokens on OpenRouter.

Sources & Further Reading

Related Models