Add vision and image generation to DeepSeek Harness: paste images for automatic transcription and continue conversation, generate images from descriptions, with fully user-configurable models and backends.
- Language
- TypeScript
- License
- MIT
- Branch
- main
Install
$ dsh plugin --profile web add github:MC5lan/dsh-multimodalRun the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial
Install via your agent
Install the DeepSeek Harness plugin MC5lan/dsh-multimodal for me: review the repository at https://github.com/MC5lan/dsh-multimodal first, then run the install command and verify the plugin loads successfully.
Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.
One-Line Pitch
Give DeepSeek Harness "eyes" and a "pen": after pasting images in the chat, your configured vision model first transcribes the image content verbatim, then passes the text to DeepSeek to continue solving the problem; when the conversation needs image generation, you can call your configured image backend to generate images and display them directly in the conversation.
Core Capabilities
- Automatically enable image attachments on the deepseek-vision conversation thread; paste/drag screenshots for automatic recognition
- Transcribe images into pure text that preserves code, error messages, and UI copy, then let DeepSeek continue solving the problem in the same turn
- Provide
generate_imagetool: enables DeepSeek to automatically call it when an illustration is needed, rendering images directly in conversation cards - Provide
extract_texttool: extract text from images in three formats: Markdown/plain text/JSON - Support any OpenAI-compatible vision endpoint (Zhipu, Alibaba Bailian, iFlytek Spark, Silicon Flow, ModelScope, Ollama, etc.) via direct configuration
- Built-in security policies post-0.2.1: API key whitelist, trusted host list, SSRF protection, sensitive field redaction, audit logs, LRU caching
Technical Implementation
- Language: TypeScript (ESM)
- Key Dependencies: @deepseek-ai/cordis (plugin injection), @deepseek-ai/dsh-llm (register providers and tools), @deepseek-ai/dsh-settings (settings page), eventsource-parser (streaming response parsing)
- Architecture Pattern: Cordis plugin, uses
apply(ctx, config)to registerdeepseek-visionprovider and user-declared extraProviders in thellmsubsystem; also registersgenerate_imageandextract_texttools in thetoolssubsystem; rewrites image messages in agent pre-step hooks - Entry Files: src/index.ts (exports
apply), src/vision.ts (vision transcription hook), src/image-gen.ts (image generation tool), src/ocr.ts (OCR tool)
Use Cases
Users who frequently need to paste error screenshots, design drafts, tables, or chat records into the dialog for the AI to continue fixing, or users who want the AI to directly generate diagrams when writing articles/plans. The plugin doesn't include any models by default; teams can directly connect existing vision/image generation services (company intranet APIs, paid platforms, local Ollama) to avoid data leakage.
Prerequisites & Compatibility
| Dependency | Min Version | Description |
|---|---|---|
| DeepSeek Harness | 0.1.0-rc.6 | Supports both web and headless profile |
| Node.js | 18+ | @types/node ^22.0.0, README recommends 18+ |
| Platform | Cross-platform | No OS restrictions |
| Native modules | None | Only pure JS packages, no node-gyp compilation |
| Runtime dependencies | cordis ^4.0.1, schemastery ^3.18.1, eventsource-parser ^3.1.1 | Declared in plugin package.json |
Installation
dsh plugin --profile web add github:MC5lan/dsh-multimodal
Configuration Options
| Config | Type | Description | Default |
|---|---|---|---|
allowedApiKeyEnvs | string[] | Environment variable/credential names allowed as vision API keys | ["DEEPSEEK_API_KEY"] |
trustedBaseUrls | string[] | Additional allowed vision API baseURLs beyond official hosts (including local Ollama) | [] |
providers.deepseek.baseURL | string | DeepSeek API address | https://api.deepseek.com |
providers.deepseek.apiKeyEnv | string | Environment variable/credential name for DeepSeek API key | DEEPSEEK_API_KEY |
providers.deepseek.models | array | Additional model list displayed under deepseek-vision route | [] |
extraProviders.<key> | object | Any OpenAI-compatible vision endpoint (with displayName/baseURL/apiKeyEnv/models) | {} |
vision.watchProvider | string | Conversation thread to enable "eyes" | deepseek-vision |
vision.transcribeProvider | string | Provider that actually performs vision transcription (empty = disabled) | "" |
vision.transcribeModel | string | Model ID for transcription (empty uses route default) | "" |
vision.fallbackProviders | string[] | Fallback providers to try in order when primary vision provider fails/rate limits | [] |
vision.transcribeMode | string | Transcription mode: auto/structured/ocr/describe/error-fix/chart-sql/design-code | auto |
vision.parallelImages | boolean | Whether to transcribe multiple images in parallel | false |
vision.costProvider/costModel/costMaxPixels | string/number | Cost routing for small images via cheap provider (≤ N pixels) | ""/""/1000000 |
vision.sceneHints | boolean | Whether to append scene hints after screenshot transcription | true |
vision.redactSensitive | boolean | Whether to redact phone numbers/ID cards/emails in transcription results | false |
vision.auditLog | boolean | Whether to write one audit log line per transcription | false |
vision.transcribePrompt | string | Transcription prompt sent to vision model | Built-in "verbatim transcription" long prompt |
vision.transcribeTimeoutMs | number | Single transcription timeout (ms) | 90000 |
ocr.provider/ocr.model | string | Provider and model used by extract_text tool | "" |
image.backends.<key> | object | Image generation backend definition (kind/baseURL/apiKeyEnv/model/defaultSize) | {} |
image.activeBackend | string | Currently active image generation backend key (empty = disabled) | "" |
image.failoverOrder | string[] | Backend key list to failover to in order when primary fails | [] |
image.verifyChineseText | boolean | Whether to use vision model to double-check if Chinese text in generated images is garbled | false |
image.verifyProvider/verifyModel | string | Vision provider and model for verification | "" |
streamIdleTimeoutMs | number | Streaming response idle timeout (ms) | 300000 |
FAQ
Q: Does this plugin come with models?
A: No. The plugin only provides integration logic and UI. Actual vision models (Zhipu, Silicon Flow, ModelScope, iFlytek Spark, local Ollama, etc.) and image generation backends (Alibaba Bailian, OpenAI protocol, custom adapters) need to be declared in the settings page or config file yourself. See extraProviders and image.backends for details.
Q: Will image bytes be uploaded to DeepSeek?
A: No. dsh-multimodal transcribes images into text using your configured vision model before sending requests to DeepSeek. Original image bytes always stay local or with your chosen vision service provider and never enter DeepSeek's context.
Q: Can vision recognition and image generation connect to local Ollama?
A: Yes. Settings → Multimodal → Platform Access has a "Local Ollama" one-click preset that automatically adds http://localhost:11434 to the trusted host list. Local loopback addresses don't require API keys.
Q: How to avoid accidentally sending other keys as vision API keys?
A: Starting from 0.2.1, the plugin maintains an API key whitelist (allowedApiKeyEnvs) and trusted host list (trustedBaseUrls): only environment variable names on the whitelist are allowed to be read as API keys, and base URLs must match known official hosts or the trusted list. Local loopback addresses are separately allowed.
Q: What happens if the backend fails to generate an image?
A: You can configure fallback backend order in image.failoverOrder. When the primary backend request fails, it will failover in order. AUTH errors or user-initiated cancellations don't trigger failover to avoid wasting quota.
Q: What happens when pasting images without a vision provider?
A: A placeholder 【Image transcription failed: reason】 will be injected into the conversation, and DeepSeek will continue solving rather than getting stuck. If you don't want images to appear, just turn off vision.transcribeProvider.
Q: What if there's no "Multimodal" entry in settings?
A: First confirm the plugin is mounted (dsh --profile web --dump-config should list dsh-multimodal), then hard refresh in browser (Ctrl+F5).
Q: How to uninstall/disable this plugin?
A: Use dsh plugin --profile web remove to uninstall (for GitHub install: remove github:MC5lan/dsh-multimodal). Uninstalling doesn't affect the dsh-multimodal section you wrote in ~/.dsh/settings.yaml. It will automatically restore when reinstalled.
Learning Curve
Beginner — Configuration options may seem numerous, but the README provides a one-click "paste key → use" workflow. Platform preset cards and paste-key auto-recognition are enabled by default. When there's no vision model, the conversation won't get stuck.
Known Issues & Limitations
- Custom image backends (
kind: custom) dynamically load user-provided.mjsfiles viaimport(). The file executes with full permissions of the host process. Both README and source code clearly state "only load files you trust" - Remote extra providers must use HTTPS and the host must match official/trusted名单; only local loopback (localhost/127.0.0.1/[::1]) is allowed to use HTTP without API key
- Image generation backends of
kind: custommust explicitly declareadapterFilein the config, otherwisevalidateConfigwill reject loading - When user-declared vision providers don't have
name/description/contextWindow, injection to model directory falls back to 128k context, butmaxTokenshas no default value (to avoid being treated as a real value). It needs to be explicitly provided in config - Vision transcription failures continue the conversation with placeholders rather than retrying. When depending on stable vision providers, it's recommended to configure at least one fallback in
fallbackProviders
English | 简体中文
Give DeepSeek a pair of eyes and a paintbrush: paste a screenshot or image into the conversation and the vision provider you configured first transcribes it verbatim (error messages, code, UI text preserved), then DeepSeek keeps solving your problem — all in the same turn, no extra steps. When an illustration is needed, DeepSeek automatically calls the image backend and the generated pictures appear right in the conversation.
Blank slate by design: this plugin ships no built-in models, providers, or backends. Vision endpoints, image backends, and models are all declared by you — plug in whatever API you already have (DeepSeek, Zhipu, Aliyun, SiliconFlow, ModelScope, Xfyun, Qianfan, local Ollama, …). Nothing is preloaded, nothing is assumed.
Compatibility: built for DeepSeek Harness
0.1.0-rc.6(Web and headless). See CHANGELOG.
Features
| Scenario | Behavior |
|---|---|
| Plain-text chat | Straight to the DeepSeek API (unchanged) |
| Image + question (e.g. error screenshot) | Your configured vision provider "looks" first → transcribes to text → DeepSeek continues from the transcription (fix code, explain, propose); hitting "stop" aborts the vision call immediately |
| Attaching an image | No more "current model does not support images" |
| User asks for an image | DeepSeek calls generate_image → the configured image backend produces pictures shown in the conversation; backend failover tries the next backend if the active one fails (AUTH/aborted skips failover — no wasted quota) |
| Any image API | A custom backend plugs any non-OpenAI/DashScope API in via a small adapter file — no plugin code changes |
| Image card | Dedicated generate_image card: thumbnail grid, click-to-zoom lightbox, one-click download, prompt & model metadata, crop-to-ask (drag a region + ask), copy params (reproducible JSON), retry button on failure (refine-aware) |
| Extracting text from an image | DeepSeek can call extract_text (OCR) → Markdown / plain text / JSON; decoupled from the watch route, usable in any session |
| Paste-key auto-connect (0.7+) | Paste any API key into the quick-add box → platform auto-detected (key fingerprint → /models probing) → endpoint + allow-list + credential + model list + feature enablement in one step |
| Vision platforms | extraProviders accepts any OpenAI-compatible vision endpoint + one-click preset cards (Zhipu, Bailian, Xfyun, ModelScope, SiliconFlow, Qianfan, local Ollama) |
| Transcription cache | Same image + same context reuses the previous transcription — no wasted vision quota (LRU, per-session) |
| Vision fallback chain | Primary vision provider rate-limited/failing → automatic switch to fallbackProviders |
| Parallel transcription | parallelImages transcribes each image in its own concurrent call (fast multi-image turns) |
| Scene modes | Built-in transcribeMode presets: error-fix (error-screenshot diagnosis), chart-sql (chart → SQL + Pandas), design-code (design mockup → HTML+CSS) |
| Cost routing | Small images (≤ costMaxPixels) automatically go to a cheap provider |
| Local vision | One-click Ollama preset keeps sensitive images off the network (see Local vision models) |
| Config migration | Settings page exports/imports the whole config as JSON (allow-listed fields only) |
Security
Since 0.2.1, the plugin hardens itself against malicious or hand-edited configurations:
- API key allow-list (
allowedApiKeyEnvs): only listed env/credential names may be read as an API key — a malicious config can no longer nameGITHUB_TOKEN(or any other secret) asapiKeyEnvto exfiltrate it - Trusted base URLs (
trustedBaseUrls): custom vision providers must target official hosts or hosts you explicitly allow — credentials cannot be silently redirected to an attacker endpoint; explicitly listed local endpoints (e.g.http://localhost:11434) are allowed over plain http and need no key - SSRF guard: generated-image downloads and
reference_imageURLs refuse loopback / private (RFC1918) / link-local addresses - No sessionId forwarding: internal harness session ids are never sent to third-party vision APIs
- Prompt-injection markers: vision & OCR outputs are wrapped in explicit "untrusted context" markers before being handed to the text model
- Sensitive-data redaction (
redactSensitive): phone numbers / 18-digit IDs / emails are masked in transcriptions (including cache hits) - Audit log (
auditLog): one line per transcription with time / image count / bytes / latency / provider - Bounded memory: transcription cache, per-session image recall, and image-kind stats are all LRU-capped — no unbounded growth
Installation
Prerequisites: official DeepSeek Harness (0.1.0-rc.6+) installed and running, Node.js 18+.
# Option 1: install straight from the GitHub repo (requires git)
dsh plugin --profile web add https://github.com/MC5lan/dsh-multimodal
# Option 2: clone and install from a local path
git clone https://github.com/MC5lan/dsh-multimodal.git
dsh plugin --profile web add /path/to/dsh-multimodal
# Also add for headless mode when needed:
dsh plugin --profile headless add /path/to/dsh-multimodal
The repo ships prebuilt artifacts (lib/), so you can install right after cloning; to rebuild from source, see Development.
After restarting dsh web, a 「多模态 / Multimodal」 page appears in the Settings sidebar — everything is configured in one place:
- Status bar (one row): 识图 ✓/✗ · 画图 ✓/✗ — computed from real capability (transcription provider set AND registered / backend exists AND active), so it never shows "configured" when the feature is actually off
- 快速接入 (quick add): the paste-any-key auto-detect box is front and center — paste an API key, the plugin figures out the platform (key fingerprint match first, then
/modelsprobing against known endpoints; key travels one-shot, never stored), and one step later the endpoint is registered, the key env allow-listed, the credential stored, the model list pre-filled, and 识图/画图 enabled. Works with any OpenAI-compatible API. Below it: platform preset cards (Zhipu, Bailian, Xfyun, ModelScope, SiliconFlow, Qianfan, local Ollama) and a custom-endpoint form - 已接入 (connected, open by default): platforms with editable model chips + image-backend management (switch / remove / add form)
- 高级 (advanced, collapsed): the two KeyCards (视觉转写 Key / 图像生成 Key), vision details (watch route, transcription provider, mode, four toggles), and config export/import
Everything in the common path is: paste key → done. The picker then shows
deepseek-vision models (e.g. DeepSeek-V4-Flash (Vision)) — that route's
sessions get the "eye".
Usage
- Fastest path (recommended, 0.7+): Settings → Multimodal → 快速接入 → paste your API key — platform auto-detected and connected in one step (endpoint + allow-list + credential + model list; 识图/画图 enabled automatically). If the key can't be recognized, the hint says why; then pick a platform card or use a custom endpoint.
- Manual path: tap a preset card (Zhipu / Bailian / Xfyun / ModelScope / SiliconFlow / Qianfan / local Ollama) and paste the key, or declare any OpenAI-compatible endpoint in
extraProviders; optionally fillproviders.deepseek.modelsto advertise models on the built-in DeepSeek route slot. - In the Web UI model picker choose a
deepseek-visionmodel — that route's sessions get the "eye". - Chat normally; paste/drag an image and ask to have it read, or just say "draw me a …" to generate one.
API Keys (or write the credentials file directly)
The Settings → Multimodal → 密钥管理 page writes both keys for you. To edit
the credentials file directly (path is $DSH_HOME/.credentials.yaml, default
~/.dsh/.credentials.yaml), it takes effect immediately:
VISION_API_KEY: <your-vision-endpoint-key> # vision transcription (see 视觉转写 Key in the UI)
IMAGE_API_KEY: <your-image-backend-key> # image generation (see 图像生成 Key in the UI)
These are your provider keys — any OpenAI-compatible endpoint you register in 平台接入 accepts its key under the env name you give it (
apiKeyEnv).
How it works
User image ──▶ host admission (model directory declares image modality → allowed, no popup)
──▶ agent/pre-step (plugin): detects images → configured vision engine transcribes
└─▶ image blocks replaced by "【图片内容转写】…" text
──▶ DeepSeek API processes the transcription + your question (image bytes never reach DeepSeek)
- The
deepseek-visionroute is the DeepSeek API with a directory that advertises image support so the UI allows attachments — images are actually transcribed by the provider you configured, never sent to DeepSeek - Image bytes are read through the dsh attachment service; the transcription instruction defaults to "transcribe verbatim + extract key error info" and can be customized in the
dsh-multimodal:section of~/.dsh/settings.yaml - On transcription failure it degrades to a
【图片转写失败: 原因】placeholder — DeepSeek keeps going instead of hanging - Settings UI: Settings → Multimodal (
settings.sectionslot, native look & feel like the Models page)
Configuration (optional, ~/.dsh/settings.yaml)
dsh-multimodal:
# 0.2.1+ security: only these env/credential names may be used as API keys
allowedApiKeyEnvs:
- DEEPSEEK_API_KEY
# - MY_PROVIDER_API_KEY # add your custom provider's key env here
# 0.2.1+ security: extra vision hosts allowed to receive credentials
trustedBaseUrls:
# - https://my-vision.example.com # add your custom provider's host here
providers:
deepseek:
models: [] # optional: advertise models on the DeepSeek route slot
vision:
watchProvider: deepseek-vision # sessions on this route get the "eye"
transcribeProvider: '' # the provider that "sees" (blank = transcription off)
fallbackProviders: [] # 0.2.2+: providers tried in order when the primary rate-limits/fails
transcribeMode: auto # auto | verbatim | structured | ocr | describe | error-fix | chart-sql | design-code
parallelImages: false # 0.2.5+: transcribe each image in its own parallel call
costProvider: '' # 0.2.5+: cheap provider used for small images (cost routing)
costModel: ''
costMaxPixels: 1000000 # images ≤ this many pixels go to costProvider
sceneHints: true # 0.2.5+: append diagnosis hint to screenshot transcriptions
customModes: {} # 0.2.4+: custom mode name → prompt pairs (plugin-in transcribers)
redactSensitive: false # 0.2.4+: mask phone/ID/email in transcription output
auditLog: false # 0.2.4+: log a line per transcription (time/images/bytes/ms/provider)
ocr:
provider: '' # OCR tool engine (blank = disabled; any registered vision provider works)
model: ''
image:
backends: {} # add backends in the UI; nothing is preloaded
activeBackend: '' # image backend key (blank = no image generation)
failoverOrder: [] # 0.2.5+: backend keys tried in order when the active backend fails
verifyChineseText: true # 0.2.3+: check generated Chinese text for garbling via a vision model
verifyProvider: '' # vision provider used for the garbling check (blank = check disabled)
transcribePrompt: ... # custom transcription instruction
transcribeTimeoutMs: 90000 # transcription timeout
# Any OpenAI-compatible vision platform (just add the key; shows up in the model picker & Models page)
extraProviders:
xfyun-vision: # Xfyun MaaS
displayName: 讯飞星辰视觉
baseURL: https://maas-api.cn-huabei-1.xf-yun.com/v2
apiKeyEnv: XFYUN_API_KEY
models:
- id: xoppaddleocrv16 # free OCR (PaddleOCR-VL-1.6)
name: PaddleOCR-VL-1.6
siliconflow-ocr: # SiliconFlow (free hosted DeepSeek-OCR)
displayName: 硅基流动 OCR
baseURL: https://api.siliconflow.cn/v1
apiKeyEnv: SILICONFLOW_API_KEY
models:
- id: deepseek-ai/DeepSeek-OCR
name: DeepSeek-OCR
modelscope-vision: # ModelScope (2000 free calls/day)
displayName: 魔搭视觉
baseURL: https://api-inference.modelscope.cn/v1
apiKeyEnv: MODELSCOPE_API_TOKEN
models:
- id: Qwen/Qwen3-VL-8B-Instruct
name: Qwen3-VL-8B
Image backends (OpenAI-compatible /images/generations and DashScope protocols) example — add them in the UI or write them here:
dsh-multimodal:
image:
backends:
z-image: # Aliyun Z-Image family (new multimodal-generation sync protocol)
kind: dashscope-v2
baseURL: https://dashscope.aliyuncs.com
apiKeyEnv: DASHSCOPE_API_KEY
model: z-image-turbo
defaultSize: 1024*1024 # supports 512*512 ~ 2048*2048
modelscope-t2i:
kind: openai-images
baseURL: https://api-inference.modelscope.cn/v1
apiKeyEnv: MODELSCOPE_API_TOKEN
model: <a lightning-badged text-to-image model in the ModelScope plaza>
activeBackend: z-image
Custom image backends (0.2.6+)
Anything that is not OpenAI-compatible or DashScope can be plugged in with a
small adapter file — no plugin code changes. Write an ES module exporting a
default object with a generate() function:
dsh-multimodal:
image:
backends:
my-api:
kind: custom
adapterFile: D:/my-adapters/my-api.mjs # absolute or relative path
baseURL: https://api.example.com/v1 # passed to the adapter
apiKeyEnv: MY_API_KEY
model: my-image-model
defaultSize: 1024*1024
activeBackend: my-api
failoverOrder: [z-image] # fall back to another configured backend
The adapter receives { prompt, size, n, negative_prompt, reference_image, apiKey, baseURL, model, signal, fetch, log } and returns
{ urls: string[], b64s: string[] }. See
scripts/adapters/example-custom.mjs
for a template. Only point adapterFile at files you trust — the adapter
runs with full process privileges. Polling backends can tune
pollIntervalMs / pollTimeoutMs per backend.
Local vision models (keep sensitive images on your machine)
The plugin routes transcription through any OpenAI-compatible endpoint, so a local
Ollama vision model keeps image data off the network. One-click
preset: Settings → Multimodal → 平台接入 → + 本地 Ollama (auto-adds
http://localhost:11434/v1 to trustedBaseUrls; local endpoints need no API key).
Model matrix (pick by VRAM):
| Model | ollama pull | VRAM | Notes |
|---|---|---|---|
llava | ollama pull llava | ~4 GB | Classic 7B, good OCR-ish |
llava-llama3 | ollama pull llava-llama3 | ~6 GB | Stronger baseline |
qwen2.5vl:7b | ollama pull qwen2.5vl:7b | ~6 GB | Better Chinese text |
minicpm-v | ollama pull minicpm-v | ~6 GB | Document/OCR focused |
qwen2.5vl:32b | ollama pull qwen2.5vl:32b | ~20 GB | High quality, big VRAM |
Compare providers before wiring them in:
node scripts/benchmark-vision.mjs shot.png \
my-endpoint=https://your-vision-api.example.com/v1,your-model,YOUR_API_KEY_ENV \
ollama=http://localhost:11434/v1,llava,SKIP
Troubleshooting
| Symptom | Cause & fix |
|---|---|
| Reply says "【图片转写失败: MISSING_CREDENTIAL…】" | Vision key missing: add it under Settings → Multimodal → 密钥管理, or write the credentials file |
| Reply says "【图片转写失败: RATE_LIMIT…】" | Free-tier rate limit; retry later or switch to a paid tier |
| Transcription is off (images ignored) | vision.transcribeProvider is blank — set it to a registered provider in Settings → Multimodal → 视觉转写 |
| No "Multimodal" entry in Settings | Verify the plugin is mounted (dsh --profile web --dump-config should list dsh-multimodal) and hard-refresh the browser (Ctrl+F5) |
| Custom vision provider skipped with a "not trusted" warning | Its base URL isn't allow-listed: add it to trustedBaseUrls (or use one of the official hosts); also make sure its key env is in allowedApiKeyEnvs |
Development
npm install
npm run build # host: tsc; client: tsc check + tsdown builds lib/client.js
License
Read the usage guide →
Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.
Listing badge
[](https://deepseek-plugin.org/plugins/MC5lan/dsh-multimodal)Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.