Add image understanding to plain text DSH models: register the image_understand tool to call free vision APIs (Qwen/Doubao/Silicon), and include built-in show_image to render images inline in the conversation flow.
- Language
- JavaScript
- License
- MIT
- Branch
- main
Install
$ dsh plugin --profile web add dsh-free-visionRun the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial
Install via your agent
Install the DeepSeek Harness plugin FuzzySoul/dsh-free-vision for me: review the repository at https://github.com/FuzzySoul/dsh-free-vision first, then run the install command and verify the plugin loads successfully.
Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.
One-Line Description
Add "vision" capability to text-only models in DeepSeek Harness: defaults to free vision APIs from Qwen, Doubao, and SiliconFlow providers, converting screenshots, errors, UI, and documents into textual evidence the model can understand, with no manual MCP service configuration required.
Core Features
- Register
image_understandtool: enables pure text models to receive local paths, HTTP(S) URLs, data URIs, or pasted image references, call vision APIs for PNG/JPEG/WebP/GIF, and return textual evidence - One-click recognition (enabled by default): replaces pasted images with cached textual descriptions during message dispatch, allowing the model to answer directly in the same turn without calling the tool first
- Register
show_imagetool: allows the model to render found/generated/captured images as inline cards in the conversation flow; images don't enter model context - Provide settings panel (Settings → Free Vision): form auto-rendered by plugin schema, supports one-click multi-provider switching, individual API Key entry, custom API addresses, saved to
~/.dsh/free-vision.jsonand effective on next call - Provide switching between 6 model providers: Qwen, Doubao, SiliconFlow (OCR specialty) are free by default; Zhipu, Tencent Hunyuan, self-hosted OpenAI-compatible endpoints are pay-per-use or self-configured
Technical Implementation
- Language: JavaScript (ESM,
"type": "module") - Key Dependencies:
@modelcontextprotocol.sdk(1.25.3, as MCP client to connect to luma-mcp),luma-mcp(1.7.1, built-in vision engine for this package),@deepseek-ai/schemastery(config schema validation) - Architecture Pattern: Host injects plugin via
cordis.patch.yml→ plugin spawns luma-mcp subprocess in-process viaapply(ctx, config)→ communicates via stdio MCP → registers general vision tools to host viactx.tools.register→ config save goes through hostwebServerGET/POST/dsh-free-vision/config; config changes trigger engine reconnect - Entry Files:
dsh/index.js(host-side ESM entry, contains schema/routing/tool registration),client/client.js(browser-side UMD bundle, injectssettings.section)
Use Cases
When users run DSH with a pure text DeepSeek model but need the model to "understand" screenshots, errors, UI, documents—for example, pasting a console error image during debugging, a design mockup when writing frontend, or needing OCR on a long screenshot—installing this plugin enables visual Q&A without switching models. Free quotas are sufficient for personal daily use.
Prerequisites & Compatibility
| Dependency | Min Version | Description |
|---|---|---|
| Node.js | >=18 | Declared in package.json:38 |
| DSH | Not declared | Injected via cordis.patch.yml, no explicit version range in package.json |
| OS | Cross-platform | Pure JS implementation, no native modules |
| Native Modules | None | No dependencies on node-pty/node:sqlite or other native extensions |
| Vision API Endpoints | Direct domestic connection | Alibaba Cloud Bailian, Volcano Ark, SiliconFlow are all domestic endpoints; subprocess strips proxy variables |
Installation
dsh plugin --profile web add dsh-free-vision
Configuration Options
| Config | Type | Description | Default |
|---|---|---|---|
| apiKey | String | Current provider's API Key; falls back to corresponding env var (e.g., DASHSCOPE_API_KEY) when empty | Empty |
| keys | Object | Per-provider API Key mapping (e.g., { qwen: 'sk-...' }) | {} |
| baseURLs | Object | Per-provider API address override (e.g., { qwen: 'https://my-proxy.example.com/v1' }); empty uses official default | {} |
| modelProvider | Enum | Provider: qwen (default) / volcengine / siliconflow / zhipu / hunyuan / custom | qwen |
| modelName | String | Optional model name override (e.g., qwen3-vl-flash); default auto-selects based on provider | Empty |
| toolName | String | Public tool name; renamed when conflicting with host's existing tools | image_understand |
| maxTokens | Number | Max generated tokens per vision call | 8192 |
| temperature | Number | Sampling temperature | 0.7 |
| multiCrop | Boolean | Auto multi-crop for large images to improve detail fidelity | true |
| toolCallTimeoutMs | Number | Single call timeout (ms) | 200000 |
| allowedDirs | String | Additional allowed image root directories (; or , separated; defaults to only working dir and user home) | Empty |
| lumaEnv | Object | Extra env vars passed to vision engine (luma-mcp) | {} |
| preservePastedImages | Boolean | Keep pasted images displayed as native thumbnails in conversation (pure text model perspective auto-rewrites to reference text) | true |
| describeAtDispatch | Boolean | Replace images with recognized description text at dispatch time for one-step answering without calling tool | true |
| describePrompt | String | Default prompt for one-step recognition | Chinese detailed description template |
| describeCacheSize | Number | In-process LRU cache size for one-step recognition descriptions (by sha256) | 64 |
| showImageEnabled | Boolean | Enable show_image tool (render model-found/generated images in conversation flow) | true |
| showImageToolName | String | Public name for show_image tool (renamed when conflicting) | show_image |
| showImageMaxBytes | Number | show_image per-image byte limit (also constrained by host attachment limit) | 26214400 |
| showImagePixels | Number | show_image pixel limit (width × height); 0 means unlimited | 40000000 |
FAQ
Q: Is it ready to use after installation?
A: One more step needed: fill in an API Key. The simplest path is to activate Alibaba Cloud Bailian which gives 500k free token quota, select model qwen3-vl-flash, and configure the obtained Key as env var DASHSCOPE_API_KEY (or paste directly in Settings → Free Vision). Saved config takes effect on next call; no restart needed.
Q: Can pure text models see images pasted into the conversation?
A: Yes, and usually no tool call needed. "One-step recognition" is enabled by default (describeAtDispatch): at dispatch time, the plugin first injects the pasted image's description text to the model, which answers directly in the same turn. image_understand remains available for precise follow-ups (e.g., OCR character-by-character transcription).
Q: Which vision models are supported? Are they all free?
A: Qwen Qwen3-VL-Flash (default, 500k free tokens), Doubao Vision (Volcano Ark new user 200k~500k tokens), SiliconFlow DeepSeek-OCR (OCR free) are the three free tiers; Zhipu GLM-4.6V, Tencent Hunyuan HY-Vision, self-hosted OpenAI-compatible endpoints are pay-per-use or self-configured.
Q: Can I use a domestic proxy or self-built gateway instead of official API address?
A: Yes. Switch to the corresponding provider card in settings panel; leave "Base URL / API Address" below empty to use official default; fill in formats like https://my-proxy.example.com/v1 and the engine will automatically splice the path and avoid duplication.
Q: Why must the subprocess "connect directly" and not use a proxy?
A: Alibaba Cloud Bailian, Volcano Ark, and SiliconFlow are all domestic endpoints; the subprocess actively strips proxy env vars like HTTP_PROXY/HTTPS_PROXY/ALL_PROXY/NO_PROXY at startup; using a proxy will result in 502 errors instead.
Q: What if image_understand tool name conflicts with another plugin?
A: Change it to something else in advanced settings via toolName; show_image can also be renamed via showImageToolName; changes take effect immediately after saving.
Q: How to completely uninstall?
A: Config file is at ~/.dsh/free-vision.json (custom path can be overridden by DSH_FREE_VISION_CONFIG_PATH env var); deleting it clears all Keys/address overrides; the plugin itself is uninstalled via dsh plugin --profile web remove dsh-free-vision.
Q: What if the engine process crashes?
A: The plugin has built-in exponential backoff auto-reconnect (dsh/index.js:648-664); it will reconnect and re-register tools after at most 30 seconds; also, the plugin's luma-mcp version is locked at 1.7.1, and the postinstall hook automatically reapplies idempotent patches during upgrades.
Learning Curve
Beginner — install the plugin, configure a free API Key, and the model can call the tool to see images on its own; advanced capabilities (one-step recognition cache, show_image inline cards, API address proxy) all have reasonable defaults, toggle in advanced settings as needed.
Known Issues & Limitations
- The subprocess actively strips eight proxy environment variables:
HTTP_PROXY/HTTPS_PROXY/ALL_PROXY/NO_PROXY/http_proxy/https_proxy/all_proxy/no_proxy(dsh/index.js:121-124); not usable for network environments that require a proxy to connect. - Only supports PNG/JPEG/WebP/GIF four image formats (
dsh/index.js:1046-1053); other formats throwUnsupported image format. - The built-in vision engine has SSRF protection and blocks fetching from loopback addresses like 127.0.0.1 (
dsh/index.js:710-711, 1199); host internal URLs must first go through pasted image reference then transfer; direct splicing is not allowed. show_imagetool explicitly does not support remote HTTP(S) image URLs (dsh/index.js:1517-1518); only accepts local paths, pasted references, or data URIs.- Image reading is restricted by whitelist: only working directory and user home are readable by default; other paths need to be added in
allowedDirsseparated by;or,(dsh/index.js:204-209). - Plugin depends on [email protected]; the
postinstallhook applies idempotent patches to this dependency during upgrades; if upstream luma-mcp changes the corresponding source code string, the patch will fail (scripts/patch-luma.mjs:46-53).
🌐 English | 中文
DSH 免费视觉插件 — 让纯文本模型获得看图能力(截图、报错、UI 分析、OCR、文档),优先使用各平台免费视觉模型,零 MCP 配置。
Free vision plugin for DeepSeek Harness (dsh) — image understanding for text-only models using free-tier vision models, with zero MCP configuration.
为什么免费 / Why free
默认使用免费额度充足的提供商,无账单惊吓:
| 提供商 | 模型 | 免费额度 | API Key 环境变量 |
|---|---|---|---|
| qwen(默认) | Qwen3-VL-Flash | 阿里云百炼限免(激活送 50万 token) | DASHSCOPE_API_KEY |
| volcengine | 豆包视觉模型 | 火山引擎豆包免费 token(20万起,可申请 50万) | VOLCENGINE_API_KEY |
| siliconflow | DeepSeek-OCR | 硅基流动 OCR 免费 | SILICONFLOW_API_KEY |
| zhipu | GLM-4.6V | 按量 | ZHIPU_API_KEY |
| hunyuan | HY-Vision | 按量 | HUNYUAN_API_KEY |
| custom | 任意 OpenAI 兼容 | — | CUSTOM_API_KEY + CUSTOM_BASE_URL + CUSTOM_MODEL_NAME |
一张 1MB 截图 ≈ 2600 token,qwen 限免额度可分析约 19 万张图。 One 1MB screenshot ≈ 2,600 tokens ≈ $0.0006 on qwen; free quota covers ~190,000 images.
特性 / Features
- 零 MCP 配置 — 不用改
cordis.patch.yml、运行时不用npx:视觉引擎(luma-mcp)作为本包依赖内置,进程内启动 - 单个通用工具 —
image_understand(可用config.toolName改名)注册到ctx.tools,每次请求模型都能看到 - 免费优先、多提供商 — 千问 / 豆包 / 硅基流动免费档开箱即用;智谱 / 混元 / custom 可切换
- 每个 Provider 可覆盖 API Base URL — 内置 Provider 可指向代理、API Gateway、本地服务或任意 OpenAI 兼容端点,无需改成 custom
- 直连 — 子进程剥离代理环境变量,国内 API 直连(带代理会导致 502)
- 任务模式 —
auto | general | ocr | ui | debug | describe;大图自动多裁剪保真 - 中英双语 — 工具描述与文档中英文都可用
安装 / Install
dsh plugin --profile web add dsh-free-vision
重启 dsh web 后,工具 image_understand 即可用。
Restart dsh web; the tool appears as image_understand.
设置界面 / Settings UI
重启 dsh web 后,打开 设置 → Free Vision 即可看到配置表单(API Key、提供商、
工具名等),由插件 schema 自动渲染。保存到 ~/.dsh/free-vision.json,下一次
调用立即生效,无需重启。
After restart, open Settings → Free Vision — a form for every config option,
saved to ~/.dsh/free-vision.json, effective on the next tool call.
配置 / Configuration
- id: free-vision
name: 'dsh-free-vision'
config:
apiKey: 'sk-xxxx' # 可选:缺省回退到提供商环境变量
baseURLs: {} # 可选:按 Provider 覆盖 API Base URL,例如 { qwen: 'https://my-proxy.example.com/v1' }
modelProvider: qwen # qwen | volcengine | siliconflow | zhipu | hunyuan | custom
modelName: qwen3-vl-flash # 可选模型覆盖
toolName: image_understand # 工具公开名(冲突时可改名)
maxTokens: 8192
temperature: 0.7
multiCrop: true
toolCallTimeoutMs: 200000
lumaEnv: {} # 传递给视觉引擎的额外环境变量
也可以只设置对应的环境变量(如 DASHSCOPE_API_KEY)。
Or just set the matching environment variable (e.g. DASHSCOPE_API_KEY).
覆盖 API 地址 / Base URL override
当 baseURLs 缺失或值为空时,继续使用该 Provider 的官方默认地址。
| Provider | 默认 Base URL |
|---|---|
| qwen | https://dashscope.aliyuncs.com/compatible-mode/v1 |
| volcengine | https://ark.cn-beijing.volces.com/api/v3 |
| siliconflow | https://api.siliconflow.cn/v1 |
| zhipu | https://open.bigmodel.cn/api/paas/v4 |
| hunyuan | https://api.hunyuan.cloud.tencent.com/v1 |
引擎会自动拼接 /chat/completions 并避免重复路径,因此以下写法都可用:
https://my-proxy.example.com/v1https://my-proxy.example.com/v1/chat/completions
也可以使用环境变量:QWEN_BASE_URL、VOLCENGINE_BASE_URL、
SILICONFLOW_BASE_URL、ZHIPU_BASE_URL、HUNYUAN_BASE_URL
(以及 custom 的 CUSTOM_BASE_URL)。
免费 Key 申请 / Free API keys
| 提供商 | 免费 Key 获取 |
|---|---|
| qwen | 阿里云百炼 bailian.console.aliyun.com — 开通即送免费额度,模型选 qwen3-vl-flash(限免) |
| volcengine | 火山引擎 volcengine.com — 豆包新用户送免费 token(20万起,可申请 50万) |
| siliconflow | 硅基流动 siliconflow.cn — DeepSeek-OCR 免费调用 |
用法 / Usage
模型调用 image_understand 时传入:
image_source(必填):本地路径、HTTP(S) URL 或 data URI(PNG/JPG/WebP/GIF,≤10MB)prompt(必填):对图片的问题 — 中英文均可task_type(可选):auto | general | ocr | ui | debug | describe
工作原理 / How it works
dsh web → cordis 加载 free-vision → 进程内启动视觉引擎(版本锁定)
→ MCP 连接 → 注册 image_understand 到 ctx.tools
→ 模型调用工具 → 引擎预处理(压缩 / 多裁剪)→ 免费视觉 API(直连)
→ 返回文字证据
开发 / Development
npm install
node test-plugin.mjs # 端到端冒烟测试(需要 API Key 环境变量)
许可证 / License
MIT — 封装 luma-mcp(MIT)与 MCP SDK(MIT)。免费额度数据来自各平台官方页面,可能变动,使用前请核实。
Read the usage guide →
Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.
Listing badge
[](https://deepseek-plugin.org/plugins/FuzzySoul/dsh-free-vision)Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.