Skip to main content

dsh-free-vision

6Stars1Forks0Issues1Watchers

Add image understanding to plain text DSH models: register the image_understand tool to call free vision APIs (Qwen/Doubao/Silicon), and include built-in show_image to render images inline in the conversation flow.

Evidence5/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki
Language
JavaScript
License
MIT
Branch
main
deepseek-harnessdshdsh-pluginfreeimage-understandingvision

Install

cmdweb profile
$ dsh plugin --profile web add dsh-free-vision

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin FuzzySoul/dsh-free-vision for me: review the repository at https://github.com/FuzzySoul/dsh-free-vision first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Line Description

Add "vision" capability to text-only models in DeepSeek Harness: defaults to free vision APIs from Qwen, Doubao, and SiliconFlow providers, converting screenshots, errors, UI, and documents into textual evidence the model can understand, with no manual MCP service configuration required.

Core Features

  • Register image_understand tool: enables pure text models to receive local paths, HTTP(S) URLs, data URIs, or pasted image references, call vision APIs for PNG/JPEG/WebP/GIF, and return textual evidence
  • One-click recognition (enabled by default): replaces pasted images with cached textual descriptions during message dispatch, allowing the model to answer directly in the same turn without calling the tool first
  • Register show_image tool: allows the model to render found/generated/captured images as inline cards in the conversation flow; images don't enter model context
  • Provide settings panel (Settings → Free Vision): form auto-rendered by plugin schema, supports one-click multi-provider switching, individual API Key entry, custom API addresses, saved to ~/.dsh/free-vision.json and effective on next call
  • Provide switching between 6 model providers: Qwen, Doubao, SiliconFlow (OCR specialty) are free by default; Zhipu, Tencent Hunyuan, self-hosted OpenAI-compatible endpoints are pay-per-use or self-configured

Technical Implementation

  • Language: JavaScript (ESM, "type": "module")
  • Key Dependencies: @modelcontextprotocol.sdk (1.25.3, as MCP client to connect to luma-mcp), luma-mcp (1.7.1, built-in vision engine for this package), @deepseek-ai/schemastery (config schema validation)
  • Architecture Pattern: Host injects plugin via cordis.patch.yml → plugin spawns luma-mcp subprocess in-process via apply(ctx, config) → communicates via stdio MCP → registers general vision tools to host via ctx.tools.register → config save goes through host webServer GET/POST /dsh-free-vision/config; config changes trigger engine reconnect
  • Entry Files: dsh/index.js (host-side ESM entry, contains schema/routing/tool registration), client/client.js (browser-side UMD bundle, injects settings.section)

Use Cases

When users run DSH with a pure text DeepSeek model but need the model to "understand" screenshots, errors, UI, documents—for example, pasting a console error image during debugging, a design mockup when writing frontend, or needing OCR on a long screenshot—installing this plugin enables visual Q&A without switching models. Free quotas are sufficient for personal daily use.

Prerequisites & Compatibility

DependencyMin VersionDescription
Node.js>=18Declared in package.json:38
DSHNot declaredInjected via cordis.patch.yml, no explicit version range in package.json
OSCross-platformPure JS implementation, no native modules
Native ModulesNoneNo dependencies on node-pty/node:sqlite or other native extensions
Vision API EndpointsDirect domestic connectionAlibaba Cloud Bailian, Volcano Ark, SiliconFlow are all domestic endpoints; subprocess strips proxy variables

Installation

dsh plugin --profile web add dsh-free-vision

Configuration Options

ConfigTypeDescriptionDefault
apiKeyStringCurrent provider's API Key; falls back to corresponding env var (e.g., DASHSCOPE_API_KEY) when emptyEmpty
keysObjectPer-provider API Key mapping (e.g., { qwen: 'sk-...' }){}
baseURLsObjectPer-provider API address override (e.g., { qwen: 'https://my-proxy.example.com/v1' }); empty uses official default{}
modelProviderEnumProvider: qwen (default) / volcengine / siliconflow / zhipu / hunyuan / customqwen
modelNameStringOptional model name override (e.g., qwen3-vl-flash); default auto-selects based on providerEmpty
toolNameStringPublic tool name; renamed when conflicting with host's existing toolsimage_understand
maxTokensNumberMax generated tokens per vision call8192
temperatureNumberSampling temperature0.7
multiCropBooleanAuto multi-crop for large images to improve detail fidelitytrue
toolCallTimeoutMsNumberSingle call timeout (ms)200000
allowedDirsStringAdditional allowed image root directories (; or , separated; defaults to only working dir and user home)Empty
lumaEnvObjectExtra env vars passed to vision engine (luma-mcp){}
preservePastedImagesBooleanKeep pasted images displayed as native thumbnails in conversation (pure text model perspective auto-rewrites to reference text)true
describeAtDispatchBooleanReplace images with recognized description text at dispatch time for one-step answering without calling tooltrue
describePromptStringDefault prompt for one-step recognitionChinese detailed description template
describeCacheSizeNumberIn-process LRU cache size for one-step recognition descriptions (by sha256)64
showImageEnabledBooleanEnable show_image tool (render model-found/generated images in conversation flow)true
showImageToolNameStringPublic name for show_image tool (renamed when conflicting)show_image
showImageMaxBytesNumbershow_image per-image byte limit (also constrained by host attachment limit)26214400
showImagePixelsNumbershow_image pixel limit (width × height); 0 means unlimited40000000

FAQ

Q: Is it ready to use after installation?

A: One more step needed: fill in an API Key. The simplest path is to activate Alibaba Cloud Bailian which gives 500k free token quota, select model qwen3-vl-flash, and configure the obtained Key as env var DASHSCOPE_API_KEY (or paste directly in Settings → Free Vision). Saved config takes effect on next call; no restart needed.

Q: Can pure text models see images pasted into the conversation?

A: Yes, and usually no tool call needed. "One-step recognition" is enabled by default (describeAtDispatch): at dispatch time, the plugin first injects the pasted image's description text to the model, which answers directly in the same turn. image_understand remains available for precise follow-ups (e.g., OCR character-by-character transcription).

Q: Which vision models are supported? Are they all free?

A: Qwen Qwen3-VL-Flash (default, 500k free tokens), Doubao Vision (Volcano Ark new user 200k~500k tokens), SiliconFlow DeepSeek-OCR (OCR free) are the three free tiers; Zhipu GLM-4.6V, Tencent Hunyuan HY-Vision, self-hosted OpenAI-compatible endpoints are pay-per-use or self-configured.

Q: Can I use a domestic proxy or self-built gateway instead of official API address?

A: Yes. Switch to the corresponding provider card in settings panel; leave "Base URL / API Address" below empty to use official default; fill in formats like https://my-proxy.example.com/v1 and the engine will automatically splice the path and avoid duplication.

Q: Why must the subprocess "connect directly" and not use a proxy?

A: Alibaba Cloud Bailian, Volcano Ark, and SiliconFlow are all domestic endpoints; the subprocess actively strips proxy env vars like HTTP_PROXY/HTTPS_PROXY/ALL_PROXY/NO_PROXY at startup; using a proxy will result in 502 errors instead.

Q: What if image_understand tool name conflicts with another plugin?

A: Change it to something else in advanced settings via toolName; show_image can also be renamed via showImageToolName; changes take effect immediately after saving.

Q: How to completely uninstall?

A: Config file is at ~/.dsh/free-vision.json (custom path can be overridden by DSH_FREE_VISION_CONFIG_PATH env var); deleting it clears all Keys/address overrides; the plugin itself is uninstalled via dsh plugin --profile web remove dsh-free-vision.

Q: What if the engine process crashes?

A: The plugin has built-in exponential backoff auto-reconnect (dsh/index.js:648-664); it will reconnect and re-register tools after at most 30 seconds; also, the plugin's luma-mcp version is locked at 1.7.1, and the postinstall hook automatically reapplies idempotent patches during upgrades.

Learning Curve

Beginner — install the plugin, configure a free API Key, and the model can call the tool to see images on its own; advanced capabilities (one-step recognition cache, show_image inline cards, API address proxy) all have reasonable defaults, toggle in advanced settings as needed.

Known Issues & Limitations

  • The subprocess actively strips eight proxy environment variables: HTTP_PROXY/HTTPS_PROXY/ALL_PROXY/NO_PROXY/http_proxy/https_proxy/all_proxy/no_proxy (dsh/index.js:121-124); not usable for network environments that require a proxy to connect.
  • Only supports PNG/JPEG/WebP/GIF four image formats (dsh/index.js:1046-1053); other formats throw Unsupported image format.
  • The built-in vision engine has SSRF protection and blocks fetching from loopback addresses like 127.0.0.1 (dsh/index.js:710-711, 1199); host internal URLs must first go through pasted image reference then transfer; direct splicing is not allowed.
  • show_image tool explicitly does not support remote HTTP(S) image URLs (dsh/index.js:1517-1518); only accepts local paths, pasted references, or data URIs.
  • Image reading is restricted by whitelist: only working directory and user home are readable by default; other paths need to be added in allowedDirs separated by ; or , (dsh/index.js:204-209).
  • Plugin depends on [email protected]; the postinstall hook applies idempotent patches to this dependency during upgrades; if upstream luma-mcp changes the corresponding source code string, the patch will fail (scripts/patch-luma.mjs:46-53).

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/FuzzySoul/dsh-free-vision)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory