Skip to main content

dsh-vision

12Stars1Forks3Issues0Watchers

Integrates a third-party image recognition model into DeepSeek Harness: the configuration panel in the webpage's bottom-right corner automatically recognizes sent images and returns results to the model, while also enabling the model to capture and recognize screenshots.

Evidence5/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki
Language
JavaScript
License
MIT
Branch
main
aideepseekdeepseek-harnessdshdsh-pluginharnessimage-recognitionllm

Install

cmdweb profile
$ dsh plugin --profile web add @linenxi-ctrl/dsh-vision

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin linenxi-ctrl/dsh-vision for me: review the repository at https://github.com/linenxi-ctrl/dsh-vision first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Sentence Positioning

dsh-vision is the vision capability extension plugin for DeepSeek Harness. It integrates any third-party vision-capable model into DSH, enabling models that don't natively support vision to understand images and screenshots: users can click to recognize images in the web panel, and models can take screenshots and call vision recognition themselves.

Core Capabilities

  • Floating whale button in the bottom-right corner of the webpage (draggable), click to open the configuration panel to fill in API address, key, model, and prompt
  • Click "📤 发送图片" (Send Image) in the panel to select an image, automatically call the third-party vision API and send the recognized text as a message back to the current session
  • Automatically detect OpenAI Chat / OpenAI Responses / Anthropic / Gemini protocols based on API address, also supports custom template for long-tail interfaces
  • Register screenshot and recognize_image tools for agents, enabling models to autonomously "screenshot → recognize image → wait for results"
  • Provide "Connection Pull" button to one-click fetch the list of models supported by the target API (including protocols) and autofill into the panel
  • Support HTTP/HTTPS proxy via proxy field to bypass direct external network restrictions

Technical Implementation

  • Language: JavaScript (Node.js ESM + native browser JS)
  • Key Dependencies: @deepseek-ai/schemastery (config schema), node:http/node:https (outbound HTTP requests), node:child_process (screenshots), @deepseek-ai/cordis (host plane injection)
  • Architecture: Three-plane injection—host plane (lib/index.js, provides settings namespace and ctx.vision service, mounts HTTP routes), client plane (lib/client.js, injects browser UI), agent plane (lib/tool.js, mounts to preset to provide tools); auto-integrates into profile layer via dsh.bundle.patch
  • Entry Files: lib/index.js (host, export name='vision') / lib/client.js (client bundle, id='@linenxi-ctrl/dsh-vision') / lib/tool.js (agent, export name='vision-tool')

Use Cases

When the DeepSeek model you're using doesn't support vision (or you want to use a cheaper vision model instead of the default), you can use this to integrate any vision-capable API like GPT-4o, Claude, Gemini, or self-hosted Qwen-VL. Typical uses: let the model interpret error screenshots from users, automatically recognize webpage content, transcribe code or tables from images into text, or use cheaper vision models to reduce main conversation costs.

Prerequisites and Compatibility

DependencyMin VersionDescription
DeepSeek Harness0.1.0-rc.6peerDependencies declaration
Node.js>=18README recommends 18+, install.sh can auto-download portable version
macOSSystem built-inscreencapture for screenshots
WindowsSystem built-inPowerShell + System.Drawing for screenshots
LinuxImageMagickNeeds to provide import command
Native npm modulesNoneOnly depends on system-provided CLI tools

Installation

dsh plugin --profile web add github:linenxi-ctrl/dsh-vision

Configuration Options

ConfigTypeDescriptionDefault
apiBasestringVision model's API address (fill in base path according to protocol)https://api.openai.com/v1
apiKeystring (key not echoed)API keyempty
modelstringModel namegpt-4o-mini
protocolstringProtocol: auto / openai-chat / openai-responses / anthropic / gemini / customauto
promptstringVision prompt (skill), customizableBuilt-in detailed Chinese prompt
proxystringOptional HTTP proxy addressempty
timeoutMsnumberSingle vision request timeout (ms)60000
requestTemplatestringCustom protocol only: request body JSON templateempty
responsePathstringCustom protocol only: dot-notation path to extract text from responseempty

FAQ

Q: What do I need to prepare to use it?

A: You need a model API that supports image input (OpenAI, Anthropic, Gemini, or a self-built compatible service). Just fill in the API address, key, and model name in the whale button panel in the bottom-right corner of the webpage.

Q: Do I need to install npm packages? Can I use it offline?

A: No need. The repository comes with install.bat / install.sh. The installation script will automatically download a portable Node.js from a domestic mirror, and it can run without pnpm, suitable for internal networks or environments where you don't want to install npm toolchains.

Q: After installation, why doesn't the model recognize images on its own?

A: The web panel button (client) and vision service (host) are mounted together by default; if you want the model to autonomously screenshot and recognize images, you still need to append the lib/tool.js mount line in the agent preset's agent.cordis.yml to bring screenshot / recognize_image into that preset's tool list.

Q: Does the "Send Image" in the panel conflict with DeepSeek's default drag-and-drop image behavior?

A: No conflict. The plugin doesn't intercept full-page drag-and-drop. DeepSeek's native image understanding process is completely unaffected. Only clicking "Send Image" in the panel goes through the external vision recognition.

Q: Should I choose auto or manually specify the protocol?

A: The default auto will automatically detect based on API address (recognizing anthropic / gemini / /responses paths). In complex scenarios, you can click "Connection Pull" to have the plugin list the models supported by the API and automatically select the protocol, saving manual input.

Q: How to adapt long-tail interfaces?

A: Switch protocol to custom. In "Custom Request Template", assemble JSON template using {{model}} / {{prompt}} / {{image}} / {{dataUrl}} / {{mime}} placeholders. In "Response Text Path", use dot notation (e.g., choices.0.message.content) to point to the returned text position. Note: placeholders must be written bare without quotes.

Q: What system components does the screenshot function need?

A: Windows needs PowerShell (built-in), macOS has screencapture built-in, Linux needs ImageMagick installed (provides import command), otherwise the screenshot tool will fail.

Q: How to completely uninstall?

A: For npm installations, first run dsh plugin --profile web remove @linenxi-ctrl/dsh-vision, then run the repository's uninstall.mjs / uninstall.bat / uninstall.sh to one-click clean up client copy, agent preset, and cordis.patch.yml mount.

Difficulty Level

Beginner — After installing the plugin, just fill in three items in the bottom-right panel (API address, key, model), and click "📤 发送图片" to use it; advanced usage (custom protocol, proxy, agent preset tool mount) are optional optimizations.

Known Issues and Limitations

  • Image size limit 20MB (after base64 decode), will throw error directly if exceeded (lib/index.js:28)
  • Custom protocol authentication defaults to Authorization: Bearer <apiKey>. Interfaces requiring special auth headers (like AWS Signature, custom HMAC) need to modify the protocol adaptation layer in lib/index.js
  • Screenshot tool depends on external CLI: Linux must have ImageMagick installed, Windows must have PowerShell available (lib/tool.js:42-60)
  • Agent tools (lib/tool.js) strongly depend on host plane's ctx.vision service; if host plugin is not mounted when calling tools, will throw "Vision service not installed" error (lib/tool.js:96)
  • recognize_image single timeout 120 seconds, screenshot single 30 seconds; external API too slow will timeout first (lib/tool.js:92, 119)
  • package.json must be saved as UTF-8 without BOM, otherwise Harness startup will fail JSON parsing (BUGFIX-NOTES.md:13-42)
  • Client bundle's __ModuleLoader__.load registered id must use full package name @linenxi-ctrl/dsh-vision, otherwise loading will be treated as unregistered by Harness (BUGFIX-NOTES.md:44-99)

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/linenxi-ctrl/dsh-vision)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory