Integrates a third-party image recognition model into DeepSeek Harness: the configuration panel in the webpage's bottom-right corner automatically recognizes sent images and returns results to the model, while also enabling the model to capture and recognize screenshots.
- Language
- JavaScript
- License
- MIT
- Branch
- main
Install
$ dsh plugin --profile web add @linenxi-ctrl/dsh-visionRun the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial
Install via your agent
Install the DeepSeek Harness plugin linenxi-ctrl/dsh-vision for me: review the repository at https://github.com/linenxi-ctrl/dsh-vision first, then run the install command and verify the plugin loads successfully.
Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.
One-Sentence Positioning
dsh-vision is the vision capability extension plugin for DeepSeek Harness. It integrates any third-party vision-capable model into DSH, enabling models that don't natively support vision to understand images and screenshots: users can click to recognize images in the web panel, and models can take screenshots and call vision recognition themselves.
Core Capabilities
- Floating whale button in the bottom-right corner of the webpage (draggable), click to open the configuration panel to fill in API address, key, model, and prompt
- Click "📤 发送图片" (Send Image) in the panel to select an image, automatically call the third-party vision API and send the recognized text as a message back to the current session
- Automatically detect OpenAI Chat / OpenAI Responses / Anthropic / Gemini protocols based on API address, also supports custom template for long-tail interfaces
- Register screenshot and recognize_image tools for agents, enabling models to autonomously "screenshot → recognize image → wait for results"
- Provide "Connection Pull" button to one-click fetch the list of models supported by the target API (including protocols) and autofill into the panel
- Support HTTP/HTTPS proxy via
proxyfield to bypass direct external network restrictions
Technical Implementation
- Language: JavaScript (Node.js ESM + native browser JS)
- Key Dependencies:
@deepseek-ai/schemastery(config schema),node:http/node:https(outbound HTTP requests),node:child_process(screenshots),@deepseek-ai/cordis(host plane injection) - Architecture: Three-plane injection—host plane (lib/index.js, provides settings namespace and ctx.vision service, mounts HTTP routes), client plane (lib/client.js, injects browser UI), agent plane (lib/tool.js, mounts to preset to provide tools); auto-integrates into profile layer via
dsh.bundle.patch - Entry Files: lib/index.js (host, export name='vision') / lib/client.js (client bundle, id='@linenxi-ctrl/dsh-vision') / lib/tool.js (agent, export name='vision-tool')
Use Cases
When the DeepSeek model you're using doesn't support vision (or you want to use a cheaper vision model instead of the default), you can use this to integrate any vision-capable API like GPT-4o, Claude, Gemini, or self-hosted Qwen-VL. Typical uses: let the model interpret error screenshots from users, automatically recognize webpage content, transcribe code or tables from images into text, or use cheaper vision models to reduce main conversation costs.
Prerequisites and Compatibility
| Dependency | Min Version | Description |
|---|---|---|
| DeepSeek Harness | 0.1.0-rc.6 | peerDependencies declaration |
| Node.js | >=18 | README recommends 18+, install.sh can auto-download portable version |
| macOS | System built-in | screencapture for screenshots |
| Windows | System built-in | PowerShell + System.Drawing for screenshots |
| Linux | ImageMagick | Needs to provide import command |
| Native npm modules | None | Only depends on system-provided CLI tools |
Installation
dsh plugin --profile web add github:linenxi-ctrl/dsh-vision
Configuration Options
| Config | Type | Description | Default |
|---|---|---|---|
apiBase | string | Vision model's API address (fill in base path according to protocol) | https://api.openai.com/v1 |
apiKey | string (key not echoed) | API key | empty |
model | string | Model name | gpt-4o-mini |
protocol | string | Protocol: auto / openai-chat / openai-responses / anthropic / gemini / custom | auto |
prompt | string | Vision prompt (skill), customizable | Built-in detailed Chinese prompt |
proxy | string | Optional HTTP proxy address | empty |
timeoutMs | number | Single vision request timeout (ms) | 60000 |
requestTemplate | string | Custom protocol only: request body JSON template | empty |
responsePath | string | Custom protocol only: dot-notation path to extract text from response | empty |
FAQ
Q: What do I need to prepare to use it?
A: You need a model API that supports image input (OpenAI, Anthropic, Gemini, or a self-built compatible service). Just fill in the API address, key, and model name in the whale button panel in the bottom-right corner of the webpage.
Q: Do I need to install npm packages? Can I use it offline?
A: No need. The repository comes with install.bat / install.sh. The installation script will automatically download a portable Node.js from a domestic mirror, and it can run without pnpm, suitable for internal networks or environments where you don't want to install npm toolchains.
Q: After installation, why doesn't the model recognize images on its own?
A: The web panel button (client) and vision service (host) are mounted together by default; if you want the model to autonomously screenshot and recognize images, you still need to append the lib/tool.js mount line in the agent preset's agent.cordis.yml to bring screenshot / recognize_image into that preset's tool list.
Q: Does the "Send Image" in the panel conflict with DeepSeek's default drag-and-drop image behavior?
A: No conflict. The plugin doesn't intercept full-page drag-and-drop. DeepSeek's native image understanding process is completely unaffected. Only clicking "Send Image" in the panel goes through the external vision recognition.
Q: Should I choose auto or manually specify the protocol?
A: The default auto will automatically detect based on API address (recognizing anthropic / gemini / /responses paths). In complex scenarios, you can click "Connection Pull" to have the plugin list the models supported by the API and automatically select the protocol, saving manual input.
Q: How to adapt long-tail interfaces?
A: Switch protocol to custom. In "Custom Request Template", assemble JSON template using {{model}} / {{prompt}} / {{image}} / {{dataUrl}} / {{mime}} placeholders. In "Response Text Path", use dot notation (e.g., choices.0.message.content) to point to the returned text position. Note: placeholders must be written bare without quotes.
Q: What system components does the screenshot function need?
A: Windows needs PowerShell (built-in), macOS has screencapture built-in, Linux needs ImageMagick installed (provides import command), otherwise the screenshot tool will fail.
Q: How to completely uninstall?
A: For npm installations, first run dsh plugin --profile web remove @linenxi-ctrl/dsh-vision, then run the repository's uninstall.mjs / uninstall.bat / uninstall.sh to one-click clean up client copy, agent preset, and cordis.patch.yml mount.
Difficulty Level
Beginner — After installing the plugin, just fill in three items in the bottom-right panel (API address, key, model), and click "📤 发送图片" to use it; advanced usage (custom protocol, proxy, agent preset tool mount) are optional optimizations.
Known Issues and Limitations
- Image size limit 20MB (after base64 decode), will throw error directly if exceeded (lib/index.js:28)
- Custom protocol authentication defaults to
Authorization: Bearer <apiKey>. Interfaces requiring special auth headers (like AWS Signature, custom HMAC) need to modify the protocol adaptation layer in lib/index.js - Screenshot tool depends on external CLI: Linux must have ImageMagick installed, Windows must have PowerShell available (lib/tool.js:42-60)
- Agent tools (lib/tool.js) strongly depend on host plane's
ctx.visionservice; if host plugin is not mounted when calling tools, will throw "Vision service not installed" error (lib/tool.js:96) - recognize_image single timeout 120 seconds, screenshot single 30 seconds; external API too slow will timeout first (lib/tool.js:92, 119)
package.jsonmust be saved as UTF-8 without BOM, otherwise Harness startup will fail JSON parsing (BUGFIX-NOTES.md:13-42)- Client bundle's
__ModuleLoader__.loadregistered id must use full package name@linenxi-ctrl/dsh-vision, otherwise loading will be treated as unregistered by Harness (BUGFIX-NOTES.md:44-99)
为 DeepSeek Harness 增加「外挂识图模型」能力:让本来不具备视觉能力的模型,通过一个可自定义地址/密钥/提示词的外部视觉模型来「看懂」图片与屏幕。
功能
- 网页配置按钮与面板:页面右下角出现一个DeepSeek 鲸鱼圆形按钮(可拖动),点击即可配置外挂识图模型的 API 地址、密钥、模型名、识图提示词(skill)、代理与超时。
- 发送图片识图并自动回传:点鲸鱼按钮打开面板,点「📤 发送图片」选图,插件会先把它发给外挂识图模型,等识别完成后把识别文本自动作为消息发回当前会话(无需手动复制粘贴),DeepSeek 基于识别文本作答。
- 模型自己截图 + 识图:插件为 agent 注入
screenshot(截屏)与recognize_image(识图)两个工具,并注入提示词,模型可自行「截图 → 识图 → 等待结果」。 - 自动适配识图 API 协议:内置 OpenAI Chat Completions、OpenAI Responses、Anthropic Messages、Google Gemini 四种协议,并按
apiBase自动探测;另有custom模板协议适配任意长尾接口。
文件结构
dsh-vision/
├── install.bat # Windows 一键安装(双击)
├── install.sh # macOS/Linux 一键安装
├── install.mjs # 安装脚本本体(npm 场景只做 agent 工具平面;目录场景全自动)
├── bootstrap-node.ps1 # Windows 引导脚本:未装 Node.js 时从国内镜像自动下载免安装版
├── package.json # 包定义(dsh.bundle + dsh.client 声明;tool 为独立子路径)
├── cordis.patch.yml # 插件挂载声明(dsh.bundle.patch 自动应用到 profile layer)
├── lib/
│ ├── index.js # host 平面插件:识图服务 + 协议适配 + settings 配置 + HTTP 路由
│ ├── tool.js # agent 工具插件:recognize_image / screenshot + 提示词注入
│ └── client.js # 客户端插件:鲸鱼按钮 / 配置面板 / 发送图片识图 / 自动回传
└── README.md
工作原理
[用户点鲸鱼按钮选图] [模型调用工具]
│ │
▼ ▼
client 转 base64 发送 screenshot 工具截屏
│ │
▼ ▼
POST /api/vision/recognize recognize_image 工具
│ │
▼ ▼
host 插件 ctx.vision 服务 ──► 协议自动适配后调用外挂识图 API
│ │
▼ ▼
识别文本 → 自动注入当前会话 识别文本返回给模型
识图请求在 host(Node)侧发起,因此不受浏览器 CORS 限制;图片请求走同源 /api/vision/recognize,同样无 CORS 问题。
安装
两种方式任选:npm 安装(标准,推荐)或手动 / 离线(下载 zip,无需 pnpm)。
方式一:npm 安装(推荐)
需要系统已装 Node.js 18+ 与 pnpm。
# 在 DSH 的 web profile 安装本插件(DSH 自动把 cordis.patch.yml 加入 profile layer)
dsh plugin --profile web add @linenxi-ctrl/dsh-vision
# (可选)配置 agent 工具平面:让模型能自己截图 + 识图
node ~/.dsh/profiles/web/node_modules/@linenxi-ctrl/dsh-vision/install.mjs
安装后无需手动改任何配置文件:package.json 的 dsh.bundle.patch 声明会被 DSH 自动 reconcile 进 profile 的 dsh.profile.bundles,cordis.patch.yml 即成为该 profile 的一个 bundle layer。
方式二:手动 / 离线(无需 pnpm,小白友好)
从 Releases 下载 zip 解压:
- Windows 双击
install.bat,macOS/Linux 运行bash install.sh——无需预装 Node.js:脚本检测不到时会自动从国内镜像(npmmirror / 华为云 / 腾讯云)下载免安装版(无需管理员权限); - 脚本会自动:复制英文副本、复制进每个 profile 的
node_modules、在cordis.patch.yml加 vision 行、创建 agent presetvision(复制随附 standard 并加入识图工具)并设为默认; - 重启 DSH(关闭后重新
dsh web)。
实测要点(DSH 0.1.0-rc.6):host + client 插件(
cordis.patch.yml)的name必须用「包名」,插件须在 profile 的node_modules下;agent 工具插件(preset)的name支持绝对路径(自动转file://);agent preset 不能叫standard(会被随附 standard 遮蔽)。
更新与卸载
更新:先卸载旧版,再安装新版即可(cordis.patch.yml 与 preset 会自动重建)。
一键卸载:
- Windows:双击
uninstall.bat - macOS/Linux:运行
bash uninstall.sh - 任意平台:
node uninstall.mjs
卸载脚本会自动:删除英文副本与 profile node_modules 里的插件、从 cordis.patch.yml 移除 vision 挂载行、删除 agent preset「vision」、恢复 settings.yaml 的默认 preset。
npm 方式安装的插件,请先用
dsh plugin --profile web remove @linenxi-ctrl/dsh-vision卸载 npm 包,再跑上面的卸载脚本清理 preset 与设置。
配置
点页面右下角鲸鱼按钮,或直接编辑 $DSH_HOME/settings.yaml 中的 vision 段:
| 字段 | 默认值 | 说明 |
|---|---|---|
apiBase | https://api.openai.com/v1 | 识图模型地址(按所选协议填到基础路径即可) |
apiKey | 空 | API 密钥(secret,不回显) |
model | gpt-4o-mini | 模型名称 |
protocol | auto | 协议:auto / openai-chat / openai-responses / anthropic / gemini / custom |
prompt | 见下 | 识图提示词(skill),可自定义 |
proxy | 空 | 可选 HTTP 代理,如 http://127.0.0.1:65532 |
timeoutMs | 60000 | 单次识图超时(毫秒) |
requestTemplate | 空 | 仅 custom:请求体 JSON 模板 |
responsePath | 空 | 仅 custom:响应文本取路径,如 choices.0.message.content |
默认识图提示词:
你是一名专业的图像识别助手。请仔细观察用户提供的图片……(详细描述 + 逐字转录文字 + 截图场景重点描述)
使用
- 发送图片识图:打开一个会话后,点右下角鲸鱼按钮 → 面板点「📤 发送图片」选图。识别期间右上角显示「外挂模型正在识图当中」,完成后自动把识别文本发回当前会话。
- 模型自主识图:直接对模型说「看看我现在屏幕上的报错」,模型会调用
screenshot截图、再调用recognize_image识图并继续。
API 协议自动适配
protocol 默认 auto,按 apiBase 自动识别;也可手动指定:
| 协议 | 识别条件 / 用法 | 请求要点 | 响应取文本 |
|---|---|---|---|
openai-chat | 默认;apiBase 填到 /v1 | POST /chat/completions,image_url 内嵌 data URL | choices[0].message.content |
openai-responses | apiBase 含 /responses | POST /responses,input_image | output[].content[].text |
anthropic | apiBase 含 anthropic | POST /v1/messages,x-api-key 头,source.base64 | content[].text |
gemini | apiBase 含 gemini/generativelanguage/googleapis | POST /models/{model}:generateContent,inline_data,x-goog-api-key 头 | candidates[0].content.parts[].text |
custom | 手动指定 | 按 requestTemplate 构造 | 按 responsePath 取路径 |
custom 模板协议
custom 用于适配上述四种之外的长尾接口:
-
requestTemplate:请求体 JSON 模板。占位符必须裸写(不带引号),替换时会自动补上 JSON 引号。支持的占位符:{{model}}→ 模型名{{prompt}}→ 识图提示词{{image}}→ 图片纯 base64(不含 data: 前缀){{dataUrl}}→ 完整data:image/...;base64,...{{mime}}→ 图片 MIME 类型
示例(等价于 OpenAI Chat):
{"model":{{model}},"messages":[{"role":"user","content":[{"type":"text","text":{{prompt}}},{"type":"image_url","image_url":{"url":{{dataUrl}}}}]}]} -
responsePath:从响应 JSON 取文本的点号路径(数字为数组下标),如choices.0.message.content、data.text、result.0.content。 -
鉴权默认走
Authorization: Bearer <apiKey>(apiKey为空则不携带);需要特殊鉴权头的接口暂不支持,可提 issue 扩展。
提示:占位符若误加了引号(写成
"{{image}}"),替换后会得到""base64""导致 JSON 非法。请保持裸写。
故障排查
| 现象 | 处理 |
|---|---|
| 识图失败:HTTP 401/403 | apiKey 未填或填错,去面板重新保存密钥 |
| 识图失败:HTTP 404 | apiBase 拼错或与协议不匹配;确认填到基础路径(如 OpenAI 填到 /v1,Anthropic 填 https://api.anthropic.com,Gemini 填到 /v1beta) |
| 识图失败:结果为空 | 协议识别不对时手动指定 protocol;custom 协议检查 responsePath 是否正确 |
| 外网直连不通 | 在 proxy 填 http://127.0.0.1:65532(或你自己的代理) |
| 点「发送图片」没反应 | 确认已打开一个会话;确认右下角有鲸鱼按钮(client 插件已挂载) |
| 模型不调用识图工具 | 确认 tool.js 已加进 preset 的 agent.cordis.yml,且该会话使用该 preset |
| 截图失败 | Windows 下需 PowerShell 可用(System.Drawing);macOS 用 screencapture;Linux 需 ImageMagick import |
Read the usage guide →
Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.
Listing badge
[](https://deepseek-plugin.org/plugins/linenxi-ctrl/dsh-vision)Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.