Add visual capabilities on demand in DSH minimal mode: drag in images/documents and invoke via folded context and Bash, keeping the minimal mode's original tool interface and system prompt unchanged.
- Language
- JavaScript
- License
- MIT
- Branch
- main
Install
$ dsh plugin --profile web add dsh-tool-visionRun the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial
Install via your agent
Install the DeepSeek Harness plugin Flora233333/dsh-minimal-vision for me: review the repository at https://github.com/Flora233333/dsh-minimal-vision first, then run the install command and verify the plugin loads successfully.
Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.
One-sentence Positioning
Adds on-demand visual capabilities to DSH's minimal mode: converts images and documents to local cached files, writes visual call instructions into collapsible context, enabling the Agent to use visual models via Bash calls as needed while preserving the original tool palette and system prompt of the minimal mode.
Core Capabilities
- Register an Agent preset named "Visual Assistant Mode", preserving the two original tools
persistent-bashandstr_replace_editor - When images are dragged in: automatically correct orientation, limit long edge to 1600 px, flatten to JPEG (mozjpeg q85), and save to local cache
- When PDF/Word/PPT are dragged in: save as-is to local file, with file path entering Agent as collapsible context
- Provide
dsh-visionCLI with 5 built-in tasks: custom, slide_review, figure_semantics, caption_grounding, flowchart_extract - Centralize API address, visual model ID, and API Key storage in "Settings → Vision", supporting both OpenAI Chat Completions and Google Gemini protocols
- Built-in
dsh-vision-doctorscript to check if launcher, preset file, and CLI path are complete
Technical Implementation
- Language: JavaScript (ES Modules,
type: "module") - Key Dependencies:
sharp(image normalization and compression),@deepseek-ai/schemastery(settings Schema),js-yaml(reading settings.yaml and credentials) - Architecture Pattern: Server-side Cordis plugin injects
webServer/settings/credentials, registers 4 HTTP routes; client-side injectsslots/conversation/sessionson web side, usesagent/pre-stephook to append collapsible context to first-round messages, visual calls executed via Bash subprocess - Entry Files:
src/index.js(server),src/client.js(browser),src/cli.js(dsh-visioncommand),src/injector.js(agent/pre-stephook)
Applicable Scenarios
Used when DSH users in minimal mode need to recognize images or read PDF/PPT: drag an image into the dialog, the plugin saves the image to local cache and inserts the path into collapsible context; when the Agent sees it needs visual evidence, it triggers the visual model via Bash by calling dsh-vision analyze. The entire chain does not pollute the minimal mode's system prompt, nor does it introduce new persistent tools. Suitable for scenarios requiring on-demand visual capability supplementation without disrupting the minimal mode's context.
Prerequisites and Compatibility
| Dependency | Minimum Version | Description |
|---|---|---|
| DSH | >= 0.1.0-rc.6 | Plugin's peerDependencies declares @deepseek-ai/dsh-llm ^0.1.0-rc.6, can use peerDependenciesMeta to mark as optional |
| Node.js | >= 20 | Declared by engines.node field |
| Platform | macOS / Linux / WSL2 | README clearly states native Windows PowerShell is not supported, tested on Ubuntu and WSL2 |
| Native Module | sharp | Image processing dependency, requires compiling platform-specific prebuilt binaries during installation |
Installation
dsh plugin --profile web add dsh-tool-vision
Configuration Options
| Config | Type | Description | Default |
|---|---|---|---|
| API Protocol | Enum (OpenAI Chat Completions / Google Gemini) | Request protocol when sending to visual service | OpenAI Chat Completions |
| API Address | URL | Interface address of visual model service provider, must be http(s) | Empty |
| Visual Model | String | Visual model ID, provided by service provider | Empty |
| API Key | String | Written to $DSH_HOME/.credentials.yaml as VISION_API_KEY, UI shows configured/unconfigured status | Unconfigured |
| Request Timeout | Number (ms) | Visual service call timeout | 120000 ms |
The
VISION_API_KEYenvironment variable has highest priority and will override the key read from the credentials file; if no key is configured locally, requests are still allowed as long as the API address points tolocalhost/127.0.0.1, facilitating local inference services.
FAQ
Q: Is it ready to use after installation?
A: No, you must first fill in the API protocol, API address, visual model ID, and API Key in "Settings → Vision" before it takes effect; clicking "Test Connection" only sends a minimal text request and does not actually upload images.
Q: What images and documents are supported?
A: Images support PNG, JPEG, WebP, GIF; documents support PDF, DOC, DOCX, PPT, PPTX. Documents themselves are not sent to the visual interface; the Agent needs to first convert them to images via Bash before passing to dsh-vision.
Q: How many images can be uploaded at once?
A: A single dsh-vision analyze call accepts a maximum of 4 images, each not exceeding 20 MB; document single file limit is 100 MB.
Q: Where are images and documents stored after upload? Are they uploaded to the cloud?
A: Files are saved to local $DSH_HOME/cache/dsh-tool-vision/ cache directory; images are only sent to the configured visual service when the Agent explicitly calls dsh-vision analyze and specifies images; documents are always processed locally only.
Q: How is it different from DSH's built-in visual capabilities?
A: This plugin does not register new chat models, nor does it rewrite the minimal mode's system prompt. Instead, it injects visual instructions and attachment paths as collapsible context, letting the Agent trigger the visual model via Bash. The original persistent-bash and str_replace_editor tools remain unchanged.
Q: Can it run directly on Windows PowerShell?
A: Not supported in current version. Tested on Ubuntu and WSL2; Windows users need to install and run DSH and this plugin inside WSL2; native Windows PowerShell support coming soon.
Q: How to check if the environment is normal after installation?
A: Run dsh-vision-doctor (included in the package), which checks if profile package, launcher, Node runtime, CLI path, and preset file are complete; any abnormal item will print an ERROR line and exit with non-zero status.
Learning Curve
Beginner — users only need to fill in three fields in "Settings → Vision" to use it, and CLI default parameters are also designed for visual Q&A; advanced users can adjust task templates and preset details.
Known Issues and Limitations
- Native Windows PowerShell not supported; Windows users must run in WSL2 (README.md:49)
- Default visual request timeout is 120 seconds; some slow-responding models may need longer wait time (src/providers.js:71 / src/config.js:13)
- Image maximum 20 MB, document maximum 100 MB; uploads exceeding these limits are directly rejected (src/upload-api.js:8 / src/file-upload-api.js:7)
- Visual capability only takes effect under the "Visual Assistant Mode" preset; other presets do not automatically inject collapsible context (src/client.js:67-70)
- Under OpenAI Chat Completions protocol, if the gateway rejects
response_format: json_object, the plugin will automatically remove that field and retry once (src/providers.js:97-99)
dsh-minimal-vision
保持极简模式的上下文边界,同时获得按需视觉能力
中文 · English
🎬 介绍视频:[开源] DSH 极简模式视觉插件 👀 轻量化+干净
该视觉插件不是新的聊天模型,也不是第三个常驻工具。它只在需要时,通过隐藏上下文和 Bash 进入 Agent 工作流。
为什么基于极简模式?
极简模式的优势不只是工具少,更重要的是首轮请求拥有干净、稳定的 system prompt 和工具面。测试中观察到要触发 DS 的灰测行为水平依赖这个起点;额外 system prompt、聊天工具或视觉模型注册都可能改变它。
本插件不改写极简模式的 system prompt,也不注册新的聊天模型。视觉说明、图片路径和文档路径作为一条默认折叠的 context 与用户消息一起进入,原有两个工具保持不变,在尽量保留极简模式行为的同时提供识图能力。
| 设计约束 | 实现方式 |
|---|---|
| 保持原始工具面 | 继续只向模型暴露 persistent-bash 和 str_replace_editor |
| 避免污染 system prompt | 用折叠 context 提供附件路径和调用说明 |
| 视觉按需触发 | Agent 判断需要视觉证据时,通过 Bash 调用 dsh-vision |
| 分离文档处理 | PDF、Word、PPT 的读取、导出、裁剪、拼图和编辑仍由 Agent 完成 |
🧭 工作流程
- 用户在
vision-assist模式中拖入图片或文档。 - 图片由
sharp自动纠正方向、铺平透明背景,并将长边限制为 1600 px;PDF、DOC、DOCX、PPT、PPTX 原样保存。 - 插件把本地附件路径放入隐藏 context,用户消息气泡保持干净。
- Agent 自行判断如何用 Bash 读取或导出文档,并在需要视觉证据时调用视觉模型。
支持 PNG、JPEG、WebP、GIF、PDF、DOC、DOCX、PPT 和 PPTX;一次最多处理 4 张图片。“测试连接”只发送最小纯文本请求,不上传附件。Vision 配置独立存在,不会出现在聊天模型列表。
附件在发送时写入 $DSH_HOME/cache/dsh-tool-vision/ 本地缓存,隐藏 context 只包含本地路径。文档本身不会直接发送到视觉接口;只有 Agent 调用 dsh-vision 时指定的图片才会发送给已配置的视觉服务。
📦 安装
⚠️ 兼容性提示:当前版本暂未适配原生 Windows PowerShell,仅支持 Bash 环境;已在 Ubuntu 和 WSL2 下通过测试。Windows 用户请暂时在 WSL2 内安装并运行 DSH 与本插件,原生 Windows PowerShell 支持即将推出。
git clone https://github.com/Flora233333/dsh-minimal-vision.git
cd dsh-minimal-vision
npm install
dsh plugin --profile web add .
dsh web
启动后选择“视觉辅助模式”,并在“设置 -> Vision”中完成配置。
配置
| 字段 | 说明 |
|---|---|
| API protocol | OpenAI Completions 或 Gemini |
| Base URL | 视觉模型服务商提供的接口地址 |
| Model | 视觉模型 ID |
| API Key | 由 DSH 凭据服务保存为 VISION_API_KEY |
插件没有写死服务商地址。CLI 每次读取 $DSH_HOME/settings.yaml 和 $DSH_HOME/.credentials.yaml;环境变量 VISION_API_KEY 优先级更高。
手动调用
"${DSH_HOME:-$HOME/.dsh}/bin/dsh-vision" \
analyze --task custom --instructions "描述图中可见内容" -- /path/to/image.png
内置任务:custom、slide_review、figure_semantics、caption_grounding、flowchart_extract。PPT/PPTX/PDF 需要先由 Agent 导出为图片。
开发
npm install
npm run check
npm test
请勿提交 node_modules/、.playwright-cli/、.env 或任何 DSH 凭据文件。
Read the usage guide →
Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.
Listing badge
[](https://deepseek-plugin.org/plugins/Flora233333/dsh-minimal-vision)Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.