Skip to main content

dsh-minimal-vision

6Stars0Forks0Issues0Watchers

Add visual capabilities on demand in DSH minimal mode: drag in images/documents and invoke via folded context and Bash, keeping the minimal mode's original tool interface and system prompt unchanged.

Evidence5/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki
Language
JavaScript
License
MIT
Branch
main
agentdeepseek-harnessdshdsh-plugindsh-pluginsjavascriptminimalvlm

Install

cmdweb profile
$ dsh plugin --profile web add dsh-tool-vision

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin Flora233333/dsh-minimal-vision for me: review the repository at https://github.com/Flora233333/dsh-minimal-vision first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-sentence Positioning

Adds on-demand visual capabilities to DSH's minimal mode: converts images and documents to local cached files, writes visual call instructions into collapsible context, enabling the Agent to use visual models via Bash calls as needed while preserving the original tool palette and system prompt of the minimal mode.

Core Capabilities

  • Register an Agent preset named "Visual Assistant Mode", preserving the two original tools persistent-bash and str_replace_editor
  • When images are dragged in: automatically correct orientation, limit long edge to 1600 px, flatten to JPEG (mozjpeg q85), and save to local cache
  • When PDF/Word/PPT are dragged in: save as-is to local file, with file path entering Agent as collapsible context
  • Provide dsh-vision CLI with 5 built-in tasks: custom, slide_review, figure_semantics, caption_grounding, flowchart_extract
  • Centralize API address, visual model ID, and API Key storage in "Settings → Vision", supporting both OpenAI Chat Completions and Google Gemini protocols
  • Built-in dsh-vision-doctor script to check if launcher, preset file, and CLI path are complete

Technical Implementation

  • Language: JavaScript (ES Modules, type: "module")
  • Key Dependencies: sharp (image normalization and compression), @deepseek-ai/schemastery (settings Schema), js-yaml (reading settings.yaml and credentials)
  • Architecture Pattern: Server-side Cordis plugin injects webServer / settings / credentials, registers 4 HTTP routes; client-side injects slots / conversation / sessions on web side, uses agent/pre-step hook to append collapsible context to first-round messages, visual calls executed via Bash subprocess
  • Entry Files: src/index.js (server), src/client.js (browser), src/cli.js (dsh-vision command), src/injector.js (agent/pre-step hook)

Applicable Scenarios

Used when DSH users in minimal mode need to recognize images or read PDF/PPT: drag an image into the dialog, the plugin saves the image to local cache and inserts the path into collapsible context; when the Agent sees it needs visual evidence, it triggers the visual model via Bash by calling dsh-vision analyze. The entire chain does not pollute the minimal mode's system prompt, nor does it introduce new persistent tools. Suitable for scenarios requiring on-demand visual capability supplementation without disrupting the minimal mode's context.

Prerequisites and Compatibility

DependencyMinimum VersionDescription
DSH>= 0.1.0-rc.6Plugin's peerDependencies declares @deepseek-ai/dsh-llm ^0.1.0-rc.6, can use peerDependenciesMeta to mark as optional
Node.js>= 20Declared by engines.node field
PlatformmacOS / Linux / WSL2README clearly states native Windows PowerShell is not supported, tested on Ubuntu and WSL2
Native ModulesharpImage processing dependency, requires compiling platform-specific prebuilt binaries during installation

Installation

dsh plugin --profile web add dsh-tool-vision

Configuration Options

ConfigTypeDescriptionDefault
API ProtocolEnum (OpenAI Chat Completions / Google Gemini)Request protocol when sending to visual serviceOpenAI Chat Completions
API AddressURLInterface address of visual model service provider, must be http(s)Empty
Visual ModelStringVisual model ID, provided by service providerEmpty
API KeyStringWritten to $DSH_HOME/.credentials.yaml as VISION_API_KEY, UI shows configured/unconfigured statusUnconfigured
Request TimeoutNumber (ms)Visual service call timeout120000 ms

The VISION_API_KEY environment variable has highest priority and will override the key read from the credentials file; if no key is configured locally, requests are still allowed as long as the API address points to localhost / 127.0.0.1, facilitating local inference services.

FAQ

Q: Is it ready to use after installation?

A: No, you must first fill in the API protocol, API address, visual model ID, and API Key in "Settings → Vision" before it takes effect; clicking "Test Connection" only sends a minimal text request and does not actually upload images.

Q: What images and documents are supported?

A: Images support PNG, JPEG, WebP, GIF; documents support PDF, DOC, DOCX, PPT, PPTX. Documents themselves are not sent to the visual interface; the Agent needs to first convert them to images via Bash before passing to dsh-vision.

Q: How many images can be uploaded at once?

A: A single dsh-vision analyze call accepts a maximum of 4 images, each not exceeding 20 MB; document single file limit is 100 MB.

Q: Where are images and documents stored after upload? Are they uploaded to the cloud?

A: Files are saved to local $DSH_HOME/cache/dsh-tool-vision/ cache directory; images are only sent to the configured visual service when the Agent explicitly calls dsh-vision analyze and specifies images; documents are always processed locally only.

Q: How is it different from DSH's built-in visual capabilities?

A: This plugin does not register new chat models, nor does it rewrite the minimal mode's system prompt. Instead, it injects visual instructions and attachment paths as collapsible context, letting the Agent trigger the visual model via Bash. The original persistent-bash and str_replace_editor tools remain unchanged.

Q: Can it run directly on Windows PowerShell?

A: Not supported in current version. Tested on Ubuntu and WSL2; Windows users need to install and run DSH and this plugin inside WSL2; native Windows PowerShell support coming soon.

Q: How to check if the environment is normal after installation?

A: Run dsh-vision-doctor (included in the package), which checks if profile package, launcher, Node runtime, CLI path, and preset file are complete; any abnormal item will print an ERROR line and exit with non-zero status.

Learning Curve

Beginner — users only need to fill in three fields in "Settings → Vision" to use it, and CLI default parameters are also designed for visual Q&A; advanced users can adjust task templates and preset details.

Known Issues and Limitations

  • Native Windows PowerShell not supported; Windows users must run in WSL2 (README.md:49)
  • Default visual request timeout is 120 seconds; some slow-responding models may need longer wait time (src/providers.js:71 / src/config.js:13)
  • Image maximum 20 MB, document maximum 100 MB; uploads exceeding these limits are directly rejected (src/upload-api.js:8 / src/file-upload-api.js:7)
  • Visual capability only takes effect under the "Visual Assistant Mode" preset; other presets do not automatically inject collapsible context (src/client.js:67-70)
  • Under OpenAI Chat Completions protocol, if the gateway rejects response_format: json_object, the plugin will automatically remove that field and retry once (src/providers.js:97-99)

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/Flora233333/dsh-minimal-vision)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory