Skip to main content

dsh-vision-toolkit/third-party/dsh-vision-toolkit

7Stars1Forks0Issues0Watchers

Give DeepSeek Harness text agent vision: ask by pasting images, visual positioning, OCR, UI reconstruction, and pixel comparison.

Evidence4/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki

ⓘ This plugin is a sub-package of the Acidmoon/DIzzy-DSH monorepo — stars and activity count the whole repository.

Language
JavaScript
Branch
master
dsh-plugin

Install

cmdweb profile
$ dsh plugin --profile web add @anionex/dsh-vision-toolkit

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin Acidmoon/DIzzy-DSH/third-party/dsh-vision-toolkit for me: review the repository at https://github.com/Acidmoon/DIzzy-DSH first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Line Description

Give text-only Agents in DeepSeek Harness a set of vision tools: paste an image and the text model can read, locate, crop, perform UI reconstruction, and do pixel-level comparison around the current question, with a switchable built-in free vision service.

Core Capabilities

  • Enable text-only models to "see" images: paste an image in DSH Web, and the text model automatically switches to the "(Vision Toolkit)" variant, preserving native thumbnails and conversation history.
  • Visual Q&A, visual grounding, element detection: answer questions about the image, provide pixel coordinates, list elements like buttons/icons in the image.
  • Cropping, vector tracing, pixel comparison, HTML screenshot: convert images to PNG crops, editable SVGs, heatmap difference reports, or render local HTML with precise viewport.
  • Long screenshot OCR, subject extraction, dominant color analysis: read long screenshots in chunks preserving Markdown and audit results; extract subject as transparent PNG; output color palette or candidate color ranking.
  • Built-in free vision service: defaults to shared Gemini 3.7 Flash vision service, returns 429 with Retry-After when shared capacity is exhausted, can switch to your own endpoint with one click.

Technical Implementation

  • Language: TypeScript (with a small amount of React/TSX client)
  • Key Dependencies: @deepseek-ai/cordis, @deepseek-ai/dsh-tools, @deepseek-ai/dsh-skill, @deepseek-ai/dsh-llm, plus saxes (for validating SVG output)
  • Architecture Pattern: Cordis plugin + Settings registration + Skill progressive exposure; core 4 tools stay resident with the session, remaining 6 tools are registered only after Agent loads the vision-tools Skill
  • Entry File: src/index.ts (exports apply(ctx, config) hook, cordis.patch.yml mounts the plugin to Profile layer stack)

Use Cases

When using text-only models like DeepSeek for screenshot troubleshooting, replicating reference pages, or extracting icon assets, the model can't see the images or answer "where is the error, where is the button". This plugin lets you paste a screenshot in DSH, enabling the text model to obtain visual evidence related to the current task, and directly output generated results to the session workspace for continued reuse.

Prerequisites & Compatibility

DependencyMin VersionDescription
DSH Host^0.1.0-rc.6Declared via 18 @deepseek-ai/dsh-* peerDependencies; Web Server integration (@deepseek-ai/dsh-host-webserver) optional
Node.js^22.19.0 or >=24.0.0Enforced by package.json#engines
Python3.11+Only when runtime.mode=managed does the plugin automatically create isolated venv; can override interpreter path in runtime.python
PlatformmacOS / Windows / Linuxpackage.json doesn't declare os/cpu; updater on Windows only "checks version" without auto-restart
BrowserChrome / Chromium / EdgeOnly required for vision_html_screenshot tool, other tools unaffected
Native ModulesNonepackage.json#dependencies only has saxes, no node-gyp dependency; vision processing handled by Python runtime

Installation

dsh plugin --profile web add @anionex/dsh-vision-toolkit

Configuration Options

ConfigTypeDescriptionDefault
provider.baseUrlstringVision model provider API root (must be http(s))https://vision.anionex.me/v1 (built-in free)
provider.credentialstringDSH Credential name (not the API key itself)ANIONEX_FREE_VISION (auto-fills shared key)
provider.modelstringMultimodal model namegemini-3.7-flash
provider.protocol'openai' | 'anthropic'Vision request protocolopenai
provider.anthropicThinking'omit' | 'disabled' | 'adaptive'Anthropic thinking field strategyomit
provider.userAgentstringOutbound User-AgentChrome 126 UA
language'zh' | 'en'Vision output languagezh
timeoutMsnumberSingle remote call timeout (ms, 1000–600000)15000
maxImageBytesnumberSingle image size limit (bytes, 1024–268435456)4194304 (4 MiB)
maxImagePixelsnumberSingle image decoded pixels limit (1–268435456)20000000
concurrencynumberConcurrent tools per session (1–16)4
runtime.mode'managed' | 'external'Runtime source: managed uses bundled snapshot + isolated venv; external uses clean pinned sourcemanaged
runtime.agentVisionToolkitPathstringUpstream checkout path required in external mode—
runtime.pythonstringOverride Python interpreterAuto-select
allowedDirsstring[]Additional directories allowed as input outside workspace[]
imageInputVariants.enabledbooleanWhether to register "(Vision Toolkit)" variant for text-only modelstrue
imageInputVariants.providersstring[]Only register variant for these provider ids; empty array = all[]
imageInputVariants.autoSwitchbooleanAuto-switch to variant when pasting imagetrue

FAQ

Q: Still prompts model doesn't support images after pasting?

A: Usually need to restart Web Profile and refresh the page to make the variant with "(Vision Toolkit)" suffix appear in model selector; alternatively, first put the image in session workspace, then call /vision-tools to have tool explicitly read the image.

Q: Free service returns 429?

A: When shared capacity is exhausted, it returns 429 with Retry-After; just wait as suggested. For stable quota, switch to your own vision endpoint in "Settings → Vision Tools" and save API Key as DSH Credential (Settings only saves reference, won't echo the key).

Q: vision_html_screenshot prompts Chrome not found?

A: Install Chrome / Chromium / Edge and retry; this only affects HTML screenshot tool, other 9 tools don't need browser.

Q: Image too large or pixels exceeded and rejected?

A: Error message will explicitly state whether it was byte or decoded pixel limit that triggered rejection. First use crop or scale to compress image within maxImageBytes (default 4 MiB) and maxImagePixels (default 20 megapixels).

Q: Will images be "uploaded" to the cloud?

A: Remote vision model calls are handled by your configured provider.baseUrl — defaults to built-in shared service. For private deployment, switch to your own endpoint; local image processing (cropping, SVG tracing, pixel comparison, dominant colors, HTML screenshot) doesn't consume vision API.

Q: How to uninstall?

A: dsh plugin --profile web remove @anionex/dsh-vision-toolkit. If you need to temporarily disable, set disabled: true in Profile patch, no need to uninstall.

Learning Curve

Beginner — default config after installation allows pasting images to ask questions; if you need to switch vision model or adjust image limits, all done through "Settings → Vision Tools" web page, no manual config needed.

Known Issues & Limitations

  • Current version focuses on screenshot understanding, visual grounding, OCR, asset extraction, UI reconstruction and pixel-level verification; does not support video, audio, camera input, nor automatic GUI clicking.
  • Not in scope: interactive annotation editing, remote service clusters, model voting, cross-session vision caching.
  • Limited self-restart capability: Windows, dynamic ports, instances managed by process manager, "auto-update and restart" can only "check version", need to manually restart Web process.
  • Shared free service has request quota limits (max 5 images per request, single image 4 MiB / 20 megapixels, single output 4096 tokens), switch to your own endpoint for higher quota.
  • First startup in offline environment needs to access Python package cache or network to prepare isolated venv; if runtime has no external network, use runtime.mode=external and pre-provision upstream source.

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/Acidmoon/DIzzy-DSH/third-party/dsh-vision-toolkit)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory