Skip to main content

dsh-vision-skill/dsh-plugins/dsh-vision-skill

7Stars1Forks0Issues0Watchers

Connect any OpenAI-compatible multimodal model to DSH: image recognition, OCR, object localization, element enumeration, dominant color detection, and long screenshot OCR. Includes paste-to-path for direct image pasting.

Evidence5/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki

ⓘ This plugin is a sub-package of the DDDFXYqiming/Agent_Extensions monorepo — stars and activity count the whole repository.

Language
JavaScript
License
MIT
Branch
main
agent-skillsai-agentdeepseek-harnessdsh-pluginprompt-engineeringpythonskillstranslation

Install

cmdweb profile
$ dsh plugin --profile web add github:DDDFXYqiming/Agent_Extensions#path:dsh-plugins/dsh-vision-skill

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin DDDFXYqiming/Agent_Extensions/dsh-plugins/dsh-vision-skill for me: review the repository at https://github.com/DDDFXYqiming/Agent_Extensions first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Line Description

Adds "vision" capability to DeepSeek Harness (DSH): wraps any OpenAI-compatible multimodal model into 8 tools + 1 runtime skill, enabling the main model to recognize images, extract text, locate targets, enumerate elements, pick colors, and perform chunked OCR—even without native image support.

Core Capabilities

  • Local Image Recognition: 6 modes (description/OCR/table/code/error/structured evidence) switchable on demand; results cached by SHA-256 content hash to avoid redundant API calls.
  • Standalone OCR: Extracts all visible text from images while preserving original layout, separate from regular image analysis.
  • Target Location: Finds specified objects (e.g., "WeChat icon") in images, returns pixel and normalized coordinate boxes; optional save of annotated preview.
  • Element Enumeration: Counts specific UI element types (buttons, links, icons, etc.) in images, numbering each with pixel coordinate boxes.
  • Dominant Color Analysis: Local pixel algorithm (no API calls), outputs theme colors with percentages; useful for color picking and palette analysis.
  • Long Screenshot Chunked OCR: Chat history and long webpage screenshots auto-chunked (with overlap) → local tesseract first → VLM fallback → merged full text.
  • Clipboard Image Recognition: Saves clipboard images to workspace before recognition, as a fallback channel for "model doesn't support pasting images."
  • Direct Paste (paste-to-path): When pasting images in Web input boxes, client intercepts, uploads to .dsh-vision/pasted/ via same-origin, and only path text enters the message—bypassing MODEL_DOES_NOT_SUPPORT_IMAGES.

Technical Implementation

  • Language: JavaScript (Node ESM, 1327 lines) + Python 3 (recognition scripts)
  • Key Dependencies: @deepseek-ai/dsh-tools / @deepseek-ai/dsh-credentials / @deepseek-ai/schemastery (peer dependencies, injected by DSH host); @deepseek-ai/dsh-client-runtime and @deepseek-ai/dsh-client-ui-conversation (optional peer, for Web client)
  • Architecture Pattern: Cordis plugin (ctx.skills.register registers runtime layer vision skill + ctx.tools.register registers tools; progressive mode exposes to agents gradually; ctx.inject(['webServer']) mounts /dsh-vision-skill/paste same-origin upload route). Recognition core runs in external Python subprocess (scripts/vision.py, Qwen official dynamic resolution preprocessing + OpenAI-compatible VLM)
  • Entry Files: lib/index.js (plugin main) / client.js (Web client paste logic) / scripts/vision.py (recognition core)

Use Cases

When the main model doesn't support direct image reading (common with pure text models), users can provide image paths, clipboard screenshots, or paste images in the Web input box to let DSH see images, read text, locate targets, extract tables, etc. Enables text-only models to have visual capabilities (screenshot Q&A, error image analysis, long chat history archiving, UI element coordinate positioning, theme color extraction).

Prerequisites and Compatibility

DependencyMin VersionDescription
DSH Host0.1.0-rc.6+Peer dependencies @deepseek-ai/dsh-client-runtime ^0.1.0-rc.6 and @deepseek-ai/dsh-client-ui-conversation ^0.1.0-rc.6; cordis 4.x
Node.js>=22.19package.json engines.node
Python3.xRuntime environment for external script scripts/vision.py, Pillow recommended
PillowAnyvision.py --check self-check; required for OCR/chunking/color analysis
tesseractAnyLocal priority path for long screenshot OCR; auto-fallback to VLM when not installed
Multimodal Model API KeyAny OpenAI-compatibleDefault MiniMax-M3 + VISION_API_KEY credential reference; also supports apiKey plain text (not recommended)
Platform SupportmacOS / Windows / LinuxServer-side and Web client cross-platform; vision_clipboard depends on PowerShell + WinForms, Windows only

Installation

dsh plugin --profile web add github:DDDFXYqiming/Agent_Extensions#path:dsh-plugins/dsh-vision-skill

Configuration Options

ConfigTypeDescriptionDefault
apiUrlstringMain vision model OpenAI-compatible endpoint (with path)https://api.minimaxi.com/v1/chat/completions
modelstringMain vision model nameMiniMax-M3
apiKeystringMain provider plain text key (not recommended, use credential)empty
credentialstringDSH Credential reference name for main provider keyVISION_API_KEY
pythonstringPython interpreter pathpython
pwshstringPowerShell path (used by clipboard recognition tool)powershell.exe
timeoutMsnumberSingle image recognition timeout (ms)180000
concurrencynumberConcurrent image recognition count (1-8)2
progressivebooleanWhether to expose progressively by agent (only one active tool globally)true
allowedDirsstring[]Image path whitelist (beyond workspace and DSH attachments directory)[]
visionProvidersobject[]Backup provider chain (order = priority; auto-switch on 429/error)[]
tesseractstringtesseract executable pathtesseract
tesseractLangsstringtesseract language pack (default Chinese+English mix)chi_sim+eng
pasteMaxBytesnumberpaste-to-path same-origin upload single image byte limit (1KB-100MB)10485760
cachebooleanEnable image content hash recognition result cachingtrue
cacheTtlSecondsnumberCache validity period (seconds, 10-86400)3600
cacheMaxEntriesnumberCache entry limit (LRU, 1-10000)200

Note: Bundle installation writes default config; works without any vision-skill lines in profile. For customization, use bare entries with id override (do NOT use insert:, otherwise triggers duplicate loader entry id startup failure).

FAQ

Q: Does this plugin require a paid vision model API?

A: Supports any OpenAI-compatible multimodal model (Qwen-VL, MiniMax-M3, Gemini, GPT-4o, etc.). Default is MiniMax-M3; change apiUrl and model in config. Keys: use DSH Credential reference VISION_API_KEY instead of plain text in config.

Q: How does a pure text model like deepseek-official "see images"?

A: Three ways: ① Send image path directly; ② Screenshot and say "look at this" to trigger clipboard recognition; ③ Paste image directly in Web input box. v0.4+ pasted images go through paste-to-path, entering message as path text only—won't trigger DSH's MODEL_DOES_NOT_SUPPORT_IMAGES admission check. Legacy pi-ai image patch kept for compatibility only.

Q: Does the second recognition of the same image still cost vision API quota?

A: By default, no. vision_analyze uses SHA-256 content hash + mode/budget/crop/prompt combination caching; within TTL, hit returns cached: true without calling vision API. Set cache to false to disable.

Q: What about ultra-tall images like long chat screenshots?

A: Use vision_long_screenshot_ocr: auto-chunks by target height (with overlap) → runs local tesseract on each chunk first, falls back to VLM on failure → merges into full text with chunk boundary info; better than single API call for ultra-tall images.

Q: What happens if the image is outside workspace or DSH attachments directory?

A: Path fence rejects with vision-skill: path exceeds allowed range. Three ways to allow: ① Put image in workspace; ② Put in ~/.dsh/attachments default allow directory; ③ Add that directory to config's allowedDirs array.

Q: After installation, startup reports "duplicate loader entry id" — how to fix?

A: Bundle already contributed vision-skill line in cordis.patch.yml; don't insert same id in your profile's patch. For customization, use bare entry to override by id (override replaces entire config line, so write all fields to preserve them).

Q: What if there's already a same-named vision skill installed at user/project level?

A: This plugin registers skill name vision at runtime layer, which may conflict with same-named skills at user/project level. Recommend choosing one; uninstall the duplicate to prevent recognition flow from being hijacked by another same-named skill.

Difficulty Level

Beginner — Default values already work; DSH 0.1.0-rc.6+ users only need to configure one vision model API to use all 8 tools—no code or patch framework modifications needed.

Known Issues and Limitations

  • vision_clipboard depends on PowerShell + System.Windows.Forms.Clipboard, Windows only; macOS/Linux clients cannot trigger this tool.
  • DSH host process may clean DSH_HOME environment variable (path fence uses homedir() fallback to ~/.dsh/attachments, but other behaviors depending on DSH_HOME may still be affected by host cleanup).
  • Same-named skill (vision) conflict: runtime layer skill may shadow user/project layer skills; resolved by DSH priority project > runtime > user.
  • After bundle installation, cordis.patch.yml already contributed id: vision-skill; profile should not additionally insert: same id, otherwise plugin loading fails.
  • Modifying lib/index.js requires DSH host restart to take effect; patch config layer supports hot reload.
  • Direct paste uses paste-to-path; max single image bytes controlled by pasteMaxBytes (default 10MB); rejected if exceeded.
  • Legacy pi-ai adapter "image→path" patch is no longer needed from v0.4; scripts/reapply-pi-ai-vision-patch.ps1 and scripts/restore-pi-ai-vision-patch* kept as compatibility scripts only, distributed with package requires user to modify hardcoded paths on local machine.

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/DDDFXYqiming/Agent_Extensions/dsh-plugins/dsh-vision-skill)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory