Skip to main content

How to use dsh-vision-skill

Connect any OpenAI-compatible multimodal model to DSH: image recognition, OCR, object localization, element enumeration, dominant color detection, and long screenshot OCR. Includes paste-to-path for direct image pasting.

This article is auto-derived from indexed fields (wiki / faq / compatibility_json), not freshly AI-generated.

This article is derived from the plugin's already-indexed fields (wiki / faq / compatibility_json / readme), not freshly generated by AI. Source field is noted at the end of each section.

Quick start

dsh-vision-skill

— source: plugin_wiki.wiki_content

Install & verify

dsh plugin --profile web add github:DDDFXYqiming/Agent_Extensions#path:dsh-plugins/dsh-vision-skill

Run the command above in your DSH Web Profile. Then enable the plugin in the plugin list.

— source: plugins.install

Key points

  • 每个子目录是独立的自包含单元,欢迎以 PR 提交新技能/插件,或提 Issue 反馈问题。
  • 贡献要求:自包含(自带脚本/模板/文档)、许可证清晰(建议 MIT)、不写死密钥与本机绝对路径。
  • 新增技能的入口文档统一为 SKILL.md,DSH 插件另附 cordis.patch.yml 与 package.json。

— source: plugin_wiki.readme_en (fallback readme_raw)

FAQ

Does this plugin require a paid vision model API?

Supports any OpenAI-compatible multimodal model (Qwen-VL, MiniMax-M3, Gemini, GPT-4o, etc.). Default is MiniMax-M3; you can change apiUrl/model/apiKey yourself; it's recommended to use DSH Credential reference (VISION_API_KEY) instead of plaintext (see package.json:54-61, lib/index.js:42-43).

How to "see images" when using deepseek-official model?

You can directly paste images, clipboard screenshots, or send image paths. The message will only contain the path text, no image blocks will appear, so it won't trigger DSH's MODEL_DOES_NOT_SUPPORT_IMAGES; the old pi-ai image patch v0.4 is kept for compatibility only and is no longer required (README.md:156-167, client.js:151-154).

Will the same image be charged against vision API quota on second recognition?

No by default. vision_analyze enables SHA-256 + content hash caching of mode/budget/crop/prompt; hits within TTL return cached:true without calling the vision API again (lib/index.js:140-176, Schemastery cacheTtlSeconds/cacheMaxEntries defaults to 3600s/200 entries).

How to recognize ultra-tall images like long chat screenshots?

Use the vision_long_screenshot_ocr tool: automatically slices (with overlap) → each slice runs local tesseract first, falls back to VLM on failure, merges and returns full text with block boundaries, more suitable for ultra-tall images than single API call (lib/index.js:583-590, 1069-1154).

What happens if the image is not in workspace or DSH attachment directory?

It will be rejected by path fence (lib/index.js:280-311). Solution: put the image in the workspace, in DSH attachment directory ~/.dsh/attachments, or add that directory to allowedDirs in config.

Getting "duplicate loader entry id" startup error after installation?

The bundle already contributes the vision-skill line in cordis.patch.yml (cordis.patch.yml:4-6), don't add insert: id: vision-skill in your own profile patch. When needing customization, use a bare entry to override config by id (last writer wins, whole line replacement); when overriding, include the reserved fields together (README.md:79-92).

What if I already have a same-named vision skill at user/project level?

This plugin registers skill name vision at runtime level, which may obscure same-named skills at user/project level (according to official priority: project > runtime > user). It's recommended to choose one; uninstall one to avoid being obscured (README.md:197).

— source: plugin_wiki.faq_json

Compatibility

  • DSH: 0.1.0-rc.6+
  • Node: >=22.19
  • Platforms: macOS, Windows, Linux

— source: plugin_wiki.compatibility_json

Pitfalls

Review the upstream repo before installing. This guide is auto-derived from indexed fields and may lag the latest release. If anything contradicts the official docs, treat the upstream source as authoritative.

— source: general rule

How to use dsh-vision-skill