How to use dsh-vision
Adds vision capability to DeepSeek Harness text
This article is auto-derived from indexed fields (wiki / faq / compatibility_json), not freshly AI-generated.
This article is derived from the plugin's already-indexed fields (wiki / faq / compatibility_json / readme), not freshly generated by AI. Source field is noted at the end of each section.
Quick start
dsh-vision
— source: plugin_wiki.wiki_content
Install & verify
dsh plugin --profile web add github:oil-oil/dsh-vision
Run the command above in your DSH Web Profile. Then enable the plugin in the plugin list.
— source: plugins.install
Key points
- macOS: built-in Vision OCR, with no extra dependency.
- Linux / Windows: Tesseract with the required language data installed.
- Original images are sent only to vision services configured by the user.
- Vision output is marked as untrusted observation data; instructions inside an image receive no system authority.
- Generated vision context affects only the current model request and does not rewrite message history.
— source: plugin_wiki.readme_en (fallback readme_raw)
FAQ
Do I need to apply for a vision model's API Key myself after installation?
Not required by default. Keep "Auto Select" and the plugin will first try models configured in Harness that declare image support, then read see private config and local OCR; only when manually specifying a cloud platform (ZenMux / Bailian / TokenDance / OpenRouter) do you need to fill in the corresponding platform's API Key.
If the current main model already supports images, will the plugin still intervene?
No. VisionBridgeAdapter will first check the input modality of the current model, and if it natively supports images, it will directly forward to the original DeepSeek adapter, with the bridging logic not involved at all; it only calls the external vision model when the original model doesn't support images.
Will multiple images be recognized separately or analyzed together?
Joint analysis. All images in a single request are bundled into the same external vision call, making it suitable for scenarios requiring multiple images simultaneously like before-after comparisons or combined evidence.
Can it recognize images on macOS without internet?
Yes. On macOS it falls back to the system's built-in Vision OCR without requiring additional installation; if system OCR is unavailable, it will try Tesseract. Other systems need to install Tesseract and language packs themselves. Note that local fallback focuses on text recognition and is not equivalent to complete multimodal semantic understanding.
How many images can be sent in a single request?
Default is 8, adjustable between 1-32 (the "Single Image Limit" on the "Vision Recognition" card in the interface). Exceeding the limit will directly throw an error, not silently truncate.
Where is the API Key saved? Is it safe?
Saved through Harness's official credential service. After writing, you can only know "whether it exists", it won't be read back in chats, settings pages, or session logs; it can also be provided via see's private config file ~/.config/see/config.env or environment variables of the same name. This repository will not write any keys.
Do I need to restart after changing the configuration?
No need. Changes to the llm-deepseek section in $DSH_HOME/settings.yaml or in the interface will take effect automatically.
How to uninstall or disable?
Uninstall with dsh plugin --profile web remove; to temporarily disable, you can set disabled: true for the dsh-vision entry in cordis.patch.yml and restart the Web Profile.
— source: plugin_wiki.faq_json
Compatibility
- DSH: =0.1.0-rc.6
- Node: >=22.19
— source: plugin_wiki.compatibility_json
Pitfalls
Review the upstream repo before installing. This guide is auto-derived from indexed fields and may lag the latest release. If anything contradicts the official docs, treat the upstream source as authoritative.
— source: general rule