How to use dsh-vision-opencode
Add a configurable image recognition model to DSH's text-only main model: automatically convert images sent in chat to text, while keeping the original multimodal model unchanged.
This article is auto-derived from indexed fields (wiki / faq / compatibility_json), not freshly AI-generated.
This article is derived from the plugin's already-indexed fields (wiki / faq / compatibility_json / readme), not freshly generated by AI. Source field is noted at the end of each section.
Quick start
dsh-vision-opencode
— source: plugin_wiki.wiki_content
Install & verify
dsh plugin --profile web add github:poiuyjie/dsh-vision-opencode
Run the command above in your DSH Web Profile. Then enable the plugin in the plugin list.
— source: plugins.install
Key points
- 聊天里发图 → 先交给视觉模型(如 MiMo-V2.5)转成文字,主模型照常回复,不用换模型
- 输入框右侧「识图模型」下拉,自动列出所有供应商中支持图片的模型
- 设置 → Vision 独立管理模型;
vision_read_image工具 /vision-image-analysisskill 支持 OCR、图表、截图理解 - 异常兜底:单次 60s 超时、失败重试 1 次、重试耗尽降级为占位文本,不拖垮回合
- 只想关掉自动转换(保留工具和选择器):
vision-opencode.autoConvert: false后重启
— source: plugin_wiki.readme_en (fallback readme_raw)
FAQ
Do I have to select a vision model after installing?
It is highly recommended to select one in Settings → Vision or the dropdown on the right side of the input box. Only after selecting can images be automatically converted to text when sending; if not selected, the plugin will show a placeholder text prompting 'Vision model not configured', and the main model can still answer normally but cannot see the image.
Can I just turn off auto-conversion and keep the tool and selector?
Yes. Change vision-opencode.autoConvert to false in settings.yaml and restart dsh; the vision_read_image tool and the vision model dropdown on the right side of the input box will still be there.
What if image conversion fails / selector doesn't appear?
Usually this happens when no vision model is selected, a non-multimodal model is selected, or the DSH version doesn't match the plugin. First confirm that provider/model are both filled in settings, then check the browser console for errors, and finally paste the console log in the repository's issue section.
Will switching to native multimodal models (like those with image input) be affected?
No. The plugin uses resolveModelInfo to identify the image input capability of the routing in real-time, and the native multimodal main model fully uses DSH's built-in image pathway, so it won't be blocked by the plugin for re-conversion.
Can turning off thinking (reasoning off) really save resources?
It depends on whether the provider really declares an 'off' setting. A few models have a real 'off' setting (hy3 off:'none'); most models can only try the 'force off' setting with three parameters: thinking:{type:disabled} / reasoning_effort:'none' / enable_thinking:false. The plugin will try each one and remember which one works, but cannot guarantee all providers will work.
What should I pay attention to before uninstalling?
First back up sessions containing images. After uninstalling, the image auxiliary information in old sessions will be reclaimed, and the pure-text main model cannot get those analysis results after restarting; you can only continue recognition by resending the image.
What image formats are supported?
Four types: PNG/JPEG/WebP/GIF. Both the vision_read_image tool and sending images in chat follow this whitelist; other formats are rejected at the attachments service layer.
Can it be used on CLI / desktop DSH?
No. The repository only declares client.platform='web', and HTTP endpoints, SSE progress push, and settings panel are all bound to the DSH web client.
— source: plugin_wiki.faq_json
Compatibility
- DSH: >=0.1.0-rc.6 <0.2.0
- Node: >=20.3
- Platforms: Web(DSH Web 客户端)
— source: plugin_wiki.compatibility_json
Pitfalls
Review the upstream repo before installing. This guide is auto-derived from indexed fields and may lag the latest release. If anything contradicts the official docs, treat the upstream source as authoritative.
— source: general rule