How to use dsh-media-skills
Provides image recognition, visual inspection, and image generation capabilities for DeepSeek Harness, and automatically registers configurable vision models.
This article is auto-derived from indexed fields (wiki / faq / compatibility_json), not freshly AI-generated.
This article is derived from the plugin's already-indexed fields (wiki / faq / compatibility_json / readme), not freshly generated by AI. Source field is noted at the end of each section.
Quick start
dsh-media-skills
— source: plugin_wiki.wiki_content
Install & verify
dsh plugin --profile web add github:akqwpeter-prog/dsh-media-skills
Run the command above in your DSH Web Profile. Then enable the plugin in the plugin list.
— source: plugins.install
Key points
- 👁️
vision-review— analyze images and screenshots, catch UI visual bugs, detect watermarks, turn images into text. - 🎨
media-tools— generate illustrations, avatars, backgrounds and banners with a free, watermark-free model. -
- Zhipu — open.bigmodel.cn → API Keys (
glm-4v-flashis free)
- Zhipu — open.bigmodel.cn → API Keys (
-
- SiliconFlow — siliconflow.cn → API Keys (Kolors is free)
-
- (optional third) Google Gemini — aistudio.google.com → Get API key; joins the vision failover chain automatically
— source: plugin_wiki.readme_en (fallback readme_raw)
FAQ
What needs to be configured after installation?
At minimum, GLM_API_KEY is required for image reading; image generation also requires SENSENOVA_API_KEY or SILICONFLOW_API_KEY. Other vision services can be used as optional fallback routes.
Can pure text models directly paste images?
The host and client patches provided by the repository need to be applied to the corresponding source code in DeepSeek Harness 0.1.0-rc.7 or rc.8. The plugin itself provides model routing and two skills; without this host capability, pure text conversations will still reject images.
Where will images be saved?
The image reading script won't save input images to the repository; pasted images in pure text conversations remain in the DSH session attachment library. Generated images are only written to the target path specified by the caller.
Will images leave the local machine?
Yes, they will. Image reading requests are sent to the currently configured third-party vision service; generation requests are also sent to SenseNova or SiliconFlow. For sensitive images, you should first consider the service provider's data terms and privacy requirements.
Which systems are supported?
The plugin code doesn't restrict the operating system and has no native modules, so it can be used in any environment that can run DSH, Python 3, and Pillow. Actual usability also depends on whether the DSH host and corresponding services are accessible.
What to do when encountering 1210 or context overflow?
GLM-4V-Flash has a single output limit of 1024 and a combined input/output limit of 16384. The plugin is configured with these limits; for long conversations, you should start a new session to retry, and check if the seed configuration has been overwritten if necessary.
How to troubleshoot image reading service exceptions?
You can run the vision-review script's --doctor option to check Python, Pillow, keys, and connectivity to each service. You can also add other vision services to the key file, or use VISION_FALLBACKS to add compatible services.
How to uninstall?
You can run dsh plugin --profile web remove dsh-media-skills. The plugin won't actively delete model entries that have already been written to DSH settings; if you no longer use that model, you need to manually delete the corresponding configuration and restart DSH.
— source: plugin_wiki.faq_json
Compatibility
- DSH: 0.1.0-rc.7~0.1.0-rc.8(贴图自动转述已验证)
- Node: 未声明
— source: plugin_wiki.compatibility_json
Pitfalls
Review the upstream repo before installing. This guide is auto-derived from indexed fields and may lag the latest release. If anything contradicts the official docs, treat the upstream source as authoritative.
— source: general rule