Skip to main content

dsh-media-skills

16Stars2Forks1Issues0Watchers

Provides image recognition, visual inspection, and image generation capabilities for DeepSeek Harness, and automatically registers configurable vision models.

Evidence5/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki
Language
Python
License
MIT
Branch
main
agent-skillsdeepseek-harnessdsh-pluginimage-generationskillvision

Install

cmdweb profile
$ dsh plugin --profile web add github:akqwpeter-prog/dsh-media-skills

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin akqwpeter-prog/dsh-media-skills for me: review the repository at https://github.com/MJorgin/dsh-media-skills first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Line Positioning

This plugin integrates image reading, visual inspection, and image generation into DeepSeek Harness. Visual models can be used in the model selector, and in adapted hosts, images can be converted to text for pure text models. Users need to prepare their own API Keys.

Core Capabilities

  • Read local images and screenshots, summarize content, extract text, and organize layout and object relationships.
  • Inspect visual anomalies in interface screenshots such as text overlapping, content overflow, element misalignment, color scheme issues, and watermarks.
  • Output structured results composed of summaries, complete OCR, reading order, object relationships, and uncertainty for easy downstream processing.
  • Automatically scale and batch process multiple images, with up to 5 images per image reading request to avoid exceeding service limits.
  • Call SenseNova or SiliconFlow to generate illustrations, avatars, backgrounds, and banners, and save them to specified files.
  • Register Zhipu (智谱) and SenseNova (商汤) vision model routes, and perform failover in the order of Zhipu, SiliconFlow, SenseNova, Gemini, and custom services when configured with keys.

Technical Implementation

  • Language: JavaScript ESM + Python 3
  • Key Dependencies: Python 3 standard library, Pillow image processing library, Node.js standard library, OpenAI-compatible service interfaces
  • Architecture Pattern: Cordis plugin entry registers skill providers and fills in missing model routes; skills load local description files and then call Python scripts to process or generate images
  • Entry File: index.js

Use Cases

Suitable for users who need to view screenshots, verify page layouts, extract text from images, or generate assets in chat. Daily single-image Q&A can directly use visual models; batch inspection, scripted processing, and having pure text models understand images are better suited for the image reading skill or patched image embedding pipeline.

Prerequisites & Compatibility

DependencyMinimum VersionDescription
DeepSeek Harness0.1.0-rc.7~0.1.0-rc.8 (full image embedding)Skills and model routes do not declare minimum DSH version; auto-transcription in pure text sessions requires host and client patches applied to corresponding version source code
Node.jsNot declaredEntry uses ESM and reads Node.js built-in fs and URL modules; repository has no npm runtime dependencies
Python 3Not declaredBoth skill scripts are launched via python3; README badge marks 3.9+ but dependencies do not enforce this version
PillowNot declaredImage reading script dynamically imports Pillow when processing JPEG compression; missing it will prompt manual installation when running vision-review diagnostic commands
Service Provider API & NetworkN/ARequires accessible image reading service and corresponding API Key at minimum; image generation additionally needs image generation service Key
PlatformCross-platformNo OS or CPU limits declared, no native modules required; actual runtime still depends on DSH host, Python, and network environment

Installation

dsh plugin --profile web add github:akqwpeter-prog/dsh-media-skills

Configuration Options

ConfigurationTypeDescriptionDefault
GLM_API_KEYSecret (string)Access credential for Zhipu image reading service; cannot use this route when not configuredNot configured
SENSENOVA_API_KEYSecret (string)Prioritized for image generation, also usable when primary image reading route failsNot configured
SILICONFLOW_API_KEYSecret (string)Used for image generation and participates in failover when primary image reading route failsNot configured
GEMINI_API_KEYSecret (string)Attempts Gemini image reading after Zhipu or other backup services failNot configured
SENSENOVA_VISION_MODELModel name (string)Overrides the model used for SenseNova image readingsensenova-6.8-flash-lite
SILICONFLOW_VISION_MODELModel name (string)Overrides the model used for SiliconFlow image readingQwen/Qwen3-VL-8B-Instruct
GEMINI_MODELModel name (string)Overrides the model used for Gemini image readinggemini-3.6-flash
GEMINI_PROXYProxy address (string)Used only for Gemini network access; falls back to system HTTPS_PROXY when not setNot explicitly set
VISION_FALLBACKSJSON arrayAdds custom image reading services; each item provides service address and model, optionally with key variable name, output limit, and structured output toggle[]

Scripts prioritize reading keys from environment variables, then from ~/.dsh/secrets/media-tools.env, and finally compatibly read from ~/.codex/secrets/media-tools.env. The plugin does not write keys to the repository; vision model routes use variables of the same name from DSH credential store.

FAQ

Q: Is additional configuration required after installation?

A: Yes. At minimum, Zhipu API Key is required for image reading; SenseNova or SiliconFlow API Key is also needed for image generation; other services can be used as needed for failover. Keys should be placed in environment variables or local credential files, not written to the repository.

Q: Can pure text models directly embed images?

A: This can be implemented in adapted DSH hosts. The repository only provides model routes and image reading/generation skills; direct image embedding also requires applying matching host and client patches to the source code of DSH 0.1.0-rc.7 or rc.8, otherwise pure text sessions may continue to reject images.

Q: Where are images saved?

A: Image reading scripts do not save input images to the repository or specified output directory. During direct image embedding, the host keeps the original image in the session attachment library; generated images are only written to the target path provided by the caller.

Q: Who are images sent to?

A: Images are sent to the current image reading service; generation tasks are sent to SenseNova or SiliconFlow. Requests to Google's free services may have their data used to improve their products; service provider terms should be confirmed before processing IDs, internal materials, or customer screenshots.

Q: Does it support macOS, Windows, and Linux?

A: The plugin has no OS or CPU limitations and no native modules. The system needs to be able to run DSH, Python 3, and Pillow, and access the selected services; other limitations come from the DSH host machine.

Q: What to do when getting error 1210 or exceeding input/output limits?

A: Zhipu image reading service requires single output not exceeding 1024, and combined input and output not exceeding 16384. The plugin is configured with these limits; for long sessions, create a new one and retry, and confirm that auto-registered model configurations have not been changed to higher output limits.

Q: How to troubleshoot when image reading service is unavailable?

A: Run the vision-review script's --doctor option to check Python, Pillow, key connectivity, and each service. You can also configure a second vision service or add custom routes via VISION_FALLBACKS.

Q: How to uninstall?

A: Run dsh plugin --profile web remove dsh-media-skills. The plugin does not actively delete model entries written to DSH settings; when no longer needed, manually delete the corresponding model configuration and restart DSH.

Difficulty Level

Advanced — Skill invocation itself is simple, but full image embedding capability requires preparing API Keys, understanding DSH model settings, and manually applying host and client patches to the corresponding version source code.

Known Issues & Limitations

  • Direct image embedding auto-transcription is not in the plugin code and still relies on DSH's visual transcription patch; repository-provided patches have only been verified for rc.7 and rc.8, and patches from different versions cannot be mixed.
  • Zhipu's single image reading request can process up to 5 images at most; the script splits more images into multiple batches, but Zhipu's output limit is 1024, and combined input and output limit is 16384.
  • The repository does not declare Python dependencies and does not automatically install Pillow; without Pillow, image reading cannot complete image compression.
  • When auto-registering model routes, it only fills in Zhipu and SenseNova configurations when the corresponding routes do not exist, does not write API Keys, and does not configure SiliconFlow or Gemini routes for users.
  • The accompanying DSH patch documentation notes that the image projection logic still lacks unit tests; existing evidence mainly comes from patch round-trip checks and type checking of specified packages.

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/akqwpeter-prog/dsh-media-skills)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory