Skip to main content

dsh-multimodal

5Stars1Forks0Issues0Watchers

Add vision and image generation to DeepSeek Harness: paste images for automatic transcription and continue conversation, generate images from descriptions, with fully user-configurable models and backends.

Evidence5/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki
Language
TypeScript
License
MIT
Branch
main
deepseek-harnessdshdsh-plugindsh-pluginsmultimodalmultimodal-aivision

Install

cmdweb profile
$ dsh plugin --profile web add github:MC5lan/dsh-multimodal

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin MC5lan/dsh-multimodal for me: review the repository at https://github.com/MC5lan/dsh-multimodal first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Line Pitch

Give DeepSeek Harness "eyes" and a "pen": after pasting images in the chat, your configured vision model first transcribes the image content verbatim, then passes the text to DeepSeek to continue solving the problem; when the conversation needs image generation, you can call your configured image backend to generate images and display them directly in the conversation.

Core Capabilities

  • Automatically enable image attachments on the deepseek-vision conversation thread; paste/drag screenshots for automatic recognition
  • Transcribe images into pure text that preserves code, error messages, and UI copy, then let DeepSeek continue solving the problem in the same turn
  • Provide generate_image tool: enables DeepSeek to automatically call it when an illustration is needed, rendering images directly in conversation cards
  • Provide extract_text tool: extract text from images in three formats: Markdown/plain text/JSON
  • Support any OpenAI-compatible vision endpoint (Zhipu, Alibaba Bailian, iFlytek Spark, Silicon Flow, ModelScope, Ollama, etc.) via direct configuration
  • Built-in security policies post-0.2.1: API key whitelist, trusted host list, SSRF protection, sensitive field redaction, audit logs, LRU caching

Technical Implementation

  • Language: TypeScript (ESM)
  • Key Dependencies: @deepseek-ai/cordis (plugin injection), @deepseek-ai/dsh-llm (register providers and tools), @deepseek-ai/dsh-settings (settings page), eventsource-parser (streaming response parsing)
  • Architecture Pattern: Cordis plugin, uses apply(ctx, config) to register deepseek-vision provider and user-declared extraProviders in the llm subsystem; also registers generate_image and extract_text tools in the tools subsystem; rewrites image messages in agent pre-step hooks
  • Entry Files: src/index.ts (exports apply), src/vision.ts (vision transcription hook), src/image-gen.ts (image generation tool), src/ocr.ts (OCR tool)

Use Cases

Users who frequently need to paste error screenshots, design drafts, tables, or chat records into the dialog for the AI to continue fixing, or users who want the AI to directly generate diagrams when writing articles/plans. The plugin doesn't include any models by default; teams can directly connect existing vision/image generation services (company intranet APIs, paid platforms, local Ollama) to avoid data leakage.

Prerequisites & Compatibility

DependencyMin VersionDescription
DeepSeek Harness0.1.0-rc.6Supports both web and headless profile
Node.js18+@types/node ^22.0.0, README recommends 18+
PlatformCross-platformNo OS restrictions
Native modulesNoneOnly pure JS packages, no node-gyp compilation
Runtime dependenciescordis ^4.0.1, schemastery ^3.18.1, eventsource-parser ^3.1.1Declared in plugin package.json

Installation

dsh plugin --profile web add github:MC5lan/dsh-multimodal

Configuration Options

ConfigTypeDescriptionDefault
allowedApiKeyEnvsstring[]Environment variable/credential names allowed as vision API keys["DEEPSEEK_API_KEY"]
trustedBaseUrlsstring[]Additional allowed vision API baseURLs beyond official hosts (including local Ollama)[]
providers.deepseek.baseURLstringDeepSeek API addresshttps://api.deepseek.com
providers.deepseek.apiKeyEnvstringEnvironment variable/credential name for DeepSeek API keyDEEPSEEK_API_KEY
providers.deepseek.modelsarrayAdditional model list displayed under deepseek-vision route[]
extraProviders.<key>objectAny OpenAI-compatible vision endpoint (with displayName/baseURL/apiKeyEnv/models){}
vision.watchProviderstringConversation thread to enable "eyes"deepseek-vision
vision.transcribeProviderstringProvider that actually performs vision transcription (empty = disabled)""
vision.transcribeModelstringModel ID for transcription (empty uses route default)""
vision.fallbackProvidersstring[]Fallback providers to try in order when primary vision provider fails/rate limits[]
vision.transcribeModestringTranscription mode: auto/structured/ocr/describe/error-fix/chart-sql/design-codeauto
vision.parallelImagesbooleanWhether to transcribe multiple images in parallelfalse
vision.costProvider/costModel/costMaxPixelsstring/numberCost routing for small images via cheap provider (≤ N pixels)""/""/1000000
vision.sceneHintsbooleanWhether to append scene hints after screenshot transcriptiontrue
vision.redactSensitivebooleanWhether to redact phone numbers/ID cards/emails in transcription resultsfalse
vision.auditLogbooleanWhether to write one audit log line per transcriptionfalse
vision.transcribePromptstringTranscription prompt sent to vision modelBuilt-in "verbatim transcription" long prompt
vision.transcribeTimeoutMsnumberSingle transcription timeout (ms)90000
ocr.provider/ocr.modelstringProvider and model used by extract_text tool""
image.backends.<key>objectImage generation backend definition (kind/baseURL/apiKeyEnv/model/defaultSize){}
image.activeBackendstringCurrently active image generation backend key (empty = disabled)""
image.failoverOrderstring[]Backend key list to failover to in order when primary fails[]
image.verifyChineseTextbooleanWhether to use vision model to double-check if Chinese text in generated images is garbledfalse
image.verifyProvider/verifyModelstringVision provider and model for verification""
streamIdleTimeoutMsnumberStreaming response idle timeout (ms)300000

FAQ

Q: Does this plugin come with models?

A: No. The plugin only provides integration logic and UI. Actual vision models (Zhipu, Silicon Flow, ModelScope, iFlytek Spark, local Ollama, etc.) and image generation backends (Alibaba Bailian, OpenAI protocol, custom adapters) need to be declared in the settings page or config file yourself. See extraProviders and image.backends for details.

Q: Will image bytes be uploaded to DeepSeek?

A: No. dsh-multimodal transcribes images into text using your configured vision model before sending requests to DeepSeek. Original image bytes always stay local or with your chosen vision service provider and never enter DeepSeek's context.

Q: Can vision recognition and image generation connect to local Ollama?

A: Yes. Settings → Multimodal → Platform Access has a "Local Ollama" one-click preset that automatically adds http://localhost:11434 to the trusted host list. Local loopback addresses don't require API keys.

Q: How to avoid accidentally sending other keys as vision API keys?

A: Starting from 0.2.1, the plugin maintains an API key whitelist (allowedApiKeyEnvs) and trusted host list (trustedBaseUrls): only environment variable names on the whitelist are allowed to be read as API keys, and base URLs must match known official hosts or the trusted list. Local loopback addresses are separately allowed.

Q: What happens if the backend fails to generate an image?

A: You can configure fallback backend order in image.failoverOrder. When the primary backend request fails, it will failover in order. AUTH errors or user-initiated cancellations don't trigger failover to avoid wasting quota.

Q: What happens when pasting images without a vision provider?

A: A placeholder 【Image transcription failed: reason】 will be injected into the conversation, and DeepSeek will continue solving rather than getting stuck. If you don't want images to appear, just turn off vision.transcribeProvider.

Q: What if there's no "Multimodal" entry in settings?

A: First confirm the plugin is mounted (dsh --profile web --dump-config should list dsh-multimodal), then hard refresh in browser (Ctrl+F5).

Q: How to uninstall/disable this plugin?

A: Use dsh plugin --profile web remove to uninstall (for GitHub install: remove github:MC5lan/dsh-multimodal). Uninstalling doesn't affect the dsh-multimodal section you wrote in ~/.dsh/settings.yaml. It will automatically restore when reinstalled.

Learning Curve

Beginner — Configuration options may seem numerous, but the README provides a one-click "paste key → use" workflow. Platform preset cards and paste-key auto-recognition are enabled by default. When there's no vision model, the conversation won't get stuck.

Known Issues & Limitations

  • Custom image backends (kind: custom) dynamically load user-provided .mjs files via import(). The file executes with full permissions of the host process. Both README and source code clearly state "only load files you trust"
  • Remote extra providers must use HTTPS and the host must match official/trusted名单; only local loopback (localhost/127.0.0.1/[::1]) is allowed to use HTTP without API key
  • Image generation backends of kind: custom must explicitly declare adapterFile in the config, otherwise validateConfig will reject loading
  • When user-declared vision providers don't have name/description/contextWindow, injection to model directory falls back to 128k context, but maxTokens has no default value (to avoid being treated as a real value). It needs to be explicitly provided in config
  • Vision transcription failures continue the conversation with placeholders rather than retrying. When depending on stable vision providers, it's recommended to configure at least one fallback in fallbackProviders

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/MC5lan/dsh-multimodal)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory