Skip to main content

How to use dsh-multimodal

Add vision and image generation to DeepSeek Harness: paste images for automatic transcription and continue conversation, generate images from descriptions, with fully user-configurable models and backends.

This article is auto-derived from indexed fields (wiki / faq / compatibility_json), not freshly AI-generated.

This article is derived from the plugin's already-indexed fields (wiki / faq / compatibility_json / readme), not freshly generated by AI. Source field is noted at the end of each section.

Quick start

dsh-multimodal

— source: plugin_wiki.wiki_content

Install & verify

dsh plugin --profile web add github:MC5lan/dsh-multimodal

Run the command above in your DSH Web Profile. Then enable the plugin in the plugin list.

— source: plugins.install

Key points

  • SSRF guard: generated-image downloads and reference_image URLs refuse loopback / private (RFC1918) / link-local addresses
  • No sessionId forwarding: internal harness session ids are never sent to third-party vision APIs
  • Prompt-injection markers: vision & OCR outputs are wrapped in explicit "untrusted context" markers before being handed to the text model
  • Sensitive-data redaction (redactSensitive): phone numbers / 18-digit IDs / emails are masked in transcriptions (including cache hits)
  • Audit log (auditLog): one line per transcription with time / image count / bytes / latency / provider

— source: plugin_wiki.readme_en (fallback readme_raw)

FAQ

Does this plugin come with models?

No. The plugin only provides access logic and UI. The actual visual models (Zhipu, SiliconFlow, ModelScope, iFlytek, local Ollama, etc.) and image generation backends (Alibaba Bailian, OpenAI protocol, custom adapters) need to be declared by you in the settings page or configuration file. See the extraProviders and image.backends sections in the README for details.

Will image data be uploaded to DeepSeek?

No. Before sending the request to DeepSeek, dsh-multimodal will first transcribe the image to text using your configured visual model. The original image bytes always stay locally or with your chosen visual service provider, and will not enter DeepSeek's context.

Can visual recognition and image generation connect to local Ollama?

Yes. In Settings → Multimodal → Platform Access, there is a one-click preset for 'Local Ollama', which will automatically add http://localhost:11434 to trusted domains. Local loopback addresses do not require an API key.

How to avoid accidentally sending other keys as visual API keys?

Since version 0.2.1, the plugin maintains an API key whitelist (allowedApiKeyEnvs) and a trusted host list (trustedBaseUrls). Only environment variable names in the whitelist can be read as API keys, and the base URL must match known official hosts or the trusted list. Local loopback addresses are allowed separately.

What happens if the backend fails to generate an image?

You can configure a failover order in image.failoverOrder for alternate backend sequences. When the primary backend request fails, it will fall back in order. Failover is not triggered for AUTH errors or user-initiated cancellations, to avoid wasting quota.

How to uninstall/disable this plugin?

Run dsh plugin --profile web remove to uninstall (if installed via GitHub, the corresponding command is remove github:MC5lan/dsh-multimodal). Uninstalling will not modify the dsh-multimodal section in your ~/.dsh/settings.yaml that you have already written.

— source: plugin_wiki.faq_json

Compatibility

  • DSH: 0.1.0-rc.6+
  • Node: >=18(@types/node ^22.0.0,README 推荐 Node 18+)

— source: plugin_wiki.compatibility_json

Pitfalls

Review the upstream repo before installing. This guide is auto-derived from indexed fields and may lag the latest release. If anything contradicts the official docs, treat the upstream source as authoritative.

— source: general rule