Skip to main content

dsh-vision-proxy

12Stars1Forks0Issues0Watchers

Enable DeepSeek Harness to send images with DeepSeek models by automatically invoking a third-party vision model to transcribe images into text before passing to DeepSeek for answering.

Evidence5/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki
Language
JavaScript
License
MIT
Branch
main
dashscopedeepseek-harnessdsh-pluginimage-understandingmultimodalocrqwenvision

Install

cmdweb profile
$ dsh plugin --profile web add dsh-vision-proxy

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin Flyvhidbwo/dsh-vision-proxy for me: review the repository at https://github.com/Flyvhidbwo/dsh-vision-proxy first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Sentence Summary

In DeepSeek Harness, when you select a DeepSeek model, GUI-attached images are natively rejected. This plugin registers a new "DeepSeek + Auto Image Understanding" route: before the request leaves outbound, it automatically sends each image to a third-party vision model to transcribe into structured text, then passes the pure text conversation to DeepSeek for answering — DeepSeek remains the "brain" of the answer, while image understanding serves as an additional capability.

Core Features

  • Automatically transcribe GUI-pasted images into text, enabling DeepSeek to answer based on image content
  • Supports any OpenAI-compatible vision endpoint: defaults to DashScope qwen3.7-flash, also compatible with QwenCloud International, Zhipu, OpenRouter, and self-hosted gateways
  • Auto-detect local Ollama — image understanding stays on-machine without keys or registration
  • Multi-model fallback chain: automatically switches to other vendors/endpoints in sequence when the primary model fails
  • Content-hash-based image caching within the process — each image is transcribed at most once per process
  • Anti-deadlock: 20-second hard timeout for free endpoints, 60-second cooldown for recently failed endpoints to avoid repeated failures

Technical Implementation

  • Language: JavaScript (ESM, with JSDoc type annotations)
  • Key Dependencies: schemastery (config schema validation), sharp (optional image downsampling), node:crypto (SHA-256 content hashing)
  • Architecture Pattern: Injects new entries via cordis.patch.yml, registers a new LLM route (default deepseek-vision) as a proxy adapter: wraps the existing DeepSeek adapter, overrides inputModalities to ['text', 'image'] in resolveModel to pass attachment pre-check, intercepts image blocks in stream by calling the vision model for transcription before forwarding
  • Entry Point: apply(ctx, config) in lib/index.js, auto-loaded by postinstall hook and cordis

Use Cases

In DeepSeek Harness, you've selected a DeepSeek model and want to directly paste screenshots, memes, UI screenshots, code screenshots, etc. to continue asking questions — DeepSeek's official API doesn't accept images, so you'd normally have to "copy to another tool to view the image, then type the description to DeepSeek." This plugin automatically transcribes images to text before the request leaves outbound, letting DeepSeek answer based on the text description, providing an experience close to native multimodal.

Prerequisites & Compatibility

DependencyMinimum VersionDescription
DSH>=0.1.0-rc.6Declared via engines.dsh in plugin package; requires cordis.patch.yml injection mechanism
Node>=22.19.0Declared via engines.node in plugin package
PlatformCross-platformmacOS / Windows / Linux all supported; on Windows, environment variable changes may not propagate to running processes, so directly writing apiKey is recommended
Native ModulesNonesharp is only an optionalDependencies; without it, large images won't be downsampled but other features work normally

Installation

dsh plugin --profile web add dsh-vision-proxy

Configuration Options

ConfigTypeDescriptionDefault
providerIdstringRoute ID displayed in model selectordeepseek-vision
innerProviderstringWrapped existing adapter route IDdeepseek-official
baseURLstringOpenAI-compatible endpoint for vision model (any vendor, including local Ollama)DashScope compatible mode
apiKeystring (secret)Vision model key; empty falls back to $VISION_API_KEY, then $DASHSCOPE_API_KEY; direct write is most reliable on Windowsempty
anonymousbooleanSkip auth header, for registration-free endpoints; subject to 20-second timeout capfalse
modelstringVision model ID (e.g., qwen3.7-flash, qwen3-vl-flash, glm-4.6v-flash)qwen3.7-flash
maxTokensnumber (1-32768)Vision model single-output limit; give ample room for reasoning models4096
timeoutMsnumber (1000-300000)Single vision model request timeout; anonymous endpoints are forcibly capped at 20 seconds regardless120000
maxImagePixelsnumber (0-100000000)Images exceeding this pixel count are automatically downsampled before transcription (requires sharp); set to 0 to disable4000000
markerstringPrefix marker before each transcribed text, for easy identification[图片转译]
failureModeenum placeholder / errorBehavior when all vision models fail: continue with placeholder text (default) or fail the entire roundplaceholder
autoLocalOllamabooleanDetect local Ollama on startup, add to fallback chain if found — images stay on machinetrue
localOllamaModelstringSpecify Ollama model ID; empty auto-selects first reported vision model from localempty
fallbackModelsarray of objectsCustom fallback chain, each can point to different vendors; non-anonymous entries without key are automatically skipped[]

FAQ

Q: How to verify it's working after installation?

A: Run dsh --profile web --dump-config | grep dsh-vision-proxy — you should see exactly one plugin entry; then restart dsh web, and the model selector will show the "DeepSeek + Auto Image Understanding" option. When you paste an image into conversation, you should see the [图片转译] marker followed by DeepSeek's answer.

Q: Can I use it without an API key?

A: Yes. The plugin auto-detects local Ollama (http://localhost:11434) by default — if found, it adds Ollama to the top of the fallback chain, keeping images on-machine without registration. If Ollama isn't installed and you have no key, transcription fails quickly within a few seconds with configuration guidance, without deadlocking the conversation.

Q: Why does it say "no API key" even though I exported environment variables on Windows?

A: Windows explorer.exe caches environment variables, so running dsh processes can't read new values. Writing apiKey directly into the plugin config in the profile's cordis.patch.yml is the only reliable method on Windows. dsh rc.6 doesn't read .env files — that's not an alternative.

Q: pnpm ≥ 10 reports "Ignored build scripts" during installation — what to do?

A: Starting with pnpm 10, dependency build scripts are blocked by default. Add allowBuilds: { dsh-vision-proxy: true, sharp: true } to the profile's pnpm-workspace.yaml, then run dsh plugin --profile web add dsh-vision-proxy again to complete bundle registration.

Q: Transcription loses small text details — what can I do?

A: This is a capability limitation of the vision model itself, not a plugin bug. For dense UI screenshots/small text scenarios, consider switching to a stronger model (like qwen3-vl-plus) or increasing maxTokens; you can also downsample images to an appropriate resolution before starting the conversation.

Q: Where is image data uploaded? Is it saved?

A: During transcription, images are sent via base64 over HTTPS to the currently active vision endpoint (configured primary model / detected local Ollama / fallback chain entry). The plugin only performs in-process SHA-256 content caching (200-entry limit), doesn't write to local files, and doesn't upload to any plugin-owned services. For sensitive images, use local Ollama or a self-hosted endpoint.

Q: If all vision endpoints fail, does the entire conversation break?

A: By default, no. With failureMode: placeholder (default), it inserts "[图片转译失败: reason]" placeholder text at the original image location and continues the conversation; failed images are recorded by content hash for 60 seconds, during which no new network requests are made. Only failureMode: error will cause the entire round to fail.

Q: Which vendors or models are available?

A: Any OpenAI-compatible vision endpoint works. The default primary model is DashScope qwen3.7-flash (cheap, fast, no rate limits), also supports QwenCloud International, Zhipu glm-4.6v-flash, OpenRouter, local Ollama, self-hosted gateways; switch the primary model via baseURL/model, and chain multiple vendors via fallbackModels.

Getting Started Difficulty

Beginner — the bundle comes with sensible defaults and works without any configuration: if you have a key, it automatically uses the paid fast lane; if not, it auto-detects local Ollama; if that fails too, it fails quickly instead of deadlocking.

Known Issues & Limitations

  • Small text details in dense UI screenshots may be missed in transcription — this is a capability limitation of the vision model itself, not a plugin bug; consider using a stronger model or increasing maxTokens (README.md:174)
  • Free anonymous endpoints have strict rate limits; when HTTP 429 occurs, the plugin fails immediately and skips (doesn't retry Retry-After), to prevent deadlocking the entire conversation (lib/index.js:248-255)
  • Recently failed endpoints enter a 60-second cooldown and are skipped, avoiding repeated hits on dead endpoints (lib/index.js:62, 390-392)
  • Transcription results are only cached in-process (SHA-256, 200-entry limit), invalidated after process restart, and not written to disk (lib/index.js:66, 333)
  • sharp is an optional dependency; when not installed, images exceeding the pixel threshold won't be downsampled, potentially increasing transcription costs (package.json:58-60)
  • No third-party anonymous free endpoints are built-in as fallback by default; when needed, you can add anonymous: true entries via fallbackModels (README.md:55, lib/index.js:117)

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/Flyvhidbwo/dsh-vision-proxy)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory