Skip to main content

dsh-vision-opencode

13Stars2Forks0Issues0Watchers

Add a configurable image recognition model to DSH's text-only main model: automatically convert images sent in chat to text, while keeping the original multimodal model unchanged.

Evidence4/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki
Language
JavaScript
License
MIT
Branch
main
deepseek-harnessdshdsh-plugindsh-plugin-marketdsh-plugins

Install

cmdweb profile
$ dsh plugin --profile web add github:poiuyjie/dsh-vision-opencode

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin poiuyjie/dsh-vision-opencode for me: review the repository at https://github.com/poiuyjie/dsh-vision-opencode first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Line Description

Equip DSH's pure text main model with a separate vision recognition model: images sent in chat are first "translated" by the vision model into text descriptions, then passed to the main model, so the main model can "see" images without switching; native multimodal main models follow the DSH native pipeline completely and won't be interfered with by the plugin.

Core Capabilities

  • Automatic chat image-to-text conversion: Pure text main models can also "see" images—images are analyzed by the vision model into "image content analysis" text and injected into context, while the original image remains in the conversation history (visible in UI)
  • Native multimodal model auto-pass-through: Plugin uses resolveModelInfo to real-time identify image input capability of the routed model; main models with built-in image capability fully retain DSH's native image pipeline
  • vision_read_image tool: Models can explicitly call it to directly read, analyze, and return text from PNG/JPEG/WebP/GIF paths (OCR, charts, screenshots scenarios)
  • "Vision Model" dropdown on right side of input box + Settings → Vision: Automatically lists all image-input-capable models from all providers; can set "disable thinking/force disable" and other reasoning strategies
  • Built-in fallback: Vision failure → 60s timeout → 1 retry attempt → degrade to placeholder text on retry exhaustion; main model turns won't be dragged down by vision failures
  • Old version modelOverrides auto-restore: For llm-pi-ai image gates written before upgrade, plugin precisely restores them on first startup based on built-in gateState ownership records

Technical Implementation

  • Language: JavaScript (ESM, type: module)
  • Key Dependencies: @deepseek-ai/dsh-llm (BlockAssembler / createUserMessage / freezeMessage), @deepseek-ai/dsh-tools (defineTool), @deepseek-ai/dsh-settings (settingsNamespace), @deepseek-ai/schemastery (runtime validation)
  • Architecture Pattern: Cordis plugin; backend index.js intercepts image-containing requests via ctx.on('llm/stream'), uses Symbol('vision-bypass') to short-circuit self-recursion, uses tools.guard to reject built-in read_image from injecting real images onto pure text main models; frontend client.js injects React component into conversation.input.right slot via window.__ModuleLoader__.load; core.js exposes pure functions for host/client reuse
  • Entry Files: index.js (backend, injects llm/stream / tools / settings / webServer), client.js (frontend, model selector recognition + settings page), core.js (image counting, replacement, model info compatibility layer)

Use Cases

When you're using DSH with a pure text main model (e.g., early DeepSeek series, open-source Qwen2.5, etc.) but want to directly send screenshots, menu screenshots, and charts in chat for the AI to see—this plugin lets you keep the main model as pure text while still "seeing" images. Any model from any provider that supports image input can be designated as an auxiliary vision recognition model; combined with "force disable thinking," you can also save first token latency. If your main model is already native multimodal (e.g., Vision series), this plugin does nothing—the native pipeline takes effect directly.

Prerequisites and Compatibility

DependencyMin VersionDescription
DSH (includes dsh-llm / dsh-tools / dsh-settings)>=0.1.0-rc.6 <0.2.0peerDependencies range, 4 DSH core packages in same range
@deepseek-ai/dsh-skillSame as aboveOptional; missing only skips skill registration, tools and auto-conversion still usable
@deepseek-ai/schemasteryAnyRuntime config validation
Node.js>=20.3Only relies on partial await / AbortSignal.any and other ES2022 features
PlatformDSH Web Clientdsh.client.platform: "web", CLI / desktop has no HTTP endpoints and SSE progress
Native ModulesNonePure JS implementation, no node-pty / sqlite / native addons

Installation

dsh plugin --profile web add github:poiuyjie/dsh-vision-opencode

Alternatively, you can use scripts/install.sh (Ubuntu one-click script, supports --vision-provider / --vision-model / --proxy and other parameters) or scripts/install.ps1 (Windows script). After installation, restart dsh and hard refresh the browser (Ctrl+Shift+R); a "Vision Model" dropdown will appear on the right side of the input box.

Configuration Options

ConfigTypeDescriptionDefault
vision-opencode.providerStringVision model provider routing id (e.g., opencode-go, my-custom)Empty
vision-opencode.modelStringVision model idEmpty
vision-opencode.visionModelsArrayPlugin-managed vision model list (includes id / provider / model / name / description / baseUrl / requestFormat / reasoning), decoupled from host provider directory[]
vision-opencode.autoConvertBooleanChat image auto-conversion toggle; after turning off, only tool and selector remaintrue
vision-opencode.visionReasoningBooleanWhether to enable thinking during image recognition (true=follow provider default; false=thinking off by default)false
vision-opencode.apiKeyString (hidden)API Key used when directly connecting to gateway (opencode-go) in force-disable mode; if empty, reads OPENCODE_GO_API_KEY environment variableEmpty
vision-opencode.mainProviderStringOld version/manual compatible main model provider; current version usually auto-identified via adapter capabilityEmpty
vision-opencode.mainModelsArrayCompatible main model list for above[]
vision-opencode.ignoredModelsArrayList of "unimported system models" actively dismissed by user in settings page, persisted to avoid re-prompting on next restart[]
vision-opencode.gateStateString (hidden)Old version modelOverrides ownership record (base64url), used for precise restoration during uninstallEmpty

Settings page (Settings → Vision) automatically generates form based on above schema; also supports direct editing of ~/.dsh/settings.yaml.

FAQ

Q: Do I have to pick a vision model after installing?

A: Strongly recommended. Pick one in Settings → Vision or the dropdown on the right side of the input box; when nothing is picked, the plugin uses placeholder text to prompt "Vision model not configured," and the main model can still answer normally but can't see the image—this is like "half-installed."

Q: Can I just turn off auto-conversion and keep the tool and selector?

A: Yes. In ~/.dsh/settings.yaml, change vision-opencode.autoConvert to false and restart dsh; the vision_read_image tool and the vision model dropdown on the right side of input box will still be there; only the chat image auto-conversion pipeline is disabled.

Q: Image conversion fails / selector doesn't appear—what to do?

A: 90% likely the vision model isn't selected, or the selected id isn't recognized as a multimodal model by the plugin. First check that both provider and model fields are filled in settings, then open browser developer tools to check Console for errors, finally go to plugin repo issue section to paste logs and ask.

Q: Will switching to a native multimodal main model (e.g., one with image input) be affected?

A: No. The plugin hooks resolveModelInfo; when it detects the current route natively supports images (inputModalities includes image), it directly passes through, following DSH's native image pipeline—it won't be intercepted for reconversion.

Q: Does "disable thinking" really save anything?

A: It depends on whether the provider really has a disable slot declared for that model. Very few models have a true "disable" slot (like hy3 off:"none"); most models can only try three wire parameters in "force disable" slot order: thinking:{type:disabled} / reasoning_effort:"none" / enable_thinking:false. The plugin tries them one by one and remembers which one marks that provider as 0 reasoning_tokens for reuse, but can't guarantee every provider can actually disable it.

Q: What should I note before uninstalling?

A: First back up conversations containing images. After uninstall, image assistant analysis in old conversations will be cleared—the pure text main model won't be able to get those analysis results after restart; if you don't resend the images, the model won't know what those images originally were.

Q: What image formats are supported?

A: PNG / JPEG / WebP / GIF four formats. Both vision_read_image tool and chat image sending use this same whitelist; other formats are rejected at the attachments service layer.

Q: Can I use it on DSH CLI / desktop?

A: No. The repo explicitly declares dsh.client.platform as "web" in package.json; HTTP endpoints (/vision-opencode/config, etc.), SSE progress push (/vision-opencode/events), settings panel, and React components are all bound to the DSH Web client.

Difficulty Level

Beginner — One command to install, restart DSH, refresh browser, then just pick a vision model from the dropdown on the right side of the input box; when you want full customization, go to Settings → Vision to adjust the schema form.

Known Issues and Limitations

  • "Disable thinking" relies on provider directory / adapter metadata: The plugin tries to distinguish "truly declared off" from "undeclared (only returns thinking slot)", but when pi-ai forms change (JSON file vs module export), it may trigger warn logs and fall back to adapter face value (index.js:594); force disable tries three parameters via direct gateway connection, can't guarantee every provider can really disable
  • If main model is not dsh-llm-pi-ai adapter (typical like dsh-llm-deepseek): Plugin warns and skips image submission gate compatibility layer installation (index.js:229); image submission on such main models may still be directly rejected by DSH
  • 60s forced timeout + 1 retry (index.js:508-512): Very long OCR or large chart recognition may be interrupted; current built-in attachments.imageLimits determines image pixel upper limit (index.js:1746-1764); large images are first cropped at service layer
  • Built-in read_image tool is forcibly intercepted on pure text main models (index.js:251-265): Returned "tool results" guide the model to use vision_read_image, but if the model insists on calling it, it keeps receiving "blocked" prompts instead of real images
  • Old version (≤0.3.2) wrote llm-pi-ai.modelOverrides: On upgrade, plugin precisely restores using gateState ownership records; if restore fails, keeps gateState waiting for next restart retry (index.js:458-462)
  • Uninstall goes through POST /vision-opencode/uninstall (needs X-Vision-Opencode-Action: uninstall custom header): Only cleans up this plugin's settings and old modelOverrides; DSH's own dependencies, profile directory, and insert entries in cordis.patch.yml need manual uninstall (index.js:1671-1674)
  • Only supports DSH Web client: dsh.client.platform: "web" explicitly declared in package.json:49; CLI / desktop / mobile and other hosts have no corresponding bundle
  • Cross-version compatibility: When DSH upgrades and changes hash prefix, frontend components can automatically adapt by scanning document.styleSheets (client.js:63-105), but if official React slot names change, components may need to relocate injection position

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/poiuyjie/dsh-vision-opencode)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory