Skip to main content

Vision & Multimodal

Sort:
shanliulingshanliuling
100

Adds in-conversation image generation to DeepSeek Harness web client, supporting Google Gemini, OpenAI, and ByteDance Seedream, with a built-in local image gallery.

$ dsh plugin --profile web add dsh-image-gen
13

Add a configurable image recognition model to DSH's text-only main model: automatically convert images sent in chat to text, while keeping the original multimodal model unchanged.

$ dsh plugin --profile web add github:poiuyjie/dsh-vision-opencode
linenxi-ctrllinenxi-ctrl
12

Integrates a third-party image recognition model into DeepSeek Harness: the configuration panel in the webpage's bottom-right corner automatically recognizes sent images and returns results to the model, while also enabling the model to capture and recognize screenshots.

$ dsh plugin --profile web add @linenxi-ctrl/dsh-vision
FlyvhidbwoFlyvhidbwo
12

Enable DeepSeek Harness to send images with DeepSeek models by automatically invoking a third-party vision model to transcribe images into text before passing to DeepSeek for answering.

$ dsh plugin --profile web add dsh-vision-proxy
dickpydickpy
8

Add AI image generation sidebar to DSH Web GUI, with OpenAI-compatible API proxy, supporting text-to-image, image-to-image, and a history/prompt template library.

$ dsh plugin --profile web add @dickpy/dsh-imagegen
7

Adds sticker support to text-only DeepSeek: registers a visual routing mechanism where images are first described as text by a VL model, then passed to DeepSeek—no model changes or official modifications required.

$ dsh plugin --profile web add dsh-deepseek-vision
54xkeee54xkeee
7

Adds visual capabilities to plain-text DeepSeek: defaults to Doubao Web free channel, supports automatic fallback across multiple channels, and enables structured evidence memory reuse across conversation turns.

$ dsh plugin --profile web add dsh-vision-web
FuzzySoulFuzzySoul
6

Add image understanding to plain text DSH models: register the image_understand tool to call free vision APIs (Qwen/Doubao/Silicon), and include built-in show_image to render images inline in the conversation flow.

$ dsh plugin --profile web add dsh-free-vision
Flora233333Flora233333
6

Add visual capabilities on demand in DSH minimal mode: drag in images/documents and invoke via folded context and Bash, keeping the minimal mode's original tool interface and system prompt unchanged.

$ dsh plugin --profile web add dsh-tool-vision
maxwell-fengmaxwell-feng
6

Enables text-only models to "see" images: uses Windows built-in OCR to recognize text from images locally, sending only the text to the model. Image data never leaves your machine.

$ dsh plugin --profile web add dsh-windows-ocr
kbpoyokbpoyo
5

Provides image bridging for DSH Web plain text model: saves pasted images to local paths and rewrites them as text markers for sending, allowing the model to independently call visual tools from its directory to view images.

$ dsh plugin --profile web add @kbpoyo/dsh-image-bridge
MC5lanMC5lan
5

Add vision and image generation to DeepSeek Harness: paste images for automatic transcription and continue conversation, generate images from descriptions, with fully user-configurable models and backends.

$ dsh plugin --profile web add github:MC5lan/dsh-multimodal
condaThinkercondaThinker
3

DSH plugin: show_image tool that renders an image inline into the web chat flow (QQ/WeChat style) — path text in the message, image bytes never enter model context.

$ dsh plugin --profile web add github:condaThinker/dsh-image-inline
xiaoxianyu-officexiaoxianyu-office
3

DSH bundle plugin: chat-image bridge + read_image deny + conversational image_recognize for text-only main models | 纯文本主模型识图桥接与识图工具

$ dsh plugin --profile web add github:xiaoxianyu-office/dsh-image-tools
OoWJZZoOOoWJZZoO
3

A plug-and-play image-reading plugin for DeepSeek Harness. After installation, DSH will no longer refuse to feed images into sessions with pure-text models; instead, the image will be mapped as [Image #N] in the session. The Agent can then invoke tools on its own to read images from the session or from specified paths.

$ dsh plugin --profile web add github:OoWJZZoO/dsh-read-image
reimu-createreimu-create
3

DSH plugin: text-only models (e.g. DeepSeek-V4) automatically see images via a vision model. Official surface-replace, cache-friendly, human transcript untouched. 纯文本模型自动识图桥

$ dsh plugin --profile web add dsh-vision
zouyuanqingzouyuanqing
3

Native interactive visual-reasoning plugin for DeepSeek Harness: precise pixel grounding (SOM grid / zoom / annotate / measure / diff / color / OCR) + MiMo V2.5 multimodal backend, zero external MCP servers.

$ dsh plugin --profile web add github:zouyuanqing/dsh-vision-primitives
NormanFxxkingRockwellNormanFxxkingRockwell
1

Automatically detects configured multimodal models and adds vision capability to the main text-only model, returning recognition results as text.

$ dsh plugin --profile web add dsh-auto-vision
xmasdongxmasdong
1

Bridges DeepSeek with image generation and vision capabilities, routing requests to compatible backends.

$ dsh plugin --profile web add --allow-build=dsh-codex-image-bridge github:xmasdong/dsh-codex-image
haimuhaimuhaimuhaimu
1

DeepSeek Harness plugin: generate images via Aliyun Bailian Qwen-Image (bl CLI). Adds a generate_image tool returning images as conversation attachments.

$ dsh plugin --profile web add dsh-image-gen
nexsjournalnexsjournal
1

Adds third-party image generation and editing capabilities to DSH, callable from chat and configured via a settings card, supporting multiple APIs.

$ dsh plugin --profile web add github:nexsjournal/dsh-imagegen-plugin
FrostLeafKEEFrostLeafKEE
1

DeepSeek Harness 插件:解除 Web GUI 图片输入限制,图片附件文本化后交给 vision skill 识图 | dsh plugin that lifts the image-input gate and hands attachments to a vision skill

$ dsh plugin --profile web add github:FrostLeafKEE/dsh-image-unlock
1

Eyes for text-only DeepSeek: view_image tool (local Ollama or any OpenAI-compatible VLM) + chat image-attachment bridge — paste/drop images in the chat and the model can see them.

$ dsh plugin --profile web add github:xsoc1/dsh-image-vision
#84·dsh-img
gmleonggmleong
1

Give text-only models eyes: analyze_image tool for DeepSeek Harness, backed by free Chinese vision APIs (GLM-4V-Flash / Qwen-VL) or any OpenAI-compatible endpoint. 给纯文本模型装上眼睛的 dsh 插件。

$ dsh plugin --profile web add dsh-img
hyls9527hyls9527
1

DSH 插件:为纯文本模型提供本地看图能力(llama.cpp / Ollama / LM Studio / vLLM 全兼容)|Local vision for text-only LLMs on any OpenAI-compatible local server

$ dsh plugin --profile web add dsh-local-vision
chenkezhen480chenkezhen480
1

Enables multimodal input and output capabilities for DeepSeek Harness, supporting images, audio, and video in conversations.

$ dsh plugin --profile web add --allow-build=dsh-plugin-multimodal github:chenkezhen480/dsh-multimodal
gloryxpnvgloryxpnv
1

Local-first vision for DeepSeek Harness: structured JSON evidence (OCR/layout/semantics) from local VLMs (LM Studio/Ollama), zero API cost, images never leave your machine.

$ dsh plugin --profile web add dsh-vision-local
sjakdhasdhsjakdhasdh
1

Vision tool plugin for DeepSeek Harness (DSH): give text-only models like deepseek-v4-flash image recognition via Alibaba Bailian / any OpenAI-compatible vision API. 给 DeepSeek Harness 无识图能力模型加识图工具。

$ dsh plugin --profile web add dsh-vision