Skip to main content

dsh-windows-ocr

6Stars2Forks2Issues0Watchers

Enables text-only models to "see" images: uses Windows built-in OCR to recognize text from images locally, sending only the text to the model. Image data never leaves your machine.

Evidence5/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki
Language
JavaScript
License
MIT
Branch
main
cordisdeepseek-harnessdshdsh-pluginocrprivacywindows

Install

cmdweb profile
$ dsh plugin --profile web add dsh-windows-ocr

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin maxwell-feng/dsh-windows-ocr for me: review the repository at https://github.com/maxwell-feng/dsh-windows-ocr first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Line Positioning

Adds "image attachment" capability to text-only models on DSH: the plugin invokes the Windows built-in OCR engine at the pre-processing stage to convert images to text, and only the recognized text is sent to the service provider with the request; without this plugin, text models continue to reject receiving images, defaulting to secure behavior.

Core Capabilities

  • Enables text-only models to receive image attachments; only the recognized text appears in requests instead of image bytes
  • Intercepts the image serialization path, rewriting image content blocks to text blocks (including nested tool-result content) before the adapter sends requests
  • Patches llm.resolveModelInfo and listModels to make models return true for "supports images", passing all three gates (admission/switching/tools)
  • Automatically wraps streams of all registered adapters and hot-registered adapters; HMR replacement of adapters triggers automatic re-wrapping
  • OCR results for each image are cached by attachment ID to avoid redundant calls across rounds
  • Each OCR run uses an independent temp directory, automatically cleaned up regardless of success/failure/timeout; also cleans up orphaned temp directories from previous crashes on startup

Technical Implementation

  • Language: TypeScript (compiles to lib/index.js)
  • Key Dependencies: @deepseek-ai/cordis (host plugin framework), node:child_process + PowerShell (executes OCR script), lib/ocr.ps1 (wraps Windows.Media.Ocr WinRT API), ctx.attachments.readImage (reads attachment bytes)
  • Architecture Pattern: Cordis plugin; cordis.patch.yml's bundle.patch inserts the windows-ocr loader entry at the loading layer; the plugin hooks into two public seams of the llm service in apply() to perform capability shimming and request rewriting; effects restore original methods on uninstall
  • Entry File: src/index.ts (compiles to lib/index.js, also bundles lib/ocr.ps1)

Use Cases

Sending screenshots, tables, PDF screenshots, code screenshots, or error screenshots to text models daily without uploading image bytes to third-party service providers; or temporarily not switching to a vision-capable model but still letting the other party "read" the text in images. Windows users can directly attach images to text-only conversations in DeepSeek Harness / DSH to use this feature.

Prerequisites & Compatibility

DependencyMinimum VersionNotes
DSH0.1.0-rc.8+ (verified in README; package.json doesn't declare dsh field)Host must provide llm/attachments services and support loader-entry injection at load time
Node>= 20From package.json#engines.node
Windows PowerShell5.1+Built into Windows 10/11; plugin invokes WinRT OCR via powershell.exe
OCR Language PackDepends on target languageE.g., Chinese recognition requires "Chinese (Simplified)" language pack (OCR-capable)
PlatformWindows 10/11Uses WinRT's Windows.Media.Ocr; no cross-platform implementation
Native ModulesNoneNo npm native dependencies; OCR completed via PowerShell process

Installation

dsh plugin --profile web add github:maxwell-feng/dsh-windows-ocr

Configuration Options

ConfigTypeDescriptionDefault
languagestringBCP-47 language tag for Windows OCR, e.g., zh-Hans, en-US; empty uses current Windows user's language preference""
passthroughbooleanDefault false, all images go through OCR; when set to true, only true vision models receive images as-is, text models still use OCRfalse
ocrScriptstringAbsolute path override for PowerShell OCR script, generally no need to changeBundled lib/ocr.ps1
timeoutMsnumberOCR timeout per image (ms); terminates child process and cleans temp dir on timeout60000
maxCacheEntriesnumberIn-process OCR result cache limit (by attachment ID)200

FAQ

Q: Can this plugin be used on macOS or Linux?

A: No. It depends on Windows' Windows.Media.Ocr (WinRT) engine and launches OCR via PowerShell. macOS / Linux have no equivalent capability. The plugin will throw errors or fail to recognize after installation on non-Windows platforms.

Q: Will images be sent to model service providers as-is by default?

A: No. With passthrough defaulting to false, every outbound request goes through local OCR first; the serialized payload only contains text content blocks, no image_url or data URI. You can verify in DevTools → Network.

Q: Chinese characters aren't being recognized. How to handle?

A: The system needs the corresponding Windows OCR language pack installed first (Settings → Time and language → Language → Add language → Chinese (Simplified)). After installation, configure language: zh-Hans to force a specific language, or keep default to let the plugin follow the current user's language preference.

Q: After installation, dsh reports duplicate loader entry id: windows-ocr. How to fix?

A: The same id can only be registered once in dsh 0.1.0-rc.8 (cordis-plugin-loader 1.0.2). If previously installed via npm bundle, dsh plugin ... add github:... will insert the same id line again, causing duplication. Choose one: either use dsh plugin remove to uninstall then reinstall, or use id-based override in profile's cordis.patch.yml (- id: windows-ocr + config:) instead of - insert:.

Q: After uninstall, will there be a gap where images are neither uploaded nor OCR'd?

A: On uninstall, the plugin's effect restores llm.resolveModelInfo / listModels and adapter streams to host originals; text models still reject receiving images per host strategy—this is fail-closed behavior; there's no state where "plugin removed but image blocks still travel naked in requests".

Q: Will OCR-failed images cause the entire conversation request to error?

A: No. Images that fail recognition are replaced with placeholder text blocks (OCR: failed to recognize this image); missing attachments are replaced with (OCR: missing attachment — image refused). Failures only affect that single image, not blocking the entire round.

Q: Do I still need this plugin with a true vision model?

A: Keep the plugin with passthrough set to true if you want true vision models to receive original images; the plugin checks before outbound if the model itself supports images, passing through the original image block if supported, only falling back to OCR if not. This way one plugin covers both text-only and vision models.

Difficulty Level

Beginner — one command to install and use; most scenarios require no config changes; only need to touch passthrough / language when switching languages or letting vision models receive original images.

Known Issues & Limitations

  • OCR language depends on system-installed language packs: When OCR script exits with code 2/3, plugin degrades to placeholder text instead of error (lib/ocr.ps1:43-55)
  • GIF animations only recognize first frame: Windows OCR engine limitation, not a bug in this plugin
  • Cache scope is single dsh process: OCR results persist throughout long sessions until process exit, constrained by maxCacheEntries limit
  • HMR and dsh full upgrade suggest restart: Plugin listens to llm/adapters-updated to automatically re-wrap new adapters, but full restart is most stable after major upgrades
  • Text models may lack "image" badge in model selector UI: listModels capability declaration is shimmed; only the visual marker in model selector UI shows inconsistency—purely cosmetic
  • No platform support outside Windows: Windows.Media.Ocr calls WinRT interfaces; macOS / Linux have no equivalent
  • port:3080占用冲突: If port 3080 is occupied when dsh web starts, EADDRINUSE occurs; use netstat -ano | findstr :3080 to find PID and taskkill /PID <pid> /F

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/maxwell-feng/dsh-windows-ocr)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory