Enables text-only models to "see" images: uses Windows built-in OCR to recognize text from images locally, sending only the text to the model. Image data never leaves your machine.
- Language
- JavaScript
- License
- MIT
- Branch
- main
Install
$ dsh plugin --profile web add dsh-windows-ocrRun the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial
Install via your agent
Install the DeepSeek Harness plugin maxwell-feng/dsh-windows-ocr for me: review the repository at https://github.com/maxwell-feng/dsh-windows-ocr first, then run the install command and verify the plugin loads successfully.
Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.
One-Line Positioning
Adds "image attachment" capability to text-only models on DSH: the plugin invokes the Windows built-in OCR engine at the pre-processing stage to convert images to text, and only the recognized text is sent to the service provider with the request; without this plugin, text models continue to reject receiving images, defaulting to secure behavior.
Core Capabilities
- Enables text-only models to receive image attachments; only the recognized text appears in requests instead of image bytes
- Intercepts the image serialization path, rewriting image content blocks to text blocks (including nested tool-result content) before the adapter sends requests
- Patches
llm.resolveModelInfoandlistModelsto make models return true for "supports images", passing all three gates (admission/switching/tools) - Automatically wraps streams of all registered adapters and hot-registered adapters; HMR replacement of adapters triggers automatic re-wrapping
- OCR results for each image are cached by attachment ID to avoid redundant calls across rounds
- Each OCR run uses an independent temp directory, automatically cleaned up regardless of success/failure/timeout; also cleans up orphaned temp directories from previous crashes on startup
Technical Implementation
- Language: TypeScript (compiles to
lib/index.js) - Key Dependencies:
@deepseek-ai/cordis(host plugin framework),node:child_process+ PowerShell (executes OCR script),lib/ocr.ps1(wraps Windows.Media.Ocr WinRT API),ctx.attachments.readImage(reads attachment bytes) - Architecture Pattern: Cordis plugin;
cordis.patch.yml'sbundle.patchinserts thewindows-ocrloader entry at the loading layer; the plugin hooks into two public seams of the llm service inapply()to perform capability shimming and request rewriting; effects restore original methods on uninstall - Entry File:
src/index.ts(compiles tolib/index.js, also bundleslib/ocr.ps1)
Use Cases
Sending screenshots, tables, PDF screenshots, code screenshots, or error screenshots to text models daily without uploading image bytes to third-party service providers; or temporarily not switching to a vision-capable model but still letting the other party "read" the text in images. Windows users can directly attach images to text-only conversations in DeepSeek Harness / DSH to use this feature.
Prerequisites & Compatibility
| Dependency | Minimum Version | Notes |
|---|---|---|
| DSH | 0.1.0-rc.8+ (verified in README; package.json doesn't declare dsh field) | Host must provide llm/attachments services and support loader-entry injection at load time |
| Node | >= 20 | From package.json#engines.node |
| Windows PowerShell | 5.1+ | Built into Windows 10/11; plugin invokes WinRT OCR via powershell.exe |
| OCR Language Pack | Depends on target language | E.g., Chinese recognition requires "Chinese (Simplified)" language pack (OCR-capable) |
| Platform | Windows 10/11 | Uses WinRT's Windows.Media.Ocr; no cross-platform implementation |
| Native Modules | None | No npm native dependencies; OCR completed via PowerShell process |
Installation
dsh plugin --profile web add github:maxwell-feng/dsh-windows-ocr
Configuration Options
| Config | Type | Description | Default |
|---|---|---|---|
language | string | BCP-47 language tag for Windows OCR, e.g., zh-Hans, en-US; empty uses current Windows user's language preference | "" |
passthrough | boolean | Default false, all images go through OCR; when set to true, only true vision models receive images as-is, text models still use OCR | false |
ocrScript | string | Absolute path override for PowerShell OCR script, generally no need to change | Bundled lib/ocr.ps1 |
timeoutMs | number | OCR timeout per image (ms); terminates child process and cleans temp dir on timeout | 60000 |
maxCacheEntries | number | In-process OCR result cache limit (by attachment ID) | 200 |
FAQ
Q: Can this plugin be used on macOS or Linux?
A: No. It depends on Windows' Windows.Media.Ocr (WinRT) engine and launches OCR via PowerShell. macOS / Linux have no equivalent capability. The plugin will throw errors or fail to recognize after installation on non-Windows platforms.
Q: Will images be sent to model service providers as-is by default?
A: No. With passthrough defaulting to false, every outbound request goes through local OCR first; the serialized payload only contains text content blocks, no image_url or data URI. You can verify in DevTools → Network.
Q: Chinese characters aren't being recognized. How to handle?
A: The system needs the corresponding Windows OCR language pack installed first (Settings → Time and language → Language → Add language → Chinese (Simplified)). After installation, configure language: zh-Hans to force a specific language, or keep default to let the plugin follow the current user's language preference.
Q: After installation, dsh reports duplicate loader entry id: windows-ocr. How to fix?
A: The same id can only be registered once in dsh 0.1.0-rc.8 (cordis-plugin-loader 1.0.2). If previously installed via npm bundle, dsh plugin ... add github:... will insert the same id line again, causing duplication. Choose one: either use dsh plugin remove to uninstall then reinstall, or use id-based override in profile's cordis.patch.yml (- id: windows-ocr + config:) instead of - insert:.
Q: After uninstall, will there be a gap where images are neither uploaded nor OCR'd?
A: On uninstall, the plugin's effect restores llm.resolveModelInfo / listModels and adapter streams to host originals; text models still reject receiving images per host strategy—this is fail-closed behavior; there's no state where "plugin removed but image blocks still travel naked in requests".
Q: Will OCR-failed images cause the entire conversation request to error?
A: No. Images that fail recognition are replaced with placeholder text blocks (OCR: failed to recognize this image); missing attachments are replaced with (OCR: missing attachment — image refused). Failures only affect that single image, not blocking the entire round.
Q: Do I still need this plugin with a true vision model?
A: Keep the plugin with passthrough set to true if you want true vision models to receive original images; the plugin checks before outbound if the model itself supports images, passing through the original image block if supported, only falling back to OCR if not. This way one plugin covers both text-only and vision models.
Difficulty Level
Beginner — one command to install and use; most scenarios require no config changes; only need to touch passthrough / language when switching languages or letting vision models receive original images.
Known Issues & Limitations
- OCR language depends on system-installed language packs: When OCR script exits with code 2/3, plugin degrades to placeholder text instead of error (lib/ocr.ps1:43-55)
- GIF animations only recognize first frame: Windows OCR engine limitation, not a bug in this plugin
- Cache scope is single dsh process: OCR results persist throughout long sessions until process exit, constrained by
maxCacheEntrieslimit - HMR and dsh full upgrade suggest restart: Plugin listens to
llm/adapters-updatedto automatically re-wrap new adapters, but full restart is most stable after major upgrades - Text models may lack "image" badge in model selector UI:
listModelscapability declaration is shimmed; only the visual marker in model selector UI shows inconsistency—purely cosmetic - No platform support outside Windows: Windows.Media.Ocr calls WinRT interfaces; macOS / Linux have no equivalent
port:3080占用冲突: If port 3080 is occupied whendsh webstarts,EADDRINUSEoccurs; usenetstat -ano | findstr :3080to find PID andtaskkill /PID <pid> /F
DeepSeek Harness (dsh) plugin that lets text-only models accept attached images: every image is recognized locally with the built-in Windows OCR engine (Windows.Media.Ocr) and only the recognized text is sent to the model API.
Privacy default: image bytes are OCR'd locally and not sent to the provider. Set passthrough: true only if you intentionally want genuine vision models to receive original image bytes.
- No configuration changes to your models — no
input: [text, image]hacks insettings.yaml. - Works with any provider/model in dsh; by default every attached image is OCR'd before the request leaves the machine.
- Vision-model passthrough is opt-in (
passthrough: true). - Fail-closed: if the plugin is not loaded, models stay text-only and image attachments are refused — nothing can silently leak. Missing attachments are replaced with a refusal text block (never left as raw
image).
Install from npm
dsh plugin --profile web add @maxwell-feng/dsh-windows-ocr
(Replace web with your profile, e.g. tui.) Prebuilt and published with Sigstore provenance — no source build or allowBuilds approval needed. Installing from source (this repo) still works via the agent guide or the manual steps below.
npm install registers the
windows-ocrrow by itself. The package ships a bundle patch (dsh.bundle+ its owncordis.patch.yml) that inserts thewindows-ocrloader entry. Do not also add a manual- insert:row with the same id to your profile — dsh0.1.0-rc.8(cordis-plugin-loader1.0.2) rejects duplicate loader entry ids anddsh webfails to boot withduplicate loader entry id: windows-ocr.
Quick install via an AI agent
Hand this repository to any AI agent, or paste the instruction below, and the agent will install and verify the plugin for you:
Please install the dsh plugin in this repository by following https://github.com/maxwell-feng/dsh-windows-ocr/blob/main/agents-install.md. Run every preflight check, choose an install mode, then complete the mandatory verification: attach an image to a text-only model session and confirm the model answers with the recognized text.
agents-install.md is a step-by-step guide written for
AI agents: preflight checks, both install modes (permanent profile patch /
temporary --patch overlay), mandatory functional verification, and
troubleshooting for the failure modes you are likely to hit. Manual install
instructions are below.
Why a plugin (not a skill)
dsh skills are Markdown instruction files injected into the model context — they cannot execute code, cannot hook the request pipeline, and cannot stop an image from being serialized. This feature needs exactly that, so it is a cordis plugin that hooks two public seams of the llm service:
- Capability shim —
ctx.llm.resolveModelInfo(alsolistModels). The host gates image attachments oninputModalities.includes("image")at three places: message admission, model switching, and theread_imagetool. The shim answers "yes", so text models admit images. - Request rewrite —
registration.adapter.stream(the single choke point bothctx.llm.streamandprepareCall().streamfunnel through). Everyimagecontent block is replaced with an OCR text block before the adapter serializes the request, so the adapter's own image check never fires, no attachment bytes are read for the wire, and noimage_urlis ever built.
you attach an image
→ admission asks ctx.llm.resolveModelInfo (shimmed: "image" ✓)
→ image stored in the local attachment store (session log, UI preview)
→ agent builds the request → adapter.stream (wrapped)
→ image block read locally (ctx.attachments.readImage) → Windows OCR
→ block replaced with <image_ocr>…text…</image_ocr>
→ adapter serializes a text-only request → provider
Requirements
- Windows 10/11 (Windows PowerShell 5.1+ ships with the OS; no install needed)
- A Windows OCR-capable language pack for your language (Settings → Time & language → Language). English is usually present; Chinese requires the Chinese language pack (OCR-capable).
dshwith a profile (tested against dsh0.1.0-rc.8)
Install
Installing via an AI agent
agents-install.md in this repository is a
step-by-step installation guide written for AI agents (and careful
humans). Give it to an agent — e.g. "install this plugin per
agents-install.md from https://github.com/maxwell-feng/dsh-windows-ocr" —
and the agent can perform the preflight checks, install, verification, and
troubleshooting on its own. The guide covers both install modes, the
mandatory functional verification (attach an image → model answers with the
OCR text), and the failure modes you are likely to hit.
Manual install
Two official ways to load this plugin, both referencing the plugin file by
absolute path (see docs/user/develop/basic). On Windows the path must be
a file:// URL — a bare C:/... path is parsed as the c: URL scheme and
the loader rejects it.
Permanent: profile patch layer
Append to your profile's cordis.patch.yml (e.g. ~/.dsh/profiles/web/cordis.patch.yml):
- insert:
- id: windows-ocr
name: 'file:///C:/absolute/path/to/windows-ocr/lib/index.js'
config:
language: ''
passthrough: false
Then restart dsh web. Remove the rows to uninstall — the plugin restores the original llm / adapter methods on unload.
Choose one way to load the plugin: the npm bundle (above) or this manual insert — never both. Both register the same
windows-ocrentry id, and dsh0.1.0-rc.8fails the boot withduplicate loader entry id: windows-ocrwhen the row exists twice. If the row is already present (for example after an npm bundle install), configure it with an id-targeted override (see Configuration below) instead of inserting a second row.
Temporary: --patch overlay
Put the same rows in an overlay file and boot with it; your profile stays untouched:
dsh --profile web --patch C:/path/to/overlay.yml
Notes
dsh webfails withEADDRINUSEon port 3080 when an older instance is still running: find it withnetstat -ano | findstr :3080and stop it (taskkill /PID/F) before starting a new one. - For a packaged install (npm / tarball /
github:user/repo), package the plugin as a bundle (dsh.bundle+cordis.patch.yml, seedocs/user/develop/basic/publish); a git install additionally needs apreparebuild script and pnpmallowBuildsconsent.
To verify the plugin loaded, look for windows-ocr in the boot logs, or check the OCR smoke test below.
Configuration
All settings live in the patch row windows-ocr (cordis.patch.yml here) and can be overridden from your profile's cordis.patch.yml:
| Key | Default | Meaning |
|---|---|---|
language | "" | BCP-47 tag for Windows OCR, e.g. zh-Hans, en-US. Empty = user profile languages. |
passthrough | false | false (default): OCR every image. true: genuine vision models receive images untouched. |
ocrScript | bundled lib/ocr.ps1 | Absolute path override for the PowerShell OCR script. |
timeoutMs | 60000 | Per-image OCR timeout. |
maxCacheEntries | 200 | Bound on the per-run OCR cache (keyed by attachment id). |
Example override in ~/.dsh/profiles/web/cordis.patch.yml — an id-targeted
row (not insert:) replaces the existing windows-ocr row's config:
- id: windows-ocr
config:
language: zh-Hans
How the model sees the image
Each image block becomes a text block (local filenames are not forwarded):
<image_ocr>
…recognized lines…
</image_ocr>
Recognition text is cached per attachment id for the lifetime of the dsh process, so repeated turns do not re-run OCR.
Temp-file hygiene
Every OCR run writes its input image and output text into a fresh temporary
directory (windows-ocr-* under the system temp dir). The directory is
removed automatically in finally — on success, on OCR error, and on timeout —
so no per-run script, image, or output file survives. At plugin start, any
orphaned windows-ocr-* directories left behind by a previously crashed
process are swept as well. Nothing is written outside the plugin's own
temporary directory and the dsh attachment store.
Smoke test (no dsh needed)
# 1x1 PNG — exercises WinRT loading, language availability, recognition
powershell.exe -NoProfile -ExecutionPolicy Bypass -File lib/ocr.ps1 -ImagePath test.png -OutFile out.txt
Get-Content out.txt
Exit code 0 with an empty/whitespace out.txt means the OCR engine works (a 1×1 image has no text). Exit 2/3 means a language pack is missing.
Verification inside dsh
- Attach an image to a text-model session and send a message — the model should answer using the recognized text.
- Confirm the image never goes out: open DevTools → Network in the web UI, inspect the request to your provider base URL, and verify the payload contains only
textcontent parts (noimage_url/ data URI).
Limitations
- OCR language availability depends on installed Windows language packs (script exits 2/3 and the plugin degrades to a placeholder text).
- GIFs: Windows OCR recognizes the first frame.
- Cache is per process; a long-lived session keeps OCR text cached, bounded by
maxCacheEntries. - Hot reload (HMR) replaces adapters; the plugin re-wraps new adapters on
llm/adapters-updated, but a full restart is the safe path after any dsh update. - The model picker may show text models without an "image" badge (cosmetic only —
listModelsis shimmed consistently). - If the OCR plugin is removed, image attachments to text models are refused again (fail-closed), not uploaded.
License
MIT
Read the usage guide →
Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.
Listing badge
[](https://deepseek-plugin.org/plugins/maxwell-feng/dsh-windows-ocr)Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.