让纯文本模型也能"看图":调用 Windows 自带 OCR 在本机把图片识别为文字,只把文字发往模型,图片字节从不出本地。
- 语言
- JavaScript
- License
- MIT
- 分支
- main
安装
$ dsh plugin --profile web add dsh-windows-ocr在终端中运行以上命令,通过 dsh CLI 安装此插件。可在右上角切换 Profile。 第一次用 dsh?看这篇新手教程
对话式安装
帮我安装 DeepSeek Harness 插件 maxwell-feng/dsh-windows-ocr:先查看仓库 https://github.com/maxwell-feng/dsh-windows-ocr 确认安全性,然后执行安装命令并验证插件加载成功。
把这段指令粘贴给 DSH Web GUI 里的助手,由它代你完成安装与验证。
一句话定位
给 DSH 上纯文本模型补上"看图附件"的能力:插件在前置环节调用 Windows 系统自带的 OCR 引擎把图片识别成文字,只有文字会随请求发到服务商;不带该插件时文本模型继续拒绝收图片,默认安全。
核心能力
- 让纯文本模型也能接收图片附件,请求里只会出现识别出的文字而不是图片字节
- 拦截图片序列化路径,在适配器发出请求前把 image 内容块改写成文字块(含嵌套的 tool-result 内容)
- 给 llm.resolveModelInfo 与 listModels 打补丁,让模型在"是否支持图片"这一项返回 true,从而通过准入/切换/工具三道闸
- 自动包装所有已注册和后续热注册的适配器 stream,HMR 替换适配器后会自动重新包装
- 每张图片的 OCR 结果按附件 id 缓存,重复轮次不重复调用
- 每次 OCR 用独立临时目录,跑完无论成功/失败/超时都自动清理;启动时还会清扫上一次崩溃遗留的孤儿临时目录
技术实现
- 语言: TypeScript(编译后输出
lib/index.js) - 关键依赖:
@deepseek-ai/cordis(宿主插件框架)、node:child_process+ PowerShell(执行 OCR 脚本)、lib/ocr.ps1(封装 Windows.Media.Ocr WinRT API)、ctx.attachments.readImage(读本机附件字节) - 架构模式: Cordis 插件,cordis.patch.yml 的
bundle.patch在加载层插入windows-ocrloader 条目;插件在 apply() 中钩住 llm 服务的两个公开接缝做能力 shim 和请求改写,effect 卸载时再恢复原方法 - 入口文件:
src/index.ts(编译产物lib/index.js,同时打包lib/ocr.ps1)
适用场景
日常给文本模型发截图、表格、PDF 截图、代码截图、报错截屏,又不想把图片字节上传到第三方服务商;或者临时不想切到带视觉能力的模型也能让对方"读"图中文本。Windows 用户在 DeepSeek Harness / DSH 里给纯文本对话直接附加图片就能用。
前置依赖与兼容性
| 依赖 | 最低版本 | 说明 |
|---|---|---|
| DSH | 0.1.0-rc.8+(README 声称在该版本验证;package.json 未声明 dsh 字段) | 宿主需提供 llm / attachments 服务,并在加载时支持 loader-entry 注入 |
| Node | >= 20 | 来自 package.json#engines.node |
| Windows PowerShell | 5.1+ | Windows 10/11 自带;插件通过 powershell.exe 调用 WinRT OCR |
| OCR 语言包 | 视所识别语言而定 | 例如识别中文需装"中文(简体)"语言包(OCR-capable) |
| 平台 | Windows 10/11 | 调 WinRT 的 Windows.Media.Ocr,无跨平台实现 |
| 原生模块 | 无 | 没有 npm 原生依赖;OCR 通过 PowerShell 进程调用完成 |
安装方式
dsh plugin --profile web add github:maxwell-feng/dsh-windows-ocr
配置项
| 配置 | 类型 | 说明 | 默认值 |
|---|---|---|---|
language | 字符串 | Windows OCR 使用的 BCP-47 语言标签,例如 zh-Hans、en-US;留空用当前 Windows 用户的语言偏好 | "" |
passthrough | 布尔 | 默认 false,所有图片一律 OCR;设为 true 后只有真正的视觉模型才会原样接收图片,文本模型仍走 OCR | false |
ocrScript | 字符串 | PowerShell OCR 脚本的绝对路径覆盖,一般无需改 | 自带 lib/ocr.ps1 |
timeoutMs | 数字 | 单张图片 OCR 超时(毫秒),超时后会终止子进程并清理临时目录 | 60000 |
maxCacheEntries | 数字 | 进程内 OCR 结果缓存(按附件 id)上限 | 200 |
常见问题
Q: 插件能用在 macOS 或 Linux 上吗?
A: 不能。它依赖 Windows 的 Windows.Media.Ocr(WinRT)引擎并通过 PowerShell 启动 OCR,macOS / Linux 没有等价能力,安装后插件启动会报错或识别不出来。
Q: 默认情况下图片会不会原样发送给模型服务商?
A: 不会。默认 passthrough 为 false,每次出站请求都会先本机 OCR,序列化 payload 里只有 text 内容块,不会出现 image_url 或 data URI;可以在 DevTools → Network 里抓包确认。
Q: 中文识别不出来,怎么处理?
A: 系统需要先装好对应语言的 Windows OCR 语言包(设置 → 时间和语言 → 语言 → 添加语言 → 中文(简体))。装好后可配 language: zh-Hans 强制指定语言,或保持默认让插件跟随当前用户的语言偏好。
Q: 安装后 dsh 启动报 duplicate loader entry id: windows-ocr 怎么解?
A: 同一 id 在 dsh 0.1.0-rc.8(cordis-plugin-loader 1.0.2)只能注册一次。如果之前用 npm bundle 装过,dsh plugin ... add github:... 会再次插入同 id 行就会重复。二选一:要么用 dsh plugin remove 卸掉再装,要么在 profile 的 cordis.patch.yml 里用按 id 覆盖(- id: windows-ocr + config:)而不是 - insert:。
Q: 卸载后会不会留下空窗期?图片会不会既不上传又不被 OCR?
A: 卸载时插件 effect 会把 llm.resolveModelInfo / listModels 和适配器 stream 恢复成宿主原版方法,文本模型仍会按宿主策略拒绝收图片——属于 fail-closed 行为;不会出现"插件卸了但 image 块还在请求里裸跑"的状态。
Q: OCR 失败的图片会不会变成请求报错整条对话?
A: 不会。识别失败的图片会被替换为占位文本块 (OCR: failed to recognize this image);缺失 attachment 的会被替换为 (OCR: missing attachment — image refused)。失败只影响这一张图,不阻断整轮对话。
Q: 我用真正的视觉模型还要这个插件吗?
A: 想让真视觉模型仍接收原图就保留插件但把 passthrough 设为 true;插件会在出站前判断模型本身是否支持图片,支持就放行原始 image 块,不支持才走 OCR。这样可以一套插件同时覆盖纯文本模型和带图模型。
上手难度
入门 — 安装一行命令即可使用,绝大多数场景无需改配置;只有需要切语言或让视觉模型收原图时才动 passthrough / language 两项。
已知问题与限制
- OCR 语言取决于系统已装的语言包:OCR 脚本退出码 2/3 时插件降级为占位文本而非报错(lib/ocr.ps1:43-55)
- GIF 动图只识别第一帧:Windows OCR 引擎行为限制,并非本插件 bug
- 缓存作用域为单 dsh 进程:长会话内 OCR 结果一直驻留直到进程退出,受
maxCacheEntries上限约束 - HMR 与 dsh 整体升级建议重启:插件会监听
llm/adapters-updated自动重新包装新适配器,但彻底升级后完整重启最稳妥 - 模型选择器 UI 上文本模型可能不带"图片"徽标:
listModels的能力声明已 shim 过,仅模型选择界面的视觉标记存在不一致,纯外观问题 - 不支持 Windows 之外的平台:Windows.Media.Ocr 调的是 WinRT 接口,macOS / Linux 无等价实现
port:3080占用冲突:dsh web启动若 3080 被占用会EADDRINUSE,需用netstat -ano | findstr :3080找到 PID 并taskkill /PID <pid> /F
DeepSeek Harness (dsh) plugin that lets text-only models accept attached images: every image is recognized locally with the built-in Windows OCR engine (Windows.Media.Ocr) and only the recognized text is sent to the model API.
Privacy default: image bytes are OCR'd locally and not sent to the provider. Set passthrough: true only if you intentionally want genuine vision models to receive original image bytes.
- No configuration changes to your models — no
input: [text, image]hacks insettings.yaml. - Works with any provider/model in dsh; by default every attached image is OCR'd before the request leaves the machine.
- Vision-model passthrough is opt-in (
passthrough: true). - Fail-closed: if the plugin is not loaded, models stay text-only and image attachments are refused — nothing can silently leak. Missing attachments are replaced with a refusal text block (never left as raw
image).
Install from npm
dsh plugin --profile web add @maxwell-feng/dsh-windows-ocr
(Replace web with your profile, e.g. tui.) Prebuilt and published with Sigstore provenance — no source build or allowBuilds approval needed. Installing from source (this repo) still works via the agent guide or the manual steps below.
npm install registers the
windows-ocrrow by itself. The package ships a bundle patch (dsh.bundle+ its owncordis.patch.yml) that inserts thewindows-ocrloader entry. Do not also add a manual- insert:row with the same id to your profile — dsh0.1.0-rc.8(cordis-plugin-loader1.0.2) rejects duplicate loader entry ids anddsh webfails to boot withduplicate loader entry id: windows-ocr.
Quick install via an AI agent
Hand this repository to any AI agent, or paste the instruction below, and the agent will install and verify the plugin for you:
Please install the dsh plugin in this repository by following https://github.com/maxwell-feng/dsh-windows-ocr/blob/main/agents-install.md. Run every preflight check, choose an install mode, then complete the mandatory verification: attach an image to a text-only model session and confirm the model answers with the recognized text.
agents-install.md is a step-by-step guide written for
AI agents: preflight checks, both install modes (permanent profile patch /
temporary --patch overlay), mandatory functional verification, and
troubleshooting for the failure modes you are likely to hit. Manual install
instructions are below.
Why a plugin (not a skill)
dsh skills are Markdown instruction files injected into the model context — they cannot execute code, cannot hook the request pipeline, and cannot stop an image from being serialized. This feature needs exactly that, so it is a cordis plugin that hooks two public seams of the llm service:
- Capability shim —
ctx.llm.resolveModelInfo(alsolistModels). The host gates image attachments oninputModalities.includes("image")at three places: message admission, model switching, and theread_imagetool. The shim answers "yes", so text models admit images. - Request rewrite —
registration.adapter.stream(the single choke point bothctx.llm.streamandprepareCall().streamfunnel through). Everyimagecontent block is replaced with an OCR text block before the adapter serializes the request, so the adapter's own image check never fires, no attachment bytes are read for the wire, and noimage_urlis ever built.
you attach an image
→ admission asks ctx.llm.resolveModelInfo (shimmed: "image" ✓)
→ image stored in the local attachment store (session log, UI preview)
→ agent builds the request → adapter.stream (wrapped)
→ image block read locally (ctx.attachments.readImage) → Windows OCR
→ block replaced with <image_ocr>…text…</image_ocr>
→ adapter serializes a text-only request → provider
Requirements
- Windows 10/11 (Windows PowerShell 5.1+ ships with the OS; no install needed)
- A Windows OCR-capable language pack for your language (Settings → Time & language → Language). English is usually present; Chinese requires the Chinese language pack (OCR-capable).
dshwith a profile (tested against dsh0.1.0-rc.8)
Install
Installing via an AI agent
agents-install.md in this repository is a
step-by-step installation guide written for AI agents (and careful
humans). Give it to an agent — e.g. "install this plugin per
agents-install.md from https://github.com/maxwell-feng/dsh-windows-ocr" —
and the agent can perform the preflight checks, install, verification, and
troubleshooting on its own. The guide covers both install modes, the
mandatory functional verification (attach an image → model answers with the
OCR text), and the failure modes you are likely to hit.
Manual install
Two official ways to load this plugin, both referencing the plugin file by
absolute path (see docs/user/develop/basic). On Windows the path must be
a file:// URL — a bare C:/... path is parsed as the c: URL scheme and
the loader rejects it.
Permanent: profile patch layer
Append to your profile's cordis.patch.yml (e.g. ~/.dsh/profiles/web/cordis.patch.yml):
- insert:
- id: windows-ocr
name: 'file:///C:/absolute/path/to/windows-ocr/lib/index.js'
config:
language: ''
passthrough: false
Then restart dsh web. Remove the rows to uninstall — the plugin restores the original llm / adapter methods on unload.
Choose one way to load the plugin: the npm bundle (above) or this manual insert — never both. Both register the same
windows-ocrentry id, and dsh0.1.0-rc.8fails the boot withduplicate loader entry id: windows-ocrwhen the row exists twice. If the row is already present (for example after an npm bundle install), configure it with an id-targeted override (see Configuration below) instead of inserting a second row.
Temporary: --patch overlay
Put the same rows in an overlay file and boot with it; your profile stays untouched:
dsh --profile web --patch C:/path/to/overlay.yml
Notes
dsh webfails withEADDRINUSEon port 3080 when an older instance is still running: find it withnetstat -ano | findstr :3080and stop it (taskkill /PID/F) before starting a new one. - For a packaged install (npm / tarball /
github:user/repo), package the plugin as a bundle (dsh.bundle+cordis.patch.yml, seedocs/user/develop/basic/publish); a git install additionally needs apreparebuild script and pnpmallowBuildsconsent.
To verify the plugin loaded, look for windows-ocr in the boot logs, or check the OCR smoke test below.
Configuration
All settings live in the patch row windows-ocr (cordis.patch.yml here) and can be overridden from your profile's cordis.patch.yml:
| Key | Default | Meaning |
|---|---|---|
language | "" | BCP-47 tag for Windows OCR, e.g. zh-Hans, en-US. Empty = user profile languages. |
passthrough | false | false (default): OCR every image. true: genuine vision models receive images untouched. |
ocrScript | bundled lib/ocr.ps1 | Absolute path override for the PowerShell OCR script. |
timeoutMs | 60000 | Per-image OCR timeout. |
maxCacheEntries | 200 | Bound on the per-run OCR cache (keyed by attachment id). |
Example override in ~/.dsh/profiles/web/cordis.patch.yml — an id-targeted
row (not insert:) replaces the existing windows-ocr row's config:
- id: windows-ocr
config:
language: zh-Hans
How the model sees the image
Each image block becomes a text block (local filenames are not forwarded):
<image_ocr>
…recognized lines…
</image_ocr>
Recognition text is cached per attachment id for the lifetime of the dsh process, so repeated turns do not re-run OCR.
Temp-file hygiene
Every OCR run writes its input image and output text into a fresh temporary
directory (windows-ocr-* under the system temp dir). The directory is
removed automatically in finally — on success, on OCR error, and on timeout —
so no per-run script, image, or output file survives. At plugin start, any
orphaned windows-ocr-* directories left behind by a previously crashed
process are swept as well. Nothing is written outside the plugin's own
temporary directory and the dsh attachment store.
Smoke test (no dsh needed)
# 1x1 PNG — exercises WinRT loading, language availability, recognition
powershell.exe -NoProfile -ExecutionPolicy Bypass -File lib/ocr.ps1 -ImagePath test.png -OutFile out.txt
Get-Content out.txt
Exit code 0 with an empty/whitespace out.txt means the OCR engine works (a 1×1 image has no text). Exit 2/3 means a language pack is missing.
Verification inside dsh
- Attach an image to a text-model session and send a message — the model should answer using the recognized text.
- Confirm the image never goes out: open DevTools → Network in the web UI, inspect the request to your provider base URL, and verify the payload contains only
textcontent parts (noimage_url/ data URI).
Limitations
- OCR language availability depends on installed Windows language packs (script exits 2/3 and the plugin degrades to a placeholder text).
- GIFs: Windows OCR recognizes the first frame.
- Cache is per process; a long-lived session keeps OCR text cached, bounded by
maxCacheEntries. - Hot reload (HMR) replaces adapters; the plugin re-wraps new adapters on
llm/adapters-updated, but a full restart is the safe path after any dsh update. - The model picker may show text models without an "image" badge (cosmetic only —
listModelsis shimmed consistently). - If the OCR plugin is removed, image attachments to text models are refused again (fail-closed), not uploaded.
License
MIT
查看使用指南 →
该插件的安装步骤、关键要点、FAQ 与兼容性说明(基于已收录字段派生)。
收录徽章
[](https://deepseek-plugin.org/plugins/maxwell-feng/dsh-windows-ocr)把这段 markdown 粘贴到你的 GitHub README,链接回本插件详情页。徽章只声明已被本站收录,不代表安全认证。