为 DSH Web UI 加上 Claude/Codex 风格的文件上传:拖拽、回形针、粘贴均可,文档自动转 Markdown,文本模型也能读懂图片。
- 语言
- TypeScript
- License
- MIT
- 分支
- main
安装
$ dsh plugin --profile web add dsh-file-upload在终端中运行以上命令,通过 dsh CLI 安装此插件。可在右上角切换 Profile。 第一次用 dsh?看这篇新手教程
对话式安装
帮我安装 DeepSeek Harness 插件 HongMing-Huang/dsh-file-upload:先查看仓库 https://github.com/HongMing-Huang/dsh-file-upload 确认安全性,然后执行安装命令并验证插件加载成功。
把这段指令粘贴给 DSH Web GUI 里的助手,由它代你完成安装与验证。
一句话定位
给 DSH Web UI 补上 Claude/Codex 风格的文件消息能力:回形针按钮、整页拖拽、剪贴板粘贴三种上传方式,文档类文件自动转成 Markdown,图片给文本模型配"自动讲解",Agent 通过 read_document 工具按需读取已上传文件。
核心能力
- 在会话输入框工具栏放一个回形针按钮,点开系统文件选择器;也支持把文件 / 文件夹拖到 Web 窗口任意位置(带全屏"释放以附件"遮罩),还支持把剪贴板里的图片或文件直接粘贴进输入框
- 上传后以 Codex 风格的
@相对路径形式插入输入框,原始字节不内联进草稿;Agent 通过read_document <path>读取,PDF/DOCX/XLSX 按需转 Markdown 并支持offset/limit分页 - 输入框里输入
@触发已上传文件的相对路径选择器,引用按官方ReferenceInsert方式插入 - 自带 MarkItDown 引擎(
markitdown-node,作为普通 npm 依赖打包进插件),无需 Python、无需下载,开箱支持 PDF/DOCX/PPTX/XLSX/HTML/CSV/JSON/XML/ZIP/Jupyter 等 20+ 格式的本地离线解析 - 图片上传时探测会话路由的模型能力:多模态模型直接走官方
read_image;文本模型(DeepSeek API 是文本-only)通过"显式端点 → 本地 Ollama VL 模型(如 DeepSeek-VL2)→ OpenAI 标准端点"的发现链自动生成"讲解图片",描述随消息发出 - 上传的文件按会话隔离存在
.dsh-uploads/<sessionId>/,loopback-only 监听、同源校验、文件名清洗、sha256 字节级去重、并发超限返 429、过期目录后台定时清扫 - 文档读取有按 (目标, 文件版本, 类型) 维度的 LRU 转换缓存(条目数和字节数双重预算),文件改了自动失效
技术实现
- 语言: TypeScript(host 端 + React 客户端),ES2020,esbuild 打 client bundle
- 关键依赖:
@deepseek-ai/dsh-fs(ctx.fs 读上传文件,继承沙箱和 fs-observation 策略)、@deepseek-ai/dsh-tools(注册read_document工具)、@deepseek-ai/dsh-credentials(解析 vision key)、markitdown-node(打包进插件的 MarkItDown TS 端口) - 架构模式: 双面 DSH 插件。host 端通过 Cordis 在
webServer上注册/api/upload的 prefix 路由,注册read_document工具到ctx.tools,向ctx.systemPrompt注入tool:read-document章节;client 端按package.json#dsh.client.platform=web注入到 DSH 浏览器运行时,通过slash/input-insert-text与slash/input-insert-reference事件把上传结果写进对话输入状态服务 - 入口文件:
src/index.ts(host 端 apply + Config schema)、src/client/index.tsx(paperclip + drag overlay + paste)、cordis.patch.yml(bundle patch 注入点)
适用场景
普通用户用 DSH Web 跟模型聊项目时经常要把本地文档丢进去:传一份 PDF 让模型总结、贴一张截图问"这是啥报错"、把整个文件夹拖进来让模型按文件名读。原生 DSH Web 没有上传通道,只能贴文本或在文件系统里 cd 进去手动 @ 路径。装上这个插件后,按回形针或直接拖,回形针按文件类型贴彩色徽标,文档以 @相对路径 引用形式进输入框,PDF/DOCX/XLSX 由内置的 MarkItDown 引擎转成 Markdown 让模型"读得到",图片给文本模型自动生成一段描述让 DeepSeek API 也能聊图。
前置依赖与兼容性
| 依赖 | 最低版本 | 说明 |
|---|---|---|
| DSH | 0.1.0-rc.6+ | peerDependencies 要求 @deepseek-ai/dsh-credentials / dsh-fs / dsh-tools 均为 ^0.1.0-rc.6 |
| Node.js | >=22.6.0 | 来自 package.json#engines.node |
| 平台 | 跨平台 | host 端跑 Node,client 端跑浏览器;上传接口仅接受 loopback 主机(127.0.0.1 / localhost / ::1) |
| 原生模块 | 无 | 仅使用 Node 内置 node:fs / node:crypto / node:path / node:child_process / node:http;浏览器侧使用标准 Fetch / Clipboard API |
安装方式
dsh plugin --profile web add github:HongMing-Huang/dsh-file-upload
配置项
所有字段都有合理默认,安装即可使用;插件注册为 Cordis row id file-upload,可按行覆盖下列任一字段。
| 配置 | 类型 | 说明 | 默认值 |
|---|---|---|---|
uploadMaxBytes | number | 单个上传体的字节上限 | 24 MiB(25165824) |
allowedExtensions | string[] | 小写扩展名白名单;空数组表示全部允许 | [] |
uploadTtlMs | number | 上传目录的存活时长,超时被清扫 | 7 天(604800000) |
sweepIntervalMs | number | 后台清扫周期;设为 0 关闭定时清扫 | 1 小时(3600000) |
maxConcurrentUploads | number | 同时进行的最大上传数,超出返 429 | 4 |
inlineTextLimit | number | 直接内联进输入框的字节上限(小文本文件) | 8 KiB(8192) |
previewTextLimit | number | 较大文本文件随响应返回的预览字节上限 | 2 KiB(2048) |
maxFileBytes | number | read_document 一次允许读取的字节上限(PDF 解析会放大内存) | 24 MiB(25165824) |
readLimit | number | read_document 单次返回的最大行数 | 2000 |
sheetRowLimit | number | XLSX 每个 sheet 保留的行数 | 200 |
maxSheets | number | XLSX 单次读取的 sheet 数;超出在尾部标注"未读取" | 5 |
cacheEntries | number | 文档解析 LRU 缓存的条目数上限 | 16 |
cacheMaxBytes | number | 文档解析 LRU 缓存的字节预算上限 | 64 MiB(67108864) |
markitdownBin | string | 可选:官方 markitdown CLI 的绝对路径;空字符串表示自动检测 PATH | "" |
markitdownTimeoutMs | number | MarkItDown CLI 单次调用的超时 | 120000 ms |
visionEndpoint | string | 图片讲解用的 OpenAI 兼容端点;空 = 自动发现(Ollama → OpenAI 标准) | "" |
visionModel | string | 图片讲解用的模型名;空 = 自动 | "" |
visionApiKeyEnv | string | 在 dsh 凭证系统中要解析的 key 名 | OPENAI_API_KEY |
visionMaxBytes | number | 传给讲解端点的图片字节上限 | 10 MiB(10485760) |
uploadDir | string | 无 sessions 服务时的兜底上传根目录 | <cwd>/uploads |
常见问题
Q: 上传后原始文件内容会进输入框吗?
A: 不会。上传后界面只插入一个 Codex 风格的 @相对路径 引用,原始字节不被内联进草稿;Agent 按需调用 read_document <path> 读取,PDF/DOCX/XLSX 会按需转 Markdown 并支持分页。
Q: 上传后文本模型(比如 DeepSeek API)为什么能"读懂"我发的图片?
A: 上传时插件会探测当前会话路由的模型:多模态模型走官方 read_image;否则触发"讲解图片"自动生成描述,按"显式 visionEndpoint → 本地 Ollama VL 模型 → OpenAI 标准端点"的顺序自动发现,描述随消息一起发出。
Q: 上传的文件存哪里?安全吗?
A: 存在当前会话工作区的 .dsh-uploads/<sessionId>/ 下,按会话隔离、未授权会话返回 403;HTTP 仅监听 loopback 主机、上传时同源校验、文件名做清洗、相同内容用 sha256 去重、并发数超限返 429,超过 TTL 的目录会被后台清扫。
Q: 必须安装 Python 或 markitdown CLI 吗?
A: 不必。MarkItDown 能力已作为 markitdown-node 依赖打进插件(覆盖 PDF/DOCX/PPTX/XLSX/HTML/CSV/JSON/XML/ZIP/Jupyter 等 20+ 格式,含图片 OCR),完全本地离线运行;若机器上已经有官方 markitdown CLI,插件会自动改用它来支持 EPUB 等更多格式。
Q: 文件上传后多久会被自动删除?
A: 默认 7 天(uploadTtlMs=604800000 毫秒)。后台每 1 小时扫一次(sweepIntervalMs=3600000 毫秒),扫描每个会话工作区下的 .dsh-uploads/<sessionId>/,整目录超时即清;把 sweepIntervalMs 设为 0 可关闭定期清扫。
Q: 安装后界面没出现回形针按钮,怎么办?
A: 客户端 bundle 只在全新页面加载时挂载。先 dsh plugin --profile web add github:HongMing-Huang/dsh-file-upload,再重启 DSH Web 服务,最后在浏览器做硬刷新(Cmd/Ctrl+Shift+R)。
Q: 单个文件最大能传多大?能限制扩展名吗?
A: 默认上限 24 MB(uploadMaxBytes),可在配置中改大;allowedExtensions 留空表示不限制,填入小写扩展名数组(如 ["pdf","docx","png"])后其他扩展名会被 415 拒绝。
Q: DSH 卸载 / 升级这个插件的命令是什么?
A: 升级:dsh plugin --profile web update github:HongMing-Huang/dsh-file-upload;卸载:dsh plugin --profile web remove github:HongMing-Huang/dsh-file-upload;之后同样要重启 Web 服务并硬刷新浏览器。
上手难度
入门 — 一条 dsh plugin 命令安装即用,无需配置文件、无需 Python/CLI,所有功能开箱默认;只需在首次安装或升级后重启 DSH Web 服务并硬刷浏览器。
已知问题与限制
- 上传 HTTP 接口只接受 loopback 主机(127.0.0.1 / localhost / ::1),外网访问会被 403 拒绝,浏览器侧发起上传须经 DSH Web 同源代理(
src/upload.ts:67-69、src/upload.ts:267-272) - 同一时刻最多 4 个上传并发,超出会立即返 429;超 24 MB 的请求体在到达之前就被 413 拒掉(
src/upload.ts:124-148) uploadDir兜底仅在 host 注入的 sessions 服务缺席时生效;正常情况下文件按 session 落在.dsh-uploads/<sessionId>/下,没有 sessions 服务的环境(src/index.ts:103)将整体落到<cwd>/uploads- GB18030 解码依赖运行时的
TextDecoder;没有 GBK decoder 的 Node 版本会回退到latin1(lossy),含中文且编码不是 UTF-8/UTF-16 的文本会被部分字符替换(src/convert.ts:65-76) - 0.5.0 起语音输入(mic 按钮、Web Speech API、MediaRecorder 回退、ASR 转写链路)被整体移除,源码层面只剩上传图片/文件路径;音频文件可作为普通附件上传但不会再被转写(
CHANGELOG.md:32-37) - DeepSeek-VL2 / qwen2.5-vl / llava / moondream / gemma3 / internvl 等视觉模型通过 Ollama 检测时按 VL 命名匹配;如果本地 Ollama 里只装了不带 VL 关键字的视觉模型,本地路径不会命中,会直接跳到 OpenAI 标准端点(
src/vision.ts:48-49) - 当显式端点、本地 Ollama、OpenAI key 三条链路全部走不通时,
describeImage直接抛错,上传响应里的imageMode会回退到ocr、图片以"路径引用 + 未生成讲解"的形式进消息(src/vision.ts:74-76、src/upload.ts:205-215) - MarkItDown CLI 解析失败会降级到
markitdown-node,再失败降级到内嵌 JS 解析器(PDF/DOCX/XLSX 最小集);CLI 调用超时由markitdownTimeoutMs控制,默认 120 秒(src/convert.ts:251-273) - 源码中未发现
TODO/FIXME/HACK/XXX注释;0.4.3 修复了"TTL 定时清扫不覆盖 session workspace"的历史 bug,0.5.0 移除了不可靠的 ASR 与语音输入
File-message plugin for DeepSeek Harness (dsh). Claude/Codex-style uploads — drag-and-drop (files and folders), paperclip picker, paste-to-attach, multi-file support; content sniffing; fully bundled document → Markdown conversion (MarkItDown engine, 20+ formats, image OCR); Codex-style @relative/path references; automatic image explanations for text-only models; and a read_document tool for agents.
English | 中文
Zero-config, install-and-use. Every feature works out of the box with sensible defaults — no Python, no downloads, no picking backends. Image explanations auto-discover a vision endpoint (local Ollama → OpenAI-compatible key from the dsh credentials seam).
Features
- Upload — composer paperclip button plus a global drag-and-drop overlay ("release to attach"), multi-file support.
- Attachment cards — color-coded type badges (PDF red / DOC blue / XLS green / TXT gray / ZIP purple / JSON gold) with name and size; removable.
- Codex-style file references — uploaded files appear in the message as
@relative/pathreferences (like OpenAI Codex), never as raw content dumped into the composer; the agent reads the file withread_document(converted to Markdown on demand). - Codex-style
@mentions — type@in the composer to pick any uploaded file by its relative path; the reference inserts as a mention. - Document → Markdown, fully bundled — the MarkItDown engine ships inside the plugin (Microsoft MarkItDown TypeScript port,
markitdown-node): PDF / DOCX / PPTX / XLSX / HTML / CSV / JSON / XML / RSS / Atom / ZIP / Jupyter / image OCR / audio transcription. No Python, no downloads, no setup. - Image explanation for text-only models — upload an image and the plugin automatically generates a description ("讲解图片") through a vision discovery chain, so the DeepSeek API (text-only) can reason about the image: explicit
visionEndpoint→ local Ollama with a VL model (e.g. DeepSeek-VL2, zero-config) → OpenAI-compatible endpoint with a dsh-credentials key. Multimodal routes / vision bridges keep the officialread_imagepath. read_documenttool for agents — line-numbered paging (offset/limit), byte-budgeted LRU cache (invalidated on file change), size pre-checks, reads throughctx.fs(inherits sandbox and fs-observation policy).- Security — loopback-only uploads, sanitized file names, session-isolated storage (
.dsh-uploads/<sessionId>), sha256 content dedup, bounded concurrency, TTL sweep.
Install
dsh plugin --profile web add dsh-file-upload
# restart dsh web
Usage
- Click the paperclip in the composer toolbar, or drag files anywhere over the window;
- Small text files land directly in the composer; documents appear as attachment cards and their path is sent with the message;
- The agent reads documents with
read_document <path>— converted to Markdown on demand, pageable withoffset/limit.
MarkItDown (fully bundled — no downloads, no setup)
The MarkItDown capability ships inside the plugin. Works out of the box: no Python, no pip, no downloads, no build-script approval.
- Bundled engine — the Microsoft MarkItDown TypeScript port (
markitdown-node) is a regular dependency covering 20+ formats: PDF, DOCX, PPTX, XLSX, HTML, CSV, JSON, XML, RSS, Atom, ZIP, Jupyter notebooks, images (OCR via Tesseract, 110+ languages), and audio transcription (via LLM, needs model credentials). - Images — OCR to text by default through the bundled engine.
- Offline — all parsing runs locally, no network calls.
Optional upgrade: if an official MarkItDown CLI already exists on your machine (or is set via
markitdownBin), the plugin prefers it (adds EPUB and more); without one the bundled engine is always available.
- id: file-upload
config:
markitdownBin: /path/to/your/markitdown # optional; empty = bundled engine only
Startup log (bundled mode):
[dsh-file-upload] Document → Markdown ready: bundled MarkItDown engine (20+ formats, image OCR) — fully packaged, no downloads, no Python.
How images are handled (auto-explained)
The plugin detects your session's model capability at upload time:
| Detected route | What happens |
|---|---|
Multimodal model (declares image input, e.g. GPT-4o / Qwen-VL / Claude / Gemini) | imageMode: native — the agent uses the official read_image tool; the image enters model context directly |
| Vision bridge installed (dsh-vision-proxy and similar) | detected automatically (they declare image input on their route) — same native path |
| Text-only model (the DeepSeek API is text-only) | an automatic image description is generated via the vision discovery chain and inserted with the message — the model reasons about the image content immediately |
Vision discovery chain (zero-config): ① explicit visionEndpoint/visionModel → ② local Ollama at http://localhost:11434 (picks a VL model such as DeepSeek-VL2 — images never leave the machine) → ③ OpenAI standard endpoint using a key from the dsh credentials seam.
Note: DeepSeek's official API does not offer vision input (the multimodal line — DeepSeek-VL2/Janus — is open-source and self-hostable); deploy DeepSeek-VL2 via Ollama for a fully local "official DeepSeek vision" experience.
The route detection mirrors the official read_image gate (ctx.llm.resolveModelInfo + inputModalities).
Configuration
All fields have sensible defaults — you can install and use the plugin without touching any of them. Tune only what you need.
| Field | Default | Description |
|---|---|---|
uploadMaxBytes | 25165824 (24 MB) | Max bytes per uploaded file |
allowedExtensions | [] | Extension allowlist; empty = all allowed |
uploadTtlMs | 604800000 (7 days) | Unreferenced upload lifetime |
sweepIntervalMs | 3600000 (1 h) | Sweep period; 0 = disabled |
maxConcurrentUploads | 4 | Concurrent upload limit |
inlineTextLimit | 8192 (8 KB) | Text inlined into the composer up to this size |
previewTextLimit | 2048 (2 KB) | Preview length for larger text files |
maxFileBytes | 25165824 | Byte cap for one document read |
readLimit | 2000 | Max lines returned by one read_document call |
sheetRowLimit | 200 | Rows kept per XLSX sheet |
maxSheets | 5 | Sheets read per workbook |
cacheEntries | 16 | Parse-cache entry count |
cacheMaxBytes | 67108864 (64 MB) | Parse-cache byte budget |
markitdownBin | '' | Optional MarkItDown CLI path; empty = auto-detect PATH |
markitdownTimeoutMs | 120000 | Timeout for one CLI invocation |
visionEndpoint | '' | Vision endpoint for image explanations; empty = auto (local Ollama → OpenAI standard) |
visionModel | '' | Vision model id; empty = auto |
visionApiKeyEnv | OPENAI_API_KEY | Credential reference for the vision key (dsh credentials seam) |
visionMaxBytes | 10485760 (10 MB) | Max image bytes sent to the vision endpoint |
Development
pnpm install
pnpm build # tsc (host) + esbuild (client bundle)
pnpm test # node --test
Architecture
src/
├── index.ts # entry: apply + Config schema + assembly
├── detect.ts # content sniffing (never trusts extensions)
├── convert.ts # MarkItDown engine + optional CLI backend
├── vision.ts # image explanations (vision discovery chain)
├── upload.ts # upload route: loopback/session/size/dedup/TTL
├── tool.ts # read_document: ctx.fs reads + paging + LRU cache
└── client/
└── index.tsx # paperclip + drag (files/folders) + paste + cards
Dual-face plugin: dsh.bundle (host) + dsh.client (web UI). No official patches — everything uses official seams (ctx.webServer, ctx.tools, ctx.systemPrompt, ctx.sessions, slash/input-insert-text, slash/input-insert-reference).
Security
- Uploads are loopback-only and same-origin checked.
- File names are sanitized (control chars, path separators, dot segments, leading dots stripped).
- Storage is session-isolated under the session's own workspace; unknown sessions get 403.
- sha256 content dedup, bounded concurrency (429 on overload), TTL sweep.
- Text extraction parses bytes, never trusts extensions; binaries are handed to the agent by path only.
License
MIT
查看使用指南 →
该插件的安装步骤、关键要点、FAQ 与兼容性说明(基于已收录字段派生)。
收录徽章
[](https://deepseek-plugin.org/plugins/HongMing-Huang/dsh-file-upload)把这段 markdown 粘贴到你的 GitHub README,链接回本插件详情页。徽章只声明已被本站收录,不代表安全认证。