为 DSH 增加"帮我批准"权限模式,由独立评审模型替代人工判定审批;其他模式行为不变。
- 语言
- TypeScript
- 分支
- main
安装
$ dsh plugin --profile web add dsh-approval-llm在终端中运行以上命令,通过 dsh CLI 安装此插件。可在右上角切换 Profile。 第一次用 dsh?看这篇新手教程
对话式安装
帮我安装 DeepSeek Harness 插件 Letter2025/dsh-approval-llm:先查看仓库 https://github.com/Letter2025/dsh-approval-llm 确认安全性,然后执行安装命令并验证插件加载成功。
把这段指令粘贴给 DSH Web GUI 里的助手,由它代你完成安装与验证。
一句话定位
本插件为 DeepSeek Harness 增加"帮我批准"权限模式:在该模式下,由独立评审模型代替人工判定 approval/request,给出 ALLOW / DENY / ESCALATE 三值决策;其他权限模式下插件完全沉默,人工审批流程不受影响。
核心能力
- 在
model-approval权限模式下,由评审模型自动判定approval/request,给出通过、拒绝或转人工 - 决策前先按白名单、黑名单、人工专属名单做确定性路由,命中的请求免去模型调用
- 模型超时、解析错误或上游故障时一律转人工(fail-to-human),不计入连续拒绝熔断
- 单会话连续拒绝达到阈值后,自动将后续请求交回人工,评审模型停止当裁判
- 评审时从会话日志回放工具真实参数,避免被 agent 自述误导
- 随包附带
configure-approval-llm技能,AI 可按"先建议再确认"流程写入评审配置
技术实现
- 语言: TypeScript(ESM,严格模式 tsconfig)
- 关键依赖:
@deepseek-ai/dsh-user-approval(审批 outcome 词汇)、@deepseek-ai/dsh-permission-presets(预设服务)、@deepseek-ai/dsh-llm(流式评审调用)、@deepseek-ai/schemastery(配置 schema 校验) - 架构模式: cordis 插件,通过
inject = ['llm','tools','permissionPresets','skills']注册并监听approval/request事件,按"确定性路由 → 熔断 → 评审模型"瀑布处理;通过package.json#dsh.bundle.patch同时插入插件行与model-approval预设 - 入口文件:
src/index.ts(apply(ctx, rawConfig))
适用场景
希望在 DSH 中实现"先让模型审、拿不准再问人"这类低风险自动化工作流的开发者。适合需要让模型在 workspace-write 模式下持续自主执行,但仍希望对删除、终端命令、job_kill 等高风险工具保留人工把关的场景;不适合无人值守的高风险任务。
前置依赖与兼容性
| 依赖 | 最低版本 | 说明 |
|---|---|---|
| @deepseek-ai/dsh-agent | ^0.1.0-rc.7 | 宿主 Agent,peerDependencies |
| @deepseek-ai/dsh-llm | ^0.1.0-rc.7 | LLM 流式接口与 BlockAssembler |
| @deepseek-ai/dsh-user-approval | ^0.1.0-rc.7 | 审批 outcome 词汇(allowed-once / rejected) |
| @deepseek-ai/dsh-permission-presets | ^0.1.0-rc.7 | 预设服务(model-approval 预设由此注入) |
| @deepseek-ai/dsh-session | ^0.1.0-rc.7 | 会话日志回放工具参数与对话路由 |
| @deepseek-ai/dsh-skill | ^0.1.0-rc.7 | bundled configure-approval-llm 技能注册 |
| @deepseek-ai/dsh-tools | ^0.1.0-rc.7 | 工具注册表(实时取工具描述) |
| @deepseek-ai/dsh-timeout | ^0.1.0-rc.7 | 评审请求超时管理 |
| Node 版本 | 未声明 | package.json 未声明 engines |
| 平台 | 跨平台 | package.json 未声明 os / cpu 限制 |
| 原生模块 | 无 | 不依赖任何原生模块 |
安装方式
dsh plugin --profile web add github:Letter2025/dsh-approval-llm
配置项
| 配置 | 类型 | 说明 | 默认值 |
|---|---|---|---|
| enabled | 布尔 | 总开关;关闭时所有请求原样交给下一层 answerer(人工审批) | true |
| modePreset | 字符串 | 触发评审模型的权限预设名;只有当前会话的预设与之匹配时才介入;设为空字符串则审查所有请求 | "model-approval" |
| provider | 字符串 | 评审模型使用的 provider;必须与 model 成对出现;未设置则回退会话中的对话路由 | 未设置 |
| model | 字符串 | 评审模型 id;必须与 provider 成对出现 | 未设置 |
| timeoutMs | 数字 | 评审请求端到端超时(毫秒);超时后转人工,不计入熔断 | 60000 |
| maxOutputTokens | 数字 | 评审模型输出 token 上限 | 256 |
| systemPrompt | 字符串 | 自定义安全策略;为空则使用内置"默认放行、关键危害才拒绝"短规则 | 内置策略 |
| allowlist | 字符串数组 | 命中即直接放行的工具名(不走模型) | [] |
| denyList | 字符串数组 | 命中即直接拒绝的工具名,优先级高于 allowlist | [] |
| humanOnlyList | 字符串数组 | 必须由人工决定的工具名,绝不自动评审 | [] |
| maxConsecutiveDenials | 数字 | 单会话连续拒绝上限;超过后该会话转人工;设为 0 关闭熔断 | 3 |
| maxArgsChars | 数字 | 渲染给评审模型的工具参数 JSON 字符上限 | 4000 |
| includeArgs | 布尔 | 是否把工具参数一起交给评审模型;参数敏感时可关闭 | true |
| notifyUser | 布尔 | 是否在每次模型通过/拒绝后向会话追加一条用户可见的决策消息 | true |
常见问题
Q: 不配置 provider/model 也能用吗?
A: 可以。未设置时会回退到本次会话最后一次 request/header 中的对话路由;如果会话日志里也没有,则评审失败并转人工,不会伪造决策。
Q: 装上以后人工审批是不是失效了?
A: 不是。插件只在 model-approval 这一个预设下应答;其他预设(含默认的 请求批准)仍然走原有的人工审批,行为与装插件前一致。
Q: 评审模型说"拿不准"或挂掉了会怎样?
A: 一律转人工:超时、解析错误、上游错误都产生 ESCALATE,并把请求交给下一个 answerer;这些失败不会被算进连续拒绝熔断,避免模型故障被误判为"反复拒绝"。
Q: 我担心模型被工具输出骗过,怎么办?
A: 这是已知风险。建议收紧 humanOnlyList(强制人工的高危工具)、denyList(黑名单)、maxConsecutiveDenials(熔断阈值),并避免在无人值守的高风险场景启用。
Q: 安装完怎么配置评审模型?
A: 让任意 agent 加载 configure-approval-llm 技能或对它说"配置审批评审模型",它会按 AI 先建议、你确认后再写盘的流程,把 approval-llm 行写进 ~/.dsh/profiles/web/cordis.patch.yml,重启 dsh web 生效。
Q: 怎么卸载?
A: 通过宿主 DSH 的插件移除命令移除并重启 dsh web;插件声明的 model-approval 预设由 bundle patch 注入,卸载后该预设将随之消失。
上手难度
进阶 — 需要理解 DSH 的权限预设、cordis patch 机制以及为评审模型选型;内置 configure-approval-llm 技能能减轻配置工作,但白名单/黑名单/熔断的策略调优仍需部署者判断。
已知问题与限制
- 暂无专门的机器可读审计事件类型:评审决策以
user/message形式出现在主链路可被持久化与重放;待宿主开放自定义事件注册通道后再补session/approval-llm-request等事件 - 暂未提供客户端徽标(浏览器半):评审决策以对话消息形式展示,而非工具卡上的盾牌图标
- 一次请求一次模型调用:批量评审(PLAN-063)需要审批服务先开放批量入口,单请求延迟受
timeoutMs与所选评审模型速度约束 - 评审模型可能被工具输出提示注入:需配置好
humanOnlyList/denyList/ 熔断参数,不要在高风险无人值守场景启用 - 列表仅精确名称匹配:
allowlist/denyList/humanOnlyList不支持通配符或参数模式匹配
English | 中文
Model-based permission approval (approve-for-me) for DeepSeek Harness.
A community plugin that adds a "model approval" permission mode (approve-for-me) to DeepSeek Harness: in that mode, approval/request asks are answered by a separate reviewer model instead of a human — the reviewer decides ALLOW / DENY / ESCALATE, and the request only reaches a human when the reviewer cannot decide or fails. In every other permission mode the plugin stays silent, so human approval is never front-run by the model.
It is the dsh equivalent of Codex's approvals_reviewer=auto_review (--approve-for-me), and it follows the review design of AGENTSCOPE-PLAN-058 / 062 / 063 (orthogonal reviewer dimension, three-way decision, routing policy, prompt isolation, fail-to-human, circuit breaker).
Warning: an AI reviewer is a policy choice, not a security guarantee. It can be fooled by prompt injection from tool output or the agent's own reason. Prefer it for low-risk workflows; keep
humanOnlyList, thedenyList, and the circuit breaker tight.
Design
| Concept (this plugin) | Codex | AGENTSCOPE-PLAN-058 |
|---|---|---|
| A dedicated permission mode activates the reviewer | approvals_reviewer: auto_review + --approve-for-me | UI preset 帮我批准 = DEFAULT + ApprovalReviewer=MODEL |
| Outside that mode the plugin delegates everything | human approval unchanged | 请求批准 = DEFAULT + HUMAN |
Answer the approval/request waterfall | approvals_reviewer: auto_review | ApprovalReviewer = MODEL (orthogonal to the permission mode) |
| Deterministic routing before any model call | "deterministic sandbox/network allowlist runs before the guardian" | SAFE_ALLOW / DENY / HUMAN_ONLY routing policy |
Reviewer model decides ALLOW / DENY / ESCALATE | Guardian subagent (Approved / Denied / TimedOut / Abort) | ALLOW / DENY / ESCALATE |
| Reviewer holds an isolated security policy | Guardian prompt isolated from the main agent | §4.5 prompt isolation |
| Tool description injected at review time | — | §4.5 dynamic tool-description injection |
| Tool arguments recovered from the session log | trust layering (arguments reviewed, not just the name) | §4.4 argument-level risk |
| Model failure → hand to human, not counted in the breaker | fail-closed guardian | §4.8 fail-to-human vs policy decision |
| Consecutive DENY threshold → hand off to a human | circuit breaker (3 consecutive) | §4.14 circuit breaker |
Explicit provider/model, else the logged conversation route | — | PLAN-062 per-agent model config + fallback |
| Decision appended to the session as a user-visible message | guardian badge in the UI | DENY event pushed to the frontend (§4.17.5) |
How it works
approval/request (waterfall)
│
├─ enabled? no ────────────────────────────────► next() (unchanged)
│
├─ mode gate: session preset ≠ modePreset ─────► next() (human approval unchanged)
│ (default modePreset = model-approval)
│
├─ RoutingPolicy (deterministic, no model call)
│ ├─ denyList hit ──────────────────────────► 'rejected'
│ ├─ allowlist hit ─────────────────────────► 'allowed-once'
│ ├─ humanOnlyList hit ─────────────────────► next() (human decides)
│ └─ else: REVIEW
│
├─ circuit breaker: consecutive DENY ≥ max ────► next() (human takes over)
│
├─ Reviewer model (isolated security-policy prompt)
│ input: tool name + description (ctx.tools)
│ + reason + tool arguments (from the session log `tool/call`)
│ decision: ALLOW ──────────────────────────► 'allowed-once' (counter resets)
│ DENY ───────────────────────────► 'rejected' (counter +1)
│ ESCALATE ───────────────────────► next()
│ timeout / parse error / provider error ─► next() (fail-to-human)
│
└─ ALLOW / DENY also append a user-visible decision message to the session,
so the main chain records why the call was approved or denied.
- ALLOW / DENY / ESCALATE map 1:1 onto the dsh outcome vocabulary (
allowed-once/rejected/ delegate). ESCALATE and model failures never fabricate a rejection — they hand the request to the next answerer (the human UI), and a deployment with no human answerer fails closed (unavailable), exactly like Codex's fail-closed guardian. - The mode gate makes the modes exclusive. In the
帮我批准preset the reviewer answers; in请求批准(and every other preset) the plugin delegates, so human approval behaves exactly as before. The two presets share sandbox/approval knobs; the recordedpermission/presetselection tells them apart. - One terminal answerer per deployment. The dsh approval chain is not a priority list of competing judges — compose one answerer. To keep human override, put a human UI answerer behind this plugin (the waterfall
next()reaches it). - Arguments are read from the session log, not from the approval request (the request deliberately carries no arguments to avoid a second rendering that could drift).
Configuration
All fields are validated by the Loader schema; defaults apply when omitted.
| Field | Default | Meaning |
|---|---|---|
enabled | true | Master switch; when false every request is delegated unchanged. |
modePreset | model-approval | The permission preset that activates the reviewer. When set, the plugin only answers asks from sessions whose effective preset equals this name; every other session delegates to the human channel. Set to '' to review every ask. |
provider / model | unset | Explicit reviewer route. Must be set together; when unset the plugin reuses the conversation route from the last request/header in the session log, and fails to human when the log has none. |
timeoutMs | 60000 | End-to-end reviewer deadline; on expiry the request is handed to a human (TIMEOUT, not counted in the breaker). |
maxOutputTokens | 256 | Reviewer output cap. |
systemPrompt | built-in policy | Custom security policy for the reviewer. The built-in policy is a short allow-by-default, deny-on-critical-harm rule set; see src/reviewer.ts. |
allowlist | [] | Tool names auto-approved without a model call (SAFE_ALLOW). |
denyList | [] | Tool names rejected outright without a model call. Wins over the allowlist. |
humanOnlyList | [] | Tool names that must be decided by a human; never auto-reviewed. |
maxConsecutiveDenials | 3 | Consecutive DENY threshold per session before the reviewer hands off to a human; 0 disables the breaker. ALLOW resets the counter. |
maxArgsChars | 4000 | Cap on tool-argument JSON rendered to the reviewer. |
includeArgs | true | Recover tool arguments from the session log for the review. |
notifyUser | true | Append a user-visible decision message (✅ 模型审批通过/❌ 模型审批拒绝 with the risk and reason) to the session after every model ALLOW/DENY, so the main chain records why. |
Example overlay (cordis.patch.yml of your profile):
- id: approval-llm
config:
provider: deepseek-official
model: deepseek-v4-flash
allowlist: [read, read_image, glob, grep]
humanOnlyList: [delete, terminal_send]
denyList: [job_kill]
maxConsecutiveDenials: 3
Install
Copy-paste for an AI agent — hand this one sentence to any AI coding agent to have it install the plugin for you: "Read https://github.com/Letter2025/dsh-approval-llm/blob/main/README.md and follow its
## Installsection to install thedsh-approval-llmbundle into the DeepSeek Harness web profile, restart thedsh webserver, and verify that the permission selector shows themodel-approval(帮我批准) preset with its shield-sparkle icon."
As an installable bundle (recommended)
This package declares dsh.bundle.patch in its package.json, so installing it activates a configuration layer that inserts the plugin row and adds the model-approval ("帮我批准") preset to the permission table — no manual preset config needed:
dsh plugin --profile web add dsh-approval-llm # installs the published npm package
Restart dsh web, then pick 帮我批准 in the permission selector (the Access chip in the input bar, which carries a shield-sparkle glyph) to switch that session's reviewer to the model. The preset table is process-level, so changing presets requires a dsh restart.
In-box bundle rows resolve from the dsh installation itself; the @deepseek-ai/* imports are peerDependencies provided by the host dsh, so pin your dsh version (the project is in developer preview with breaking changes). Installing from a local checkout instead: pnpm run build, then dsh plugin --profile web add ./dsh-approval-llm from the parent directory.
Bundled skill: configure the reviewer
The package ships one bundled skill (configure-approval-llm, source bundled), so installing the plugin also puts a configuration guide in the skill catalog. Ask any agent to "configure the approval reviewer", or load the skill directly — it walks an AI-proposes / user-confirms flow: probe the current model and provider settings, write the approval-llm overlay into ~/.dsh/profiles/web/cordis.patch.yml, then present the full config for your confirmation before a restart takes effect. The guide covers choosing a reviewer model (same provider preferred, contextWindow ≥ the main model), and tightening allowlist / denyList / humanOnlyList / maxConsecutiveDenials for your deployment.
As a source overlay (dev)
- insert:
- id: approval-llm
name: './src/index.ts' # path to this package's entry, or an absolute path
config:
provider: deepseek-official
model: deepseek-v4-flash
Run dsh with the overlay (dsh web --patch ./cordis.patch.yml), or merge the row into your profile's cordis.patch.yml. The source overlay inserts only the plugin row, not the preset — either also install the bundle layer above, or add the model-approval preset to the permission row yourself (a patch replaces the whole row config, so restate every preset):
- id: permission
config:
presets:
read-only:
sandbox: read-only
approval: ask
name: 只读
workspace-write:
sandbox: workspace-write
approval: ask
name: 请求批准
model-approval:
sandbox: workspace-write
approval: ask
name: 帮我批准
description: 审批由独立的评审模型决定;拿不准或模型故障时转人工。
danger-full-access:
sandbox: danger-full-access
approval: never
name: 完全放开
Build & test
The plugin lives inside the DeepSeek Harness checkout at custom_plugin/dsh-approval-llm; @deepseek-ai/* resolves against the checkout's own node_modules (built lib declarations + @types), so keep the harness built (pnpm run build at the repo root). The node_modules junction into the harness is provided by the checkout.
pnpm run typecheck # tsc --noEmit (strict)
pnpm run test # vitest: 39 unit tests, no network
pnpm run build # tsc emit to lib/ (ESM, relative imports rewritten)
Roadmap
- Client-side badge & toggle: a browser half (
dsh.clientin this package) can render a shield icon on tool cards whose ask the reviewer decided, and a settings row that writes the plugin'senabled/modePresetto a hot-reloaded settings namespace. The host loader already discoversdsh.clientpackages from the same row, so the install path is unchanged. - Batch review (PLAN-063): one reviewer call for several pending tools needs a batch entry point in the approval service.
- Wildcard / argument-pattern routing in the deterministic policy.
Security model
- Mode gating keeps human approval unchanged: outside the configured preset the plugin delegates every request, so
请求批准behaves exactly as before the plugin existed. - Prompt isolation: the reviewer prompt is assembled by this plugin from its own config; the main agent never sees the security policy, so it cannot tailor asks to pass review. Tool descriptions come from the live registry (
ctx.tools.schemas()), arguments from the durable log — the reviewer judges the real call, not the agent's claim. - Fail-to-human:
TIMEOUT,PARSE_ERROR, and provider errors produce ESCALATE (delegate), never a fabricated denial, and are not counted in the circuit breaker (PLAN-058 §4.8 separation of model failure from policy decision). - Fail-closed by composition: a deployment with no human answerer resolves
unavailable, which callers treat as denial. - Circuit breaker:
maxConsecutiveDenialsconsecutive DENY on one session hands the rest of the session's asks to a human — the reviewer stops being the judge when it keeps saying no. - Sensitive data: review input (tool arguments, reason) is used only for the reviewer request and structured logs; it is not stored beyond the ordinary session log and console output. Turn
includeArgsoff if arguments are sensitive.
Known limitations
- No dedicated log-only audit event. Model decisions are visible in the main chain as
user/messagenotices (durable and replayable), and the built-inapproval/asked+approval/decidedpair records the ask/outcome. A dedicated machine-readable audit event (like asession/approval-llm-requestwith the full review context) is still blocked by the harness persistence policy for out-of-repo event types; when the harness ships a registration surface, this plugin should add one. - No client-side badge (yet). The decision notice appears as a transcript message; a Codex-style shield icon anchored to the tool card needs a client plugin (the package ships no browser half yet). See the roadmap below.
- One request per model call. PLAN-063's batch review (one call for several pending tools) needs a batch entry point in the approval service; the dsh seam processes one request at a time. Per-request latency is bounded by
timeoutMsand the reviewer model choice. - AI-reviewer trust is a deployment decision. The reviewer can be prompt-injected through tool output. Keep
humanOnlyList,denyList, and the breaker configured; do not enable this for high-risk, unattended workflows. - Exact-name routing only. Lists match whole tool names; there is no wildcard or argument-pattern matching. Add patterns as a follow-up if needed.
License
MIT
查看使用指南 →
该插件的安装步骤、关键要点、FAQ 与兼容性说明(基于已收录字段派生)。
收录徽章
[](https://deepseek-plugin.org/plugins/Letter2025/dsh-approval-llm)把这段 markdown 粘贴到你的 GitHub README,链接回本插件详情页。徽章只声明已被本站收录,不代表安全认证。