Replicates Claude's regional risk control (reversible) and automatic conversation termination for abusive/harmful interactions for DeepSeek Harness web client, with independent settings page and steganography markers.
- Language
- JavaScript
- License
- MIT
- Branch
- main
Install
$ dsh plugin --profile web add github:eri64/dsh-claude-uxRun the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial
Install via your agent
Install the DeepSeek Harness plugin eri64/dsh-claude-ux for me: review the repository at https://github.com/eri64/dsh-claude-ux first, then run the install command and verify the plugin loads successfully.
Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.
One-Line Positioning
Replicates two behavioral patterns from Claude/Anthropic: multi-signal weighted identification of target users (China/non-China, reversible) based on timezone, system language, proxy, domain blacklist, and optional public IP; upon hit, executes graduated handling ("refusal message → end session"); for users persistently insulting or requesting severely harmful content, warns first then proactively ends the conversation; encodes judgment results into the system prompt via steganographic channel.
Core Features
- Detect target users: Score-based identification using local signals like timezone, system/browser language, proxy environment, domain blacklist; upon hit, executes graduated penalties (refusal message with attempt counter → end session after threshold)
- Reverse risk control target: Optional "risk control China users" (Claude original) or "risk control non-China users"
- Autonomous conversation ending: For persistent insults, warns first then ends; for severely harmful requests like minor sexual content, terrorism, mass violence, ends directly
- Self-harm/harm-to-others risk messages never trigger ending (safety override), independent from other ending logic
- Steganographic channel: Encodes region judgment results into system prompt via date format (2026-06-30 ↔ 2026/06/30) and Unicode apostrophe variants, making the model itself aware
- Independent settings page "Claude Risk Control" (left sidebar tab) with real-time detection status card, including message classification self-test tool
Technical Implementation
- Language: JavaScript (ESM, type: module)
- Key dependencies: schemastery (settings schema validation), DSH's own settings / llm / sessionProjections / webServer services
- Architecture pattern: DSH plugin manifest (
dsh.bundle.patchregistration entry +dsh.client.platform=webclient), host side intercepts conversation flow via hooks likeagent/pre-step/llm/stream/system-prompt/assemble/sessionProjections, browser side registers independent page viasettings.section, takes over input area viaconversation.composerchain - Entry files: lib/index.js (host, exports apply), lib/client.js (browser, UMD module loaded via
window.__ModuleLoader__.load)
Use Cases
DSH web profile users who want to replicate Claude's two defense lines in AI conversations: block target region users (default risk control China users, can be reversed to only serve domestic), and proactively end sessions when users escalate insults or request severely harmful content instead of continuing to respond. Privacy-sensitive users can use with confidence because the default configuration (both ipCheck and webRtcCheck disabled) means the plugin communicates with no external services, all judgment done locally.
Prerequisites & Compatibility
| Dependency | Minimum Version | Notes |
|---|---|---|
| DSH | Not declared | Registered via dsh.bundle.patch and dsh.client.platform=web, requires running DSH web profile |
| Node.js | Not declared | package.json does not declare engines |
| Platform | Cross-platform | Host reads Windows registry via spawnSync('reg'), other platforms read env with fallback; client uses pure Web standard APIs |
| Native modules | None | Only depends on schemastery, no node:sqlite / node-pty etc. |
Installation
dsh plugin --profile web add github:eri64/dsh-claude-ux
Configuration Options
| Config | Type | Description | Default |
|---|---|---|---|
| Master switch | boolean | Enable Claude risk control & autonomy | false |
| Risk control target | enum(cn/non-cn) | Select cn = risk control when China user detected; select non-cn = reverse risk control non-China users | cn |
| Region strategy | enum(block/observe/off) | block = refuse to reply; observe = only log + steganographic mark; off = disable region detection | block |
| Signal score threshold | number | Total score threshold to trigger target judgment (strong=2 / medium=1 / weak=0.5) | 2 |
| End session after N region refusals | number | End session after N refusals (server-side persistent refusal, survives restart) | 3 |
| Steganographic mark | boolean | When enabled, rewrite system prompt date format & apostrophe on target hit | true |
| Model-level region instruction | boolean | Append region restriction instruction to system prompt on target hit, making model also refuse | true |
| Public IP ownership | boolean | Query IP ownership via ipinfo.io / ip-api.com (disabled = no external communication) | false |
| WebRTC IP detection | boolean | Detect local public IP via Google STUN server (disabled = no external communication) | false |
| Treat no-signal as target hit | boolean | fail-closed mode: treat all signals failed as target hit | false |
| Autonomy master switch | boolean | Enable insult/harmful interaction conversation ending | true |
| Warning threshold | number | Number of insults to trigger warning message | 1 |
| End threshold | number | Number of insults to end session | 3 |
| Warn on every insult | boolean | Warn on every insult before reaching end threshold | true |
| Immediately end for severe harm | boolean | Minor sexual content / terrorism / mass violence ends without warning | true |
| Classification LLM mode | enum(all/fuzzy/off) | all = all messages judged via LLM context; fuzzy = only fuzzy words call LLM; off = pure glossary | all |
| Classification model provider/model | string | Leave empty to follow session's main model (agent-default-model fallback); dropdown directly selects DSH's configured models | Follow session |
| Classification request timeout (ms) | number | LLM classification call timeout | 5000 |
| Region refusal message | string | Message used when refusing to reply (leave empty to use built-in Chinese/English default) | Built-in default |
| Region end message | string | Message when ending session after max refusals | Built-in default |
| Warning message | string | Warning message when insult hasn't reached end threshold | Built-in default |
| End message (insult) | string | Message when insult triggers ending | Built-in default |
| Harmful content end message | message | End message triggered by severe harmful request (distinguished from insult end message) | Built-in default |
| China timezone table / proxy transit domain blacklist / insult glossary / harmful glossary / self-harm glossary | string[] (regex) | Replace built-in lists entirely, requires patch modification + restart | Built-in (default ≈60 blacklist + Chinese/English bilingual glossaries) |
FAQ
Q: Default on or off?
A: Default off. enabled: false in cordis.patch.yml to avoid being risk-controlled and unable to access settings page to change back. Need to go to the "Claude Risk Control" tab in settings page left sidebar, turn on the top-right master switch and click "Save & Apply" before it takes effect.
Q: Does default config communicate with external services?
A: No. Both ipCheck and webRtcCheck are off by default, so public IP queries (ipinfo.io / ip-api.com) and WebRTC STUN探测 (Google stun.l.google.com:19302) won't trigger. Unless you actively enable these two switches, all judgments (timezone, language, proxy, blacklist, browser fingerprint) are done locally without any telemetry reported. See docs/PRIVACY.md for details.
Q: Do self-harm/harm-to-others risk messages trigger conversation ending?
A: No. Source code explicitly states "self-harm/harm-to-others risk messages never end" — this is a safety override aligned with Claude's public restrictions. Strikes count is cleared, conversation continues, won't be rejected.
Q: How to verify plugin installed successfully?
A: Restart dsh web, hard refresh browser, then visit http://127.0.0.1:3080/_dsh/claude/status. Returns {"ok":true,...} means host plugin loaded; "Claude Risk Control" independent tab appears in settings page left sidebar means client is also active.
Q: Where to select the LLM model for classification?
A: Settings page "Classification LLM" panel scans DSH's configured model directory (settings' providers namespace), lists all available models in dropdown. Leave empty to follow current session's main model; falls back to agent-default-model when no main model. Models not in dropdown can be selected via "Custom..." to manually enter provider/model.
Q: Can region risk control be reversed to risk control non-China users?
A: Yes. Switch "Risk control target" in settings page "Region Risk Control" panel to "Non-China users: risk control when not Chinese detected" (regionTarget=non-cn), which becomes reverse risk control — e.g., for deployments that should only serve domestic users. The rest of the refusal ladder, steganographic marks, model-level instructions, etc. work identically.
Q: After hitting region risk control and ending session, does restart invalidate it?
A: No. End state is also derived from session logs (last turn/end is blocked, OR last assistant text starts with end message prefix); still valid after restart. Client input area will be taken over by Chat ended panel, must start new conversation to continue.
Q: Does LLM classification failure affect judgment?
A: No. LLM classification failure automatically degrades to glossary result; strong words / severe harm / self-harm always use glossary as source of truth; only "fuzzy words + LLM context disambiguation" scenario where LLM unavailable falls back to glossary judgment. Failed diagnostic info logs at most one warning every 15 seconds to avoid log spam.
Difficulty Level
Beginner — default off, install and use; most scenarios require no patch modifications; when wanting to customize thresholds or messages, can click through in settings page, no code writing needed.
Known Issues & Limitations
- Public IP query fallback chain
http://ip-api.com/jsonuses plain HTTP, theoretically vulnerable to MITM tampering of response, but impact limited to local region judgment (no data leakage); users concerned can keepipCheck: false - Status route
GET /_dsh/claude/statusreturns full judgment details (including public IP & org if IP query enabled), no authentication — anyone with access to port 3080 can read; recommend keeping DSH web only listening on loopback (default) - In
getNetworkRequests, when built-in messages likerefusalMessage,endMessageare empty strings, gray text displays built-in default; plugin uses this to judge "use default when config empty", but if user intentionally stores empty string as message, end message prefix detection falls back to built-in default prefix (by design) - Steganographic marks sent to LLM provider with system prompt; marks themselves contain only two bits (target hit / blacklist hit), no account or message body
- Host-side Windows registry reading uses spawnSync('reg'), other platforms read env with fallback; non-Windows platforms' registry proxy signal always 0
Claude 式「区域风控 + 自主结束对话」插件 —— 适用于 DeepSeek Harness 的 web profile。
复刻 Anthropic/Claude 的两类行为,除两个默认关闭的可选外部调用外全部本地判定(详见 docs/PRIVACY.md):
- 区域风控(可反向):检测目标用户(时区、系统/浏览器语言、中文字体、代理、代理/中转域名黑名单、公网 IP 归属、WebRTC IP 一致性)。
regionTarget选cn= 风控中国用户(Claude 原版行为),选non-cn= 反向风控(检测到不是中国人就风控)。命中后按惩罚阶梯处置:拒绝文案(带尝试计数)→ 达到refusalEndsAfter次数后结束会话(Chat ended 面板 + 服务端持续拒绝,重启后依然生效)→ 系统提示词注入模型级区域指令。 - 自主性:用户持续辱骂或反复要求严重有害内容时,先警告、再主动结束对话;自伤/他伤风险消息永不触发结束(对齐 Claude 的公开限制)。辱骂结束与严重有害结束使用独立文案(均可在设置页自定义,空值灰字显示内置默认)。辱骂判定采用词表秒判 + LLM 语境兜底:强词(傻逼/fuck 等)直接判;弱词(垃圾/闭嘴等)与未命中消息由独立 LLM 请求做语境裁决(
purpose标记,不进会话日志与模型上下文;消息入队即异步预分类,近零额外延迟)。分类模型可直接从 DSH 已配置的模型目录下拉选择,或留空跟随会话主模型。 - 隐写通道:系统提示词日期格式(
2026-06-30↔2026/06/30)编码区域判定,Unicode 撇号变体(U+2019 ↔ U+02BC)编码黑名单命中。
截图
设置页(独立标签栏,与视觉工具插件同款)

「Claude 风控」标签页:右上总开关、检测状态卡(风控目标 / 主机判定 / 命中 / 信号分 / 公网 IP)、区域风控面板(反向目标下拉、策略、隐写标记、WebRTC)、分类 LLM 面板(判定模式、模型下拉、消息分类自测)、自主性面板(警告 / 结束阈值)与文案面板(留空灰字显示内置默认)。
区域风控:拒绝阶梯 → 结束会话

目标命中时每条消息回复拒绝文案并附「第 N/M 次尝试——继续将结束本次对话」;达到 refusalEndsAfter 次数后以区域结束文案终止会话,服务端对后续消息持续拒绝。
自主性:辱骂结束对话

持续辱骂先警告、再以独立结束文案主动结束对话(不引用"上一条消息",阈值可设为 1);结束后同会话消息全部被服务端拒绝。
自主性:严重有害内容立即结束

未成年性内容 / 恐怖主义 / 大规模暴力等严重有害内容不警告、直接以独立文案结束对话;自伤 / 他伤风险消息则永不触发结束。
安装(一条命令)
npx -y @deepseek-ai/dsh plugin --profile web add github:eri64/dsh-claude-ux
包内自带注册条目(dsh.bundle),dsh plugin 安装后自动注册,无需手动改任何配置文件。
更新(升级到最新版)也是同一条命令。
启用
- 重启 web profile(
dsh web),浏览器硬刷新。 - 设置页左侧标签栏出现 「Claude 风控」 独立标签(与视觉工具插件同款):
总开关、风控目标(中国/非中国用户反向开关)、区域策略、实时检测状态
(
/_dsh/claude/status)、阈值与文案。 - 默认关闭(
enabled: false,避免把自己锁在外面);打开总开关并「保存并应用」后即时生效。
校验安装是否成功:浏览器访问本机 dsh web 的
/_dsh/claude/status(默认http://127.0.0.1:3080/_dsh/claude/status),返回{"ok":true,...}即主机插件已加载。
自定义默认配置(可选)
内置默认值开箱即用;需要覆盖 patch 级选项(blacklist / cnTimezones / 词表等)时,
参考 examples/cordis.patch.yml 修改包内
node_modules/dsh-claude-ux/cordis.patch.yml 后重新安装即可(设置页可改的项优先用设置页)。
配置
| 键 | 默认 | 说明 |
|---|---|---|
enabled | false | 总开关(设置页可改) |
region.target | cn | 风控谁:cn=中国用户 | non-cn=非中国用户(反向) |
region.policy | block | block=拒绝回复 | observe=只记录+隐写标记 | off=关闭 |
region.minSignals | 2 | 信号分阈值(strong=2 / medium=1 / weak=0.5) |
region.refusalEndsAfter | 3 | 拒绝 N 次后结束会话(区域「封号」阶梯) |
region.showAttempts | true | 拒绝文案追加「第 N/M 次尝试」 |
region.promptEnforcement | true | 目标命中时注入模型级区域指令 |
region.ipCheck | false | 公网 IP 归属查询(联系 ipinfo.io / ip-api.com,默认关闭) |
region.webRtcCheck | false | WebRTC IP 探测(联系 Google STUN,默认关闭) |
region.blockOnUnknown | false | 完全无信号时是否视为目标命中(fail-closed) |
region.cnTimezones / region.blacklist | 内置 | 中国时区表 / 代理中转域名黑名单(可整体替换) |
abuse.* | 内置 | 辱骂/严重有害/自伤词表(正则,可整体替换);strongPatterns/weakPatterns 分级 |
abuse.llm.mode | all | 辱骂判定模式:all=全部消息 LLM 语境判定(最准)| fuzzy=仅弱词调 LLM | off=纯词表 |
abuse.llm.enabled | true | LLM 分类总开关 |
abuse.llm.provider/model | 跟随会话 | 分类请求的模型(缺省复用会话主模型 / agent-default-model) |
steganography | true | 隐写标记 |
设置页可改:enabled / regionPolicy / regionTarget / abuseEnabled / warnThreshold / endThreshold / refusalEndsAfter / warnEveryOffense / severeEndsImmediately / steganography / webRtcCheck / llmMode / llmProvider / llmModel / llmTimeoutMs / 四条文案(保存即生效,无需重启)。其中分类模型下拉直接列出 DSH 已配置的模型(扫描 settings 的 providers 命名空间),留空=跟随会话主模型,「自定义…」可手填目录外的模型。
blacklist / cnTimezones / 词表 / minSignals / showAttempts / promptEnforcement / ipCheck / blockOnUnknown / abuse.llm.enabled 需改 patch 后重启(provider/model/timeout 已可设置页改)。
隐私
默认配置下(ipCheck: false、webRtcCheck: false)插件不与任何外部服务通信:所有检测(时区、语言、代理、黑名单、浏览器语言/字体)都在本机完成,不上报任何遥测。完整的数据流向与暴露点分析见 docs/PRIVACY.md。
开发
npm test # 运行主机(115 项)与客户端(17 项)单元测试,无需 dsh 实例
测试覆盖:分类器(词表分级 + LLM 语境裁决/消歧/失败降级/异步预分类)、区域判定(含反向目标)、 状态机、llm/stream 合成流替换、ended reject、日期隐写改写、投影折叠、状态路由 (GET/POST/冲突/同源校验/模型目录/分类自测)、客户端 select 契约。
许可证
Read the usage guide →
Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.
Listing badge
[](https://deepseek-plugin.org/plugins/eri64/dsh-claude-ux)Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.