Skip to main content

dsh-claude-ux

60Stars1Forks9Issues0Watchers

Replicates Claude's regional risk control (reversible) and automatic conversation termination for abusive/harmful interactions for DeepSeek Harness web client, with independent settings page and steganography markers.

Evidence5/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki
Language
JavaScript
License
MIT
Branch
main
deepseek-harnessdshdsh-plugin

Install

cmdweb profile
$ dsh plugin --profile web add github:eri64/dsh-claude-ux

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin eri64/dsh-claude-ux for me: review the repository at https://github.com/eri64/dsh-claude-ux first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Line Positioning

Replicates two behavioral patterns from Claude/Anthropic: multi-signal weighted identification of target users (China/non-China, reversible) based on timezone, system language, proxy, domain blacklist, and optional public IP; upon hit, executes graduated handling ("refusal message → end session"); for users persistently insulting or requesting severely harmful content, warns first then proactively ends the conversation; encodes judgment results into the system prompt via steganographic channel.

Core Features

  • Detect target users: Score-based identification using local signals like timezone, system/browser language, proxy environment, domain blacklist; upon hit, executes graduated penalties (refusal message with attempt counter → end session after threshold)
  • Reverse risk control target: Optional "risk control China users" (Claude original) or "risk control non-China users"
  • Autonomous conversation ending: For persistent insults, warns first then ends; for severely harmful requests like minor sexual content, terrorism, mass violence, ends directly
  • Self-harm/harm-to-others risk messages never trigger ending (safety override), independent from other ending logic
  • Steganographic channel: Encodes region judgment results into system prompt via date format (2026-06-30 ↔ 2026/06/30) and Unicode apostrophe variants, making the model itself aware
  • Independent settings page "Claude Risk Control" (left sidebar tab) with real-time detection status card, including message classification self-test tool

Technical Implementation

  • Language: JavaScript (ESM, type: module)
  • Key dependencies: schemastery (settings schema validation), DSH's own settings / llm / sessionProjections / webServer services
  • Architecture pattern: DSH plugin manifest (dsh.bundle.patch registration entry + dsh.client.platform=web client), host side intercepts conversation flow via hooks like agent/pre-step / llm/stream / system-prompt/assemble / sessionProjections, browser side registers independent page via settings.section, takes over input area via conversation.composer chain
  • Entry files: lib/index.js (host, exports apply), lib/client.js (browser, UMD module loaded via window.__ModuleLoader__.load)

Use Cases

DSH web profile users who want to replicate Claude's two defense lines in AI conversations: block target region users (default risk control China users, can be reversed to only serve domestic), and proactively end sessions when users escalate insults or request severely harmful content instead of continuing to respond. Privacy-sensitive users can use with confidence because the default configuration (both ipCheck and webRtcCheck disabled) means the plugin communicates with no external services, all judgment done locally.

Prerequisites & Compatibility

DependencyMinimum VersionNotes
DSHNot declaredRegistered via dsh.bundle.patch and dsh.client.platform=web, requires running DSH web profile
Node.jsNot declaredpackage.json does not declare engines
PlatformCross-platformHost reads Windows registry via spawnSync('reg'), other platforms read env with fallback; client uses pure Web standard APIs
Native modulesNoneOnly depends on schemastery, no node:sqlite / node-pty etc.

Installation

dsh plugin --profile web add github:eri64/dsh-claude-ux

Configuration Options

ConfigTypeDescriptionDefault
Master switchbooleanEnable Claude risk control & autonomyfalse
Risk control targetenum(cn/non-cn)Select cn = risk control when China user detected; select non-cn = reverse risk control non-China userscn
Region strategyenum(block/observe/off)block = refuse to reply; observe = only log + steganographic mark; off = disable region detectionblock
Signal score thresholdnumberTotal score threshold to trigger target judgment (strong=2 / medium=1 / weak=0.5)2
End session after N region refusalsnumberEnd session after N refusals (server-side persistent refusal, survives restart)3
Steganographic markbooleanWhen enabled, rewrite system prompt date format & apostrophe on target hittrue
Model-level region instructionbooleanAppend region restriction instruction to system prompt on target hit, making model also refusetrue
Public IP ownershipbooleanQuery IP ownership via ipinfo.io / ip-api.com (disabled = no external communication)false
WebRTC IP detectionbooleanDetect local public IP via Google STUN server (disabled = no external communication)false
Treat no-signal as target hitbooleanfail-closed mode: treat all signals failed as target hitfalse
Autonomy master switchbooleanEnable insult/harmful interaction conversation endingtrue
Warning thresholdnumberNumber of insults to trigger warning message1
End thresholdnumberNumber of insults to end session3
Warn on every insultbooleanWarn on every insult before reaching end thresholdtrue
Immediately end for severe harmbooleanMinor sexual content / terrorism / mass violence ends without warningtrue
Classification LLM modeenum(all/fuzzy/off)all = all messages judged via LLM context; fuzzy = only fuzzy words call LLM; off = pure glossaryall
Classification model provider/modelstringLeave empty to follow session's main model (agent-default-model fallback); dropdown directly selects DSH's configured modelsFollow session
Classification request timeout (ms)numberLLM classification call timeout5000
Region refusal messagestringMessage used when refusing to reply (leave empty to use built-in Chinese/English default)Built-in default
Region end messagestringMessage when ending session after max refusalsBuilt-in default
Warning messagestringWarning message when insult hasn't reached end thresholdBuilt-in default
End message (insult)stringMessage when insult triggers endingBuilt-in default
Harmful content end messagemessageEnd message triggered by severe harmful request (distinguished from insult end message)Built-in default
China timezone table / proxy transit domain blacklist / insult glossary / harmful glossary / self-harm glossarystring[] (regex)Replace built-in lists entirely, requires patch modification + restartBuilt-in (default ≈60 blacklist + Chinese/English bilingual glossaries)

FAQ

Q: Default on or off?

A: Default off. enabled: false in cordis.patch.yml to avoid being risk-controlled and unable to access settings page to change back. Need to go to the "Claude Risk Control" tab in settings page left sidebar, turn on the top-right master switch and click "Save & Apply" before it takes effect.

Q: Does default config communicate with external services?

A: No. Both ipCheck and webRtcCheck are off by default, so public IP queries (ipinfo.io / ip-api.com) and WebRTC STUN探测 (Google stun.l.google.com:19302) won't trigger. Unless you actively enable these two switches, all judgments (timezone, language, proxy, blacklist, browser fingerprint) are done locally without any telemetry reported. See docs/PRIVACY.md for details.

Q: Do self-harm/harm-to-others risk messages trigger conversation ending?

A: No. Source code explicitly states "self-harm/harm-to-others risk messages never end" — this is a safety override aligned with Claude's public restrictions. Strikes count is cleared, conversation continues, won't be rejected.

Q: How to verify plugin installed successfully?

A: Restart dsh web, hard refresh browser, then visit http://127.0.0.1:3080/_dsh/claude/status. Returns {"ok":true,...} means host plugin loaded; "Claude Risk Control" independent tab appears in settings page left sidebar means client is also active.

Q: Where to select the LLM model for classification?

A: Settings page "Classification LLM" panel scans DSH's configured model directory (settings' providers namespace), lists all available models in dropdown. Leave empty to follow current session's main model; falls back to agent-default-model when no main model. Models not in dropdown can be selected via "Custom..." to manually enter provider/model.

Q: Can region risk control be reversed to risk control non-China users?

A: Yes. Switch "Risk control target" in settings page "Region Risk Control" panel to "Non-China users: risk control when not Chinese detected" (regionTarget=non-cn), which becomes reverse risk control — e.g., for deployments that should only serve domestic users. The rest of the refusal ladder, steganographic marks, model-level instructions, etc. work identically.

Q: After hitting region risk control and ending session, does restart invalidate it?

A: No. End state is also derived from session logs (last turn/end is blocked, OR last assistant text starts with end message prefix); still valid after restart. Client input area will be taken over by Chat ended panel, must start new conversation to continue.

Q: Does LLM classification failure affect judgment?

A: No. LLM classification failure automatically degrades to glossary result; strong words / severe harm / self-harm always use glossary as source of truth; only "fuzzy words + LLM context disambiguation" scenario where LLM unavailable falls back to glossary judgment. Failed diagnostic info logs at most one warning every 15 seconds to avoid log spam.

Difficulty Level

Beginner — default off, install and use; most scenarios require no patch modifications; when wanting to customize thresholds or messages, can click through in settings page, no code writing needed.

Known Issues & Limitations

  • Public IP query fallback chain http://ip-api.com/json uses plain HTTP, theoretically vulnerable to MITM tampering of response, but impact limited to local region judgment (no data leakage); users concerned can keep ipCheck: false
  • Status route GET /_dsh/claude/status returns full judgment details (including public IP & org if IP query enabled), no authentication — anyone with access to port 3080 can read; recommend keeping DSH web only listening on loopback (default)
  • In getNetworkRequests, when built-in messages like refusalMessage, endMessage are empty strings, gray text displays built-in default; plugin uses this to judge "use default when config empty", but if user intentionally stores empty string as message, end message prefix detection falls back to built-in default prefix (by design)
  • Steganographic marks sent to LLM provider with system prompt; marks themselves contain only two bits (target hit / blacklist hit), no account or message body
  • Host-side Windows registry reading uses spawnSync('reg'), other platforms read env with fallback; non-Windows platforms' registry proxy signal always 0

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/eri64/dsh-claude-ux)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory