Provides model tier routing for DSH: routes auxiliary requests and subtasks to lightweight models by session, while main dialogues and complex tasks use strong-tier models, effective across providers.
ⓘ This plugin is a sub-package of the biociao/dsh-science monorepo — stars and activity count the whole repository.
- Language
- JavaScript
- License
- MIT
- Branch
- main
Install
$ dsh plugin --profile web add dsh-model-tierRun the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial
Install via your agent
Install the DeepSeek Harness plugin biociao/dsh-science/packages/dsh-model-tier for me: review the repository at https://github.com/biociao/dsh-science first, then run the install command and verify the plugin loads successfully.
Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.
One-Sentence Description
A session-level model tier routing plugin for DeepSeek Harness: automatically routes auxiliary requests (title/compaction) and sub-tasks in a session to lightweight models, escalates complex tasks to strong-tier models, main conversation follows the tier scheme, and each tier can point to different providers.
Core Features
- Register a virtual provider "Smart Tiering" in the session model selector, converting multiple tiering schemes into selectable "models", enabling routing per-session via opt-in; unaffected sessions remain completely untouched
- Auto-decide tiers by request type: auxiliary requests (session title, compaction) go to light tier, sub-tasks go to light tier, main conversation goes to default tier
- Escalate to strong tier by rules: optional "recent user message exceeds N characters" / "sub-agent chain depth ≥ N" triggers strong tier, automatically routing long inputs and deep-chain reasoning to stronger models
- Optional LLM pre-classifier: uses a small model to classify tasks as light/default/strong on top of structural rules, cached by (sessionId, message hash), classifies only once per round, falls back to structural tier on failure/timeout
- Multi-scheme presets + hot reload: settings page can save multiple tiering schemes, config file hot-reloads by mtime, no profile restart needed
- Cross-provider routing: three tiers can point to different providers/models respectively; when a tier is unconfigured, falls back in order default → light → strong, then passes through global default model
Technical Implementation
- Language: JavaScript (ESM, Node built-in modules)
- Key Dependencies: Zero third-party dependencies; only uses
node:fs,node:path,node:os,node:url - Architecture Pattern: Registers process-level virtual LLM adapter via
ctx.llm.registerAdapter(["model-tier"], ...); subscribes toagent/requestevent for global default model fallback guard; client injects "settings page section" and "chat page bottom dock" viaslots.inject - Entry Files:
engines/model-tier.mjs(routing engine, exportsname = "dsh-model-tier",inject = ["llm"]) +engines/model-tier-ui.mjs(settings page HTTP routes) +client/model-tier-ui/src/index.js(settings page + routing metrics UI)
Use Cases
Users who want to simultaneously use strong models from different providers (e.g., DeepSeek, Zhipu) and cheap/local lightweight models, wanting to automatically route models by request type within the same session to save token costs and latency. Suitable for multi-model collaboration (flagship for default tier, local small models for light tier, reasoning-strong models for strong tier) in development, research, and automation workflow scenarios. Claude Code users migrating to DSH can seamlessly reuse the "small model handles auxiliary requests" tiering approach.
Prerequisites & Compatibility
| Dependency | Min Version | Description |
|---|---|---|
| Node | >= 18 | Declared in package.json engines.node |
| DSH | Not declared | No dsh version field in package.json, runs on currently released DSH host |
| Platform | Cross-platform | No platform-specific native modules on server; client declares dsh.client.platform: "web" |
| Native Modules | None | Only depends on Node built-in modules (fs/path/os/url) |
Installation
dsh plugin --profile web add dsh-model-tier
Configuration Options
| Config | Type | Description | Default |
|---|---|---|---|
tiers.strong | Object | Strong tier {provider, model, reasoningEffort?}, for complex tasks (deep-chain sub-tasks / super long inputs) | Unconfigured → falls back to default tier |
tiers.default | Object | Default tier, main conversation in selected sessions goes through this tier | Unconfigured → falls back to light/strong tier |
tiers.light | Object | Light tier, for auxiliary requests (title/compaction) + sub-tasks | Unconfigured → falls back to default tier |
routing.auxiliary | String Array | purpose list treated as auxiliary requests, auto go to light tier | ["session-title", "compaction"] |
routing.subagents | String | "light" means sub-tasks route to light tier; other values don't route | "light" |
routing.subagentDepthStrong | Number/null | Escalate to strong tier when sub-agent delegationDepth ≥ N | null (off) |
routing.escalateOnChars | Number/null | Escalate to strong tier when latest user message text length ≥ N (after stripping harness-injected <system-reminder> segments) | null (off) |
routing.classify | Object/null | LLM pre-classifier: true/{} enables (defaults to light tier model for classification), or explicitly specify {provider, model, timeoutMs?, maxChars?, maxTokens?}; non-thinking model recommended | null (off) |
enabled | Boolean | Master switch: when off, "Smart Tiering" no longer appears in model selector | true |
When all three tiers are unconfigured, no tiering scheme appears in the model selector, so sample config can be safely shipped with the package without polluting environments that don't expect the provider.
FAQ
Q: What's the relationship between this plugin and Claude Code's Opus/Sonnet/Haiku strategy?
A: It's the DSH implementation of the same idea: within the same session, automatically route auxiliary requests (title/compaction) and sub-tasks to light tier, main conversation goes to default tier, complex tasks escalate to strong tier by rules, corresponding to Claude Code's smallModel/--model-small and task complexity judgment.
Q: Will all sessions be automatically tiered after installation?
A: No. The plugin only registers a virtual provider "Smart Tiering" in the session model selector. You need to manually select "Smart Tiering / some scheme" in each session to enable routing. Unselected sessions remain completely unaffected (opt-in per session, no global routing).
Q: Where is the config file saved? Do I need to restart after changes?
A: Saved to $DSH_HOME/model-tier.json (default ~/.dsh/model-tier.json). Routing engine caches by file mtime, changes take effect hot, no need to restart profile or reopen session.
Q: Is LLM pre-classifier (routing.classify) required?
A: No. It's an optional enhancement: when enabled, every user question or sub-task dispatch will use a small model to classify the task as light/default/strong, overriding the structural tier. Defaults to light tier model for classification; failure/timeout/garbage response automatically falls back to structural tier (won't block main request).
Q: Can I use a thinking model for the classifier target?
A: Not recommended. Reasoning content from thinking models also consumes maxTokens quota (default 512, adjustable via maxTokens). In testing, thinking has filled the quota and produced empty strings where classification words couldn't be obtained. README explicitly recommends non-thinking models.
Q: What's the relationship with dsh-science?
A: dsh-model-tier is a companion package to dsh-science. Installing dsh-science automatically brings dsh-model-tier as a dependency; but it's completely independent and can be used in any profile alone via dsh plugin add dsh-model-tier.
Difficulty Level
Beginner — configuration only involves three-tier provider/model YAML, after writing yaml you can see "Smart Tiering" group in model selector and enable it in sessions; advanced features (classifier, deep-chain escalation) have clear switches and defaults.
Known Issues & Limitations
- Selection window race condition (known limitation): Since platform persists "session model selection" as global default model, engine uses
agent/requestlistener for best-effort fallback; however, sessions created in the extremely short window between "selecting scheme" and "next request" will still inherit the tiering scheme (source:README.zh.md:48-49,engines/model-tier.mjs:651-681) - Cross-tier thinking compatibility: When history mixes assistant messages from other tiers (non-thinking / no reasoning capture) producing tool calls, and target tier is a thinking model like DeepSeek that requires
reasoning_contentreturn, it triggers 400 INVALID_REQUEST. Engine automatically adds placeholder reasoning content to messages missing reasoning blocks and retries transparently once (generates warn log, but only retries once, doesn't replay already produced content; source:engines/model-tier.mjs:344-369) - Classifier cache limit 200 items: Classification results cached by (sessionId, message hash), items evicted by insertion order when limit exceeded; occasional duplicate classification in long-session dense scenarios is normal (source:
engines/model-tier.mjs:318-323) - Light tier model reasoning effort: Explicit
reasoningEffortin tier is forwarded as-is; model's reasoning effort in selector is only forwarded when target model declares support, otherwise host hard-fails withUNSUPPORTED_REASONING_EFFORT; light tier models typically don't support reasoning, auxiliary requests like title/compaction will silently fall back (source:engines/model-tier.mjs:622-639) - Unconfigured tier schemes don't go live: If a scheme saved in settings page has all three tiers incompletely configured with provider/model, corresponding tiers fall back in TIER_FALLBACK order, and when still no target available, throws error "No available model" (source:
engines/model-tier.mjs:294-303,engines/model-tier.mjs:597-606)
A Claude Science–style research workbench for DeepSeek Harness — for genomics / pathogens / human health / bioinformatics projects.
One-liner: dsh-science — Claude Science-style research workbench for DSH: ReAct research-loop engine (research_* tools), versioned artifacts with provenance (artifact_* tools), an SSH remote-compute engine (remote_* tools, mirroring Claude Science's Computer / Remote compute clusters), and 11 science skills for genomics / pathogens / bioinformatics.
- ReAct research loop engine —
research_init/research_state/research_hypothesis/research_experiment/research_findings/research_phase/research_review/research_report, persisted in aresearch-manifest.jsonstate machine (Question → Hypothesis → Experiment → Observe → Analyze → Conclude → Next Question). - Versioned artifacts with provenance —
artifact_save/artifact_list/artifact_show/artifact_diff/artifact_verify/artifact_deprecate/artifact_reproduce: every result saved asartifacts/<name>/v<N>/with per-file SHA-256,artifact.jsonprovenance (command / inputs / environment / envFile) and an append-onlyprovenance.md. - Remote compute engine (SSH / HPC clusters) — 16 tools:
remote_host_add/remote_host_probe/remote_host_notes/remote_run/remote_status/remote_logs/remote_pull/remote_cancel/remote_execetc. Connect lab workstations or HPC clusters via~/.ssh/configaliases (nothing installed on the host, zero third-party deps). Long bioinformatics jobs run as detached processes on workstations or viasbatchon SLURM — they survive connection loss; submission asks for approval by default;remote_statusbatch-monitors and auto-transitions state (running → succeeded/failed/killed);remote_pullfetches outputs back (files over the size threshold stay on the host with their paths recorded). - Remote Hosts config UI (bundle/profile-level) — a Settings > 远程主机 page (the analog of Claude Science's Settings > Compute > SSH hosts): list/add/probe/edit/remove hosts, plus each project's access allowlist and job summary. Host-side REST API (
webServerroute/dsh-science/remote-hosts/*,engines/remote-hosts-ui.mjs) + client bundle (client/remote-hosts-ui/, built byscripts/build-client-bundle.mjs) sharing the same data files as the remote engine. Requires a web-process restart to activate (see docs/remote-hosts-ui.md). - Model Tier router (tiered, cross-provider) — via the companion bundle
dsh-model-tier: within one session, automatically routes auxiliary requests (session titles, compaction summaries) and subagent/background tasks to a light tier, keeps the main conversation on the default tier, and escalates complex work (deep subagent chains, very long inputs) to a strong tier — each tier may point at a different provider (e.g. strong GLM-5.3 / default deepseek-v4-flash / light minimax-M3), mirroring Claude Code's Opus/Sonnet/Haiku strategy. Built on DSH's nativeagent/request+llm/streamwaterfall extension points; a no-op when the tier's provider is unregistered. Installed automatically with dsh-science, but also standalone-installable into any profile (dsh plugin add dsh-model-tier). - 11 science skills — research-loop, science-project-setup, artifact-provenance, scientific-reviewer, literature-connector, parallel-delegation, manuscript-writing, bioinformatics-toolkit, conda-environments, data-inventory, remote-compute.
All in-repo engine plugins are zero-dependency (Node built-ins + the system OpenSSH binaries, sharing engines/core.mjs) and register plain cordis tools; the companion dsh-model-tier router is likewise zero-dependency. Installable either as a profile bundle (dsh plugin add) or as an agent preset (科学模式).
v0.2.0: Model Tier router (new, companion bundle)
Mirrors Claude Code's Opus/Sonnet/Haiku tiering: within one session, auxiliary requests (purpose ∈ {session-title, compaction}) and subagents (session.meta.origin === 'subagent') are routed to the light tier; the main conversation keeps its own per-session model selection (never overridden); deep subagent chains (delegationDepth ≥ subagentDepthStrong) and very long inputs (escalateOnChars, opt-in) escalate to the strong tier. An optional LLM pre-classifier (routing.classify) grades each user prompt / subtask dispatch by complexity (light / default / strong) before routing. Each tier is {provider, model, reasoningEffort?} and may span providers.
Ships as the standalone bundle dsh-model-tier — dsh-science depends on it and mounts it in its cordis.patch.yml, but it can equally be installed on its own into any profile (dsh plugin add dsh-model-tier):
- id: model-tier
name: dsh-model-tier
config:
tiers:
strong: { provider: zai-coding-cn, model: glm-5.3 }
default: { provider: deepseek-official, model: deepseek-v4-flash }
light: { provider: opencode-go, model: minimax-m2.7 }
routing:
auxiliary: [session-title, compaction]
subagents: light
subagentDepthStrong: 3
- Host plane — mounted in the profile bundle (
cordis.patch.yml), not the agent preset, so it applies to every session and subagent on the profile. - Safety rails — no
tiersconfigured → inert no-op; target provider unregistered → no routing; a failing light-tier call automatically falls back to the original route (auxiliary features never break). - Verified —
node packages/dsh-model-tier/test/model-tier.test.mjs(zero-dependency unit matrix) +bash packages/dsh-model-tier/scripts/test-model-tier.sh(E2E: light tier pointed at a local mock LLM; asserts the title request is actually routed).
v0.2.0: Remote compute (new)
Mirrors Claude Science's Remote compute clusters / Computer capability, following its documented mechanism:
- Host registration + read-only probe —
remote_host_addtakes a~/.ssh/configalias (oruser@host; ProxyJump etc. handled by OpenSSH), with optional port/identityFile overrides; probing records CPUs, memory, GPUs, CUDA driver, conda/module/Apptainer presence, scratch dirs,sbatchand SLURM partitions (remote_host_probere-runs it). Host registry:$DSH_HOME/remotes/hosts.json. - Job submission —
remote_runcopies script + inputs into<scratch>/<jobId>/(default~/dsh-scratch); workstations run it as a detachednohup+setsidprocess (connection-loss safe), SLURM clusters getsbatch(with--time); default job timeout 30 min; submission asks for approval by default (the analog of Claude Science's "Run this job on?" card). - Monitoring & reaction —
remote_statusbatch-probes (ps / squeue+sacct / done+exitcode markers) and auto-transitions state;remote_logstails logs;remote_pullfetches outputs and writespulled-manifest.json(files > 100 MB stay on the host with recorded paths);remote_cancelkills (process group / scancel). Job registry:<project>/.dsh/remotes/jobs.json, persists across sessions. - Host Details document —
remote_host_notesmaintains per-host notes (environment activation, partitions/account, conventions) that the model reads before submitting jobs. - Per-project access allowlist (allowed servers, isolated per project) — every host-connecting action (add/probe,
remote_host_probe,remote_exec,remote_run) requires the host to be in the project's allowlist (.dsh/remotes/allowlist.json) by default; first use pops an approval dialog and, on approval, persists the grant at project scope (the analog of Claude Science's "This project" approval scope). The project root resolves by priority:research-manifest.json(research project) →.dsh(workspace) →.git→ session cwd — multiple research projects in one workspace keep separate allowlists; grants never leak across projects (authorization paths fail closed when no session cwd is available). Review withremote_host_allowlist, revoke withremote_host_revoke, pre-grant withremote_host_allow(approval-gated); disable withrequireHostAccess: falsefor unattended runs.
v0.1.1 hardening (robustness update)
- Concurrency-safe state: all manifest/artifact writes go through a lightweight file lock (O_EXCL + stale reclaim) and atomic tmp+rename — parallel subagents can no longer corrupt or lose updates on
research-manifest.json/artifacts.json. - Structured error codes (
ERR_NOT_INIT/ERR_NOT_FOUND/ERR_VALIDATION/ERR_PATH/ERR_QUOTA/ERR_LOCK_TIMEOUT/ERR_IO) instead of opaque strings. - Hypothesis state machine (proposed → testing → supported/refuted/inconclusive) and forward-only phase transitions (rewind requires config).
- manifest ↔ artifacts linked:
research_statemerges the artifact index;artifact_savewrites back to the manifest. - Manifest schema v1→v2 migration on load, persisted on next write.
- Artifact upgrades: streaming SHA-256 (big files), identical-content dedup via hardlink,
artifact_diff/artifact_verify/artifact_deprecate, envFile + input hashes in provenance. - Structured JSON outputs (
research_report,artifact_diff,artifact_verify) and an audit log at<root>/.science.log.
Install
Option A — profile bundle (community standard)
dsh plugin --profile web add dsh-science # after npm publish
# or straight from GitHub:
dsh plugin --profile web add "github:biociao/dsh-science"
Restart the profile (or refresh the Web GUI). The bundle inserts the three engines
into the profile layer stack; the research_* / artifact_* / remote_* tools
become available to every agent on that profile.
Option B — agent preset (full 科学模式 experience, per-agent)
git clone https://github.com/biociao/dsh-science ~/.dsh/.agent-presets/science
# or from a local checkout:
bash scripts/install.sh # copy (or: bash scripts/install.sh link)
Then create a session in the DSH Web GUI and pick the 科学模式 preset — the preset carries the research persona + engines with per-agent scoping.
Skills
The 11 skills are discovered automatically from a project's .dsh/skills/
(drop this repo's skills/ into your project), or install them machine-wide:
bash scripts/install-skills.sh # -> ~/.dsh/skills (respects $DSH_HOME)
Quick start (first session)
research_init— createresearch-manifest.json+ the project skeleton (experiments/ literature/ artifacts/ analyses/ figures/ manuscript/ reviews/ data/ envs/).- Read
research_stateat the start of every session; the loop state persists across sessions. - Run the loop:
research_hypothesis(H1/H2/…) →research_experiment(E01/…, createsexperiments/<id>/{design.md,log.md,code/,results/}) → run code →research_findings(appends to log.md, updates hypothesis status, advances the loop) →artifact_savefor anything worth citing or reproducing. - When GPU/cluster/specialized environments are needed:
remote_host_addthe host →remote_runa background job (approval required) → pollremote_status/remote_logs→remote_pulloutputs when done →artifact_saveto archive. See docs/remote-compute.md and theremote-computeskill. - For key claims: extract the claim, have a review subagent check it against the
execution records (see the
scientific-reviewerskill), archive withresearch_review(writesreviews/R0n/report.md).
Repository layout
dsh-science/
├── package.json # dsh.bundle.patch -> ./cordis.patch.yml (+ dsh.client + exports)
├── cordis.patch.yml # bundle patch: inserts the engines by subpath export + mounts dsh-model-tier
├── packages/
│ └── dsh-model-tier/ # 配套独立 bundle:模型分档路由(可单独 dsh plugin add)
├── engines/ # canonical engine sources (bundle form)
│ ├── core.mjs # shared core: locks, atomic writes, error codes, streaming sha256, structured tools, audit
│ ├── research-loop.mjs
│ ├── artifact-registry.mjs
│ ├── remote-compute.mjs# SSH/local transports, host registry + probe, job submit/monitor/pull/cancel
│ └── remote-hosts-ui.mjs# Remote Hosts 设置页的宿主 REST API(webServer 路由)
├── client/ # client 插件(设置页 UI,bundle/profile 级)
│ └── remote-hosts-ui/ # src/index.js 源码 · lib/client.js 打包产物(build-client-bundle.mjs)
├── preset/ # agent-preset form (mirrors engines/ via sync-engines.sh)
│ ├── agent.cordis.yml # references ./engines/*.mjs (relative, preset mount)
│ ├── preset.yml
│ └── engines/ # mirror — keep in sync: bash scripts/sync-engines.sh
├── skills/ # 11 SKILL.md skills
├── scripts/
│ ├── install.sh # install preset -> ~/.dsh/.agent-presets/science
│ ├── install-skills.sh # install skills -> ~/.dsh/skills
│ ├── sync-engines.sh # mirror engines/ -> preset/engines/
│ ├── init-project.sh # project skeleton without a science session
│ ├── build-client-bundle.mjs # wrap client src -> __ModuleLoader__ bundle (lib/client.js)
│ ├── smoke-test.mjs # 125 checks against a temp workspace (node >= 18)
│ └── stability-test.mjs# 25 concurrency/atomicity/stress checks (locks, lost-update, soak, migration)
└── test/verify-bundle.sh # isolated end-to-end bundle install + boot + client scan check
Verification
node scripts/smoke-test.mjs # engine logic + end-to-end loop + error codes + migration
node scripts/stability-test.mjs # concurrency / atomicity / lock / stress stability checks
bash test/verify-bundle.sh # pnpm pack -> isolated profile -> install -> boot check
All are part of the release checklist and are safe to run in CI (both test scripts
write only to a temp workspace; the bundle test uses an isolated $DSH_HOME).
FAQ
Why subpath exports and not relative paths in the bundle?
dsh plugin add installs the package into the profile and its cordis.patch.yml
rows join the profile composition. The profile loader resolves a row name
relative to the profile directory (not the package), so ./engines/x.mjs
fails with ERR_MODULE_NOT_FOUND. Referencing dsh-science/engines/x.mjs
(subpath export, exports in package.json) resolves from the profile's
node_modules and works — verified experimentally on dsh 0.1.0-rc.6.
The agent-preset mount, by contrast, resolves relative names from the preset
directory, which is why preset/agent.cordis.yml can use ./engines/*.mjs.
Bundle or preset — which should I use?
- Bundle: tools available to every agent on the profile; one command to install.
- Preset: the full 科学模式 experience (research persona, per-agent scoping).
The persona row in
cordis.patch.ymlis commented out because a profile-wide persona would apply to all agents — uncomment it before publishing if that is what you want.
Where do the skills come from?
A project's .dsh/skills/ is auto-discovered; scripts/install-skills.sh puts
them machine-wide in ~/.dsh/skills (respecting $DSH_HOME).
Development
Branching model & release workflow (main = release, dev = integration, feat/* = features,
tag-triggered npm publish + GitHub Release via Actions): see
docs/branching.md.
bash scripts/sync-engines.sh # after editing engines/*.mjs — keeps preset/engines in sync
node scripts/smoke-test.mjs # logic + static package checks
node scripts/stability-test.mjs # concurrency / atomicity / lock stability checks
bash test/verify-bundle.sh # end-to-end bundle install + boot
Workspace-isolation patch — DSH's New Session used to fall back to the most
recently used Workspace, letting automation spawn sessions into unrelated
projects. scripts/patch-session-isolation.mjs applies the one-line guard
(idempotent, backs up first; re-apply after every dsh upgrade). See
docs/workspace-isolation.md.
node scripts/patch-session-isolation.mjs apply # idempotent, backs up first
node scripts/patch-session-isolation.mjs status # check current state
node scripts/patch-session-isolation.mjs revert # restore pristine file
Community
- Topic: github.com/topics/dsh-plugin
- Curated lists: awesome-dsh-plugin · awesome-deepseek-harness
License
MIT — see LICENSE.
Read the usage guide →
Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.
Listing badge
[](https://deepseek-plugin.org/plugins/biociao/dsh-science/packages/dsh-model-tier)Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.