Skip to main content

dsh-llm-fallbacks

13Stars2Forks0Issues0Watchers

When LLM requests persistently fail (authentication, quota, rate limiting), automatically switch provider/model along the fallback chain, preventing DSH proxy tasks from being interrupted by model issues.

Evidence5/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki
Language
TypeScript
License
MIT
Branch
main
deepseek-harnessdeepseek-harness-plugindshdsh-pluginfallbackssubagents

Install

cmdweb profile
$ dsh plugin --profile web add dsh-llm-fallbacks

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin omdsh-dev/dsh-llm-fallbacks for me: review the repository at https://github.com/omdsh-dev/dsh-llm-fallbacks first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Sentence Positioning

When DSH agent's LLM requests repeatedly fail (authentication errors, quota exhausted, 429 rate limits), this plugin automatically switches along the fallback chain to the next available provider/model, allowing the current step to continue directly on the new model, preventing model issues from interrupting the entire task; both web and dsh-tui share the same configuration.

Core Capabilities

  • Automatic Fault Degradation: When configured failure codes are detected in the agent/request-error tap, it walks down the fallback chain resolved by role, picks the next available model not in cooldown, and recovers automatically (no longer using host's default retry)
  • Time Slot Routing: Switches the "effective all-day chain" based on clock windows of fallbacks.tz (default Asia/Shanghai); different models for peak/off-peak, after crossing the window the next root request automatically switches chains (not counted as failure switch, not subject to cooldown)
  • Virtual Auto Model: After plugin is enabled, it injects a line FallbacksChain / Auto into the host LLM adapter directory; selecting it as the main model uses the entire current effective fallback chain as the root agent's primary model
  • Three-Stage Role Resolution: On first request of sub-agents, resolve role in order: explicit agentPreset → declared rules → optional LLM auto-match; inject the resolved chain head model into the first request
  • Cooling & Reversion: Models that are switched away or failed are not selected again within cooldownMs; after expiry, automatically revert to main model based on revertPolicy; each step also has maxSwitchesPerStep safety valve to prevent infinite chain walking
  • Zero Config = No-Op: With enabled: false and no chains, the plugin is completely no-op, won't intercept or switch anything

Technical Implementation

  • Language: TypeScript (host side ESM, client side React + TSX)
  • Key Dependencies: @deepseek-ai/cordis (Cordis 4 plugin framework) / @deepseek-ai/schemastery (settings schema & defaults) / @deepseek-ai/dsh-settings (installSettingsSection to register settings namespace) / react (client side FallbacksCard & ConversationFallbackSwitch)
  • Architecture Pattern: Pure mount plugin—inserts a line in profile's bundle stack via bundle/cordis.patch.yml (must be registered after dsh-base's llm-retry so the plugin's agent/request-error listener takes over after llm-retry's retry budget is exhausted); host side apply() connects decision logic at agent/request-error and agent/request taps, client side injects settings card into web settings sidebar via dsh.client.inject; config read/write goes through plugin's own /api/fallbacks/get|set|reset RPC channel
  • Entry Files: src/index.ts (host side: apply(ctx, config) entry, FallbacksService registration, gateway registration, commands registration) / src/client/index.ts (client side: settings card FallbacksCard, conversation switch badge ConversationFallbackSwitch)

Applicable Scenarios

Suitable for DSH users using multiple providers simultaneously: single provider's occasional authentication, quota, and rate limit issues are the most common failure causes for long-running tasks; after enabling, agents can silently switch to fallback models when the primary model is temporarily abnormal; the time-based effective chain switching feature is also suitable for placing cheaper models during peak hours and flagship models during off-peak hours. Note that the rootChain tail must be deepseek-official/deepseek-v4-flash or deepseek-official/deepseek-v4-pro as the final fallback; other models are not supported as the chain endpoint.

Prerequisites & Compatibility

DependencyMinimum VersionDescription
DSH^0.1.1-rc.1peerDependencies declared @deepseek-ai/dsh-* series; at runtime provided by dsh host as bundle, local npm registry pulls via autoInstallPeers (README badge DSH-0.1.1--rc.1)
Node.js>= 22package.json engines.node; build scripts need pnpm ≥ 10 (only needed for local directory install)
PlatformCross-platformPure TypeScript, no OS / CPU restriction fields
Native ModulesNoneNo node-gyp native dependencies, no usage of built-in native APIs like node:sqlite

Installation

dsh plugin --profile web add github:omdsh-dev/dsh-llm-fallbacks

Configuration Options

All configurations are under DSH shared settings document's fallbacks: namespace; enabling means giving up DSH's built-in fallback:

ConfigTypeDescriptionDefault
enabledBooleanMaster switch. When false, plugin doesn't intercept any requests, behavior identical to not installedfalse
triggerCodesString ArrayFailure codes that trigger fallback decision (AUTH / QUOTA / RATE_LIMIT); requests with codes not in this list pass through unchanged['AUTH', 'QUOTA', 'RATE_LIMIT']
rootChainString ArrayMain agent's all-day fallback chain; last item must be exactly one of deepseek-official/deepseek-v4-flash or deepseek-official/deepseek-v4-pro (fallback model), preceding items are fallback models in order[]
timeSlotsArrayTime slot rows; preset rows (kind: preset) use frozen windows (梁文峰 / 梁文谷 / GLM峰 / GLM谷, fixed UTC+8 not adjustable), custom rows (kind: custom) can set start / end / days; first window matching current time becomes effective chain for next root request[]
roles.listArrayDeclared role collection; each item id (unique lowercase alphanumeric hyphen ≤32 chars, inherit is reserved word cannot be used as id) + persona (persona description, free text) + optional chain (dedicated fallback chain) + optional fallback ('inherit-root' / 'none')[]
roles.rulesArraySub-agent rules: match first sub-agent by provider / model pattern, route it to a role in roles.list or built-in inherit. root never matches rules, goes directly to rootChain[]
cooldownMsNumber (ms)How many ms a failed/switched-away model is not selected again; after expiry handled according to revertPolicy300000 (5 minutes)
revertPolicyEnumReversion strategy after cooldown expiry: 'cooldown-expiry' reverts to main model on expiry, 'never' stays on last fallback for the entire session'cooldown-expiry'
maxSwitchesPerStepNumberMaximum fallback switch attempts allowed within a single step; after exceeding, stop switching and preserve original error semantics, preventing fallback chain from amplifying latency8
alwaysModeRetryCapNumberWhen provider is retryPolicy.mode: 'always', how many retries before forcing switch away; 0 means disabled5
presetsEnumWhether to auto-declare 7 built-in preset roles on apply (task / sonic / scout / designer / librarian / reviewer / security-reviewer)'bundled'
roleAutoMatchBooleanWhether to enable third-stage LLM auto-match in sub-agent's three-stage role resolution; when false, only explicit agentPreset + rules two paths remaintrue
tzString IANA TimezoneTimezone used for time slot matching; locks to Asia/Shanghai when any preset time slot exists'Asia/Shanghai'

FAQ

Q: After installation, agent behavior hasn't changed at all. Is it not taking effect?

A: No, it's in "zero-config no-op" state. Default enabled: false and rootChain is empty array, plugin actively skips intervention on all requests. Need to set enabled: true in Fallbacks settings card or YAML, plus add at least one rootChain with V4 official model as the last item.

Q: When does configuration take effect after modification?

A: Takes effect on save. Host side reads latest config through settings' onChange hook, hot-reloads role id set, maxSwitchesPerStep and all other settings; no need to restart dsh or create new session.

Q: Is it too rigid that rootChain tail must use a specific model?

A: It's intentionally designed: rootChain tail is the final fallback of the entire fallback chain, plugin binds it with virtual Auto row, expecting it to point to an official flagship/flash model that can cover all scenarios. Using other models as tail will only warn, go legacy compatibility path, but web settings card and gateway will reject saving chains with non-official tail.

Q: How do I use the new Auto in the model selector?

A: It's a virtual line registered to host LLM adapter directory (provider is FallbacksChain, model id is Auto), appears immediately when plugin is enabled. After selecting, the root agent's primary model uses the first item in rootChain for current effective window; selecting a real model maintains traditional "primary model + fallback on failure" behavior. Auto appears only when enabled: true, displays regardless of whether chain is configured; when chain is non-compliant, it only refuses to take over, won't hide.

Q: Do sub-agent sessions also need fallback chain configuration?

A: Not mandatory, but recommended: sub-agents not matching any rules default to built-in role, equivalent to directly using rootChain; to give specific sub-agents specific chains, declare a role in roles.list (with chain), then route to it via roles.rules by provider/model pattern. LLM auto-match is enabled by default, after adding roles.list will automatically pick the most suitable role for sub-agents.

Q: How to disable certain failure types from triggering fallback?

A: Remove corresponding failure codes from triggerCodes. Default only has AUTH / QUOTA / RATE_LIMIT three; 5xx class retryable errors are first handled by dsh's built-in llm-retry with fallback budget, only entering this fallback chain after exceeding, no additional declaration needed.

Q: How do dsh-tui terminal users edit configuration?

A: Find the fallbacks section in /settings screen (requires dsh-tui ≥ v0.8.5). Booleans (enabled, roleAutoMatch) render as toggles, enums (presets, revertPolicy) render as selectors, numbers render as number inputs, complex structures (rootChain, timeSlots, roles.list, roles.rules) are JSON text boxes. Invalid drafts (bad JSON, non-compliant chain, bad time slot rows) prevent saving, won't write bad config.

Q: Will upgrading to a new version lose configuration?

A: No, configuration is stored in DSH shared settings document, plugin only reads/writes its own fallbacks namespace, doesn't touch other fields. Note: after upgrading, if your old config doesn't have explicit enabled field, schema parses to false by default, plugin will suddenly become no-op, please proactively add enabled: true.

Difficulty Level

Advanced — requires understanding concepts like fallback chains, cooldown, role matching, and writing a reasonable fallbacks: YAML (or clicking through web card); "install and it works" isn't enough, must configure at least one chain to exit no-op mode.

Known Issues & Limitations

  • Sessions upgraded from <0.2.2 may fail to load: old versions persistently wrote fallbacks/switch session events, new version dsh's session reading module can't recognize (plugin and host parse different module instances), new plugin has stopped writing such events, but historical sessions need to run pnpm repair:fallbacks-switch-logs -- --apply --backup (stop dsh first) to mark old events as ignorable, will keep a .bak backup
  • rootChain tail must be official V4 model (deepseek-official/deepseek-v4-flash or deepseek-official/deepseek-v4-pro one of two), web settings card and gateway reject any other tail nodes on save; to keep non-official tail can only hand-write in YAML, will have one warning on startup, will go legacy compatibility path through fallback chain at runtime
  • As long as any kind: preset time slot exists, tz is locked to Asia/Shanghai, preset time windows are hardcoded UTC+8 constants, cannot be adjusted in UI; custom time slot rows are not subject to this restriction
  • Fallback switching within a single step cannot exceed maxSwitchesPerStep, after exceeding stops switching and preserves original error semantics; this is intentional safety valve to avoid infinite fallback loop amplifying latency
  • No config field like rootMode: whether root request goes "chain as primary model" or "primary model then chain" is determined by selecting Auto or real model in model selector, callers shouldn't look for such config in YAML
  • roleAutoMatch enabled by default causes plugin to make one extra limited LLM call on sub-agent first request (5s timeout, 32 token cap; src/automatch.ts:52-55) to pick role, network failure/timeout/no declared roles will automatically fall back to inherit, won't throw error
  • roles.rules don't match root agent sessions (PR #62 feedback): root requests can only go through rootChain; to give root agent specific chain, can only modify rootChain itself

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/omdsh-dev/dsh-llm-fallbacks)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory