Provides Rapid-MLX local model routing for DeepSeek Harness, automatically reading server model metadata (context window, inference/tool parsers), eliminating manual settings.yaml configuration.
- Language
- JavaScript
- License
- Apache-2.0
- Branch
- main
Install
$ dsh plugin --profile web add @raullenchai/dsh-providerRun the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial
Install via your agent
Install the DeepSeek Harness plugin raullenchai/rapid-mlx-dsh-provider for me: review the repository at https://github.com/raullenchai/rapid-mlx-dsh-provider first, then run the install command and verify the plugin loads successfully.
Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.
One-Line Pitch
Adds local Rapid-MLX model routing to DeepSeek Harness (DSH), reading model facts (context window, reasoning/tool-call support, etc.) directly from the server, eliminating manual settings.yaml maintenance; includes 5 model management tools and the /rapid-mlx overview command, enabling Agents to view, download, and delete local models within a session.
Core Features
- Auto-discover model metadata: On registration, call the local Rapid-MLX server's
/v1/modelsand write context window, reasoning/tool parsers, MoE/hybrid architecture, vision capabilities, etc. into the routing - Prioritize memory-fit capacity upper limit: Use the server's
max_model_len(capacity estimated based on Apple Silicon unified memory), falling back tocontext_windowfor older server versions, ensuring correct timing for dsh-compaction-basic compression - Make reasoning tiers "tell the truth": Models where the server reports
reasoning_parser: nullno longer show off/low/medium/high selectors (they wouldn't work anyway) - Provide 5 model management tools:
rapid_mlx_serving(view models being served),rapid_mlx_cached(view download cache),rapid_mlx_pull(download model, cancellable),rapid_mlx_remove(delete cache, forced -y),rapid_mlx_health(independent API and CLI health reporting) - Register
/rapid-mlxoverview command: Single-line output of server/CLI health, current served model facts, and total cache usage
Technical Implementation
- Language: JavaScript (ESM, pure JS no build step)
- Key dependencies:
@deepseek-ai/dsh-llm(LlmAdapter base class & LlmError contract),@deepseek-ai/schemastery(Config schema validation & defaults),@deepseek-ai/dsh-tools(defineTool tool registration),@deepseek-ai/dsh-subprocess(subprocess seam for CLI execution) - Architecture pattern: Injected via
cordis.patch.ymlas a profile bundle layer; plugin injects['llm', 'tools', 'subprocess', 'commands']four services, depending ondsh.bundlefield being loaded by profile;Configdeclared via schemastery + environment variable fallbacks (RAPID_MLX_BASE_URL/RAPID_MLX_CLI) - Entry file:
lib/index.js(apply(ctx, config)is the cordis hook entry)
Use Cases
Use this plugin when DSH users want Agents to directly connect to a Rapid-MLX server running locally (typically MLX models on Apple Silicon). The most direct pain point: the generic openai-completions routing forces you to manually fill in baseURL, context window, and reasoning tiers—changing models requires editing settings; this plugin pulls metadata from the server, requires no maintenance, and follows server-side model switches while triggering dsh-compaction-basic compression at the right moment.
Prerequisites & Compatibility
| Dependency | Min Version | Notes |
|---|---|---|
| DeepSeek Harness | ^0.1.0-rc.8 (peerDependencies) | End-to-end verified on 0.1.0-rc.7; rc.8's LlmAdapter contract bytes match rc.7 |
| Node | >= 22.15.0 | README and GitHub Actions CI both enforce 22.15 (dsh uses Node Zstd stream API) |
| Cordis | ^4.0.1 | peerDependency |
| Rapid-MLX Server | Any version that can run OpenAI-compatible /v1/models | max_model_len is a vLLM/SGLang standard field; older versions with only context_window also work |
| Runtime Platform | macOS (Apple Silicon) | Rapid-MLX itself depends on Apple Silicon MLX; plugin code is pure JS cross-platform, but actual usage requires a local Rapid-MLX server |
| Native Modules | None | Only depends on Node built-in node:os.tmpdir |
Installation
dsh plugin --profile web add @raullenchai/dsh-provider
Configuration Options
| Config | Type | Description | Default |
|---|---|---|---|
baseURL | string | Root URL for local Rapid-MLX OpenAI-compatible API. Can be overridden via RAPID_MLX_BASE_URL env var | http://localhost:8000/v1 |
cliCommand | string | Executable name (PATH resolved) or absolute path for rapid-mlx CLI used for downloading/deleting/querying cache. Can be overridden via RAPID_MLX_CLI env var | rapid-mlx |
FAQ
Q: Do I need to write model config in settings.yaml after installation?
A: No need to write context window, reasoning tiers, vision capabilities, etc. Just set the agent-default-model's provider to rapid-mlx and model to the actual model name running on the server (short alias or Hugging Face repo ID both work). Other metadata is pulled by the plugin from the server.
Q: What's the difference between max_model_len and context_window?
A: context_window is the "theoretical max context" defined during model training; max_model_len is the "actual capacity limit" estimated by the server based on the machine's unified memory (weights + KV cache). The plugin prioritizes feeding the latter to the compression module to avoid compression timing being later than what the machine can actually handle.
Q: Can it run on non-Apple Silicon machines?
A: The plugin code is pure JS and cross-platform. But actual usability requires a local Rapid-MLX server (MLX backend limits to Apple Silicon). Installing the plugin on Intel Mac or Linux is fine—there's just no available server to connect to.
Q: What happens if the server responds slowly or returns 502?
A: In stream() path: non-2xx HTTP throws LlmError (401/403 → credential error codes, 429 → QUOTA, 413 → context exceeded, 5xx → PROVIDER_ERROR). When /v1/models fetch fails, it silently returns an empty list, letting DSH take the "unknown model" branch instead of breaking the entire session.
Q: Does it support tool calls and thinking chains?
A: Yes. In streaming output, tool-call increments are concatenated via argumentsDelta, and complete ToolCallBlock is given at block-end in one go; reasoning_content goes through the reasoning channel separately, not mixed into the main text. But image content throws UNSUPPORTED—it won't be silently dropped.
Q: How to uninstall?
A: dsh plugin --profile web remove @raullenchai/dsh-provider. After removal, the rapid-mlx routing, 5 tools, and /rapid-mlx command all go offline without affecting your existing settings.yaml content.
Q: Will it break if dsh upgrades to a new version?
A: As of rc.8, the LlmAdapter contract matches rc.7 byte-for-byte, so the plugin is forward compatible. dsh is a developer preview and evolves quickly—recommend watching the inject field in cordis.patch.yml (currently ['llm', 'tools', 'subprocess', 'commands']) and the Config schema for potential additions.
Onboarding Difficulty
Beginner — if you have an existing Rapid-MLX server running locally, just run dsh plugin add, and only change two lines in settings.yaml (provider and model).
Known Issues & Limitations
recommended_sampling(per-model recommended sampling parameters) is only read, not auto-appliedtool_call_parseris only read, not used for "fast-fail when model doesn't support tool calls"—may still enter loopsis_hybrid/is_moe/capabilitiesare only read, not yet used to change behavior at the routing layer- True memory-aware capacity is not yet connected:
resolveModel()currently returns the server-reported max_model_len, which is already an "estimated based on machine" value, but finer-grained "current available capacity" requires Rapid-MLX side to first expose ausable-capacityfield - Image input currently directly throws
LlmError('UNSUPPORTED'), not silently dropped; if you need image support, use the generic openai-completions provider - Routing name is fixed to
rapid-mlx—if your settings.yaml has already declared a provider with the same name underllm-pi-ai.providers,registerAdapter's provider exclusivity will cause a conflict—pick one or rename
A native Rapid-MLX provider for
DeepSeek Harness — so dsh
gets its model facts from the server instead of from whatever you typed into
settings.yaml.
Status: published to npm as
@raullenchai/dsh-provider. The end-to-enddshrun in Verified was on an M3 Ultra againstdsh 0.1.0-rc.7;dsh 0.1.0-rc.8is API-compatible — theLlmAdaptercontract is byte-identical and the only changes are additive — and the adapter is re-verified against rc.8 at the protocol and unit-test level. DSH is still a developer preview that moves fast, so treat this as tracking a moving target, not a frozen compatibility promise.
What it does for you
DSH can already talk to a local Rapid-MLX server through its generic
openai-completions provider. That route works — but it knows nothing about
your model beyond what you hand-wrote:
# what the generic route makes you maintain, by hand, per model
llm-pi-ai:
providers:
rapid-mlx:
baseURL: http://localhost:8000/v1
defaultContextWindow: 262144 # you looked this up. is it still right?
models:
- id: qwen3.6-35b-8bit
contextWindow: 262144
reasoningEfforts: {off: none, low: low, medium: medium, high: high}
Rapid-MLX's /v1/models already publishes all of that and more. This adapter
reads it, so:
1. Nothing to hand-write, and nothing to re-write when you switch models.
Swap what rapid-mlx serve is running and dsh follows. No re-running setup,
no stale numbers.
2. The reasoning control tells the truth. Rapid-MLX reports whether a model actually has a reasoning parser. A model that can't reason no longer shows an off/low/medium/high selector that does nothing.
3. Compaction is timed with the capacity that actually fits this Mac, not a
number that drifted. This is the one that quietly costs you.
dsh-compaction-basic asks the provider for the route's capacity and compacts
at thresholdRatio × capacity (0.8 by default). The provider prefers the
server's max_model_len — Rapid-MLX's memory-fitted ceiling (what fits in
unified memory: weights + KV cache), in the vLLM/SGLang-standard field — over
the native context_window, and falls back to context_window on an older
server that doesn't report it. So compaction is timed to what the machine can
actually hold, not the model's advertised window (which it may not have room
for) and not a hand-written number copied from another model.
Install
Needs Node ≥ 22.15 (dsh imports Node's Zstd stream API without declaring it) and a running Rapid-MLX server.
# From npm:
dsh plugin --profile web add @raullenchai/dsh-provider
# …or straight from source — the package ships plain JS with no build step:
dsh plugin --profile web add github:raullenchai/rapid-mlx-dsh-provider
export RAPID_MLX_BASE_URL=http://localhost:8000/v1 # optional; this is the default
dsh web
Then point the agent at the route:
# $DSH_HOME/settings.yaml
agent-default-model:
provider: rapid-mlx
model: qwen3.6-35b-8bit
Verified: that command installs and activates as a profile layer against
dsh 0.1.0-rc.7. To hack on it locally instead, see
Local development.
Model management (v0.2.0)
Beyond the provider route, the plugin registers five tools and a /rapid-mlx
command so the agent can see and manage models without leaving the session.
The split follows which surface actually owns each fact: served-model facts
come from the structured HTTP /v1/models; the download cache and pull/remove
are CLI-only, so those — and only those — shell out to rapid-mlx through the
harness subprocess seam.
| Tool | Source | What it does |
|---|---|---|
rapid_mlx_serving | HTTP /v1/models | The model(s) served right now, deduped, with context window, reasoning/tool parsers, MoE/hybrid, and modalities. |
rapid_mlx_cached | rapid-mlx models --cached | Downloaded models and their on-disk size. |
rapid_mlx_pull | rapid-mlx pull <name> | Download a model (alias or HF repo id). Cancellable; no fixed deadline. |
rapid_mlx_remove | rapid-mlx rm -y <name> | Delete a cached model to free disk. |
rapid_mlx_health | HTTP + rapid-mlx --version | API up? CLI reachable? Reported as two independent facts. |
/rapid-mlx prints a one-shot overview: health, the served model and its facts,
and total cache disk usage.
The CLI is resolved from the cliCommand config (default rapid-mlx on
PATH, or $RAPID_MLX_CLI); set it to an absolute path if the binary is not on
the harness's PATH. rapid_mlx_pull/rapid_mlx_remove are the only tools
that change anything on disk, and they run non-interactively (rm is forced
with -y because the subprocess seam ignores stdin).
Verified
Against dsh 0.1.0-rc.7 on an M3 Ultra:
- Installs and activates as a profile layer (no "declares no
dsh.bundle" warning; the entry shows up indsh --profile headless --dump-config). - Registers the
rapid-mlxroute withctx.llmand serves real queries. - Plain chat, a single tool call, and the multi-step bug-fix task that gates
Rapid-MLX releases — the last one fixed the bug and made the target repo's own
test pass, verified independently, in 36 s on
qwen3.6-35b-8bit.
Not done yet
Being explicit, because the point of the adapter is to use what the server says and some of it is still only read:
recommended_sampling— should be applied automatically per model.tool_call_parser— should letdshfail fast on a model that cannot emittool_calls, instead of looping.is_hybrid/is_moe/capabilities— read, not yet acted on.- Memory-aware capacity. Today
resolveModel()reports the model's advertised context window. On a Mac the real ceiling is unified memory, and reporting that instead is the biggest remaining win — it needs Rapid-MLX to expose a usable-capacity figure first. - Images are not carried through
stream()— but they now refuse withLlmError(..., 'UNSUPPORTED')rather than being dropped, per the cookbook. Text, reasoning and tool calls are carried. - The route is registered as
rapid-mlx. If yoursettings.yamlalso declares arapid-mlxprovider underllm-pi-ai, the two compete for one route name (registerAdapterowns provider exclusivity). Use one or rename ours.
Conformance with the official adapter contract
Built against
docs/cookbook/adding-an-llm-adapter.md
and its "protocol obligations" section. Each item has a test:
| Obligation | How it is met |
|---|---|
usage before finish, nothing after finish | usage is buffered and flushed at end-of-stream, so a trailing usage-only chunk cannot reorder it |
Tool-call arguments are raw JSON strings end to end | fragments stream as argumentsDelta and reassemble unparsed |
| Block indexes in first-seen order, reused per block | verified across a reasoning-then-text response |
| Errors take exactly two sanctioned paths | transport/protocol failures throw LlmError with a stable code; nothing ends the stream quietly |
Honor options.signal | passed to fetch and to the SSE reader; an AbortError is re-thrown unchanged, not reclassified |
A field the provider cannot honor throws UNSUPPORTED | image content refuses instead of being narrowed away |
| Config is a schemastery schema with env fallback | export const Config, fed from cordis.patch.yml via !!js process.env.RAPID_MLX_BASE_URL |
finish.replayState is not emitted: Rapid-MLX needs no native response ids
or signatures on follow-up calls, so there is nothing lossless to project.
Three things worth knowing before you edit this
Each of these cost real debugging time:
dsh.bundleinpackage.jsonis what makes this a plugin. Without it the package installs as an inert dependency anddshonly warns. It is also what gets it appended to the profile'sdsh.profile.bundles. CI fails if it goes missing.LlmReasoningEffortInfo.nameis required. Returning{id}alone fails the whole model withINVALID_MODEL_REASONING— an error that names the model, not the missing field.- DSH has no
toolrole.Message.roleis only system|user|assistant; a tool result is a user-role message whosesource.kind === 'tool'carries thecallIdand whose content holds aToolResultBlock. Flatten those into plain user text and the model reissues the same call forever — the symptom is an empty answer and a non-zero exit, with nothing on stderr.
Local development
pnpm links a local path outside the profile tree, so Node's parent-walk
never reaches $DSH_HOME/profiles/node_modules and the peer deps fail to
resolve. Symlink them in — dev only, node_modules is gitignored and excluded
from the published files:
mkdir -p node_modules/@deepseek-ai
ln -sfn <dsh-install>/node_modules/@deepseek-ai/dsh-llm node_modules/@deepseek-ai/dsh-llm
ln -sfn <dsh-install>/node_modules/@deepseek-ai/cordis node_modules/@deepseek-ai/cordis
export DSH_HOME=/tmp/dsh-dev # never your real ~/.dsh
dsh plugin --profile headless add "$PWD"
export RAPID_MLX_BASE_URL=http://127.0.0.1:8000/v1
dsh --profile headless "say hello"
A real npm install needs none of this: the package lands inside the profile
tree, where the flat fallback resolves bare names normally.
When testing agent behaviour, use a strong 8-bit model. A multi-step task here
failed on qwen3.5-9b-4bit and passed on qwen3.6-35b-8bit — 4-bit confounds
"weak model" with "broken integration".
The engine side guards these fields
Living in its own repo means a rename in Rapid-MLX would break this package
silently — nothing there imports it and this CI does not run there. So the
fields are pinned on the side that owns them, by
tests/test_model_card_client_contract.py in
Rapid-MLX, which names this package
as its reason. It pins the wire shape: field names, nullability, and the fact
that ModelInfo does not set exclude_none — which is what makes
"reasoning_parser": null distinguishable from an older server that omits the
key entirely.
If you start reading a new /v1/models field here, add it there too.
Otherwise the guard silently stops covering what this package actually uses.
License
Apache-2.0, matching Rapid-MLX.
Read the usage guide →
Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.
Listing badge
[](https://deepseek-plugin.org/plugins/raullenchai/rapid-mlx-dsh-provider)Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.