Skip to main content

rapid-mlx-dsh-provider

20Stars19Forks0Issues19Watchers

Provides Rapid-MLX local model routing for DeepSeek Harness, automatically reading server model metadata (context window, inference/tool parsers), eliminating manual settings.yaml configuration.

Evidence5/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki
Language
JavaScript
License
Apache-2.0
Branch
main
apple-siliconcoding-agentdeepseek-harnessdshdsh-pluginllmlocal-llmmlx

Install

cmdweb profile
$ dsh plugin --profile web add @raullenchai/dsh-provider

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin raullenchai/rapid-mlx-dsh-provider for me: review the repository at https://github.com/raullenchai/rapid-mlx-dsh-provider first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Line Pitch

Adds local Rapid-MLX model routing to DeepSeek Harness (DSH), reading model facts (context window, reasoning/tool-call support, etc.) directly from the server, eliminating manual settings.yaml maintenance; includes 5 model management tools and the /rapid-mlx overview command, enabling Agents to view, download, and delete local models within a session.

Core Features

  • Auto-discover model metadata: On registration, call the local Rapid-MLX server's /v1/models and write context window, reasoning/tool parsers, MoE/hybrid architecture, vision capabilities, etc. into the routing
  • Prioritize memory-fit capacity upper limit: Use the server's max_model_len (capacity estimated based on Apple Silicon unified memory), falling back to context_window for older server versions, ensuring correct timing for dsh-compaction-basic compression
  • Make reasoning tiers "tell the truth": Models where the server reports reasoning_parser: null no longer show off/low/medium/high selectors (they wouldn't work anyway)
  • Provide 5 model management tools: rapid_mlx_serving (view models being served), rapid_mlx_cached (view download cache), rapid_mlx_pull (download model, cancellable), rapid_mlx_remove (delete cache, forced -y), rapid_mlx_health (independent API and CLI health reporting)
  • Register /rapid-mlx overview command: Single-line output of server/CLI health, current served model facts, and total cache usage

Technical Implementation

  • Language: JavaScript (ESM, pure JS no build step)
  • Key dependencies: @deepseek-ai/dsh-llm (LlmAdapter base class & LlmError contract), @deepseek-ai/schemastery (Config schema validation & defaults), @deepseek-ai/dsh-tools (defineTool tool registration), @deepseek-ai/dsh-subprocess (subprocess seam for CLI execution)
  • Architecture pattern: Injected via cordis.patch.yml as a profile bundle layer; plugin injects ['llm', 'tools', 'subprocess', 'commands'] four services, depending on dsh.bundle field being loaded by profile; Config declared via schemastery + environment variable fallbacks (RAPID_MLX_BASE_URL / RAPID_MLX_CLI)
  • Entry file: lib/index.js (apply(ctx, config) is the cordis hook entry)

Use Cases

Use this plugin when DSH users want Agents to directly connect to a Rapid-MLX server running locally (typically MLX models on Apple Silicon). The most direct pain point: the generic openai-completions routing forces you to manually fill in baseURL, context window, and reasoning tiers—changing models requires editing settings; this plugin pulls metadata from the server, requires no maintenance, and follows server-side model switches while triggering dsh-compaction-basic compression at the right moment.

Prerequisites & Compatibility

DependencyMin VersionNotes
DeepSeek Harness^0.1.0-rc.8 (peerDependencies)End-to-end verified on 0.1.0-rc.7; rc.8's LlmAdapter contract bytes match rc.7
Node>= 22.15.0README and GitHub Actions CI both enforce 22.15 (dsh uses Node Zstd stream API)
Cordis^4.0.1peerDependency
Rapid-MLX ServerAny version that can run OpenAI-compatible /v1/modelsmax_model_len is a vLLM/SGLang standard field; older versions with only context_window also work
Runtime PlatformmacOS (Apple Silicon)Rapid-MLX itself depends on Apple Silicon MLX; plugin code is pure JS cross-platform, but actual usage requires a local Rapid-MLX server
Native ModulesNoneOnly depends on Node built-in node:os.tmpdir

Installation

dsh plugin --profile web add @raullenchai/dsh-provider

Configuration Options

ConfigTypeDescriptionDefault
baseURLstringRoot URL for local Rapid-MLX OpenAI-compatible API. Can be overridden via RAPID_MLX_BASE_URL env varhttp://localhost:8000/v1
cliCommandstringExecutable name (PATH resolved) or absolute path for rapid-mlx CLI used for downloading/deleting/querying cache. Can be overridden via RAPID_MLX_CLI env varrapid-mlx

FAQ

Q: Do I need to write model config in settings.yaml after installation?

A: No need to write context window, reasoning tiers, vision capabilities, etc. Just set the agent-default-model's provider to rapid-mlx and model to the actual model name running on the server (short alias or Hugging Face repo ID both work). Other metadata is pulled by the plugin from the server.

Q: What's the difference between max_model_len and context_window?

A: context_window is the "theoretical max context" defined during model training; max_model_len is the "actual capacity limit" estimated by the server based on the machine's unified memory (weights + KV cache). The plugin prioritizes feeding the latter to the compression module to avoid compression timing being later than what the machine can actually handle.

Q: Can it run on non-Apple Silicon machines?

A: The plugin code is pure JS and cross-platform. But actual usability requires a local Rapid-MLX server (MLX backend limits to Apple Silicon). Installing the plugin on Intel Mac or Linux is fine—there's just no available server to connect to.

Q: What happens if the server responds slowly or returns 502?

A: In stream() path: non-2xx HTTP throws LlmError (401/403 → credential error codes, 429 → QUOTA, 413 → context exceeded, 5xx → PROVIDER_ERROR). When /v1/models fetch fails, it silently returns an empty list, letting DSH take the "unknown model" branch instead of breaking the entire session.

Q: Does it support tool calls and thinking chains?

A: Yes. In streaming output, tool-call increments are concatenated via argumentsDelta, and complete ToolCallBlock is given at block-end in one go; reasoning_content goes through the reasoning channel separately, not mixed into the main text. But image content throws UNSUPPORTED—it won't be silently dropped.

Q: How to uninstall?

A: dsh plugin --profile web remove @raullenchai/dsh-provider. After removal, the rapid-mlx routing, 5 tools, and /rapid-mlx command all go offline without affecting your existing settings.yaml content.

Q: Will it break if dsh upgrades to a new version?

A: As of rc.8, the LlmAdapter contract matches rc.7 byte-for-byte, so the plugin is forward compatible. dsh is a developer preview and evolves quickly—recommend watching the inject field in cordis.patch.yml (currently ['llm', 'tools', 'subprocess', 'commands']) and the Config schema for potential additions.

Onboarding Difficulty

Beginner — if you have an existing Rapid-MLX server running locally, just run dsh plugin add, and only change two lines in settings.yaml (provider and model).

Known Issues & Limitations

  • recommended_sampling (per-model recommended sampling parameters) is only read, not auto-applied
  • tool_call_parser is only read, not used for "fast-fail when model doesn't support tool calls"—may still enter loops
  • is_hybrid / is_moe / capabilities are only read, not yet used to change behavior at the routing layer
  • True memory-aware capacity is not yet connected: resolveModel() currently returns the server-reported max_model_len, which is already an "estimated based on machine" value, but finer-grained "current available capacity" requires Rapid-MLX side to first expose a usable-capacity field
  • Image input currently directly throws LlmError('UNSUPPORTED'), not silently dropped; if you need image support, use the generic openai-completions provider
  • Routing name is fixed to rapid-mlx—if your settings.yaml has already declared a provider with the same name under llm-pi-ai.providers, registerAdapter's provider exclusivity will cause a conflict—pick one or rename

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/raullenchai/rapid-mlx-dsh-provider)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory