Skip to main content

dsh-web-search-pro

29Stars1Forks4Issues0Watchers

Extends web search capabilities for DeepSeek Harness with multi-engine routing and fusion, 20+ platform search, SQLite persistent caching, per-site extraction like ScriptCat, and Playwright rendering.

Evidence5/5methodologySourceInstallMaintenanceDSH versionSecurity scan
Machine-auditedInstall commandRepo verifieddsh-plugin topicLicenseREADMEAI wiki
Language
TypeScript
License
MIT
Branch
master
deepseek-harnessdshdsh-pluginpluginweb-search

Install

cmdweb profile
$ dsh plugin --profile web add dsh-web-search-pro

Run the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial

Install via your agent

Install the DeepSeek Harness plugin anweat/dsh-web-search-pro for me: review the repository at https://github.com/anweat/dsh-web-search-pro first, then run the install command and verify the plugin loads successfully.

Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.

One-Line Positioning

Adds a unified "web search + scraping + snapshot" toolkit to DeepSeek Harness, replacing the single-backend native search with multi-engine routing + persistent cache + 20 platform search + ScriptCat-style per-site extraction + Playwright fallback rendering.

Core Capabilities

  • Multi-engine routing and fusion: Sequential fallback between DeepSeek native, Exa, DuckDuckGo, Bing, Jina, with RRF inverse ranking fusion for parallel queries
  • 20 platform search: GitHub / Bilibili / YouTube / V2EX / Xiaohongshu / Twitter / Reddit / Instagram / Facebook / RSS, plus Zhihu / Weibo / Douban / Tieba / Douyin / Kuaishou / arXiv / PubMed
  • Persistent cache: Search results, page snapshots, and per-site extraction rules stored in SQLite, reusable across restarts, hot data in-process via LRU
  • Readable scraping: Three-tier approach with Jina Reader → HTTP + rule extraction → Playwright fallback, automatic escalation on failure
  • Persistent page snapshots: Headless browser screenshots + HTML + text saved to local directory
  • Query history and playback: Filter history by keyword/engine/platform, replay by queryId, supports JSON export
  • Custom platforms: Add new sites via URL template + result selector in settings.yaml without code changes

Technical Implementation

  • Language: TypeScript
  • Key Dependencies: @deepseek-ai/cordis (Hook container), @deepseek-ai/dsh-tools & dsh-web (tool definitions and ctx.web abstraction), @deepseek-ai/dsh-settings (settings.yaml hot-reload segment), jsdom (HTML parsing)
  • Architecture Pattern: Cordis bundle plugin, mounts this plugin line and dsh-browser line simultaneously via cordis.patch.yml; injects three services: tools / systemPrompt / browser
  • Entry File: src/index.ts:apply(ctx, config)

Use Cases

Scenarios where you want DSH to truly "go online for research": looking up latest API docs when coding, searching open-source project issues, tracking academic papers, doing topic research, browsing Chinese community discussions. For regular DSH users encountering "model answers are outdated/can't find latest info", this plugin significantly improves information density and freshness.

Prerequisites & Compatibility

DependencyMin VersionDescription
DSH Harness0.1.0-rc.6+All @deepseek-ai/* runtime dependencies are rc.6
@anweat/dsh-browser^0.1.2Companion plugin providing browser service (this plugin's cordis.patch auto-mounts)
Node.js>=22.19.0From package.json#engines
PlatformCross-platformNo specific platform restrictions declared
Native ModulesNoneOnly Node built-ins and pure JS dependencies

Installation

dsh plugin --profile web add github:anweat/dsh-web-search-pro

Configuration Options

ConfigTypeDescriptionDefault
enginesstring arrayMulti-engine sequential fallback (or parallel fusion), available ids: seam, exa, ddg, bing, jina, github, bilibili, v2ex, youtube, arxiv, pubmed["ddg","bing","exa","seam","jina"]
parallelEnginesbooleanWhen enabled, queries all engines in parallel and fuses results with RRFfalse
ttlSecondsnumberValidity period for search results and page snapshots (seconds)3600
memoryCacheEntriesnumberIn-process LRU cache entry limit128
searchMaxResultsnumberMax results per single search8
timeoutMsnumberSingle call coordination timeout (milliseconds)30000
rrfConstantnumberRRF fusion constant k60
freshnessBoost / freshnessDaysnumberExtra scoring weight (0..1) and decay days for recent content0.2 / 30
authorityBoost / authorityDomainsnumber / string arrayExtra scoring for authoritative domains (.edu/.gov/.org etc.)0.25 / []
exaApiKey / jinaApiKeystringFill when enabling Exa / Jina; can also use env vars or .credentials.yamlNot set
exaApiKeyEnv / jinaApiKeyEnvstringEnvironment variable name referenceEXA_API_KEY / JINA_API_KEY
registerProvider / providerIdboolean / stringAfter enabling, takes over DSH's built-in web_search tool, routing through this pluginfalse / web-search-pro
enableCliBackendsbooleanWhether to allow calling external CLIs like gh / bili / yt-dlp / opencli / agent-reachtrue
opencliEnabled / agentReachEnabledbooleanIndividual toggles for opencli and agent-reach backendstrue
platformRulesdictionaryOverride result selectors by platform id (no code changes needed after site redesigns){}
customPlatformsdictionaryUser-defined sites: URL template + selector + optional Cookie{}
playwright.enabledbooleanWhether to enable Playwright fallback scraping and snapshotstrue
playwright.snapshotDirstringSnapshot save directory<dbDir>/snapshots
dbPathstringSQLite database path$DSH_HOME/data/web-search-pro/store.db
verbosebooleanAppends this load info to apply.log next to dbPath on startupfalse

FAQ

Q: What's the difference from DSH's built-in web_search?

A: The built-in tool only uses ctx.web as a single backend; this plugin upgrades it to multi-engine routing, automatic fallback, fusion ranking, plus SQLite persistent cache, query history, 20 platform search, ScriptCat-style per-site extraction, and Playwright fallback scraping with snapshots.

Q: Do I need an API Key?

A: Not required by default. Default engines ddg / bing / seam don't need keys; only when enabling Exa or Jina do you need one, which can be filled in settings.yaml or via environment variables EXA_API_KEY / JINA_API_KEY.

Q: How to search Zhihu, Xiaohongshu, Bilibili, etc.?

A: General search engines can find results, but Chinese communities (Zhihu/Weibo/Douban/Tieba/Douyin/Kuaishong/Xiaohongshu) go through Playwright-driven logged-in browser; first use scripts/save-login.mjs to log in and save cookies, then set playwright.storageStatePath in settings.yaml.

Q: What external tools need to be installed after installation?

A: Out of the box you can search general web pages. To use dedicated backends like GitHub / Bilibili / YouTube / agent-reach, first use the web_deps tool to detect, then use the detected install commands to supplement.

Q: Where is data stored?

A: Default is under $DSH_HOME/data/web-search-pro/: store.db stores history and cache, snapshots/ stores web page snapshots. Change dbPath or snapshot directory in settings.yaml's web-search-pro section.

Q: How to uninstall?

A: Use dsh plugin --profile web remove dsh-web-search-pro. SQLite database and snapshots remain in the repo directory by default, manually delete as needed.

Learning Curve

Beginner — out of the box after installation you can search general web pages; only need further configuration when enabling Exa/Jina keys, scraping Chinese communities, or customizing platforms.

Known Issues & Limitations

  • Selectors are "best effort": Chinese community platform result selectors are hardcoded, may fail to retrieve results after site redesigns, can be overridden via platformRules in settings.yaml without code changes.
  • Login state depends on external script: Zhihu/Weibo/Douban/Tieba/Douyin/Kuaishou/Xiaohongshu require users to first run scripts/save-login.mjs to log in and save cookies, otherwise returns empty results.
  • Backends relying on external CLIs need self-installation: GitHub / Bilibili / YouTube / agent-reach backends require system installation of gh / bili-cli / yt-dlp / agent-reach; opencli and playwright are bundled with dsh-browser plugin, no manual installation needed.
  • Unregister addon disabled by default: registerProvider defaults to false, DSH's built-in web_search won't automatically switch to this plugin; to take over, manually set to true or set env var DSH_WEB_SEARCH_PROVIDER=web-search-pro.

Read the usage guide →

Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.

Listing badge

Listed on deepseek-plugin.org
[![Listed on deepseek-plugin.org](https://img.shields.io/badge/listed_on-deepseek--plugin.org-007EC6)](https://deepseek-plugin.org/plugins/anweat/dsh-web-search-pro)

Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.

← Back to plugin directory