Extends web search capabilities for DeepSeek Harness with multi-engine routing and fusion, 20+ platform search, SQLite persistent caching, per-site extraction like ScriptCat, and Playwright rendering.
- Language
- TypeScript
- License
- MIT
- Branch
- master
Install
$ dsh plugin --profile web add dsh-web-search-proRun the command above in your terminal to install this plugin via the dsh CLI. You can switch Profile in the top-right corner. New to dsh? Read the beginner tutorial
Install via your agent
Install the DeepSeek Harness plugin anweat/dsh-web-search-pro for me: review the repository at https://github.com/anweat/dsh-web-search-pro first, then run the install command and verify the plugin loads successfully.
Paste this instruction to the DSH Web GUI assistant — it will install and verify for you.
One-Line Positioning
Adds a unified "web search + scraping + snapshot" toolkit to DeepSeek Harness, replacing the single-backend native search with multi-engine routing + persistent cache + 20 platform search + ScriptCat-style per-site extraction + Playwright fallback rendering.
Core Capabilities
- Multi-engine routing and fusion: Sequential fallback between DeepSeek native, Exa, DuckDuckGo, Bing, Jina, with RRF inverse ranking fusion for parallel queries
- 20 platform search: GitHub / Bilibili / YouTube / V2EX / Xiaohongshu / Twitter / Reddit / Instagram / Facebook / RSS, plus Zhihu / Weibo / Douban / Tieba / Douyin / Kuaishou / arXiv / PubMed
- Persistent cache: Search results, page snapshots, and per-site extraction rules stored in SQLite, reusable across restarts, hot data in-process via LRU
- Readable scraping: Three-tier approach with Jina Reader → HTTP + rule extraction → Playwright fallback, automatic escalation on failure
- Persistent page snapshots: Headless browser screenshots + HTML + text saved to local directory
- Query history and playback: Filter history by keyword/engine/platform, replay by queryId, supports JSON export
- Custom platforms: Add new sites via URL template + result selector in settings.yaml without code changes
Technical Implementation
- Language: TypeScript
- Key Dependencies: @deepseek-ai/cordis (Hook container), @deepseek-ai/dsh-tools & dsh-web (tool definitions and ctx.web abstraction), @deepseek-ai/dsh-settings (settings.yaml hot-reload segment), jsdom (HTML parsing)
- Architecture Pattern: Cordis bundle plugin, mounts this plugin line and dsh-browser line simultaneously via cordis.patch.yml; injects three services:
tools / systemPrompt / browser - Entry File: src/index.ts:apply(ctx, config)
Use Cases
Scenarios where you want DSH to truly "go online for research": looking up latest API docs when coding, searching open-source project issues, tracking academic papers, doing topic research, browsing Chinese community discussions. For regular DSH users encountering "model answers are outdated/can't find latest info", this plugin significantly improves information density and freshness.
Prerequisites & Compatibility
| Dependency | Min Version | Description |
|---|---|---|
| DSH Harness | 0.1.0-rc.6+ | All @deepseek-ai/* runtime dependencies are rc.6 |
| @anweat/dsh-browser | ^0.1.2 | Companion plugin providing browser service (this plugin's cordis.patch auto-mounts) |
| Node.js | >=22.19.0 | From package.json#engines |
| Platform | Cross-platform | No specific platform restrictions declared |
| Native Modules | None | Only Node built-ins and pure JS dependencies |
Installation
dsh plugin --profile web add github:anweat/dsh-web-search-pro
Configuration Options
| Config | Type | Description | Default |
|---|---|---|---|
| engines | string array | Multi-engine sequential fallback (or parallel fusion), available ids: seam, exa, ddg, bing, jina, github, bilibili, v2ex, youtube, arxiv, pubmed | ["ddg","bing","exa","seam","jina"] |
| parallelEngines | boolean | When enabled, queries all engines in parallel and fuses results with RRF | false |
| ttlSeconds | number | Validity period for search results and page snapshots (seconds) | 3600 |
| memoryCacheEntries | number | In-process LRU cache entry limit | 128 |
| searchMaxResults | number | Max results per single search | 8 |
| timeoutMs | number | Single call coordination timeout (milliseconds) | 30000 |
| rrfConstant | number | RRF fusion constant k | 60 |
| freshnessBoost / freshnessDays | number | Extra scoring weight (0..1) and decay days for recent content | 0.2 / 30 |
| authorityBoost / authorityDomains | number / string array | Extra scoring for authoritative domains (.edu/.gov/.org etc.) | 0.25 / [] |
| exaApiKey / jinaApiKey | string | Fill when enabling Exa / Jina; can also use env vars or .credentials.yaml | Not set |
| exaApiKeyEnv / jinaApiKeyEnv | string | Environment variable name reference | EXA_API_KEY / JINA_API_KEY |
| registerProvider / providerId | boolean / string | After enabling, takes over DSH's built-in web_search tool, routing through this plugin | false / web-search-pro |
| enableCliBackends | boolean | Whether to allow calling external CLIs like gh / bili / yt-dlp / opencli / agent-reach | true |
| opencliEnabled / agentReachEnabled | boolean | Individual toggles for opencli and agent-reach backends | true |
| platformRules | dictionary | Override result selectors by platform id (no code changes needed after site redesigns) | {} |
| customPlatforms | dictionary | User-defined sites: URL template + selector + optional Cookie | {} |
| playwright.enabled | boolean | Whether to enable Playwright fallback scraping and snapshots | true |
| playwright.snapshotDir | string | Snapshot save directory | <dbDir>/snapshots |
| dbPath | string | SQLite database path | $DSH_HOME/data/web-search-pro/store.db |
| verbose | boolean | Appends this load info to apply.log next to dbPath on startup | false |
FAQ
Q: What's the difference from DSH's built-in web_search?
A: The built-in tool only uses ctx.web as a single backend; this plugin upgrades it to multi-engine routing, automatic fallback, fusion ranking, plus SQLite persistent cache, query history, 20 platform search, ScriptCat-style per-site extraction, and Playwright fallback scraping with snapshots.
Q: Do I need an API Key?
A: Not required by default. Default engines ddg / bing / seam don't need keys; only when enabling Exa or Jina do you need one, which can be filled in settings.yaml or via environment variables EXA_API_KEY / JINA_API_KEY.
Q: How to search Zhihu, Xiaohongshu, Bilibili, etc.?
A: General search engines can find results, but Chinese communities (Zhihu/Weibo/Douban/Tieba/Douyin/Kuaishong/Xiaohongshu) go through Playwright-driven logged-in browser; first use scripts/save-login.mjs to log in and save cookies, then set playwright.storageStatePath in settings.yaml.
Q: What external tools need to be installed after installation?
A: Out of the box you can search general web pages. To use dedicated backends like GitHub / Bilibili / YouTube / agent-reach, first use the web_deps tool to detect, then use the detected install commands to supplement.
Q: Where is data stored?
A: Default is under $DSH_HOME/data/web-search-pro/: store.db stores history and cache, snapshots/ stores web page snapshots. Change dbPath or snapshot directory in settings.yaml's web-search-pro section.
Q: How to uninstall?
A: Use dsh plugin --profile web remove dsh-web-search-pro. SQLite database and snapshots remain in the repo directory by default, manually delete as needed.
Learning Curve
Beginner — out of the box after installation you can search general web pages; only need further configuration when enabling Exa/Jina keys, scraping Chinese communities, or customizing platforms.
Known Issues & Limitations
- Selectors are "best effort": Chinese community platform result selectors are hardcoded, may fail to retrieve results after site redesigns, can be overridden via platformRules in settings.yaml without code changes.
- Login state depends on external script: Zhihu/Weibo/Douban/Tieba/Douyin/Kuaishou/Xiaohongshu require users to first run scripts/save-login.mjs to log in and save cookies, otherwise returns empty results.
- Backends relying on external CLIs need self-installation: GitHub / Bilibili / YouTube / agent-reach backends require system installation of gh / bili-cli / yt-dlp / agent-reach; opencli and playwright are bundled with dsh-browser plugin, no manual installation needed.
- Unregister addon disabled by default: registerProvider defaults to false, DSH's built-in web_search won't automatically switch to this plugin; to take over, manually set to true or set env var DSH_WEB_SEARCH_PROVIDER=web-search-pro.
增强型、可持久化的扩展网页搜索插件 for DeepSeek Harness(DSH)。
一个 DSH bundle 插件,把多引擎网页搜索、平台搜索、持久化缓存、脚本猫式按站提取、Playwright 渲染打包成模型可直接调用的 9 个工具。灵感来自 MediaCrawler、Agent-Reach、脚本猫/油猴 userscript、opencli 与 playwright。
安装
dsh plugin --profile web add dsh-web-search-pro # 自动装 dsh-browser(dependency)+ 自动挂载 browser 行(本 patch)
# 或本地目录 / tarball:
dsh plugin --profile web add ./dsh-web-search-pro
# 重启(web profile 关闭了 HMR):
dsh --profile web
dsh-browser 需先发布到 npm(本地测试可用
dsh plugin --profile web add ../dsh-browser ../dsh-web-search-pro一条命令显式列两个)。 依赖@deepseek-ai/*已发布到 npm(^0.1.0-rc.6,与社区 dsh-cc-tui 一致)。 若你的 harness 是本地源码 checkout(如0.1.0-rc.5),版本号可能有出入——用dsh plugin --profile web add ./<path>并在 profile 的pnpm-workspace.yaml里对齐版本后重装即可。
工具(9 个)
| 工具 | 作用 |
|---|---|
web_search_pro | 多引擎搜索 + RRF 融合 + 内存/SQLite 双层缓存 + 历史 |
web_fetch_pro | 可读化抓取(Jina → HTTP+规则抽取 → Playwright 兜底)+ 快照缓存 |
web_platform_search | 20 平台:GitHub/B站/YouTube/V2EX/小红书/Twitter/Reddit/IG/FB/RSS + 知乎/微博/豆瓣/贴吧/抖音/快手(Playwright 登录态) |
web_snapshot | Playwright 全页截图 + HTML + 文本落盘 |
web_history / web_cache_clear / web_search_stats | 持久历史 / 清缓存 / 存储统计 |
web_rule | 持久化按站提取规则(脚本猫式,list/upsert/remove) |
web_deps | 检测/安装外部依赖(gh/bili/yt-dlp/opencli/agent-reach/mcporter/playwright) |
配置
三层,越靠前越日常:
-
$DSH_HOME/settings.yaml→web-search-pro:段(热重载,改完即生效):web-search-pro: exaApiKey: 'sk-...' # 或环境变量 EXA_API_KEY / .credentials.yaml jinaApiKey: 'jina_...' # 或环境变量 JINA_API_KEY engines: [ddg, bing, exa, seam, jina] parallelEngines: false ttlSeconds: 3600 searchMaxResults: 8 -
cordis.yml
config:(部署级默认值,见cordis.patch.yml)。 -
环境变量 / 凭据:
$EXA_API_KEY、$JINA_API_KEY(exaApiKeyEnv/jinaApiKeyEnv引用)。
外部依赖(按需)
多数后端需要系统额外安装的工具;插件提供 web_deps 工具检测与安装:
| 依赖 | 用途 | 安装 |
|---|---|---|
| gh | GitHub 后端 | winget install GitHub.cli / choco install gh |
| bili-cli | B站后端 | uv tool install bili-cli / pipx install bili-cli |
| yt-dlp | YouTube 后端 | uv tool install yt-dlp / pip install yt-dlp |
| opencli | 小红书/Twitter/Reddit/IG/FB | npm i -g opencli |
| agent-reach | agent-reach 后端 | uv tool install agent-reach / pip install agent-reach |
| playwright | 渲染/截图后端 | npm i -g playwright && playwright install chromium |
平台与引擎
seam(ctx.web/DeepSeek 原生)· exa · ddg · bing · jina · github · bilibili · v2ex · youtube。默认顺序 ddg, bing, exa, seam, jina(免费优先),失败自动回退;multi 并行融合。
开发
pnpm install
pnpm build # tsc src → lib
源码在 src/;lib/ 为发布产物(已提交)。
License
MIT
中文社区平台登录态
zhihu / weibo / douban / tieba / douyin / kuaishou 的免登录公开接口都被风控, 所以走 Playwright 驱动登录态浏览器(借鉴 MediaCrawler 思路、MIT 独立实现,未用其签名算法):
- 登录一次保存登录态:
node scripts/save-login.mjs all login-state.json - 在
$DSH_HOME/settings.yaml里设playwright.storageStatePath - 站点改版时无需改代码,用
platformRules按平台覆盖结果选择器
详见 LOGIN.md。
历史管理
web_history 支持:kind/query/engine/platform 过滤、replay(用 queryId 回放已存结果)、 export(把过滤后的历史+结果写成 JSON 文件)。
自定义平台(解析 cookie 去搜索)
在 settings.yaml 里定义任意站点(URL 模板 + 结果选择器 + 可选 Cookie), web_platform_search 就能直接搜它——不需要改代码:
web-search-pro:
customPlatforms:
mybili:
name: '我的B站'
url: 'https://search.bilibili.com/all?keyword={query}'
item: '.bili-video-card'
title: '.bili-video-card__info--tit'
link: 'a'
# 需要登录的站点补 cookie(a=b; c=d,自动应用到 url 域名)
myforum:
name: '某论坛'
url: 'https://forum.example.com/search?q={query}'
item: '.thread'
title: '.thread-title a'
link: '.thread-title a'
cookie: 'sessionid=abc123; csrftoken=xyz'
Read the usage guide →
Install steps, key points, FAQ and compatibility for this plugin — auto-derived from indexed fields.
Listing badge
[](https://deepseek-plugin.org/plugins/anweat/dsh-web-search-pro)Paste this markdown into your GitHub README to link back to this listing. The badge only states the listing — not a security endorsement.